1. Setup and Installation
Before you start, ensure you have:
Python: Version 3.7 or higher installed.
Libraries: InstallrequestsandBeautifulSoup.
pip install requests beautifulsoup4
Editor: Any Python-supported IDE, such as VS Code or PyCharm.
2. Analyzing GitHub HTML Structure
To scrape GitHub folders, you need to understand the HTML structure of a repository page. On a GitHub repository page:
Folders are linked with paths like/tree/<branch>/<folder>.
Files are linked with paths like/blob/<branch>/<file>.
Each item (folder or file) is inside a <div> with the attribute role="rowheader" and contains an <a> tag. For example:
<div role="rowheader">
<a href="/owner/repo/tree/main/folder-name">folder-name</a>
</div>
3. Implementing the Scraper
3.1. Recursive Crawling Function
The script will recursively scrape folders and print their structure. To limit the recursion depth and avoid unnecessary load, we’ll use a depth parameter.
import requests
from bs4 import BeautifulSoup
import time
def crawl_github_folder(url, depth=0, max_depth=3):
"""
Recursively crawls a GitHub repository folder structure.
Parameters:
- url (str): URL of the GitHub folder to scrape.
- depth (int): Current recursion depth.
- max_depth (int): Maximum depth to recurse.
"""
if depth > max_depth:
return
headers = {"User-Agent": "Mozilla/5.0"}
response = requests.get(url, headers=headers)
if response.status_code != 200:
print(f"Failed to access {url} (Status code: {response.status_code})")
return
soup = BeautifulSoup(response.text, 'html.parser')
# Extract folder and file links
items = soup.select('div[role="rowheader"] a')
for item in items:
item_name = item.text.strip()
item_url = f"https://github.com{item['href']}"
if '/tree/' in item_url:
print(f"{' ' * depth}Folder: {item_name}")
crawl_github_folder(item_url, depth + 1, max_depth)
elif '/blob/' in item_url:
print(f"{' ' * depth}File: {item_name}")
# Example usage
if __name__ == "__main__":
repo_url = "https://github.com/<owner>/<repo>/tree/<branch>/<folder>"
crawl_github_folder(repo_url)
4. Features Explained
Headers for Request: Using aUser-Agentstring to mimic a browser and avoid blocking.
Recursive Crawling:
- Detects folders (
/tree/) and recursively enters them. - Lists files (
/blob/) without entering further.
- Detects folders (
Indentation: Reflects folder hierarchy in the output.
Depth Limitation: Prevents excessive recursion by setting a maximum depth (max_depth).
5. Enhancements
5.1. Exporting Results
Save the output to a structured JSON file for easier usage.
import json
def crawl_to_json(url, depth=0, max_depth=3):
"""Crawls and saves results as JSON."""
result = {}
if depth > max_depth:
return result
headers = {"User-Agent": "Mozilla/5.0"}
response = requests.get(url, headers=headers)
if response.status_code != 200:
print(f"Failed to access {url}")
return result
soup = BeautifulSoup(response.text, 'html.parser')
items = soup.select('div[role="rowheader"] a')
for item in items:
item_name = item.text.strip()
item_url = f"https://github.com{item['href']}"
if '/tree/' in item_url:
result[item_name] = crawl_to_json(item_url, depth + 1, max_depth)
elif '/blob/' in item_url:
result[item_name] = "file"
return result
if __name__ == "__main__":
repo_url = "https://github.com/<owner>/<repo>/tree/<branch>/<folder>"
structure = crawl_to_json(repo_url)
with open("output.json", "w") as file:
json.dump(structure, file, indent=2)
print("Repository structure saved to output.json")
5.2. Error Handling
Add robust error handling for network errors and unexpected HTML changes:
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
except requests.exceptions.RequestException as e:
print(f"Error fetching {url}: {e}")
return
5.3. Rate Limiting
To avoid being rate-limited by GitHub, introduce delays:
import time
def crawl_with_delay(url, depth=0):
time.sleep(2) # Delay between requests
# Crawling logic here
6. Ethical Considerations
Compliance: Adhere to GitHub’s Terms of Service.
Minimize Load: Respect GitHub’s servers by limiting requests and adding delays.
Permission: Obtain permission for extensive crawling of private repositories.
7. Complete Code
Here’s the consolidated script with all features included:
import requests
from bs4 import BeautifulSoup
import json
import time
def crawl_github_folder(url, depth=0, max_depth=3):
result = {}
if depth > max_depth:
return result
headers = {"User-Agent": "Mozilla/5.0"}
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
except requests.exceptions.RequestException as e:
print(f"Error fetching {url}: {e}")
return result
soup = BeautifulSoup(response.text, 'html.parser')
items = soup.select('div[role="rowheader"] a')
for item in items:
item_name = item.text.strip()
item_url = f"https://github.com{item['href']}"
if '/tree/' in item_url:
print(f"{' ' * depth}Folder: {item_name}")
result[item_name] = crawl_github_folder(item_url, depth + 1, max_depth)
elif '/blob/' in item_url:
print(f"{' ' * depth}File: {item_name}")
result[item_name] = "file"
time.sleep(2) # Avoid rate-limiting
return result
if __name__ == "__main__":
repo_url = "https://github.com/<owner>/<repo>/tree/<branch>/<folder>"
structure = crawl_github_folder(repo_url)
with open("output.json", "w") as file:
json.dump(structure, file, indent=2)
print("Repository structure saved to output.json")
Key Notes
Recursive Depth: Thedepthparameter prevents infinite recursion.
Rate Limiting: Avoid rapid requests to prevent IP bans. Use time delays (time.sleep) if necessary.
HTML Updates: Adjust selectors (div[role="rowheader"] a) as GitHub updates its structure.
Enhancements
Export Results: Save crawled data to JSON or CSV.
Parallel Requests: Use libraries likeasyncioorconcurrent.futuresfor faster crawling.
Error Handling: Handle network issues and retries gracefully.
By following this guide, you can efficiently crawl folder structures on GitHub repositories programmatically. Adapt the solution for your specific requirements while adhering to ethical practices.
SOCIAL SHARE CARD GENERATOR