Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Sichere ProgrammierungWhy Claude Code keeps writing shell commands that fail on your Mac(20.09.2026 um 21:06 Uhr)
Sichere Programmierungllms.txt v2: What the Spec Says, and What 137,000 Domains Show(20.09.2026 um 21:17 Uhr)
Sicherheitslücken (CVE)NiceTryGPT: Less pattern matching. More actual hacking.(20.09.2026 um 21:19 Uhr)
IT Security VideoActivities BoF (kde2026)(20.09.2026 um 00:00 Uhr)
IT Security Toolsirdoc-app(20.09.2026 um 20:33 Uhr)
Sichere ProgrammierungWhy Claude Code keeps writing shell commands that fail on your Mac(20.09.2026 um 21:06 Uhr)
Sichere Programmierungllms.txt v2: What the Spec Says, and What 137,000 Domains Show(20.09.2026 um 21:17 Uhr)
Sicherheitslücken (CVE)NiceTryGPT: Less pattern matching. More actual hacking.(20.09.2026 um 21:19 Uhr)
IT Security VideoActivities BoF (kde2026)(20.09.2026 um 00:00 Uhr)
IT Security Toolsirdoc-app(20.09.2026 um 20:33 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Scraping YellowPages with Python in 2026: What Actually Works (and What Doesn't)

Reagiere als Erste:r — dein Feedback zählt!

I spent a week trying to scrape YellowPages.com with Python. Most tutorials online are from 2021-2023 and silently fail now. Here's what I found.

The tutorials are all broken

If you Google "scrape yellowpages python", you'll find guides using requests + BeautifulSoup. They look clean, they make sense, and they don't work.

import requests
from bs4 import BeautifulSoup

url = "https://www.yellowpages.com/search?search_terms=plumbers&geo_location_terms=Austin%2C+TX"
resp = requests.get(url)
# resp.status_code == 403. Every time.

YellowPages.com moved behind Cloudflare sometime in 2023. Every request now passes through a JavaScript challenge. requests can't execute JavaScript, so it gets a 403 or an empty challenge page. Same with httpx, urllib3, or any pure HTTP library.

What about Selenium/Playwright?

Headless browsers can execute the JS challenge:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://www.yellowpages.com/search?search_terms=plumbers&geo_location_terms=Austin%2C+TX")
    page.wait_for_selector(".search-results")
    html = page.content()

This works... sometimes. Three problems kill it at scale:

1. IP blocks. Default Chromium uses your IP. After 10-20 requests, Cloudflare blocks you. You need residential proxies, and specifically US ones -- non-US IPs get rejected.

2. Browser fingerprinting. Headless Chrome has detectable fingerprint differences. Cloudflare catches navigator.webdriver=true, missing plugins, and other signals. You need stealth patches.

3. Session reuse kills extraction. If you navigate to 50 business detail pages in one browser session, Cloudflare starts returning challenge pages instead of data after the 8th or 9th page. You need fresh browser contexts per page.

The proxy problem

Even with Playwright + stealth, you need proxies. Not just any proxies -- residential US proxies specifically.

Free proxy lists? Dead within hours. Datacenter proxies? Blocked instantly. Shared residential proxies? Rate-limited because 100 other scrapers are on the same pool.

You'll spend $50-100/month on a residential proxy service just to keep the scraper running. Then you need rotation logic, error handling for burned IPs, and retry logic.

What I actually use now

After fighting this for a week, I switched to using a pre-built scraper that handles all of this: Yellow Pages Scraper on Apify.

It handles:

  • Cloudflare bypass with residential proxies (built in)
  • Stealth browser automation (no fingerprint detection)
  • Fresh browser sessions per detail page
  • Automatic retries on blocked requests
  • Structured output (JSON, CSV, Excel)

You put in search terms + locations and get back clean data:

{
  "name": "Joe's Plumbing LLC",
  "phone": "(512) 555-0142",
  "email": "[email protected]",
  "website": "https://joesplumbing.com",
  "address": "1234 Main St, Austin, TX 78701",
  "rating": 4.5,
  "reviewCount": 47,
  "categories": ["Plumbers", "Water Heaters"],
  "yearsInBusiness": 12
}

Cost: about $0.005 per result. A run of 100 businesses costs $0.60. Cheaper than maintaining your own proxy subscription.

When to roll your own

Building your own scraper makes sense if:

  • You need custom post-processing that can't be done after export
  • You're scraping at massive scale (10K+ results/day) and want to optimize costs
  • You need real-time streaming rather than batch results
  • You want to learn how browser automation and anti-bot bypass work

If you just need business leads from YellowPages, the pre-built tool saves you a week of proxy debugging.

Key technical lessons if you do build it yourself

  1. Use Go or Playwright, not Selenium. Selenium is slower and has more detectable fingerprints.
  2. Residential US proxies are mandatory. Budget $50-100/month minimum.
  3. Fresh browser context per detail page. Session reuse drops email extraction from ~22% to ~3%.
  4. Parse JSON-LD, not HTML. YellowPages embeds structured data as application/ld+json. It's cleaner and more stable than CSS selectors.
  5. Respect rate limits. 1-3 seconds between requests. Faster than that triggers immediate blocks.

I'm a data engineer building automation tools for lead generation. If you have questions about scraping public business directories, drop a comment.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Scraping YellowPages with Python in 2026: What Actually Works (and What Doesn't)

Thematisch verwandte Begriffe: Scraping, YellowPages, with, Python · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-93956 | A flaw has been found in olivier-ls PHP-FTS up to 1.1.2. Affected by thi…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick