🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 3 Min Lesezeit
0

Building a Universal Property Listing Scraper with Python and JSON-LD

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




Building a Universal Property Listing Scraper with Python and JSON-LD



Ever wanted to extract structured data from real estate listings across multiple sites without writing custom parsers for each one? I built a Property Listing Scraper that works with any property website — from Zillow to Rightmove to Imobiliare — using a combination of JSON-LD, OpenGraph metadata, and smart HTML pattern matching.






The Problem



Every real estate site structures its data differently. Zillow uses one schema, Rightmove uses another, and smaller sites like Imobiliare.ro have their own format entirely. Building separate scrapers for each is a maintenance nightmare.






The Solution: Multi-Layer Extraction



The scraper uses three extraction layers, falling back gracefully:






1. JSON-LD (Structured Data)



Many modern property sites embed structured data using Schema.org vocabulary. This is the gold standard — clean, machine-readable, and standardised.




CODE
for script in soup.find_all("script", type="application/ld+json"):
data = json.loads(script.string)
if "Product" in str(data.get("@type", "")):
# Extract price, address, coordinates, images









2. OpenGraph Meta Tags



When JSON-LD isn't available, we fall back to OpenGraph metadata that most sites provide for social sharing.




CODE
og_title = soup.find("meta", property="og:title")
og_image = soup.find("meta", property="og:image")









3. Regex Pattern Matching



As a final fallback, we scan the page text for common price and property patterns:




CODE
price_match = re.search(r"([$£€]\s*[\d,]+(?:\.\d+)?)", text)
bed_match = re.search(r"(\d+)\s*(?:bed|bedroom|camera)", text)









Extracted Data Fields



Each property listing yields:





  • Price & Currency — parsed from any format (USD, EUR, GBP, RON)


  • Property Type — apartment, house, studio, land, office


  • Bedrooms & Bathrooms — multi-language support (English, Italian, Romanian)


  • Area — square metres detection


  • Address Components — street, city, region, postal code, country


  • Coordinates — latitude/longitude from geo data


  • Images — up to 20 high-quality images per listing






Search Page Support



The scraper automatically detects whether a URL is:





  1. An individual listing — extracts data directly


  2. A search results page — discovers listing links, then scrapes each one concurrently (5 parallel requests)






Try It



You can use the scraper right now:





The Apify actor uses pay-per-event pricing at $0.01 per property extracted — you only pay for actual results.






Use Cases





  • Market Analysis: Track property prices across multiple markets


  • Lead Generation: Build databases of available properties


  • Price Monitoring: Watch specific neighbourhoods for price changes


  • Investment Research: Compare yields across countries and currencies






Conclusion



By combining JSON-LD extraction with meta tags and regex fallbacks, we achieve universal compatibility without sacrificing data quality. The scraper handles the messy reality of real estate websites so you can focus on analysing the data.






Built with Python, BeautifulSoup4, httpx, and the Apify platform.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Building a Universal Property Listing Scraper with Python and JSON-LD

Thematisch verwandte Begriffe: Building, Universal, Property, Listing · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...