🔧 AI Nachrichten Debian is Voting on Whether to Allow AI-Assisted Contributions(23.08.2026 um 09:34 Uhr)
🔧 AI Nachrichten The Linux Kernel Is Approaching 2,000 CVEs Per Release(29.08.2026 um 20:00 Uhr)
⚠️ Malware / Trojaner / VirenCitrix Adds a Linux-Powered Escape Hatch For Compromised Windows PCs(30.08.2026 um 17:34 Uhr)
🔧 AI Nachrichten Debian is Voting on Whether to Allow AI-Assisted Contributions(23.08.2026 um 09:34 Uhr)
🔧 AI Nachrichten The Linux Kernel Is Approaching 2,000 CVEs Per Release(29.08.2026 um 20:00 Uhr)
⚠️ Malware / Trojaner / VirenCitrix Adds a Linux-Powered Escape Hatch For Compromised Windows PCs(30.08.2026 um 17:34 Uhr)

🔧 Programmierung 🕛 vor 1 Monat 6 Min Lesezeit
0

Your HTML is fine. The CDN still blocks the bot.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

This started as a comment

AI-crawler response codes. Source: Cloudflare Radar, AI Insights, 7-day view, 18 Jul 2026.



Read that again. More than one in fifteen requests from AI crawlers is being actively refused at the edge — 403 or 429 — before it ever touches the content. Not deprioritised. Refused.



And most of that isn't a decision. It's a default — a managed WAF ruleset, a "block AI bots" toggle flipped in 2024, a bot-fight score set too high. The same dashboard's robots.txt tracker shows GPTBot, CCBot and ClaudeBot as the most-disallowed crawlers across the top domains — much of it inherited, not authored.





AI bot transparency tracking. Source: Cloudflare Radar, AI Insights, as of 18 Jul 2026.



So real verification isn't the name. It's the origin. Two ways:



1. Verify by published IP ranges. OpenAI, Anthropic and the rest publish their crawler IPs — openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json, and their equivalents. Reverse-DNS the requesting IP, confirm the hostname belongs to the vendor, then forward-resolve it back to the same IP. If either step fails, it's a spoofer wearing the name.



2. Verify by cryptographic signature. This is where it's heading. Web Bot Auth — a Cloudflare-led IETF draft (draft-meunier-web-bot-auth-architecture, built on RFC 9421) — has crawlers sign each request with an Ed25519 key, so identity is proven, not claimed. It's integrated into Cloudflare's Verified Bots program with a dedicated Signed Agents directory, and ChatGPT's agent was in the first signed cohort in 2025. It isn't only Cloudflare, either — AWS WAF added Web Bot Auth support in late 2025, auto-allowing verified agents.



Note the status honestly: Web Bot Auth is an active IETF draft, not a ratified standard. But with Cloudflare, OpenAI, Anthropic and AWS moving in lockstep, it's already the de-facto direction.



The point for you: "block the bad bots" and "let the AI in" are the same problem, and the User-Agent solves neither. Verify by origin or by signature — otherwise you block the crawler you want and admit the one you don't.






What to actually do



Check the logs, not the browser. Pull a week of access logs, filter to the AI User-Agents, look at the status codes. That's the whole audit — an afternoon, not a project.



If you see 403/429, find the rule. It's almost always a managed ruleset, a "block AI bots" setting, or an over-eager bot-fight score — not a line you wrote. Allowlist the crawlers you want by verified IP or Web Bot Auth, never by User-Agent alone.






The rule I took away



Your HTML is the last thing in the request's path, not the first. Before a bot reads a single byte of it, the request has to clear DNS, the CDN, the WAF, the rate limiter, the bot-management score. Any one of them can end the request with a status code you never see — on a page that renders flawlessly for you.



"It works in my browser" was always a weak claim. For crawlers it isn't even the right question. The right question is: what status code did the bot get? And the only honest answer is in the logs.






Related reading:





Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Bits und so #1021 (Passwort für Laufwerk)
1 Quelle
Bits und so #1022 (Wie Weißbier)
1 Quelle
KI-Agenten entdecken deutsches Wiki als Kommunikationskanal
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Your HTML is fine. The CDN still blocks the bot.

Thematisch verwandte Begriffe: Your, HTML, fine, still · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...