You know what's worse than a bug in production?
A bug you introduced while fixing another bug — deployed at midnight, right when your biggest traffic spike ever is happening.
The Setup: Our Best Day Ever
I run featured us in their newsletter.
Traffic exploded:
| Day | Active Users | New Users |
|---|---|---|
| March 19 (feature day) | 448 | 445 |
| March 20 (day after) | 154 | 139 |
448 users. For a solo indie SaaS, that felt massive.
But there was a problem I wouldn't discover until 24 hours later.
The "Fix" That Broke Everything
Earlier on March 19, I noticed that some video generations were failing because Remotion Lambda (our video renderer on AWS) couldn't download images from fal.ai's temporary CDN URLs fast enough. The URLs were expiring or timing out.
So at midnight (00:20 JST, March 20), I shipped what I thought was a solid improvement:
Pre-fetch all media from fal.ai CDN → Supabase Storage before rendering
Add retry logic — 3 attempts with exponential backoff
Throw early if all clips failed to prefetch, instead of sending unreliable URLs to the renderer
That third point was the killer. Here's the diff:
// video-prefetch.ts — the "improvement"
const cached = results.filter((r, i) => r.url !== videos[i].url).length;
if (cached === 0 && videos.length > 0) {
throw new Error(
`All ${videos.length} video clips failed to prefetch from CDN`
);
}
My reasoning: "If we can't cache any clips locally, Remotion will probably fail anyway. Let's fail fast and give the user a clear error instead of wasting 5 minutes on a doomed render."
It sounded so reasonable.
The Data I Should Have Checked First
The next evening, I ran a query against our projects table:
SELECT status, content_mode, COUNT(*) as cnt
FROM projects
WHERE created_at >= '2026-03-19T15:00:00+00:00' -- after deploy
AND created_at < '2026-03-20T15:00:00+00:00'
GROUP BY status, content_mode;
| Status | Mode | Count |
|---|---|---|
| failed | video_short | 10 |
| completed | image | 3 |
Zero. Not a single video succeeded after my deploy.
Every single failure had the same message:
All 3 video clips failed to prefetch from CDN — rendering would likely fail
Meanwhile, before my fix on the same day:
| Status | Mode | Count |
|---|---|---|
| completed | video_short | 22 |
| failed | video_short | 7 |
| completed | image | 6 |
| failed | image | 8 |
76% video success rate before. 0% after. My "safety check" didn't just fail to help — it made things infinitely worse.
Why My Assumption Was Wrong
Here's what I didn't understand about my own infrastructure:
The prefetch path (broken):
Inngest (Vercel Edge) → fal.media CDN → frequently times out
The render path (working fine):
Remotion Lambda (AWS us-east-1) → fal.media CDN → usually succeeds
The prefetch runs on Vercel's serverless infrastructure. Remotion Lambda runs on AWS in us-east-1. The network path from AWS to fal.ai's CDN was far more reliable than from Vercel.
Before my fix, when prefetch failed, the code silently fell back to the original fal.media URLs. Remotion Lambda would then download them directly — and it usually worked. The 7 pre-fix failures were mostly unrelated issues (prompt rejections, Lambda crashes), not CDN problems.
By adding the throw, I cut off the fallback path entirely. The pipeline would die at the prefetch step without ever giving Remotion a chance to try.
The Human Cost
6 unique users were affected. All on the free plan, all trying video generation for the first time:
| User | Failed Attempts |
|---|---|
turns GitHub repos into promotional videos. It works again now. Probably. 😅 Have you ever shipped a "fix" that made things worse? I'd love to hear your story in the comments. Vollständiger Original-Bericht Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to. Wie bewertest du diesen Beitrag? 1 Klick Feedback Teilen mit Netzwerk & Team: Hat Ihnen dieser Tipp / Anleitung geholfen? Community-Analysen & Experten-Meinungen 0Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog. Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf „ Eigene Analyse verfassen“! Community Pulse: Relevanz-Einschätzung 1 Klick Experten-Votum 🔴 Akute Relevanz 0% 🟡 In Evaluierung 0% 🟢 Keine Auswirkung 0% Spannende Innovation 0% Verwandte Story-Cluster & Quellen (Vektor-KI) Tipp: Mit Pfeiltasten [ ← ] und [ → ] blättern
Ähnliche Beiträge
🔍 Verwandte News
Auch interessante Nachrichten My Safety Check Killed 100% of Video Generations — Right When Traffic Spiked 3xThematisch verwandte Begriffe: Safety, Check, Killed, Video · 6 Treffer 🔧 AI Nachrichten DZone.com Feed Building Agentic RAG, Step by Step: From Static Retrieval to Reasoning Pipelines ⚠️ Malware / Trojaner / Viren Elastic Security Labs Stopping Vulnerable Driver Attacks
Laden...
Videos werden geladen ...
Laden...
Beiträge werden geladen ...
Laden...
Videos werden geladen ...
Laden...
Beiträge werden geladen ...
Laden...
Videos werden geladen ...
Laden...
Beiträge werden geladen ...
Laden...
Videos werden geladen ...
Laden...
Beiträge werden geladen ...
Laden...
Videos werden geladen ... 🔖 Gespeicherte Artikel
📂
Keine gespeicherten Artikel vorhanden.
📂 News
⏱️ 3 Min
vor 10 Min
Artikeldaten werden geladen...
tsecurity.de AppOffline-Lesen, Eilmeldungen & 0ms Ladezeit
Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.
Nächster Beitrag
🤖
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster:
Security Explorer
Match:
lädt…
Aktivitäten deiner Analystenlädt…
Neues Thema oder Eilmeldung einreichenReiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung. Heiß diskutierte Einreichungen |
SOCIAL SHARE CARD GENERATOR