Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Unix & Linux ServerPeppermintOS Is Moving From Xorg to XLibre to Avoid Wayland(22.09.2026 um 19:58 Uhr)
Sicherheitslücken (CVE)USN-8803-1: Sudo vulnerability(22.09.2026 um 16:15 Uhr)
Sichere ProgrammierungClaude Opus 5.5 is now available in GitHub Copilot(22.09.2026 um 19:10 Uhr)
Sichere ProgrammierungColab is now part of your Google AI plan(22.09.2026 um 20:51 Uhr)
Sichere ProgrammierungThe Hidden Production Risks of Third-Party SDKs(22.09.2026 um 20:00 Uhr)
Unix & Linux ServerPeppermintOS Is Moving From Xorg to XLibre to Avoid Wayland(22.09.2026 um 19:58 Uhr)
Sicherheitslücken (CVE)USN-8803-1: Sudo vulnerability(22.09.2026 um 16:15 Uhr)
Sichere ProgrammierungClaude Opus 5.5 is now available in GitHub Copilot(22.09.2026 um 19:10 Uhr)
Sichere ProgrammierungColab is now part of your Google AI plan(22.09.2026 um 20:51 Uhr)
Sichere ProgrammierungThe Hidden Production Risks of Third-Party SDKs(22.09.2026 um 20:00 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Failure Engineering Explained by Uncle to Nephew — Episode 5: Recovery — How Systems Heal Themselves

Episode 4 covered handling — containing the damage. Episode 5 answers what comes after: the danger is contained, but is the system actually healthy again? Saturday, Round 5 👦 Nephew: We've detected it. We've handled it. The …

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

Episode 4 covered handling — containing the damage. Episode 5 answers what comes after: the danger is contained, but is the system actually healthy again?










Saturday, Round 5



👦 Nephew: We've detected it. We've handled it. The danger's contained. Now what?



Uncle: Imagine your Node.js server crashes.



👦 Nephew: Okay.



👨‍🦳 Uncle: Your API is down. No requests. No users. How does it come back?



👦 Nephew: Someone SSHs into the server... runs npm start?



👨‍🦳 Uncle: Would you wake an engineer at 3 AM every time that happened?



👦 Nephew: ...hopefully not.



👨‍🦳 Uncle: Exactly. The goal of recovery is to need humans less often — not never.









Part 1 — Recovery Isn't Handling



👨‍🦳 Uncle: House catches fire. Firefighters arrive, put it out. Is the house usable now?



👦 Nephew: No... it's still burnt. Someone still has to rebuild it.



👨‍🦳 Uncle: Putting out the fire and rebuilding the house are two completely different jobs.



👦 Nephew: So handling is the fire department. Recovery is the rebuild.



👨‍🦳 Uncle: Exactly that.




Failure
|
Detection
|
Handling → stops the damage from spreading
|
System still unhealthy
|
Recovery → brings the system back to healthy
|
Healthy again






👦 Nephew: So a system can be "handled" and still be completely broken.



👨‍🦳 Uncle: Every time. Handling contains. Recovery restores. Two different jobs, and a lot of engineers stop at the first one.









Part 2 — Small Recovery



👨‍🦳 Uncle: Your Node process crashed. What should happen?



👦 Nephew: Restart it.



👨‍🦳 Uncle: Who restarts it?



👦 Nephew: ...



👨‍🦳 Uncle: Go on.



👦 Nephew: ...me?



👨‍🦳 Uncle: At 3 AM?



👦 Nephew: ...oh. No. That's exactly the thing you asked me at the start, isn't it.



👨‍🦳 Uncle: Same question, different angle. So if not you — who?



👦 Nephew: Something has to be watching the process, ready to bring it back up the second it dies.



Uncle: That's what PM2 does. Or systemd. Or Docker's restart policy. Different tools, same job.




PM2       →  watches your Node process, restarts it on crash
systemd → watches a system service, restarts it on crash
Docker → watches a container, restarts it on crash

All solving the exact same problem: someone has to notice, and act, without you.






Nephew: So none of these tools are actually special — they're all just different hands doing the same job.



👨‍🦳 Uncle: Exactly. The tool doesn't matter yet. The idea does.









Part 3 — Recovery Isn't Always Restart



👨‍🦳 Uncle: Your database got deleted. Restart it?



👦 Nephew: ...no. That won't bring the data back.



👨‍🦳 Uncle: Right. Different failures need completely different recovery actions.




Node crash            →  Restart
Database deleted → Restore from backup
Redis disconnected → Reconnect
Server died → Failover to another server
Worker died → Spin up another worker






👦 Nephew: So "recovery" isn't one action. It's a whole category of different responses, and picking the right one depends entirely on what actually broke.



👨‍🦳 Uncle: That's the level-up moment for today. Beginners think recovery means restart. Real recovery means matching the fix to the failure — same lesson as handling, one lifecycle stage later.









Part 4 — Recovery Without a Human, and How It Actually Knows



👨‍🦳 Uncle: If Node crashes 100 times today, do you want 100 phone calls?



👦 Nephew: No. Obviously not.



👨‍🦳 Uncle: So who should be doing the restarting, 100 times, without you knowing each one happened?



👦 Nephew: The system itself.



👨‍🦳 Uncle: That's self-healing — recovery with no human in the loop. Now — a pod dies in Kubernetes. What happens?



👦 Nephew: Kubernetes... notices, and creates another pod?



👨‍🦳 Uncle: How did it know the pod died?



👦 Nephew: ...I actually don't know. I just assumed it magically knew.



👨‍🦳 Uncle: Nothing magic about it. Remember Episode 3?



👦 Nephew: Health checks.



👨‍🦳 Uncle: The pod stops answering its health check. Kubernetes sees that, and only then decides to act.




Pod stops responding
|
Health check fails ← this is DETECTION, from Episode 3
|
Kubernetes notices
|
Creates another pod ← this is RECOVERY, today's episode
|
Traffic resumes






👦 Nephew: So self-healing isn't its own separate magic trick. It's detection and recovery, wired directly into each other.



👨‍🦳 Uncle: That's the whole insight. Recovery doesn't start with "fix it." It starts with "notice it's broken" — which means every recovery system is quietly standing on top of a detection system.









Part 5 — The Recovery Ladder



👨‍🦳 Uncle: Put everything today in order, smallest fix to biggest.




Small Failure
|
Reconnect
|
Restart
|
Replace
|
Restore
|
Failover
|
Disaster Recovery






👦 Nephew: So recovery isn't one thing — it's a ladder, and how bad the failure is decides how high up you have to climb.



👨‍🦳 Uncle: Most days, you're at the bottom rung. The system heals itself quietly, nobody notices. The higher you climb, the fewer people have ever actually had to.









Part 6 — Recovery Isn't Instant



👨‍🦳 Uncle: Can every system recover?



Nephew: ...I want to say yes, but I feel like you're about to prove me wrong.



👨‍🦳 Uncle: Database corruption. You restore from last night's backup. What happened to everything written in the last ten minutes before the corruption?



👦 Nephew: ...gone. Lost.



👨‍🦳 Uncle: So the system recovered. Is it perfect?



👦 Nephew: No — it recovered, but not to exactly where it was.




Database corruption
|
Restore backup
|
Lose last 10 minutes
|
Recovered — but not perfect






👨‍🦳 Uncle: That gap has a name — how much time it took you to recover, and how much data you lost getting there. We'll properly name both of those in the Disaster Recovery module. For now, just hold onto the idea: climbing higher up that ladder costs you something, and the cost isn't always zero.









Part 7 — Final Exercise



👨‍🦳 Uncle: Server crashed.



👦 Nephew: Restart.



👨‍🦳 Uncle: Redis disconnected.



Nephew: Reconnect.



👨‍🦳 Uncle: Database deleted.



👦 Nephew: Restore backup.



👨‍🦳 Uncle: Entire AWS region gone.



👦 Nephew: ...multi-region?



👨‍🦳 Uncle: And what if there isn't another region?



👦 Nephew: ...then you can't recover?



👨‍🦳 Uncle: Then you're no longer recovering. You're surviving. That's Disaster Recovery.



👦 Nephew: Saturday?



👨‍🦳 Uncle: Saturday.









What we covered in Episode 5




  • Recovery is not the same job as handling — handling contains, recovery restores

  • The simplest recovery: restart, done automatically by PM2, systemd, or Docker

  • Recovery isn't always a restart — the action has to match the failure

  • Self-healing isn't magic — it's detection (Episode 3) wired directly into recovery

  • The Recovery Ladder: reconnect → restart → replace → restore → failover → disaster recovery

  • Recovery isn't always perfect or instant — a first look at what becomes RTO and RPO

  • The line between recovering and surviving: what happens when the next rung on the ladder doesn't exist



Next up — Module 3: Resilience Patterns, Episode 6: "Retry" — the pattern you've already used casually for two episodes, now built properly, with all the edge cases that break it in production.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Failure Engineering Explained by Uncle to Nephew — Episode 5: Recovery — How Systems Heal Themselves

Thematisch verwandte Begriffe: Failure, Engineering, Explained, Uncle · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-75517 | Novu provides an API for sending notifications through multiple channels…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick