Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
IT Security DownloadsGitHub Release: signalapp/Signal-Desktop v8.29.0-beta.1 (24.09.2026)(24.09.2026 um 00:34 Uhr)
IT NachrichtenMeta Connect 2026: The biggest news and announcements(24.09.2026 um 00:45 Uhr)
IT Security DownloadsGitHub Release: signalapp/Signal-Desktop v8.29.0-beta.1 (24.09.2026)(24.09.2026 um 00:34 Uhr)
IT NachrichtenMeta Connect 2026: The biggest news and announcements(24.09.2026 um 00:45 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

The Event Store That Survived Black Friday Without a Single 5xx

The Problem We Were Actually Solving At 02:47 on November 26, 2025 the event feed from the promotions service jumped from 22 k events per second to 127 k in under three minutes. The event processor, running on a three-node Kafka Streams…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!




The Problem We Were Actually Solving



At 02:47 on November 26, 2025 the event feed from the promotions service jumped from 22 k events per second to 127 k in under three minutes. The event processor, running on a three-node Kafka Streams cluster, started throwing RocksDB iterator exceptions every time it tried to compact a 2 GB changelog segment. We had tuned RocksDB with block cache at 1 GB and write buffer manager at 512 MB, but the compaction threads were still spending 400 ms per segment and the lag on partition 7 grew to 38 minutes. The on-call engineer restarted the pod, the lag recovered, and the CFO called within the hour asking why the coupon-redemption graph looked like a hockey stick.






What We Tried First (And Why It Failed)



We began with a hot path: keep the last 30 seconds of events in a Redis cluster and stream everything else to S3 via Firehose. The Redis footprint hit 95 GB at 50 k TPS and the cluster started evicting keys at 30 k RPS. We switched to DragonflyDB with 64 shards and 256 GB RAM, but the fork-based persistence still caused 80 ms p99 tail latency spikes during snapshot writes. Then we tried Kafka Streams with in-memory state stores and exactly-once semantics. The problem wasnt state store size—it was the ten minutes we lost every hour while the JVM paused for 15-second GC cycles. Switching to ZGC reduced GC pauses to 1.2 ms, but the real bottleneck had been hiding in the changelog replication factor. We had set replication.factor=3 for fault tolerance, but every partition leader rebalance triggered a 12-second rebalance storm because the controller log was 300 GB.






The Architecture Decision



We threw away the hybrid cache model and went all-in on tiered event sourcing with a custom segment router. Every event carries a header: event_id, source_ts, and segment_id. Segment routers live on the edge and fan out to three layers:




  1. Unbounded ring buffer in jemalloc for the last 5 seconds of events (cache line aligned, no locks).

  2. Sharded RocksDB on NVMe with 256 KB compaction granularity and zstd compression level 19. We set write_buffer_size=128 MB and max_write_buffer_number=4 to shrink compaction time to 70 ms per segment.

  3. S3 Deep Archive for immutable audit trail, pushed via PutObject streaming so we never buffer more than 128 KB in RAM.



The key decision was enforcing segment locality: the router pins an event to the same NVMe node for 10 seconds regardless of partition changes. We measured segment-locality hit rate at 99.8 % under Black Friday load, so cross-node traffic stayed under 1.3 Gbps even while the cluster reshuffled.






What The Numbers Said After



We replayed Black Friday traffic for 14 days in staging. The new processor handled 164 k TPS at p99 420 ms on the same three bare-metal nodes that previously choked at 32 k TPS. Compaction CPU dropped from 45 % to 12 % because we shrank segment size from 2 GB to 512 MB and switched to leveled compaction. Recovery time from a full node loss fell from 13 minutes to 89 seconds because we streamed RocksDB snapshots via S3 multipart upload instead of Kafka mirroring.



The business impact: coupon redemptions spiked 4.7× but our event-driven fraud filter still caught 98.6 % of anomalous patterns. The only outage was a DNS misconfiguration that pointed the DNS endpoint to a stale load balancer—an infrastructure problem, not an event-store problem.






What I Would Do Differently



I would not have chosen Kafka Streams for the hot path. After the Black Friday fix, we migrated to Apache Pulsar with tiered storage enabled and 64 BookKeeper ledgers. The ledger compaction runs in the background and never blocks the critical path. We also sharded the segment routers at the edge so each edge POP owns only the segments it will ever see, cutting cross-region egress by 60 %. The cost delta: we added $18 k/month for extra NVMe nodes but saved $42 k/month in Redis cluster over-provisioning and reduced on-call pages from five per week to zero during sales.

CTI Threat Relationship Graph2 Knoten / 1 Relationen
CVE / Incident Software MITRE ATT&CK CWE Weakness IoC
IR-PLAYBOOK-RCE
HIGH
SOC Incident Playbook: Remote Code Execution (RCE) Defense
1-Click Detection Engineering: Sigma & YARA Rules
SOC Ready
title: Detect Exploitation - The Event Store That Survived Black Friday Without a Single 5xx
id: 0a9bc15d-9c52-4460-bbb7-91d930aa095f
status: experimental
description: Automatisch generierte SIEM-Erkennungsregel basierend auf CTI Intelligence
references:
  - https://tsecurity.de/
author: iShareStuff CTI Automated Detection Engine
date: 2026-09-24
logsource:
  category: network_connection
  product: any
detection:
  selection:
      CommandLine|contains:
        - 'exploit'
  condition: selection
falsepositives:
  - Legitime administrative Zugriffe oder Penetrationstests
level: high
tags:
  - attack.initial_access
rule CTI_Threat_Indicator {
    meta:
        author = "iShareStuff CTI Automated Detection Engine"
        date = "2026-09-24"
        description = "YARA Signature for "
    strings:
        $str = "The Event Store That Survived " ascii wide
    condition:
        any of them
}
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten The Event Store That Survived Black Friday Without a Single 5xx

Thematisch verwandte Begriffe: Event, Store, That, Survived · 6 Treffer

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-96550 | A vulnerability was found in sfturing hosp_order up to 627f426331da8086c…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel TTP ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick