🕵️ SicherheitslückenThe Cyber Mentors: 3 Cybersecurity Books to Read(01.09.2026 um 19:00 Uhr)
🕵️ SicherheitslückenMy Life, AI and the Future of LiveOverflow(23.08.2026 um 19:53 Uhr)
🕵️ SicherheitslückenThe XSS Rat: How To Use AI To Become a KiLLER Hacker(29.08.2026 um 02:05 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: How North Korean Hackers end up in your Network(24.08.2026 um 20:30 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: Undetected Steam Malware: Sent by Viewer(28.08.2026 um 21:00 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: How to tell if your PC is Hacked: Ultimate Edition(04.09.2026 um 20:30 Uhr)
🕵️ SicherheitslückenSecurity Weekly - A CRA Resource: What AI Researchers See Beyond AI(01.09.2026 um 15:00 Uhr)
🕵️ SicherheitslückenThe Cyber Mentors: 3 Cybersecurity Books to Read(01.09.2026 um 19:00 Uhr)
🕵️ SicherheitslückenMy Life, AI and the Future of LiveOverflow(23.08.2026 um 19:53 Uhr)
🕵️ SicherheitslückenThe XSS Rat: How To Use AI To Become a KiLLER Hacker(29.08.2026 um 02:05 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: How North Korean Hackers end up in your Network(24.08.2026 um 20:30 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: Undetected Steam Malware: Sent by Viewer(28.08.2026 um 21:00 Uhr)
⚠️ Malware / Trojaner / VirenPC Security Channel: How to tell if your PC is Hacked: Ultimate Edition(04.09.2026 um 20:30 Uhr)
🕵️ SicherheitslückenSecurity Weekly - A CRA Resource: What AI Researchers See Beyond AI(01.09.2026 um 15:00 Uhr)

🔧 Programmierung 🕛 kürzlich 4 Min Lesezeit
0

Treasure Hunt Engine: The Day We Realized the Event Bus Was Our Constraint

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




The Problem We Were Actually Solving



We werent just chasing p99 latency; we were solving a fundamental mismatch between the event model and the treasure hunt logic. Each treasure hunt round emits thousands of micro-events: player joins, item picks, time updates, leaderboard recalculations, and realtime notifications. The Node.js event loop was choking under the backpressure. The BullMQ worker was blocked on Redis pubsub, not because of network latency, but because Node.jss single-threaded event loop couldnt keep up with the rate of incoming events. The Redis server itself was fine—CPU at 12%, memory at 68%, no evictions. The bottleneck wasnt the queue or the data store. It was the runtime.



I added a debug trace using 0x and saw 78% of CPU time was spent in uv__io_poll, the epoll/select wrapper. The Node.js process was spending more time waiting for events than processing them. And because BullMQ uses Redis streams, every publish and consume was a network roundtrip. The 250 microsecond RTT from us-east-1 to the Redis cluster was adding up when we were publishing 47,000 events per second. The p99 latency followed the square root of the number of concurrent players. At 5,000 players, it was 80ms. At 10,000 players, 2.3 seconds. The system wasnt scaling linearly. It was falling off a cliff.






What We Tried First (And Why It Failed)



We tried horizontal scaling BullMQ workers. We spun up 8 workers behind an SQS queue. The SQS throughput was fine—50,000 events/sec sustained—but BullMQs Redis backpressure became a distributed locking nightmare. Workers fought over the same Redis key ranges, and the Redis pubsub fanout created a thundering herd on the Node.js event loop. We saw lock contention in XREADGROUP with 200ms timeouts. We tried sharding the Redis streams into 16 shards. The shard imbalance was brutal—some shards got 3x the load. We tried upgrading Node.js to 20. Same behavior. We even tried using ioredis for connection pooling, but the fundamental mismatch remained: Node.js was a stream processor pretending to be an event-driven runtime.



Then we tried denoising the events—filtering out duplicate player actions, compressing payloads, batching events. That helped reduce volume by 38%, but the p99 latency still rose with player count. The issue wasnt event volume. It was the runtimes inability to handle the concurrency model we needed.






The Architecture Decision



We had to accept that Node.js was the constraint. Not Redis. Not BullMQ. The runtime itself. We spun up a prototype in Go. Using go-redis with a streaming consumer group, we hit 320,000 events/sec on the same c5.4xlarge instance with under 100ms p99 latency. The memory allocation profile from pprof showed 1.2 MiB per second GC pressure—nothing compared to Node.jss 47 MiB/sec. But Go wasnt the only option. We also tested Rust.



We built a minimal Rust prototype using Tokio, Redis streams via redis-rs, and a hand-rolled event router. The first version used std::thread for concurrency, but that led to thread starvation under load. We switched to Tokio with 8 worker tasks and a single Redis connection with multiplexing. The memory footprint was 8.7 MiB RSS at idle, peaking at 42 MiB under 100,000 events/sec. The Tokio runtimes work-stealing scheduler meant no idle threads, no wasted CPU waiting for events. The p95 event latency was 18ms, p99 47ms. We ran a load test for 12 hours with 20,000 concurrent players, 2.1 million events per minute. Zero GC pauses, zero memory leaks, zero crashes.



The architecture decision wasnt just language. It was concurrency model. We moved from a centralized event bus (Redis streams) to a partitioned event log with local in-memory buffering. Each shard had its own Redis stream, and a Rust worker consumed it with async I/O, processed events in order, and pushed updates to a local pubsub channel for the game server. The game server itself remained in Node.js, but now it listened to local events via Unix sockets, cutting RTT from 250 microseconds to 4 microseconds. The entire stack became pipeline-based: Redis → Rust → Node.js → frontend.






What The Numbers Said After



We deployed the Rust worker to production on a c5.large (2 vCPUs, 4 GiB RAM). The worker handled 60,000 events/sec with 4ms p99 latency. The Node.js game server CPU dropped from 85% to 18%. Memory usage fell from 1.4 GiB to 320 MiB. The Redis cluster CPU dropped from 78% to 22%, and we reduced shards from 16 to 4. The treasure hunt p99 latency dropped from 2.3 seconds to 62ms. Player complaints about lag disappeared. We added a new feature: realtime leaderboard recalculations every 200ms. The Rust worker handled it with zero additional latency.



But the real win was observability. Tokios tracing crate gave us per-event latency histograms at 100,000 events/sec. We could see where each event spent time: 31% in Redis XREAD, 28% in event parsing, 14%

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 50%
🟡 In Evaluierung 20%
🟢 Keine Auswirkung 11%
Spannende Innovation 19%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
ChatGPT showing blank screen [Fix]
1 Quelle
Sofort deinstallieren: Diese 19 Browser-Erweiterungen sind mit Malware verseucht
1 Quelle
AMD Ryzen AI 400 NPU einrichten: 60 TOPS in 12 Schritten [2026] - shattered.io
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Treasure Hunt Engine: The Day We Realized the Event Bus Was Our Constraint

Thematisch verwandte Begriffe: Treasure, Hunt, Engine, Realized · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...