Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
Windows Tipps & SecurityGrafikkarte vor Überhitzung schützen: So geht’s(25.09.2026 um 08:00 Uhr)
••••••••••
Windows Tipps & SecurityGrafikkarte vor Überhitzung schützen: So geht’s(25.09.2026 um 08:00 Uhr)
••••••••••
Intelligence View
⚡ tsecurity.de Intelligence

OpenAI Is Running PostgreSQL at Millions of QPS (Here’s How It Didn’t Explode)

PostgreSQL has a reputation. Great database. Rock‑solid. But “web‑scale”? Eh… OpenAI just published how they’re running Postgres at millions of queries per second for ~800 million users. With: One primary ~50 read replicas And no shar…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

PostgreSQL has a reputation.



Great database.


Rock‑solid.


But “web‑scale”? Eh…



OpenAI just published how they’re running Postgres at millions of queries per second for ~800 million users.



With:




  • One primary

  • ~50 read replicas

  • And no sharding (for Postgres)



Let’s talk about what actually made this work.









The Architecture (Surprisingly Simple)




  • Single primary PostgreSQL

  • Writes go to the primary

  • Reads go almost everywhere else

  • ~50 read replicas across regions



This only works because the workload is:




  • Extremely read‑heavy

  • Ruthlessly optimized



Postgres wasn’t the bottleneck.

Bad assumptions were.









The Failure Mode That Kept Biting Them



Every major incident followed the same pattern:




  1. Cache fails / misses spike

  2. Expensive queries flood Postgres

  3. Latency rises

  4. Requests start timing out

  5. Retries kick in

  6. Load gets worse



Congratulations, you’ve built a feedback loop from hell.









Writes Are the Real Problem (Thanks, MVCC)



Postgres MVCC means:




  • Updates copy the entire row

  • Dead tuples pile up

  • Autovacuum becomes a full‑time job



At scale:




  • Reads are cheap

  • Writes are dangerous



So OpenAI made a call most teams avoid.









Rule #1: Don’t Let Postgres Do Write‑Heavy Work




  • Shardable, write‑heavy workloads → CosmosDB

  • PostgreSQL stays unsharded

  • No new tables allowed on the main Postgres cluster



Why?

Because retroactively sharding a massive OLTP system is how roadmaps go to die.









ORMs Will Absolutely Hurt You If You Let Them



One real incident:




  • A query joining 12 tables

  • Traffic spike

  • CPU pegged

  • SEV triggered



Lessons learned:




  • ORMs generate wild SQL

  • Multi‑table joins at scale are a footgun

  • Complex joins often belong in the app layer



Also:




  • Kill idle transactions

  • Set idle_in_transaction_session_timeout

  • Always read the SQL your ORM emits









PgBouncer Is Not Optional



They hit:




  • Connection storms

  • Exhausted connection limits

  • Cascading failures



Fix:




  • PgBouncer in transaction / statement pooling

  • Deployed via Kubernetes per replica

  • Co‑located with clients



Results:




  • Connection time: 50ms → 5ms

  • Way fewer idle connections

  • Fewer “why is Postgres on fire” moments









Cache Misses Can Take You Down



A cache miss storm is just a DDoS you did to yourself.



Solution:




  • Cache locking / leasing

  • One request fetches from Postgres

  • Everyone else waits



This single pattern prevented multiple outages.









Rate Limiting at Every Layer



They rate‑limit:




  • Endpoints

  • Connection pools

  • Proxies

  • Individual query digests



They’ll even:




  • Block specific queries at the ORM layer



Retries are treated as dangerous, not helpful.









One Primary Is a Risk — Reduce the Blast Radius



If the primary dies:




  • Writes fail

  • But reads still work



How:




  • Critical read paths are replica‑only

  • Primary runs in HA with a hot standby

  • Fast, reliable failovers



This turns “everything is down” into “writes are temporarily unavailable.”



Huge difference.









Noisy Neighbors Are Real



New feature launches caused:




  • CPU spikes

  • Latency for unrelated traffic



Fix:




  • Workload isolation

  • High‑priority vs low‑priority instances

  • Product‑level separation



One bad feature should not take down ChatGPT.









Schema Changes Are Basically Treated as Incidents



Rules:




  • No table rewrites

  • No long migrations

  • 5‑second schema change timeout

  • Indexes only with CONCURRENTLY

  • Backfills are rate‑limited and can take weeks



Speed is optional.

Stability isn’t.









The Results




  • Millions of QPS

  • Low double‑digit ms p99 latency

  • ~50 global read replicas

  • Five‑nines availability

  • One SEV‑0 in a year



Postgres didn’t fail.

Engineering discipline scaled it.









The Takeaway



PostgreSQL can scale insanely far if:




  • Your workload is read‑heavy

  • Writes are treated with respect

  • Caching is defensive

  • Retries are controlled

  • ORMs are kept on a leash



Most Postgres horror stories aren’t about Postgres.



They’re about letting the system do work it was never meant to do.

1. Sofort-Triage & Abwehrmaßnahmen

SOC Incident Playbook: Vulnerability Remediation & Verification
Syntax validiert (0 Fehler)
title: Detect Exploitation - OpenAI Is Running PostgreSQL at Millions of QPS (Here’s How It Didn’t Explode)
id: b1be6d1d-15a7-423b-b2a1-5726f06b4afa
status: experimental
description: Automatisch generierte SIEM-Erkennungsregel basierend auf CTI Intelligence
references:
  - https://tsecurity.de/
author: iShareStuff CTI Automated Detection Engine
date: 2026-09-26
logsource:
  category: network_connection
  product: any
detection:
  selection:
      CommandLine|contains:
        - 'exploit'
  condition: selection
falsepositives:
  - Legitime administrative Zugriffe oder Penetrationstests
level: high
tags:
  - attack.initial_access
Syntax validiert (0 Fehler)
rule CTI_Threat_Indicator {
    meta:
        author = "iShareStuff CTI Automated Detection Engine"
        date = "2026-09-26"
        description = "YARA Signature for "
    strings:
        $str = "OpenAI Is Running PostgreSQL a" ascii wide
    condition:
        any of them
}
Syntax validiert (0 Fehler)
index=security sourcetype IN ("cisco:asa", "pan:traffic", "zeek_conn", "suricata", "WinEventLog:Security")
("OpenAI Is Running PostgreSQL at Millions")
| stats count earliest(_time) as first_seen latest(_time) as last_seen by src_ip, dest_ip, dest_host, signature
| eval first_seen=strftime(first_seen, "%Y-%m-%d %H:%M:%S"), last_seen=strftime(last_seen, "%Y-%m-%d %H:%M:%S")
| sort - count
Syntax validiert (0 Fehler)
message: "*OpenAI Is Running PostgreSQL at Millions*"
Syntax validiert (0 Fehler)
CommonSecurityLog
| where Message has "OpenAI Is Running PostgreSQL at Millions"
| summarize EventCount = count(), FirstSeen = min(TimeGenerated), LastSeen = max(TimeGenerated) by SourceIP, DestinationIP, DestinationPort, Activity
| extend DetectionRule = "iShareStuff-CTI-Compiled"
| sort by EventCount desc

2. Cyber Threat Intelligence & Forensik

🎯
MITRE ATT&CK Matrix Navigator 14 Taktiken
Reconnaissance
-
Resource Development
-
Initial Access
Execution
Persistence
-
Privilege Escalation
Defense Evasion
Credential Access
-
Discovery
-
Lateral Movement
-
Collection
-
Command and Control
Exfiltration
-
Impact
tsecurity.de Cognitive Threat RAG
Fokus-Vektor:

Kognitive Analyse für identifizierte Bedrohung: Erhöhte Bedrohungslage im Bereich OpenAI Is Running PostgreSQL at Millions.... Basierend auf 368k Vektor-Korrelationen werden sofortige Isolationsmaßnahmen für betroffene Endpunkte empfohlen.

🛡️ Angriffsfläche & Exposure

Netzwerk/Remote-Zugriff ohne Vorauthentifizierung möglich.

⚡ Empfohlene Sofortmaßnahmen
  • 1. Perimeter-Inspektion: Relevante Portfreigaben und exponierte Endpunkte unverzüglich scannen.
  • 2. Patch-Applikation: Hersteller-Hotfix einspielen oder betroffene Daemons in isolierte DMZ-Segmente überführen.
  • 3. Telemetrie & EDR-Alerts: Prozessaufrufe und Child-Processes auf anomale Shell-Spawns überwachen.
🔗 Semantisch verwandte Zero-Days MariaDB 11.7 VEC
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten OpenAI Is Running PostgreSQL at Millions of QPS (Here’s How It Didn’t Explode)

Thematisch verwandte Begriffe: OpenAI, Running, PostgreSQL, Millions · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-100537 | OpenClaw (npm package 'openclaw') before 2026.8.1 fails to apply the or…
Advisory →
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag