Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
YouTube Security VideosAndroid Police: Samsung is smashing records! #shorts #tech #phones(21.09.2026 um 13:55 Uhr)
YouTube Security Videosheise & c't: Bundesnetzagentur wollte diesen Futterautomaten verbieten(21.09.2026 um 13:53 Uhr)
YouTube Security VideosNeil Patel: Your Google Traffic Isn't An Asset It's A Loan #shorts(21.09.2026 um 14:05 Uhr)
Windows Tipps & SecurityF-14 A Tomcat Top Gun endlich als Revell Klemmbausteinmodell erhältlich(21.09.2026 um 14:27 Uhr)
Sichere ProgrammierungShow the Hand-Back Sample Before Approving an Agent Score(21.09.2026 um 14:15 Uhr)
Sichere ProgrammierungHybrid retrieval in one Postgres query: RRF over tsvector + pgvector(21.09.2026 um 14:15 Uhr)
YouTube Security VideosAndroid Police: Samsung is smashing records! #shorts #tech #phones(21.09.2026 um 13:55 Uhr)
YouTube Security Videosheise & c't: Bundesnetzagentur wollte diesen Futterautomaten verbieten(21.09.2026 um 13:53 Uhr)
YouTube Security VideosNeil Patel: Your Google Traffic Isn't An Asset It's A Loan #shorts(21.09.2026 um 14:05 Uhr)
Windows Tipps & SecurityF-14 A Tomcat Top Gun endlich als Revell Klemmbausteinmodell erhältlich(21.09.2026 um 14:27 Uhr)
Sichere ProgrammierungShow the Hand-Back Sample Before Approving an Agent Score(21.09.2026 um 14:15 Uhr)
Sichere ProgrammierungHybrid retrieval in one Postgres query: RRF over tsvector + pgvector(21.09.2026 um 14:15 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Where Tensor-Parallel Inference Hits the NVLink Wall

Where tensor-parallel inference hits the NVLink wall 2026-05-31 · GPU / distributed systems Tensor parallelism splits each layer across GPUs, so every forward pass pays for an all-reduce over the network fabric. On a single node that …

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!




Where tensor-parallel inference hits the NVLink wall



2026-05-31 · GPU / distributed systems



Tensor parallelism splits each layer across GPUs, so every forward pass pays for an

all-reduce over the network fabric. On a single node that fabric is NVLink/NVSwitch — and

how close you get to its theoretical budget decides whether TP helps or hurts. This post

measures it on 4× H100 and explains where the wall is.



Repo with the full harness and CSVs:

nccl-collectives-bench.






What was measured



A bandwidth sweep (message size 8 B → 8 GB) of the three collectives that bound distributed

LLM work — all-reduce, all-gather, reduce-scatter — driving the canonical

nvidia/nccl-tests and adding a parser + analysis layer on top. The headline number:





  • All-reduce bus bandwidth ≈ 366 GB/s, about 77 % of the per-GPU NVLink uni-directional
    budget
    on this box. That 77 % is the practical ceiling TP communication runs into; the
    remaining gap is protocol overhead and the algorithm's traffic multiplier.

  • Algorithm ranking at large messages: NVLS > Ring > Tree. NVLink SHARP (NVLS) offloads
    the reduction into the switch, which is why it pulls ahead once messages are big enough to
    amortise setup.

  • A protocol study (Simple / LL / LL128) showing the small-message latency floor — the
    regime that actually matters for decode, where each token's all-reduce is tiny.






Why it matters for inference



Training all-reduces gradients on big tensors, so it lives in the bandwidth-bound regime

where 366 GB/s is good news. Decode is the opposite: one token at a time means small

messages, so you're pinned against the latency floor, not the bandwidth ceiling. That is the

real "TP wall" — past a certain TP degree, the per-token all-reduce latency dominates and

adding GPUs makes decode slower, not faster.



The repo also includes an eager-vs-CUDA-Graph comparison of that decode latency wall:

capturing the per-token step as a graph removes launch overhead that would otherwise be

indistinguishable from communication cost — a reminder to measure the right thing before

blaming the fabric.






Takeaway



"Use tensor parallelism" is not free advice. Measure the all-reduce on your fabric, know

your 77 %, and know that the number that decides decode latency is the small-message floor —

not the big-message bandwidth everyone quotes.



→ Methodology, raw CSVs, and the roofline analysis:

github.com/waynehacking8/nccl-collectives-bench

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Where Tensor-Parallel Inference Hits the NVLink Wall

Thematisch verwandte Begriffe: Where, TensorParallel, Inference, Hits · 6 Treffer

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-94097 | A vulnerability was determined in Netcore NBR200V2 1.3.241127.071246. Th…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick