Zum Hauptinhalt springen
🪟 Windows TippsHow to install and set up DBeaver on Windows 11(18.09.2026 um 02:31 Uhr)
🪟 Windows TippsHow to install and set up DBeaver on Windows 11(18.09.2026 um 02:31 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Llama 3.1 405B accelerated to almost a thousand tokens per second

Cerebras finally found enough of their CS-3 to launch Llama 405B, applied Speculative Decoding to it, which they used to speed up 70B up to 2k tokens, and outperformed SambaNova by almost 6 times. It will cost $6 input/$12 output per million tokens and is already available in beta. All users will be given access in the first quarter of 2025.

You have to wait so long because of the extremely poor availability of hardware - in order to run Llama 405B, you need 20-30 CS-3. By comparison, Condor Galaxy, a supercomputer powered by Cerebras chips, has only 64 CS-3s. And it costs more than one hundred million dollars. I hope that if they manage to switch to mass production, the cost of their systems will drop significantly. Otherwise, the profitability of such an API is questionable.

It’s not just Cerebras that has problems with availability—Groq also has them, which have been promising API 405B for more than three months, but apparently there just aren’t enough chips (about four thousand Groq chips are needed to run 405B). In the meantime, they have almost caught up with Cerebras on the Llama 70B inference - 1669 tokens per second, while promising that the next generation of chips will be much faster.

Unfortunately, access to all users via chat was not given this time. And the context length is only 8k so far, but they promise to make 128k available at release. The speed in this context, however, sags, but still more than half a thousand tokens per second. Hopefully for a full release R1 they will dig up another supercomputer, and we will have a model that thinks in seconds instead of minutes.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Llama 3.1 405B accelerated to almost a thousand tokens per second

Thematisch verwandte Begriffe: Llama, 405B, accelerated, almost · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-61591 | djust provides Phoenix LiveView-style reactive server-side rendering for…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
News ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

↗ Original-Quelle