Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Windows Tipps & SecurityNighthawk M7 Pro im Test: Flexibler, aber teurer 5G-Router(21.09.2026 um 10:30 Uhr)
Sichere ProgrammierungNeue Gmail-Funktion: So sparst du jetzt Zeit bei Einmalcodes(21.09.2026 um 10:00 Uhr)
Sichere ProgrammierungYour GIF exporter is fine — the container is the problem(21.09.2026 um 10:01 Uhr)
Sichere ProgrammierungCSS, Motion, or GSAP? I Choose by Who Owns the Animation(21.09.2026 um 10:12 Uhr)
Windows Tipps & SecurityNighthawk M7 Pro im Test: Flexibler, aber teurer 5G-Router(21.09.2026 um 10:30 Uhr)
Sichere ProgrammierungNeue Gmail-Funktion: So sparst du jetzt Zeit bei Einmalcodes(21.09.2026 um 10:00 Uhr)
Sichere ProgrammierungYour GIF exporter is fine — the container is the problem(21.09.2026 um 10:01 Uhr)
Sichere ProgrammierungCSS, Motion, or GSAP? I Choose by Who Owns the Animation(21.09.2026 um 10:12 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

A 35-billion-parameter agent that punches like a trillion-parameter model

A 35-billion-parameter model called Agents-A1 matches trillion-parameter models on multi-step agent tasks, according to a new paper from Shanghai AI Lab. The key insight: instead of scaling parameter count, the researchers scaled the…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

A 35-billion-parameter model called Agents-A1 matches trillion-parameter models on multi-step agent tasks, according to a new paper from Shanghai AI Lab. The key insight: instead of scaling parameter count, the researchers scaled the "horizon" — the length and variety of action sequences the model trains on — producing a small model that sustains plans across long sequences of tool use as well as giants do. The work is on arXiv, and its title captures the thesis: scaling the horizon, not the parameters.






Key facts





  • What: Shanghai AI Lab argues you can reach giant-model performance on long tasks not by adding parameters, but by training on much longer chains of real work.


  • When: 2026-06-30


  • Primary source: read the source (arXiv 2606.30616)



Agents-A1 has 35 billion total parameters — small by frontier standards — yet matches trillion-parameter models on agent tasks: long, multi-step jobs where the AI must use tools, take actions, observe results, and keep working toward a goal across many turns. A giant model has vast raw knowledge, but agent work depends less on knowing more facts and more on sustaining a plan across a long sequence of actions without losing the thread. So instead of scaling parameter count, the researchers scaled the horizon — the length and variety of the action sequences the model learns from.



Concretely, they built an infrastructure that connects external knowledge, actions, observations, and checks on whether each action worked, and used it to generate training examples that average around forty-five thousand words per task. The model learns from full, extended episodes of real problem-solving, not short snippets. Training on long trajectories teaches the specific skill agents need: carrying context and a goal across dozens of steps, the difference between studying finished essays and watching someone work through an entire project from start to finish.



The training structure leans on distillation, an idea we cover in distillation. Rather than making one model good at everything at once, the team first trained separate specialist teacher models, each expert in one domain, then distilled all of them into a single student model — routing the student to learn from whichever teacher was most relevant for a given kind of task. This lets one modestly sized model absorb the strengths of several specialists. It is also built as a mixture-of-experts model, so only part of it activates at any moment, keeping running costs down; our lesson on mixture of experts explains why that design is everywhere now.



The reported results are strong across a spread of demanding agent and science benchmarks — the paper claims leading or highly competitive numbers on tasks involving tool use, web browsing, and scientific reasoning, holding its own against trillion-parameter systems on the long-horizon work it was built for. If that holds up under independent testing, the implication is meaningful: the path to capable agents runs partly through better, longer training data rather than only through ever-larger and more expensive models — good news for anyone who cannot afford to train a trillion-parameter system.



The honest caveat is the standard one for a self-reported paper: these are the authors' own benchmark numbers, and benchmark performance and real-world reliability are not the same thing — a point our lesson on how AI gets benchmarked makes at length. Matching a giant model on a curated test set is impressive but does not guarantee matching it on the messy, open-ended tasks people actually throw at agents, where, as a wave of new benchmarks this week showed, even the best frontier models still struggle badly. There is also a selection effect: it is easier to reach parity on the exact kinds of tasks you designed your training data around. Still, the core argument is a healthy corrective to size-worship. Bigger is one way to get better, but it is not the only way — and for the specific challenge of agents that have to think across a long stretch of work, teaching a smaller model on longer examples may be the smarter bet.






Originally published on Ground Truth, where every claim is checked against the primary source.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten A 35-billion-parameter agent that punches like a trillion-parameter model

Thematisch verwandte Begriffe: 35billionparameter, agent, that, punches · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-94030 | A security vulnerability has been detected in SerenityOS up to 3d83e4509…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick