🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 8 Min Lesezeit
0

I Raised Gemma 4's Token Cap. The Dense Model Stopped Refusing.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

TL;DR: Last week I argued Gemma 4 Dense regressed on a grounded-retrieval scenario under tightened prompts — and called the MoE-vs-Dense divergence architecture-mediated. The comment thread, led by Robin Converse on her sovereign Ollama stack, proposed an alternative: my max_tokens: 400 cap was starving Gemma's reasoning layer before the visible reply completed. I re-ran the same six scenarios with one variable changed — budget raised from 400 to 4096. Dense recovered on every scenario, including the false-refusal headline that anchored the original article. MoE did too. The original MoE-vs-Dense divergence largely disappears when reasoning has room to finish. The cap was doing the work. Walking it back publicly.









Why I Re-Ran It



Last week I published ( and asked: what does the same test look like on the managed Gemini API side? She also separately filed two upstream Ollama bugs — walked back her framing on was confirmed by multiple users and resolved in a later release. That kind of fair-witness practice is what made me trust the hypothesis enough to test it.



The hypothesis itself, sharp: my max_tokens: 400 cap was starving Gemma 4's reasoning layer before the visible reply completed. Capability ceiling and orchestration pressure look identical from the outside, as mapped a separate substitution-vs-decision boundary that sharpened where to look. , each named structural reframes I had to engage. The collective shape of the thread was: the cap is one variable, the architecture is a category, and you've been calling the result by the wrong name.



The cap was the part I could test fastest. One variable. Re-run the same six scenarios on the same two architectures with the only change being the budget.









The Experiment



Single-variable change against the original v2 conditions:











































Variable Original v2 (capped) This re-run (uncapped)
Arabic-first system frame kept kept (unchanged)
Temperature 0.3 0.3 (unchanged)

max_tokens floor
400 4096
Scenarios 6 Arabic e-commerce 6 (same set)
Models 26B MoE + 31B Dense 26B MoE + 31B Dense
Calls 12 12


The implementation was four lines in the chat-models server file called it from her stack on Tuesday. The cross-validation now exists: same architecture, two deployment contexts (sovereign Ollama, managed Gemini API), one finding — uncap the budget and the failure mode evaporates.



What still partially holds from the original article:





  • Reluctance, not hallucination, as the dominant failure mode for grounded Arabic chat on open models when the budget is too tight to complete reasoning.


  • Variant-specific prompt tuning is real — but the variant-specific thing isn't an architectural slot; it's a token budget shaped to the variant's reasoning footprint.


  • Latency on the Google API is still a chasm for interactive chat — that part wasn't a cap artifact.



What I retract:




  • The "architecture has slots for sequential sub-behavior that dense doesn't" hypothesis was too strong. Dense plus budget gets the same grounded behavior.

  • The matrix's "31B Dense regression" reads now as "31B Dense under-budgeted regression."

  • The closing line — "I think I was tuning architecture, not size" — should have been "I think I was tuning budget, mediated by architecture."









What's Still Open





  • Temperature dimension untested in this run. I held temperature at 0.3 to keep this one-variable. proposed making preconditions first-class through his NEXUS protocol; shared that JAMES's per-stage DEFAULT_MAX_TOKENS defaults (200/400/400/400 across four cognitive stages) match this pathology — their 2026-05-18 internal eval reported empty responses on gemma4:e4b at exactly those four stages. Uncapped replication is queued this week. If it reproduces, that's three independent deployment contexts (sovereign Ollama, managed Gemini API, JAMES production) on the same cap pathology before any cross-experiment swap runs.









The Lesson



The community ran my falsifier for me. The cap is one variable; architecture is a category. If your model is failing under a tight token budget, raise the budget before you reach for architectural explanations — and if a thoughtful commenter offers you a single-variable test against your strongest claim, run it.



The article you can write a week later, with the comment thread folded in, is stronger than the one you ship alone. Robin Converse is the right person to share the next round with.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten I Raised Gemma 4's Token Cap. The Dense Model Stopped Refusing.

Thematisch verwandte Begriffe: Raised, Gemma, Token, Dense · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...