, an agent workstation for open-weight models. The core claim behind it is that the harness — retrieval, tool surface, control loop — is what lets an open-weight model perform, not just the model's raw weights. This post is me trying to falsify that claim with a controlled run, and publishing every output file so you can check it.
Short version: on Terminal-Bench 2.0, single attempt, Atlarix resolved 42/89 and opencode resolved 39/89 on the same model. That 3-task gap is within k=1 noise — I'm not claiming a win. What it shows is that the harness isn't bottlenecking the model. Details and caveats below; raw files at the end.
The experiment
The only variable is the harness. Everything else is pinned identical across both agents.
Benchmark:terminal-bench/terminal-bench-2— all 89 tasks, one isolated container each, automated verifiers.
Model:minimax/minimax-m3, routed through OpenRouter, pinned to a single provider at fp8 — identical for both harnesses.
Infrastructure:
What's next
- More open-weight models, so no claim rests on one.
- The official Terminal-Bench (k=5) submission — on the roadmap.
- More benchmarks beyond terminal tasks.
If you spot something wrong in the result files, that's the point — tell me.
Built in Nairobi.
↗ Original-Artikel auf dev.to lesenVollständiger Original-BerichtAusführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
SOCIAL SHARE CARD GENERATOR