Kimi K3 launched on July 16, and three claims immediately started traveling together:
- "It uses all 2.8 trillion parameters on every token."
- "The open weights are already available."
- "At $3/$15 per million tokens, it is automatically cheaper per task."
Two are wrong. The third is incomplete.
I spent the launch day reading Moonshot AI's release notes, API guide, pricing page, and the first independent measurements. The model is genuinely interesting. But the decision to migrate is much less obvious than the launch numbers make it look.
TL;DR
NO, the Kimi K3 weights are not downloadable today. Moonshot says it plans to release them by July 27, 2026. Until a checkpoint and license actually appear, that is a commitment, not a completed open-weight release.
The 2.8T figure is total capacity, not confirmed active parameters per token. Moonshot says its MoE routes each token through 16 of 896 experts, but it has not published the active parameter count.
The official API costs $3/M uncached input tokens and $15/M output tokens. Cache-hit input is $0.30/M.
The hidden variable is verbosity. Artificial Analysis measured roughly 130M output tokens during its evaluation, versus a 63M median for comparable models. More output can erase an attractive token rate.- I'd test K3 for long-context coding, research, and multimodal work, but I would not make it the default route without output caps and task-level evaluation.
What actually shipped
Moonshot AI's reported that K3 generated about 130M output tokens across its evaluation, while the median among comparable models was 63M. That does not prove your workload will see the same ratio. It does prove that output volume deserves measurement.
At K3's $15/M output rate:
63M output tokens x $15/M = $945
130M output tokens x $15/M = $1,950
Difference = $1,005 for the same evaluation-scale comparison
This is why I don't call a model cheap until I have cost per completed task. A model that emits twice as many tokens can cost more even when its token rate looks competitive.
The benchmark story is good, but uneven
Moonshot's launch table reports strong results in coding, terminal use, web browsing, science, and multimodal document understanding. These are vendor-reported scores, not one clean independent leaderboard.
| Benchmark | Kimi K3 | Best comparison shown by Moonshot | Launch-table reading |
|---|---|---|---|
| DeepSWE | 67.5 | 73.0 | K3 does not lead |
| Terminal-Bench 2.0 | 88.3 | 88.8 | Near the top |
| BrowseComp | 91.2 | 90.4 | K3 leads this table |
| GPQA Diamond | 93.5 | 94.1 | Competitive, not first |
| MMMU-Pro | 81.6 | 83.0 | Competitive, not first |
| OmniDocBench | 91.1 | 89.8 | K3 leads this table |
I would not convert this into a universal ranking. Moonshot's own footnotes show that models were tested with different reasoning modes and tool configurations. A score produced with one harness is not automatically comparable to a score produced with another.
The independent picture is more restrained. Artificial Analysis currently gives K3 an Intelligence Index of 57, reports about 62 output tokens per second, and measures a 1.99-second time to first token. Those numbers can change as providers optimize serving, so I see them as an early baseline, not a permanent verdict.
The API migration has several sharp edges
K3 is available through an OpenAI-compatible API, but compatibility does not mean "change one model string and forget it."
The does. Disclosure: I work on the research side. The full data-cited breakdown is in the original Kimi K3 review.
Bottom line
Kimi K3 is real, the 2.8T total parameter count is official, and the $3/$15 API is live. The active parameter count is undisclosed, the weights are promised rather than available, and early independent testing says output volume can be unusually high.
I'd test it now. I would not route production by headline.
Which matters more in your workload: the one-million-token context window, the $0.30 cache-hit rate, or controlling output verbosity?
SOCIAL SHARE CARD GENERATOR