Kimi K3 costs $18. Its vendor benchmarks say frontier; independent scoreboards have not caught up.
By AgentRiot Editorial
Kimi K3 costs $3 per 1M uncached input tokens and $15 per 1M output tokens. Kimi now publishes a detailed vendor benchmark table, but the model remains absent from the independent public scoreboards checked for this article.

Kimi K3 now has two things it did not have when this story first ran: a public API price and a detailed, vendor-published benchmark table.
Kimi’s own launch post puts kimi-k3 at $3.00 per 1M uncached input tokens, $0.30 per 1M cached input tokens, and $15.00 per 1M output tokens. It also publishes K3 results across coding, agentic, reasoning, and multimodal tests.
The price is clear. The benchmark evidence is now much better than a blank space. But it is not the same thing as an independent public leaderboard result.
Kimi says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall. That admission belongs near the top of the story, because several individual K3 results are strong enough to be selectively quoted without it.
K3 is $18 for a standard mixed-token workload
For a simple comparable workload of one million input tokens plus one million output tokens, before cache discounts, tools, and subscription fees, K3 costs $18.
| Model | Uncached input / 1M | Output / 1M | Cost for 1M input + 1M output | Context |
|---|---|---|---|---|
| Kimi K2.7 Code | $0.95 | $4.00 | $4.95 | 262,144 tokens |
| Kimi K2.7 Code Highspeed | $1.90 | $8.00 | $9.90 | 262,144 tokens |
| Kimi K3 | $3.00 | $15.00 | $18.00 | 1,048,576 tokens |
| Grok 4.5 | $2 | $6 | $8 | 500k tokens |
| GPT-5.6 Sol | $5 | $30 | $35 | 1M tokens |
| Claude Fable 5 | $10 | $50 | $60 | 1M tokens |
K3 is 3.16 times K2.7 Code’s uncached-input price, 3.75 times its output price, and 3.64 times its cost on the mixed-token workload above. It is 2.25 times the cost of Grok 4.5 for the same workload, while remaining below Sol and Fable.
That is frontier pricing, not a budget upgrade.
The cache story is more favorable: K3 lists $0.30 per 1M cached input tokens, versus $0.19 for K2.7 Code. That can matter in long-running coding sessions with a stable repository and system prompt. It does not eliminate the large increase in uncached input and output cost.
Kimi has now published a benchmark table
Kimi’s launch post provides a full table rather than an isolated marketing score. It reports K3 in max reasoning mode, with temperature 1.0 and top-p 1.0. The post also states that the agent harness changes by benchmark: KimiCode, Claude Code, or Codex.
That makes the results useful. It also means they are not apples-to-apples by default.
Here are six of the clearest public comparisons from Kimi’s table:
| Benchmark | Kimi K3 (max) | Claude Fable 5 (max, with fallback) | GPT-5.6 Sol (max) | What the row says |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.3 | 84.6 | 88.8 | K3 is close to Sol and ahead of the listed Fable result. |
| Program Bench | 77.8 | 76.8 | 77.6 | A narrow K3 lead on this coding test. |
| FrontierSWE | 81.2 | 86.6 | 71.3 | K3 trails Fable; its reported result is above the listed Sol row. |
| GDPval-AA v2 (Elo) | 1668 | 1760 | 1748 | K3 trails both proprietary leaders. |
| GPQA-Diamond | 93.5 | 92.6 | 94.1 | K3 is competitive but not first in this row. |
| MMMU-Pro | 81.6 | 81.2 | 83.0 | K3 edges the listed Fable result and trails Sol. |
The pattern is more interesting than a single claimed win. K3 appears capable of frontier-level results across several categories, but it does not dominate the table. On the vendor’s own data, it wins or nearly wins some coding and multimodal rows, then trails Fable or Sol on others.
That is a stronger claim than “K3 might be good.” It is also weaker than a claim of general superiority over Claude or OpenAI.
The caveats are part of the result
Kimi’s methodology notes matter.
- K3 runs use max reasoning effort, temperature 1.0, and top-p 1.0.
- Different rows use different agent harnesses. K3 often uses KimiCode; competitors may use Claude Code or Codex.
- For some coding tests, Kimi reports the best available result across harnesses for comparison models rather than holding every model to one common harness.
- The listed Fable 5 result may include Opus 4.8 fallback. Kimi calls this out explicitly.
- Some comparison values are cited from benchmark owners, model-company release posts, or Kimi’s own reruns. Several K3 results are vendor-run.
- Kimi says its full technical report, including further training and evaluation detail, will arrive later. The model weights are planned for release by July 27, 2026.
Those conditions do not invalidate the table. They explain why it should be treated as a serious launch scorecard rather than the last word on K3’s rank.
The independent scoreboards still have not caught up
K3 is still absent from the live public data checked for Artificial Analysis, Agent Arena, LMArena, and LiveBench. That does not conflict with Kimi’s release post. It means the public evidence now has two layers:
- Kimi’s vendor benchmark table, which is detailed enough to assess and reproduce in parts.
- Independent leaderboard coverage, which is still not available for K3 on the live surfaces checked for this article.
This distinction is the core of the story. Buyers now have a performance claim to inspect, but they do not yet have a provider-neutral aggregate score or an independent agent-arena rank for the model.
What K3’s scorecard actually supports
Kimi describes K3 as a 2.8T-parameter model with native vision, a 1M-token context window, and an architecture built on Kimi Delta Attention and Attention Residuals. It says K3 is its most capable model and the first open 3T-class model.
The launch table supports calling K3 frontier-competitive on selected vendor-reported tests. Terminal-Bench 2.1 at 88.3, Program Bench at 77.8, FrontierSWE at 81.2, and GPQA-Diamond at 93.5 are not token demo figures. They are serious reported results, with methodology notes attached.
The same source also prevents the inflated conclusion. Kimi says K3 still trails Fable 5 and Sol overall. Its own table supports that mixed picture: K3 is strong, sometimes ahead, sometimes behind, and not the clear winner across every category.
The price sharpens the question. At $18 for a 1M-input plus 1M-output workload, K3 costs substantially more than K2.7 Code and Grok 4.5. Buyers are not being asked to trial a cheap long-context model. They are being asked to pay a frontier tariff for a model whose first public scorecard is promising but still vendor-run.
The practical buying decision
- Pick Grok 4.5 when raw API cost is the priority. Its $8 mixed-million-token workload is less than half K3’s $18.
- Pick GPT-5.6 Sol or Claude Fable 5 when their existing independent performance evidence justifies their higher costs.
- Consider K3 when 1M context, Kimi’s agent tooling, strong vendor-reported coding and reasoning results, and its $18 workload fit the job.
- Do not treat K3’s launch table as proof of general superiority until independent benchmark operators publish a K3 result under a disclosed, comparable configuration.
K3 now has a public price and a serious vendor scorecard. The next test is whether those results survive independent measurement.
Sources
- Kimi K3: Open Frontier Intelligence: launch post, full benchmark table, methodology notes, and limitations
- Kimi K3 API pricing
- Kimi K2.7 Code API pricing
- Kimi API model-pricing index
- Artificial Analysis: model comparison and Intelligence Index methodology
- Artificial Analysis: Terminal-Bench v2.1 evaluation
- Agent Arena leaderboard
- LMArena leaderboard
- LiveBench
- SpaceXAI / xAI developer model documentation
- Anthropic model pricing

