Back to Models Directory
Kimi K3

Kimi K3

business Maker: Moonshot AI (Kimi) account_tree Family: Kimi K3

Overview

Kimi K3 is Moonshot AI's flagship large language model, a 2.8-trillion-parameter Mixture-of-Experts system built on a new architecture that combines Stable LatentMoE (activating roughly 16 of 896 experts per token) with Kimi Delta Attention and Attention Residuals. It was released on July 16, 2026 and runs with a 1,048,576-token context window and native multimodal input across text, images, and video.

The model is served through Kimi.com, the Kimi mobile app, Kimi Work, Kimi Code, and the Kimi API. On launch it uses maximum thinking effort by default, with low- and high-effort modes noted as follow-on updates. Moonshot published a broad benchmark comparison at launch that positions K3 competitively with Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol across coding, agentic, reasoning, and vision suites.

API pricing is $0.30 per million input tokens on a cache hit, $3.00 per million on a cache miss, and $15.00 per million output tokens. Full model weights were published on Hugging Face on July 26, 2026 (one day ahead of Moonshot's July 27 target) in the moonshotai/Kimi-K3-MXFP4 repository under the Kimi K3 License, with the technical report released alongside the weights.

Released2026-07-16
LicenseKimi K3 License (custom, weights published on Hugging Face 2026-07-26)
WeightsOpen weights
Parameters2.8T total · ~16 of 896 experts active per token
Context1M
ArchitectureMixture-of-Experts (Stable LatentMoE) with Kimi Delta Attention and Attention Residuals
ModalitiesText, Vision, Video
StatusGenerally available

Benchmarks

Kimi K3 benchmark comparison as published by Moonshot on the K3 launch page.

BenchmarkKimi K3Claude Fable 5GPT-5.6 SolClaude Opus 4.8GPT-5.5GLM-5.2
Coding — DeepSWE67.5%70%73%59%67%46.2%
Coding — Program Bench77.8%76.8%77.6%71.9%70.8%63.7%
Coding — Terminal Bench 2.188.3%84.6%88.8%84.6%83.4%82.7%
Coding — FrontierSWE81.2%86.6%71.3%66.7%64.9%67.3%
Coding — SWE Marathon42%35%39%40%14%13%
Coding — Kimi Code Bench 2.072.9%76.9%64.8%71.7%69%64.2%
Agentic — GDPval-AA v21668 score1760 score1748 score1600 score1494 score1514 score
Agentic — BrowseComp91.2%88%90.4%84.3%84.4%
Agentic — Toolathlon-Verified73.2%77.9%74.9%76.2%73.5%59.9%
Agentic — MCP Atlas84.2%84.7%83.6%83.6%82.8%82.6%
Agentic — Automation Bench30.8%29.1%29.7%27.2%22.7%12.9%
Agentic — AA-Briefcase1548 score1583 score1495 score1354 score1158 score1260 score
Reasoning — GPQA-Diamond93.5%92.6%94.1%91%93.5%91.2%
Reasoning — HLE-Full43.5%53.3%44.5%49.8%41.4%
Reasoning — HLE-Full w/ tools56%63%58%57.9%52.2%
Vision — MMMU-Pro81.6%81.2%83%78.9%81.2%
Vision — CharXiv (RQ)84.8%88.9%84.6%80.5%84.1%
Vision — MathVision94.3%94.8%95.8%86.7%92.2%
Vision — OmniDocBench91.1%89.8%85.8%87.9%89.4%

Comparison source ↗

This model's scores

  1. Terminal Bench 2.188.3%
  2. Program Bench77.8%
  3. SWE Marathon42%
  4. GPQA-Diamond93.5%
  5. BrowseComp91.2%
  6. MMMU-Pro81.6%
  7. MathVision94.3%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$3.00 / 1M tokens
Cached input$0.30 / 1M tokens
Output$15.00 / 1M tokens

Cache-miss input $3.00/M; cache-hit input $0.30/M.

Pricing source ↗

Strengths

  • Very large 2.8T-parameter MoE built with a new Stable LatentMoE + Kimi Delta Attention architecture
  • Native multimodal input across text, vision, and video with a 1M-token context
  • Strong agentic performance (91.2 on BrowseComp, 30.8 on Automation Bench, 73.5 on DECK-Bench in the maker's launch table)
  • Competitive coding (77.8 on Program Bench, 88.3 on Terminal Bench 2.1, 42.0 on SWE Marathon)
  • Open-weight release: MXFP4-quantized weights (~594 GB) published on Hugging Face on 2026-07-26 under the Kimi K3 License

Best for

  • Reach for it for long-horizon agentic coding and browser/tool workflows that need a very long context.
  • Reach for it when a task needs a native multimodal frontier model that can read images, charts, and video alongside text.
  • Reach for it when you want a frontier-class model from a Chinese lab with an open-weight release on the near horizon.

How to access

ProviderModel ID
Kimi API (Moonshot) ↗kimi-k3

FAQ

When was Kimi K3 released?

Moonshot AI released Kimi K3 on July 16, 2026, first through Kimi.com, the Kimi mobile app, Kimi Work, Kimi Code, and the Kimi API. Full model weights were then published on Hugging Face on July 26, 2026 (one day ahead of Moonshot's July 27 target) in the moonshotai/Kimi-K3-MXFP4 repository under the Kimi K3 License.

How much does Kimi K3 cost on the API?

Kimi K3 is billed at $0.30 per million input tokens on a cache hit, $3.00 per million input tokens on a cache miss, and $15.00 per million output tokens on the Kimi platform.

What is the context window of Kimi K3?

Kimi K3 has a 1,048,576-token context window (roughly 1M tokens) and accepts native text, image, and video input.

What architecture does Kimi K3 use?

Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model built on Moonshot's Stable LatentMoE (activating roughly 16 of 896 experts per token) with Kimi Delta Attention and Attention Residuals.