New release · Jul 2026

Kimi K3

Moonshot's 2.8T-parameter open-weight multimodal agentic model — frontier reasoning, native vision, and a 1M-token context window.

Choose an Elastic endpoint or a Dedicated deployment — both OpenAI-compatible. Tell us your workload and we'll get you a key.

By Moonshot · Multimodal agentic · Kimi K3 License
Model details

Specs & substance

Spec sheet

Developed byMoonshot
Model familyKimi
Use caseMultimodal agentic
ModalityLLM
Context window1M tokens
ArchitectureMoE + KDA
VersionK3
LicenseKimi K3 License
Pricing$3.00 input/M · $15.00 output/M · $0.30 cache read/M
ReleasedJul 2026

Kimi K3 is a frontier open-weight model for long-horizon autonomy, combining reasoning with native text+image understanding in one model. It is designed for coding with minimal oversight across large repositories and terminal tool orchestration, as well as agentic knowledge work.

K3 has 2.8T total parameters with 104B activated. It is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) with a Stable LatentMoE framework activating 16 of 896 experts across 93 layers composed of 69 KDA + 24 Gated MLA. Its MoonViT-V2 vision encoder has 401M parameters. K3 uses MXFP4 weights with MXFP8 activations and was trained with quantization-aware training from the SFT stage onward, with a 1M-token context window.

Kimi K3 runs on Parasail via Elastic endpoints and Dedicated deployments, with OpenAI-compatible access and capacity sized to your workload.

Benchmarks

Measured on the work that matters

Kimi K3 model-card results alongside selected comparison models.

Reasoning
Benchmark
Kimi K3Moonshot · Open
GPT-5.6 SolOpenAI · Closed
Claude Fable 5Anthropic · Closed
Claude Opus 4.8Anthropic · Closed
GPT-5.5OpenAI · Closed
GLM-5.2Z.AI · Open
GPQA Diamond%
93.5
94.1
92.6
91.0
93.5
91.2
HLE-Full%
43.5
44.5
53.3
49.8
41.4
AA-LCR%
74.7
73.7
70.0
67.7
74.3
71.3
CritPt%
23.4
32.3
28.6
20.9
27.1
20.9
Coding
Benchmark
Kimi K3Moonshot · Open
GPT-5.6 SolOpenAI · Closed
Claude Fable 5Anthropic · Closed
Claude Opus 4.8Anthropic · Closed
GPT-5.5OpenAI · Closed
GLM-5.2Z.AI · Open
Terminal-Bench 2.1%
88.3
88.8
88.0
84.6
83.4
82.7
FrontierSWE%
81.2
71.3
86.6
66.7
64.9
67.3
ProgramBench%
77.8
77.6
76.8
71.9
70.8
63.7
DeepSWE%
67.5
73.0
70.0
59.0
67.0
46.2
SWE-Marathon%
42.0
39.0
35.0
40.0
14.0
13.0
SciCode%
58.7
56.1
60.2
53.5
56.1
50.5
Agentic
Benchmark
Kimi K3Moonshot · Open
GPT-5.6 SolOpenAI · Closed
Claude Fable 5Anthropic · Closed
Claude Opus 4.8Anthropic · Closed
GPT-5.5OpenAI · Closed
GLM-5.2Z.AI · Open
MCPMark-Verified%
94.5
92.9
87.4
76.4
92.9
BrowseComp%
91.2
90.4
88.0
84.3
84.4
OSWorld-Verified%
84.8
83.0
85.0
83.4
79.0
MCP-Atlas%
84.2
83.6
84.7
83.6
82.8
82.6
Toolathlon-Verified%
76.5
74.9
77.9
76.2
73.5
59.9
APEX-Agents%
41.0
39.9
43.3
39.4
38.5
35.6
GDPval-AA v2Elo
1686
1736
1747
1593
1491
1510
AA-BriefcaseElo
1548
1495
1583
1354
1158
1260
Vision
Benchmark
Kimi K3Moonshot · Open
GPT-5.6 SolOpenAI · Closed
Claude Fable 5Anthropic · Closed
Claude Opus 4.8Anthropic · Closed
GPT-5.5OpenAI · Closed
GLM-5.2Z.AI · Open
OmniDocBench%
91.1
85.8
89.8
87.9
89.4
Video-MME (w. sub)%
90.0
89.5
86.0
89.3
MathVision%
94.3
95.8
94.8
86.7
92.2
CharXiv (RQ)%
84.8
84.6
88.9
80.5
84.1
MMMU-Pro%
81.6
83.0
81.2
78.9
81.2
WorldVQA ForceAnswer%
51.0
41.8
56.7
39.1
38.5

All scores are self-reported by Moonshot in the Kimi K3 model card on Hugging Face; the comparison columns are in turn cited there from Artificial Analysis, Vals AI, and each benchmark’s official leaderboard. Where a benchmark is reported both with and without tool augmentation, we show the lower, no-tools number.

Bring Kimi K3 to your workload.

Request access and tell us what you're building — we'll help get you set up.