Kimi K3
Moonshot's 2.8T-parameter open-weight multimodal agentic model — frontier reasoning, native vision, and a 1M-token context window.
Choose an Elastic endpoint or a Dedicated deployment — both OpenAI-compatible. Tell us your workload and we'll get you a key.
Specs & substance
Spec sheet
Kimi K3 is a frontier open-weight model for long-horizon autonomy, combining reasoning with native text+image understanding in one model. It is designed for coding with minimal oversight across large repositories and terminal tool orchestration, as well as agentic knowledge work.
K3 has 2.8T total parameters with 104B activated. It is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) with a Stable LatentMoE framework activating 16 of 896 experts across 93 layers composed of 69 KDA + 24 Gated MLA. Its MoonViT-V2 vision encoder has 401M parameters. K3 uses MXFP4 weights with MXFP8 activations and was trained with quantization-aware training from the SFT stage onward, with a 1M-token context window.
Kimi K3 runs on Parasail via Elastic endpoints and Dedicated deployments, with OpenAI-compatible access and capacity sized to your workload.
Measured on the work that matters
Kimi K3 model-card results alongside selected comparison models.
All scores are self-reported by Moonshot in the Kimi K3 model card on Hugging Face; the comparison columns are in turn cited there from Artificial Analysis, Vals AI, and each benchmark’s official leaderboard. Where a benchmark is reported both with and without tool augmentation, we show the lower, no-tools number.