RepoMoonshot AI (Kimi)Moonshot AI (Kimi)published Jul 27, 2026seen 10h

MoonshotAI/Kimi-K3

Open original ↗

Captured source

source ↗
published Jul 27, 2026seen 10hcaptured 10hhttp 200method plain

MoonshotAI/Kimi-K3

Description: Open Frontier Intelligence

License: NOASSERTION

Stars: 2392

Forks: 179

Open issues: 6

Created: 2026-07-27T08:01:37Z

Pushed: 2026-07-28T02:40:06Z

Default branch: main

Fork: no

Archived: no

README:

📰 Tech Blog | 📄 Full Report

1. Model Introduction

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Key Features

  • New Architecture: Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE sparsity with a Stable LatentMoE framework that activates 16 out of 896 experts — yielding an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.
  • Long-Horizon Coding: Operating with minimal human oversight, Kimi K3 sustains long engineering sessions, navigates massive repositories, and orchestrates terminal tools — from GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and even chip design.
  • Agentic Knowledge Work: Kimi K3 advances end-to-end knowledge work, producing deep research with interactive visualizations, widgets and dashboards, and motion design and video editing, powered by its native multimodal architecture.
  • Native Multimodality & Long Context: Kimi K3 understands text, images, and video within the same model, and supports a 1-million-token context window.
  • Open Frontier Weights: We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further innovation.

2. Model Summary

3. Evaluation Results

Benchmark Kimi K3 (max) Claude Fable 5 (max, w/ fallback) GPT-5.6 Sol (max) Claude Opus 4.8 (max) GPT-5.5 (xhigh) GLM-5.2 (max)

Reasoning & Knowledge

GPQA Diamond 93.5 92.6 94.1 91.0 93.5 91.2

CritPt 23.4 28.6 32.3 20.9 27.1 20.9

AA-LCR 74.7 70.0 73.7 67.7 74.3 71.3

HLE-Full 43.5 / 56.0 53.3 / 63.0 44.5 / 58.0 49.8 / 57.9 41.4 / 52.2 —

Coding

DeepSWE 67.5 70.0 73.0 59.0 67.0 46.2

ProgramBench 77.8 76.8 77.6 71.9 70.8 63.7

Terminal-Bench 2.1 88.3 88.0 88.8 84.6 83.4 82.7

FrontierSWE 81.2 86.6 71.3 66.7 64.9 67.3

SWE-Marathon 42.0 35.0 39.0 40.0 14.0 13.0

PostTrainBench 36.6 41.4 34.6 34.1 28.4 34.3

MLS-Bench-Lite 48.3 49.9 46.2 42.8 35.5 40.4

SciCode 58.7 60.2 56.1 53.5 56.1 50.5

Kimi Code Bench 2.0 72.9 76.9 64.8 71.7 69.0 64.2

Agentic

BrowseComp 91.2 88.0 90.4 84.3 84.4 —

DeepSearchQA (F1) 95.0 94.2 — 93.1 — —

ResearchRubrics 76.2 — 73.8 73.5 64.0 71.1

GDPval-AA v2 (Elo) 1686 1747 1736 1593 1491 1510

Toolathlon-Verified 76.5 77.9 74.9 76.2 73.5 59.9

MCPMark-Verified 94.5 87.4 92.9 76.4 <td align="center" style="vertical-align: middle; text-align: center"

Excerpt shown — open the source for the full document.