Kimi K2.5

Moonshot AI llm

Natively multimodal agentic model continually pretrained on roughly 15T mixed visual-text tokens, with an Agent Swarm mode that orchestrates up to 100 sub-agents per prompt.

overall ⌄

58.8

24 of 93 ranked

price · $/M tokens

0.6 / 3

in / out, as of 2026-07-25

to run it yourself ⌄

600 GB

minimum · 2.4 TB recommended

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
MMLU-Pro87.1 thinking
GPQA87.6 thinking
AIME 202596.1 thinking
SWE-bench76.8 70.8
SWE-bench Pro50.7
HLE31.5
BrowseComp60.6

Cheaper, and at least as good

On a 1,000-in, 1,000-out call these cost less and rank no lower — a hard comparison to argue with, and a narrow one. Value covers what it misses.

DeepSeek V4-Flash DeepSeek$0.00042overall 59.3
DeepSeek V4-Pro DeepSeek$0.00130overall 68.3
MiniMax M3 MiniMax$0.00150overall 65.1

Lineage

↑ finetune Kimi K2
view in full graph →

Metadata

orgMoonshot AI
released2026-01-27
params1T-A32B
licensemodified-mit
hf idmoonshotai/Kimi-K2.5