Claude Sonnet 4.5

Anthropic llm

Anthropic's coding-focused model that set a new state of the art on SWE-bench Verified at launch. The 82.0% figure reported with additional test-time compute is deliberately not recorded.

overall ⌄

54.4

30 of 93 ranked

price · $/M tokens

3 / 15

in / out, as of 2026-07-25

to run it yourself ⌄

API only

no weights to run

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
SWE-bench77.2 71.4 70.6

Cheaper, and at least as good

On a 1,000-in, 1,000-out call these cost less and rank no lower — a hard comparison to argue with, and a narrow one. Value covers what it misses.

DeepSeek V4-Flash DeepSeek$0.00042overall 59.3
DeepSeek V4-Pro DeepSeek$0.00130overall 68.3
MiniMax M3 MiniMax$0.00150overall 65.1
Kimi K2.5 Moonshot AI$0.00360overall 58.8
Qwen3.5 397B A17B Alibaba$0.00420overall 63.7

Lineage

No recorded lineage — root or standalone model.

view in full graph →

Metadata

orgAnthropic
released2025-09-29
licenseproprietary