Claude Sonnet 4

Anthropic llm

Anthropic's mid-tier Claude 4 model, succeeding Claude 3.7 Sonnet with improved coding and instruction-following. Scores are extended-thinking, single attempt, no tools.

overall ⌄

37.5

64 of 93 ranked

price · $/M tokens

3 / 15

in / out, as of 2026-07-25

to run it yourself ⌄

API only

no weights to run

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
GPQA75.4 thinking
AIME 202570.5 thinking
SWE-bench72.7 none64.9

Cheaper, and at least as good

On a 1,000-in, 1,000-out call these cost less and rank no lower — a hard comparison to argue with, and a narrow one. Value covers what it misses.

Devstral Small 2 Mistral AI$0.00040overall 42.7
Ministral 3 14B Reasoning Mistral AI$0.00040overall 48.1
DeepSeek V4-Flash DeepSeek$0.00042overall 59.3
DeepSeek V4-Pro DeepSeek$0.00130overall 68.3
MiniMax M2.1 MiniMax$0.00150overall 50.9

Lineage

No recorded lineage — root or standalone model.

view in full graph →

Metadata

orgAnthropic
released2025-05-22
licenseproprietary