Claude 3.7 Sonnet

Anthropic llm

Anthropic's first hybrid reasoning model, letting users toggle standard and extended-thinking modes. Anthropic publishes no single-attempt extended-thinking GPQA figure for this model, so the standard-mode score is recorded.

overall ⌄

35.9

70 of 93 ranked

price · $/M tokens

no published price

to run it yourself ⌄

API only

no weights to run

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
GPQA68
SWE-bench63.7 52.8

Lineage

No recorded lineage — root or standalone model.

view in full graph →

Metadata

orgAnthropic
released2025-02-24
licenseproprietary