Grok 4

xAI llm

xAI's reasoning-native flagship model, which uses tools during its chain of thought by default. xAI publishes its benchmark results only as chart images, so no sourced scores are recorded here yet.

overall ⌄

nothing published

price · $/M tokens

no published price

to run it yourself ⌄

API only

no weights to run

Benchmarks

What its maker published, and what anyone else measured. Every figure links the document it came from.

benchmarkpublishedmeasured
MMLU-Pro86.6
GPQA87.7
AIME 202592.7
HLE23.9

Lineage

↓ finetune Grok 4.1
view in full graph →

Metadata

orgxAI
released2025-07-09
licenseproprietary