Model Comparisons
Head-to-head comparisons of the models on the AI leaderboard: benchmark scores and API pricing side by side, refreshed weekly. Each page answers the two questions that matter: which is better, and which is cheaper for your workload.
18 matchups · refreshed weekly
Claude Fable 5
92
vs
o3
95.1
Compare in depth →
Claude Fable 5
92
vs
Gemini 2.5 Pro
89.7
Compare in depth →
o3
95.1
vs
Gemini 2.5 Pro
89.7
Compare in depth →
Claude Opus 4.8
89.3
vs
o3
95.1
Compare in depth →
Claude Opus 4.8
89.3
vs
GPT-4.1
88.4
Compare in depth →
Claude Opus 4.8
89.3
vs
Gemini 2.5 Pro
89.7
Compare in depth →
Claude Opus 4.8
89.3
vs
Claude Sonnet 4.6
86.5
Compare in depth →
Claude Sonnet 4.6
86.5
vs
GPT-4o
85.2
Compare in depth →
Claude Sonnet 4.6
86.5
vs
Gemini 2.5 Flash
85.9
Compare in depth →
DeepSeek R1
88.2
vs
o3
95.1
Compare in depth →
DeepSeek R1
88.2
vs
Claude Opus 4.8
89.3
Compare in depth →
DeepSeek R1
88.2
vs
DeepSeek V3
83.9
Compare in depth →
Llama 4 Maverick
87.3
vs
GPT-4o
85.2
Compare in depth →
Gemini 2.5 Pro
89.7
vs
Gemini 2.5 Flash
85.9
Compare in depth →
GPT-4.1
88.4
vs
GPT-4o
85.2
Compare in depth →
Claude Haiku 4.5
79.8
vs
GPT-4o mini
79.8
Compare in depth →
Gemini 2.5 Flash
85.9
vs
GPT-4o mini
79.8
Compare in depth →
Grok 3
84.2
vs
GPT-4.1
88.4
Compare in depth →