Which Model Should I Use?
Four quick questions, one ranked answer. Recommendations combine live benchmark scores from the weekly leaderboard with published API prices, so the answer comes from real numbers.
1 · What's the main task?
2 · What matters more?
3 · Open weights (self-hosting)?
4 · Very long documents (200K+ tokens)?
Recommendation
How it ranks: your task picks the benchmark column (or the combined index); "balanced" and "lowest cost" subtract a cost penalty proportional to the log of a typical monthly API bill (10K requests at 2K input / 500 output tokens); the open-weights and long-context answers filter the field. Retired models are excluded. Same data as the leaderboard and cost calculator. Shortlist here, then test the winner on your actual workload.