AI Model Benchmark Scores

Compare 67 LLMs by MMLU, HumanEval, MATH, and Arena Elo scores. Find the best model for your use case — ranked by actual benchmark performance.

MMLU
HumanEval
MATH
Arena Elo
Composite

MMLU Scores — All Models

Premium
Mid
Budget
# Model Tier MMLU HumanEval MATH Arena Elo Composite Cost $/1M

Head-to-Head Benchmark Comparison

Understanding AI Model Benchmarks

Benchmark scores provide a standardized way to compare AI model capabilities. Here's what each benchmark measures:

Important: Benchmark scores are estimates based on published results and community data. Actual performance varies by task, prompt, and use case. Always test with your specific workload before committing to a model.

This was a snapshot. What about next month?
Prices change. New models launch. Our tools catch what a one-time calculation can't — and saves you money every month.
Free Tools → 🔍 Free audit first

All Tools Are Free

No signup required to 67-model comparison, migration code snippets, PDF reports, price alerts, and cost monitoring. ✅ All tools free.

Free Tools →
Free Tools — Optimize Your Costs

Get model routing, caching strategies, and save 40%+ on API costs. 100% free — no signup required.

Run Free Cost Audit → Free Tools →

No signup required · 100% free