AI Benchmark Leaderboard
State-of-the-art results across key evaluation benchmarks.
Updated 2026-09-21
Full leaderboards on Papers With Code ↗Results reflect published evaluations and may use different prompting strategies or few-shot settings — direct comparison across rows should be made cautiously. Scores auto-update weekly via the Papers With Code API.
Massive Multitask Language Understanding
57 diverse subjects spanning STEM, humanities, social sciences, and professional domains. Tests breadth of knowledge and reasoning.
reasoningFull leaderboard ↗
Metric: Accuracy (%)
Higher is better: Yes
