LLM benchmarks

Compare 100 LLM benchmarks by score and price

Tokenscost tracks 100 public LLM evaluation benchmarks across 8 categories — reasoning, math, coding, multimodal, multilingual, instruction following, and agent multi-turn — paired with live $/1M token pricing for every model.

Benchmarks100
Categories8
Models scored13

All 100 tracked benchmarks

A crawler-readable index. Each row links to the live card on /benchmarks with full leaderboards and best score-per-dollar.

BenchmarkCategoryTop modelTop score
GPQA DiamondReasoning and LogicGemini 3.1 Pro94.1
Humanity's Last ExamReasoning and LogicClaude Fable 553.3
ARC-AGI-2Reasoning and LogicGemini 3.1 Pro31.1
LCR (Long-Context Retrieval)Reasoning and LogicGPT-5.5 Pro75.7
AIME 2025Mathematical Problem SolvingGemini 3.1 Pro98.4
MATH-500Mathematical Problem SolvingClaude Fable 598.2
AIME 2024Mathematical Problem SolvingGemini 3.1 Pro98.1
SWE-Bench VerifiedComputer Science and ProgrammingClaude Fable 579.4
LiveCodeBenchComputer Science and ProgrammingGemini 3 Pro91.7
SciCodeComputer Science and ProgrammingClaude Fable 560.2
HumanEval+Computer Science and ProgrammingClaude Fable 596.3
MMLU-ProGeneral KnowledgeGemini 3.1 Pro89.8
SimpleQAGeneral KnowledgeGPT-5.5 Pro62.5
τ²-BenchMulti-turnClaude Fable 578.5
TerminalBenchMulti-turnClaude Fable 562.9
IFBenchInstruction FollowingClaude Fable 586.2
MMLU-ProXMultilingualGemini 3.1 Pro84.6
MMMUMultimodalGemini 3.1 Pro83.6
MathVistaMultimodalGemini 3.1 Pro82.1
ChartQAMultimodalGemini 3.1 Pro92.8
DROPReasoning and LogicGPT-5.5 Pro93.2
BIG-Bench HardReasoning and LogicClaude Fable 594.8
MuSRReasoning and LogicClaude Fable 582.5
GSM8KMathematical Problem SolvingClaude Fable 598.6
FrontierMathMathematical Problem SolvingGPT-5.5 Pro32.4
MBPP+Computer Science and ProgrammingClaude Fable 592.1
Aider PolyglotComputer Science and ProgrammingClaude Fable 584.7
TruthfulQAGeneral KnowledgeClaude Fable 578.4
MGSMMultilingualGemini 3.1 Pro94.6
MT-BenchMulti-turnClaude Fable 596.4
ARC-AGI-1Reasoning and LogicGPT-5.5 Pro87.5
ZebraLogicReasoning and LogicClaude Fable 581.2
FOLIOReasoning and LogicClaude Fable 589.3
LogiQA 2.0Reasoning and LogicGPT-5.5 Pro84.6
ReClorReasoning and LogicClaude Fable 588.4
ProofWriterReasoning and LogicClaude Fable 596.2
AGIEvalReasoning and LogicGPT-5.5 Pro87.8
NoChaReasoning and LogicGemini 3.1 Pro72.3
OlympicArenaReasoning and LogicGPT-5.5 Pro62.4
NPHardEvalReasoning and LogicGPT-5.5 Pro68.2
MATHMathematical Problem SolvingClaude Fable 598.1
OlympiadBenchMathematical Problem SolvingGPT-5.5 Pro74.6
MathBenchMathematical Problem SolvingClaude Fable 589.4
ASDivMathematical Problem SolvingClaude Fable 598.6
SVAMPMathematical Problem SolvingClaude Fable 597.5
MultiArithMathematical Problem SolvingClaude Fable 599.5
TheoremQAMathematical Problem SolvingGPT-5.5 Pro70.2
PutnamBenchMathematical Problem SolvingGPT-5.5 Pro18.4
MATH Level 5Mathematical Problem SolvingClaude Fable 591.3
GSM-HardMathematical Problem SolvingClaude Fable 588.7
HumanEvalComputer Science and ProgrammingClaude Fable 599.4
MBPPComputer Science and ProgrammingClaude Fable 595.8
CodeContestsComputer Science and ProgrammingGPT-5.5 Pro58.7
APPSComputer Science and ProgrammingClaude Fable 562.8
CRUXEvalComputer Science and ProgrammingClaude Fable 590.2
BigCodeBenchComputer Science and ProgrammingClaude Fable 561.4
RepoBenchComputer Science and ProgrammingClaude Fable 572.6
SWE-Bench LiteComputer Science and ProgrammingClaude Fable 578.4
ClassEvalComputer Science and ProgrammingClaude Fable 584.2
DS-1000Computer Science and ProgrammingClaude Fable 587.5
MMBenchMultimodalGemini 3.1 Pro91.4
SEED-BenchMultimodalGemini 3.1 Pro78.9
RealWorldQAMultimodalGemini 3.1 Pro78.2
AI2DMultimodalGemini 3.1 Pro95.1
DocVQAMultimodalGemini 3.1 Pro96.4
InfographicVQAMultimodalGemini 3.1 Pro87.2
OCRBenchMultimodalGemini 3.1 Pro92.8
BLINKMultimodalGemini 3.1 Pro72.4
MM-VetMultimodalGemini 3.1 Pro84.6
ScreenSpotMultimodalGemini 3.1 Pro88.4
MMLUGeneral KnowledgeClaude Fable 592.6
TriviaQAGeneral KnowledgeClaude Fable 594.7
Natural QuestionsGeneral KnowledgeClaude Fable 578.3
OpenBookQAGeneral KnowledgeClaude Fable 596.4
ARC-ChallengeGeneral KnowledgeClaude Fable 597.2
HellaSwagGeneral KnowledgeClaude Fable 596.1
WinoGrandeGeneral KnowledgeClaude Fable 594.5
CommonsenseQAGeneral KnowledgeClaude Fable 590.7
PIQAGeneral KnowledgeClaude Fable 592.5
Social IQaGeneral KnowledgeClaude Fable 587.4
IFEvalInstruction FollowingClaude Fable 592.4
AlpacaEval 2.0Instruction FollowingClaude Fable 578.6
MT-Bench-101Instruction FollowingClaude Fable 589.7
FLASKInstruction FollowingClaude Fable 587.2
InfoBenchInstruction FollowingClaude Fable 590.5
FollowBenchInstruction FollowingClaude Fable 582.4
Arena-HardInstruction FollowingClaude Fable 594.2
XCOPAMultilingualGemini 3.1 Pro94.5
XNLIMultilingualGemini 3.1 Pro89.4
FLORES-200MultilingualGemini 3.1 Pro62.8
M3ExamMultilingualGemini 3.1 Pro76.8
Global-MMLUMultilingualGemini 3.1 Pro85.7
BelebeleMultilingualGemini 3.1 Pro91.4
MINTMulti-turnClaude Fable 572.5
BFCLMulti-turnClaude Fable 587.6
AgentBenchMulti-turnClaude Fable 565.2
GAIAMulti-turnClaude Fable 574.6
WebArenaMulti-turnClaude Fable 558.4
OSWorldMulti-turnClaude Fable 552.7
SWE-LancerMulti-turnClaude Fable 546.8

Reasoning and Logic (17)

Reasoning and Logic

GPQA Diamond

Google-proof graduate questions in biology, physics, chemistry. Top: Gemini 3.1 Pro 94.1 · Claude Fable 5 92.6 · GPT-5.5 Pro 89.4.

Reasoning and Logic

Humanity's Last Exam

Frontier multidisciplinary stress test from Center for AI Safety. Top: Claude Fable 5 53.3 · Gemini 3.1 Pro 44.7 · GPT-5.5 Pro 41.6.

Reasoning and Logic

ARC-AGI-2

Abstract visual reasoning puzzles trivial for humans, hard for LLMs. Top: Gemini 3.1 Pro 31.1 · Claude Fable 5 17.6 · GPT-5.5 Pro 15.9.

Reasoning and Logic

LCR (Long-Context Retrieval)

Find and use info buried inside very long documents. Top: GPT-5.5 Pro 75.7 · GPT-5.5 75.6 · GPT-5.4 75.0.

Reasoning and Logic

DROP

Discrete reasoning over paragraphs — numbers, dates, sorting. Top: GPT-5.5 Pro 93.2 · Claude Fable 5 92.6 · Gemini 3.1 Pro 91.8.

Reasoning and Logic

BIG-Bench Hard

23 BIG-Bench tasks where prior models trailed humans. Top: Claude Fable 5 94.8 · GPT-5.5 Pro 94.1 · Gemini 3.1 Pro 93.5.

Reasoning and Logic

MuSR

Multistep soft reasoning over natural-language narratives. Top: Claude Fable 5 82.5 · GPT-5.5 Pro 80.1 · Gemini 3.1 Pro 78.4.

Reasoning and Logic

ARC-AGI-1

Abstraction and Reasoning Corpus — visual pattern puzzles. Top: GPT-5.5 Pro 87.5 · Claude Fable 5 85.2 · Gemini 3.1 Pro 83.0.

Reasoning and Logic

ZebraLogic

Logic-grid puzzles testing constraint-satisfaction reasoning. Top: Claude Fable 5 81.2 · GPT-5.5 Pro 78.4 · Gemini 3.1 Pro 74.6.

Reasoning and Logic

FOLIO

First-order logic natural-language entailment. Top: Claude Fable 5 89.3 · GPT-5.5 Pro 87.6 · Gemini 3.1 Pro 86.1.

Reasoning and Logic

LogiQA 2.0

Civil-service style logical reading comprehension. Top: GPT-5.5 Pro 84.6 · Claude Fable 5 83.1 · Gemini 3.1 Pro 80.7.

Reasoning and Logic

ReClor

LSAT-style logical reasoning multiple choice. Top: Claude Fable 5 88.4 · GPT-5.5 Pro 87.2 · Gemini 3.1 Pro 85.9.

Reasoning and Logic

ProofWriter

Synthetic multi-step natural-language proofs. Top: Claude Fable 5 96.2 · GPT-5.5 Pro 95.4 · Gemini 3.1 Pro 94.1.

Reasoning and Logic

AGIEval

Human-centric exams (SAT, LSAT, GRE, Gaokao). Top: GPT-5.5 Pro 87.8 · Claude Fable 5 86.2 · Gemini 3.1 Pro 84.5.

Reasoning and Logic

NoCha

Novel-length claim verification over 100k-token books. Top: Gemini 3.1 Pro 72.3 · Claude Fable 5 69.8 · GPT-5.5 Pro 65.4.

Reasoning and Logic

OlympicArena

Multi-discipline Olympiad-level reasoning across 7 subjects. Top: GPT-5.5 Pro 62.4 · Claude Fable 5 60.7 · Gemini 3.1 Pro 58.9.

Reasoning and Logic

NPHardEval

NP-complete/hard algorithmic problem solving. Top: GPT-5.5 Pro 68.2 · Claude Fable 5 66.5 · o3 62.1.

Multi-turn (10)

Multi-turn

τ²-Bench

Multi-turn agent tasks in airline and retail domains. Top: Claude Fable 5 78.5 · Claude Opus 4.8 74.0 · GPT-5.5 Pro 72.1.

Multi-turn

TerminalBench

Long-horizon shell tasks — agents controlling a real terminal. Top: Claude Fable 5 62.9 · GPT-5.5 Pro 57.6 · Claude Opus 4.8 54.5.

Multi-turn

MT-Bench

GPT-4-judged two-turn conversational quality benchmark. Top: Claude Fable 5 96.4 · GPT-5.5 Pro 95.8 · Gemini 3.1 Pro 95.1.

Multi-turn

MINT

Multi-turn tool use and natural language feedback. Top: Claude Fable 5 72.5 · GPT-5.5 Pro 71.2 · Gemini 3.1 Pro 69.4.

Multi-turn

BFCL

Berkeley Function Calling Leaderboard. Top: Claude Fable 5 87.6 · GPT-5.5 Pro 86.4 · Gemini 3.1 Pro 84.8.

Multi-turn

AgentBench

Multi-environment evaluation of LLM agents. Top: Claude Fable 5 65.2 · GPT-5.5 Pro 63.8 · Gemini 3.1 Pro 61.7.

Multi-turn

GAIA

General AI Assistants benchmark — real-world tasks. Top: Claude Fable 5 74.6 · GPT-5.5 Pro 72.8 · Gemini 3.1 Pro 70.4.

Multi-turn

WebArena

Realistic web-task execution in a sandbox. Top: Claude Fable 5 58.4 · GPT-5.5 Pro 56.7 · Gemini 3.1 Pro 54.2.

Multi-turn

OSWorld

Real desktop OS computer-use tasks. Top: Claude Fable 5 52.7 · GPT-5.5 Pro 50.4 · Gemini 3.1 Pro 48.6.

Multi-turn

SWE-Lancer

Real Upwork freelance software engineering tasks. Top: Claude Fable 5 46.8 · GPT-5.5 Pro 44.5 · Gemini 3.1 Pro 41.7.

Mathematical Problem Solving (15)

Mathematical Problem Solving

AIME 2025

American Invitational Math Exam — competition-grade problems. Top: Gemini 3.1 Pro 98.4 · GPT-5.5 Pro 97.0 · Claude Fable 5 96.7.

Mathematical Problem Solving

MATH-500

500 competition math problems across algebra, geometry, number theory. Top: Claude Fable 5 98.2 · Gemini 3.1 Pro 97.9 · GPT-5.5 Pro 97.4.

Mathematical Problem Solving

AIME 2024

Prior-year AIME, widely cited in model release notes. Top: Gemini 3.1 Pro 98.1 · Claude Fable 5 97.5 · GPT-5.5 Pro 97.1.

Mathematical Problem Solving

GSM8K

Grade-school math word problems requiring chain-of-thought. Top: Claude Fable 5 98.6 · Gemini 3.1 Pro 98.3 · GPT-5.5 Pro 98.1.

Mathematical Problem Solving

FrontierMath

Research-level mathematics from Epoch AI — extremely hard. Top: GPT-5.5 Pro 32.4 · Claude Fable 5 28.7 · Gemini 3.1 Pro 26.5.

Mathematical Problem Solving

MATH

Hendrycks 12,500 competition math problems. Top: Claude Fable 5 98.1 · GPT-5.5 Pro 97.8 · Gemini 3.1 Pro 97.2.

Mathematical Problem Solving

OlympiadBench

Bilingual Olympiad math and physics problems. Top: GPT-5.5 Pro 74.6 · Claude Fable 5 72.8 · Gemini 3.1 Pro 70.4.

Mathematical Problem Solving

MathBench

Hierarchical math from primary to college level. Top: Claude Fable 5 89.4 · GPT-5.5 Pro 88.7 · Gemini 3.1 Pro 87.5.

Mathematical Problem Solving

ASDiv

Diverse grade-school arithmetic word problems. Top: Claude Fable 5 98.6 · GPT-5.5 Pro 98.4 · Gemini 3.1 Pro 98.0.

Mathematical Problem Solving

SVAMP

Math word problems with sensitivity perturbations. Top: Claude Fable 5 97.5 · GPT-5.5 Pro 97.1 · Gemini 3.1 Pro 96.4.

Mathematical Problem Solving

MultiArith

Multi-step arithmetic word problems. Top: Claude Fable 5 99.5 · GPT-5.5 Pro 99.4 · Gemini 3.1 Pro 99.2.

Mathematical Problem Solving

TheoremQA

Theorem-application questions across STEM. Top: GPT-5.5 Pro 70.2 · Claude Fable 5 68.7 · Gemini 3.1 Pro 66.4.

Mathematical Problem Solving

PutnamBench

Formalized Putnam competition problems in Lean/Coq/Isabelle. Top: GPT-5.5 Pro 18.4 · Claude Fable 5 16.9 · Gemini 3.1 Pro 14.2.

Mathematical Problem Solving

MATH Level 5

Hardest subset of the MATH benchmark. Top: Claude Fable 5 91.3 · GPT-5.5 Pro 90.7 · Gemini 3.1 Pro 88.6.

Mathematical Problem Solving

GSM-Hard

GSM8K with adversarially large numbers. Top: Claude Fable 5 88.7 · GPT-5.5 Pro 87.5 · Gemini 3.1 Pro 85.9.

Computer Science and Programming (16)

Computer Science and Programming

SWE-Bench Verified

Resolve real GitHub issues in real Python repos. Top: Claude Fable 5 79.4 · Claude Opus 4.8 74.5 · GPT-5.5 Pro 72.8.

Computer Science and Programming

LiveCodeBench

Continuously updated coding contests — leak-resistant. Top: Gemini 3 Pro 91.7 · Gemini 3.1 Pro 90.8 · Claude Fable 5 89.3.

Computer Science and Programming

SciCode

Scientific computing tasks: numerics, simulations, ML pipelines. Top: Claude Fable 5 60.2 · Gemini 3.1 Pro 58.9 · GPT-5.5 Pro 57.4.

Computer Science and Programming

HumanEval+

Classic Python function synthesis with hardened tests. Top: Claude Fable 5 96.3 · GPT-5.5 Pro 95.7 · Gemini 3.1 Pro 95.1.

Computer Science and Programming

MBPP+

Mostly Basic Python Problems with hardened test cases. Top: Claude Fable 5 92.1 · GPT-5.5 Pro 91.3 · Gemini 3.1 Pro 90.4.

Computer Science and Programming

Aider Polyglot

Multi-language code editing benchmark from the Aider project. Top: Claude Fable 5 84.7 · GPT-5.5 Pro 81.2 · Claude Opus 4.8 79.5.

Computer Science and Programming

HumanEval

OpenAI 164 Python coding problems (pass@1). Top: Claude Fable 5 99.4 · GPT-5.5 Pro 99.1 · Gemini 3.1 Pro 98.8.

Computer Science and Programming

MBPP

Mostly Basic Python Problems (pass@1). Top: Claude Fable 5 95.8 · GPT-5.5 Pro 95.2 · Gemini 3.1 Pro 94.4.

Computer Science and Programming

CodeContests

Competitive programming problems from DeepMind. Top: GPT-5.5 Pro 58.7 · Claude Fable 5 56.9 · Gemini 3.1 Pro 54.3.

Computer Science and Programming

APPS

Coding interview problems across three difficulty tiers. Top: Claude Fable 5 62.8 · GPT-5.5 Pro 61.4 · Gemini 3.1 Pro 59.7.

Computer Science and Programming

CRUXEval

Code understanding via input and output prediction. Top: Claude Fable 5 90.2 · GPT-5.5 Pro 89.4 · Gemini 3.1 Pro 87.6.

Computer Science and Programming

BigCodeBench

Practical coding tasks across 139 Python libraries. Top: Claude Fable 5 61.4 · GPT-5.5 Pro 60.2 · Gemini 3.1 Pro 57.8.

Computer Science and Programming

RepoBench

Repository-level code completion benchmark. Top: Claude Fable 5 72.6 · GPT-5.5 Pro 71.4 · Gemini 3.1 Pro 69.7.

Computer Science and Programming

SWE-Bench Lite

300 curated real GitHub issue fixes. Top: Claude Fable 5 78.4 · GPT-5.5 Pro 76.9 · Gemini 3.1 Pro 74.2.

Computer Science and Programming

ClassEval

Class-level Python code generation. Top: Claude Fable 5 84.2 · GPT-5.5 Pro 83.4 · Gemini 3.1 Pro 81.6.

Computer Science and Programming

DS-1000

Data-science coding problems across 7 libraries. Top: Claude Fable 5 87.5 · GPT-5.5 Pro 86.4 · Gemini 3.1 Pro 84.6.

General Knowledge (13)

General Knowledge

MMLU-Pro

Harder, deduplicated successor to MMLU across 57 subjects. Top: Gemini 3.1 Pro 89.8 · Claude Fable 5 89.5 · GPT-5.5 Pro 88.2.

General Knowledge

SimpleQA

Short factual questions — hallucination stress test. Top: GPT-5.5 Pro 62.5 · Gemini 3.1 Pro 58.1 · Claude Fable 5 53.9.

General Knowledge

TruthfulQA

Questions designed to elicit common human misconceptions. Top: Claude Fable 5 78.4 · GPT-5.5 Pro 76.2 · Gemini 3.1 Pro 74.8.

General Knowledge

MMLU

Massive multitask language understanding across 57 subjects. Top: Claude Fable 5 92.6 · GPT-5.5 Pro 92.1 · Gemini 3.1 Pro 91.5.

General Knowledge

TriviaQA

Open-domain trivia question answering. Top: Claude Fable 5 94.7 · GPT-5.5 Pro 94.2 · Gemini 3.1 Pro 93.5.

General Knowledge

Natural Questions

Real Google search queries answered from Wikipedia. Top: Claude Fable 5 78.3 · GPT-5.5 Pro 77.5 · Gemini 3.1 Pro 76.4.

General Knowledge

OpenBookQA

Elementary-science QA requiring multi-hop reasoning. Top: Claude Fable 5 96.4 · GPT-5.5 Pro 95.8 · Gemini 3.1 Pro 95.1.

General Knowledge

ARC-Challenge

AI2 Reasoning Challenge — grade-school science. Top: Claude Fable 5 97.2 · GPT-5.5 Pro 96.8 · Gemini 3.1 Pro 96.3.

General Knowledge

HellaSwag

Commonsense sentence completion. Top: Claude Fable 5 96.1 · GPT-5.5 Pro 95.7 · Gemini 3.1 Pro 95.2.

General Knowledge

WinoGrande

Adversarial Winograd-style commonsense. Top: Claude Fable 5 94.5 · GPT-5.5 Pro 94.0 · Gemini 3.1 Pro 93.4.

General Knowledge

CommonsenseQA

Commonsense QA grounded in ConceptNet. Top: Claude Fable 5 90.7 · GPT-5.5 Pro 90.1 · Gemini 3.1 Pro 89.4.

General Knowledge

PIQA

Physical commonsense reasoning. Top: Claude Fable 5 92.5 · GPT-5.5 Pro 92.0 · Gemini 3.1 Pro 91.4.

General Knowledge

Social IQa

Social commonsense reasoning about everyday interactions. Top: Claude Fable 5 87.4 · GPT-5.5 Pro 86.8 · Gemini 3.1 Pro 86.1.

Instruction Following (8)

Instruction Following

IFBench

Adherence to nuanced writing constraints and formats. Top: Claude Fable 5 86.2 · Gemini 3.1 Pro 84.7 · GPT-5.5 Pro 83.4.

Instruction Following

IFEval

Verifiable instruction-following constraints. Top: Claude Fable 5 92.4 · GPT-5.5 Pro 91.6 · Gemini 3.1 Pro 90.5.

Instruction Following

AlpacaEval 2.0

Length-controlled win-rate vs GPT-4-Turbo. Top: Claude Fable 5 78.6 · GPT-5.5 Pro 76.2 · Gemini 3.1 Pro 73.4.

Instruction Following

MT-Bench-101

Fine-grained multi-turn dialogue capability. Top: Claude Fable 5 89.7 · GPT-5.5 Pro 88.4 · Gemini 3.1 Pro 87.0.

Instruction Following

FLASK

Fine-grained skill-set evaluation. Top: Claude Fable 5 87.2 · GPT-5.5 Pro 86.0 · Gemini 3.1 Pro 84.7.

Instruction Following

InfoBench

Decomposed-requirement instruction following. Top: Claude Fable 5 90.5 · GPT-5.5 Pro 89.7 · Gemini 3.1 Pro 88.4.

Instruction Following

FollowBench

Multi-level fine-grained constraint following. Top: Claude Fable 5 82.4 · GPT-5.5 Pro 81.2 · Gemini 3.1 Pro 79.8.

Instruction Following

Arena-Hard

500 challenging real-user prompts from Chatbot Arena. Top: Claude Fable 5 94.2 · GPT-5.5 Pro 92.7 · Gemini 3.1 Pro 90.4.

Multilingual (8)

Multilingual

MMLU-ProX

MMLU-Pro translated into 14 languages. Top: Gemini 3.1 Pro 84.6 · Claude Fable 5 83.2 · GPT-5.5 Pro 82.0.

Multilingual

MGSM

GSM8K translated into 10 typologically diverse languages. Top: Gemini 3.1 Pro 94.6 · Claude Fable 5 93.8 · GPT-5.5 Pro 92.5.

Multilingual

XCOPA

Causal commonsense reasoning across 11 languages. Top: Gemini 3.1 Pro 94.5 · Claude Fable 5 93.7 · GPT-5.5 Pro 92.8.

Multilingual

XNLI

Natural language inference in 15 languages. Top: Gemini 3.1 Pro 89.4 · Claude Fable 5 88.5 · GPT-5.5 Pro 87.7.

Multilingual

FLORES-200

Translation quality across 200 languages. Top: Gemini 3.1 Pro 62.8 · Claude Fable 5 61.4 · GPT-5.5 Pro 60.2.

Multilingual

M3Exam

Real human exam questions across 9 languages. Top: Gemini 3.1 Pro 76.8 · Claude Fable 5 75.4 · GPT-5.5 Pro 74.2.

Multilingual

Global-MMLU

Culturally-translated MMLU across 42 languages. Top: Gemini 3.1 Pro 85.7 · Claude Fable 5 84.6 · GPT-5.5 Pro 83.5.

Multilingual

Belebele

Multilingual reading comprehension in 122 languages. Top: Gemini 3.1 Pro 91.4 · Claude Fable 5 90.5 · GPT-5.5 Pro 89.6.

Multimodal (13)

Multimodal

MMMU

College-level multimodal questions across 30 subjects. Top: Gemini 3.1 Pro 83.6 · Claude Fable 5 81.4 · GPT-5.5 Pro 80.7.

Multimodal

MathVista

Visual mathematical reasoning across charts, diagrams, geometry. Top: Gemini 3.1 Pro 82.1 · Claude Fable 5 80.4 · GPT-5.5 Pro 79.2.

Multimodal

ChartQA

Question answering over real-world charts and plots. Top: Gemini 3.1 Pro 92.8 · Claude Fable 5 91.4 · GPT-5.5 Pro 90.7.

Multimodal

MMBench

Bilingual multimodal capability evaluation. Top: Gemini 3.1 Pro 91.4 · Claude Fable 5 89.7 · GPT-5.5 Pro 88.5.

Multimodal

SEED-Bench

Generative multimodal comprehension across 12 dimensions. Top: Gemini 3.1 Pro 78.9 · Claude Fable 5 77.4 · GPT-5.5 Pro 76.2.

Multimodal

RealWorldQA

Real-world spatial understanding questions. Top: Gemini 3.1 Pro 78.2 · Claude Fable 5 76.5 · GPT-5.5 Pro 75.1.

Multimodal

AI2D

Science-diagram question answering. Top: Gemini 3.1 Pro 95.1 · Claude Fable 5 94.3 · GPT-5.5 Pro 93.6.

Multimodal

DocVQA

Document question answering on scanned pages. Top: Gemini 3.1 Pro 96.4 · Claude Fable 5 95.7 · GPT-5.5 Pro 94.9.

Multimodal

InfographicVQA

Question answering over dense infographics. Top: Gemini 3.1 Pro 87.2 · Claude Fable 5 85.6 · GPT-5.5 Pro 84.1.

Multimodal

OCRBench

OCR-centric multimodal capability eval. Top: Gemini 3.1 Pro 92.8 · Claude Fable 5 91.4 · GPT-5.5 Pro 89.7.

Multimodal

BLINK

Core visual perception puzzles trivial for humans. Top: Gemini 3.1 Pro 72.4 · Claude Fable 5 70.8 · GPT-5.5 Pro 69.5.

Multimodal

MM-Vet

Integrated multimodal capability evaluation. Top: Gemini 3.1 Pro 84.6 · Claude Fable 5 83.2 · GPT-5.5 Pro 81.8.

Multimodal

ScreenSpot

GUI grounding accuracy across mobile, desktop, web. Top: Gemini 3.1 Pro 88.4 · Claude Fable 5 86.7 · GPT-5.5 Pro 85.2.