GPQA Diamond
Google-proof graduate questions in biology, physics, chemistry. Top: Gemini 3.1 Pro 94.1 · Claude Fable 5 92.6 · GPT-5.5 Pro 89.4.
LLM benchmarks
Tokenscost tracks 100 public LLM evaluation benchmarks across 8 categories — reasoning, math, coding, multimodal, multilingual, instruction following, and agent multi-turn — paired with live $/1M token pricing for every model.
A crawler-readable index. Each row links to the live card on /benchmarks with full leaderboards and best score-per-dollar.
| Benchmark | Category | Top model | Top score |
|---|---|---|---|
| GPQA Diamond | Reasoning and Logic | Gemini 3.1 Pro | 94.1 |
| Humanity's Last Exam | Reasoning and Logic | Claude Fable 5 | 53.3 |
| ARC-AGI-2 | Reasoning and Logic | Gemini 3.1 Pro | 31.1 |
| LCR (Long-Context Retrieval) | Reasoning and Logic | GPT-5.5 Pro | 75.7 |
| AIME 2025 | Mathematical Problem Solving | Gemini 3.1 Pro | 98.4 |
| MATH-500 | Mathematical Problem Solving | Claude Fable 5 | 98.2 |
| AIME 2024 | Mathematical Problem Solving | Gemini 3.1 Pro | 98.1 |
| SWE-Bench Verified | Computer Science and Programming | Claude Fable 5 | 79.4 |
| LiveCodeBench | Computer Science and Programming | Gemini 3 Pro | 91.7 |
| SciCode | Computer Science and Programming | Claude Fable 5 | 60.2 |
| HumanEval+ | Computer Science and Programming | Claude Fable 5 | 96.3 |
| MMLU-Pro | General Knowledge | Gemini 3.1 Pro | 89.8 |
| SimpleQA | General Knowledge | GPT-5.5 Pro | 62.5 |
| τ²-Bench | Multi-turn | Claude Fable 5 | 78.5 |
| TerminalBench | Multi-turn | Claude Fable 5 | 62.9 |
| IFBench | Instruction Following | Claude Fable 5 | 86.2 |
| MMLU-ProX | Multilingual | Gemini 3.1 Pro | 84.6 |
| MMMU | Multimodal | Gemini 3.1 Pro | 83.6 |
| MathVista | Multimodal | Gemini 3.1 Pro | 82.1 |
| ChartQA | Multimodal | Gemini 3.1 Pro | 92.8 |
| DROP | Reasoning and Logic | GPT-5.5 Pro | 93.2 |
| BIG-Bench Hard | Reasoning and Logic | Claude Fable 5 | 94.8 |
| MuSR | Reasoning and Logic | Claude Fable 5 | 82.5 |
| GSM8K | Mathematical Problem Solving | Claude Fable 5 | 98.6 |
| FrontierMath | Mathematical Problem Solving | GPT-5.5 Pro | 32.4 |
| MBPP+ | Computer Science and Programming | Claude Fable 5 | 92.1 |
| Aider Polyglot | Computer Science and Programming | Claude Fable 5 | 84.7 |
| TruthfulQA | General Knowledge | Claude Fable 5 | 78.4 |
| MGSM | Multilingual | Gemini 3.1 Pro | 94.6 |
| MT-Bench | Multi-turn | Claude Fable 5 | 96.4 |
| ARC-AGI-1 | Reasoning and Logic | GPT-5.5 Pro | 87.5 |
| ZebraLogic | Reasoning and Logic | Claude Fable 5 | 81.2 |
| FOLIO | Reasoning and Logic | Claude Fable 5 | 89.3 |
| LogiQA 2.0 | Reasoning and Logic | GPT-5.5 Pro | 84.6 |
| ReClor | Reasoning and Logic | Claude Fable 5 | 88.4 |
| ProofWriter | Reasoning and Logic | Claude Fable 5 | 96.2 |
| AGIEval | Reasoning and Logic | GPT-5.5 Pro | 87.8 |
| NoCha | Reasoning and Logic | Gemini 3.1 Pro | 72.3 |
| OlympicArena | Reasoning and Logic | GPT-5.5 Pro | 62.4 |
| NPHardEval | Reasoning and Logic | GPT-5.5 Pro | 68.2 |
| MATH | Mathematical Problem Solving | Claude Fable 5 | 98.1 |
| OlympiadBench | Mathematical Problem Solving | GPT-5.5 Pro | 74.6 |
| MathBench | Mathematical Problem Solving | Claude Fable 5 | 89.4 |
| ASDiv | Mathematical Problem Solving | Claude Fable 5 | 98.6 |
| SVAMP | Mathematical Problem Solving | Claude Fable 5 | 97.5 |
| MultiArith | Mathematical Problem Solving | Claude Fable 5 | 99.5 |
| TheoremQA | Mathematical Problem Solving | GPT-5.5 Pro | 70.2 |
| PutnamBench | Mathematical Problem Solving | GPT-5.5 Pro | 18.4 |
| MATH Level 5 | Mathematical Problem Solving | Claude Fable 5 | 91.3 |
| GSM-Hard | Mathematical Problem Solving | Claude Fable 5 | 88.7 |
| HumanEval | Computer Science and Programming | Claude Fable 5 | 99.4 |
| MBPP | Computer Science and Programming | Claude Fable 5 | 95.8 |
| CodeContests | Computer Science and Programming | GPT-5.5 Pro | 58.7 |
| APPS | Computer Science and Programming | Claude Fable 5 | 62.8 |
| CRUXEval | Computer Science and Programming | Claude Fable 5 | 90.2 |
| BigCodeBench | Computer Science and Programming | Claude Fable 5 | 61.4 |
| RepoBench | Computer Science and Programming | Claude Fable 5 | 72.6 |
| SWE-Bench Lite | Computer Science and Programming | Claude Fable 5 | 78.4 |
| ClassEval | Computer Science and Programming | Claude Fable 5 | 84.2 |
| DS-1000 | Computer Science and Programming | Claude Fable 5 | 87.5 |
| MMBench | Multimodal | Gemini 3.1 Pro | 91.4 |
| SEED-Bench | Multimodal | Gemini 3.1 Pro | 78.9 |
| RealWorldQA | Multimodal | Gemini 3.1 Pro | 78.2 |
| AI2D | Multimodal | Gemini 3.1 Pro | 95.1 |
| DocVQA | Multimodal | Gemini 3.1 Pro | 96.4 |
| InfographicVQA | Multimodal | Gemini 3.1 Pro | 87.2 |
| OCRBench | Multimodal | Gemini 3.1 Pro | 92.8 |
| BLINK | Multimodal | Gemini 3.1 Pro | 72.4 |
| MM-Vet | Multimodal | Gemini 3.1 Pro | 84.6 |
| ScreenSpot | Multimodal | Gemini 3.1 Pro | 88.4 |
| MMLU | General Knowledge | Claude Fable 5 | 92.6 |
| TriviaQA | General Knowledge | Claude Fable 5 | 94.7 |
| Natural Questions | General Knowledge | Claude Fable 5 | 78.3 |
| OpenBookQA | General Knowledge | Claude Fable 5 | 96.4 |
| ARC-Challenge | General Knowledge | Claude Fable 5 | 97.2 |
| HellaSwag | General Knowledge | Claude Fable 5 | 96.1 |
| WinoGrande | General Knowledge | Claude Fable 5 | 94.5 |
| CommonsenseQA | General Knowledge | Claude Fable 5 | 90.7 |
| PIQA | General Knowledge | Claude Fable 5 | 92.5 |
| Social IQa | General Knowledge | Claude Fable 5 | 87.4 |
| IFEval | Instruction Following | Claude Fable 5 | 92.4 |
| AlpacaEval 2.0 | Instruction Following | Claude Fable 5 | 78.6 |
| MT-Bench-101 | Instruction Following | Claude Fable 5 | 89.7 |
| FLASK | Instruction Following | Claude Fable 5 | 87.2 |
| InfoBench | Instruction Following | Claude Fable 5 | 90.5 |
| FollowBench | Instruction Following | Claude Fable 5 | 82.4 |
| Arena-Hard | Instruction Following | Claude Fable 5 | 94.2 |
| XCOPA | Multilingual | Gemini 3.1 Pro | 94.5 |
| XNLI | Multilingual | Gemini 3.1 Pro | 89.4 |
| FLORES-200 | Multilingual | Gemini 3.1 Pro | 62.8 |
| M3Exam | Multilingual | Gemini 3.1 Pro | 76.8 |
| Global-MMLU | Multilingual | Gemini 3.1 Pro | 85.7 |
| Belebele | Multilingual | Gemini 3.1 Pro | 91.4 |
| MINT | Multi-turn | Claude Fable 5 | 72.5 |
| BFCL | Multi-turn | Claude Fable 5 | 87.6 |
| AgentBench | Multi-turn | Claude Fable 5 | 65.2 |
| GAIA | Multi-turn | Claude Fable 5 | 74.6 |
| WebArena | Multi-turn | Claude Fable 5 | 58.4 |
| OSWorld | Multi-turn | Claude Fable 5 | 52.7 |
| SWE-Lancer | Multi-turn | Claude Fable 5 | 46.8 |
Google-proof graduate questions in biology, physics, chemistry. Top: Gemini 3.1 Pro 94.1 · Claude Fable 5 92.6 · GPT-5.5 Pro 89.4.
Frontier multidisciplinary stress test from Center for AI Safety. Top: Claude Fable 5 53.3 · Gemini 3.1 Pro 44.7 · GPT-5.5 Pro 41.6.
Abstract visual reasoning puzzles trivial for humans, hard for LLMs. Top: Gemini 3.1 Pro 31.1 · Claude Fable 5 17.6 · GPT-5.5 Pro 15.9.
Find and use info buried inside very long documents. Top: GPT-5.5 Pro 75.7 · GPT-5.5 75.6 · GPT-5.4 75.0.
Discrete reasoning over paragraphs — numbers, dates, sorting. Top: GPT-5.5 Pro 93.2 · Claude Fable 5 92.6 · Gemini 3.1 Pro 91.8.
23 BIG-Bench tasks where prior models trailed humans. Top: Claude Fable 5 94.8 · GPT-5.5 Pro 94.1 · Gemini 3.1 Pro 93.5.
Multistep soft reasoning over natural-language narratives. Top: Claude Fable 5 82.5 · GPT-5.5 Pro 80.1 · Gemini 3.1 Pro 78.4.
Abstraction and Reasoning Corpus — visual pattern puzzles. Top: GPT-5.5 Pro 87.5 · Claude Fable 5 85.2 · Gemini 3.1 Pro 83.0.
Logic-grid puzzles testing constraint-satisfaction reasoning. Top: Claude Fable 5 81.2 · GPT-5.5 Pro 78.4 · Gemini 3.1 Pro 74.6.
First-order logic natural-language entailment. Top: Claude Fable 5 89.3 · GPT-5.5 Pro 87.6 · Gemini 3.1 Pro 86.1.
Civil-service style logical reading comprehension. Top: GPT-5.5 Pro 84.6 · Claude Fable 5 83.1 · Gemini 3.1 Pro 80.7.
LSAT-style logical reasoning multiple choice. Top: Claude Fable 5 88.4 · GPT-5.5 Pro 87.2 · Gemini 3.1 Pro 85.9.
Synthetic multi-step natural-language proofs. Top: Claude Fable 5 96.2 · GPT-5.5 Pro 95.4 · Gemini 3.1 Pro 94.1.
Human-centric exams (SAT, LSAT, GRE, Gaokao). Top: GPT-5.5 Pro 87.8 · Claude Fable 5 86.2 · Gemini 3.1 Pro 84.5.
Novel-length claim verification over 100k-token books. Top: Gemini 3.1 Pro 72.3 · Claude Fable 5 69.8 · GPT-5.5 Pro 65.4.
Multi-discipline Olympiad-level reasoning across 7 subjects. Top: GPT-5.5 Pro 62.4 · Claude Fable 5 60.7 · Gemini 3.1 Pro 58.9.
NP-complete/hard algorithmic problem solving. Top: GPT-5.5 Pro 68.2 · Claude Fable 5 66.5 · o3 62.1.
Multi-turn agent tasks in airline and retail domains. Top: Claude Fable 5 78.5 · Claude Opus 4.8 74.0 · GPT-5.5 Pro 72.1.
Long-horizon shell tasks — agents controlling a real terminal. Top: Claude Fable 5 62.9 · GPT-5.5 Pro 57.6 · Claude Opus 4.8 54.5.
GPT-4-judged two-turn conversational quality benchmark. Top: Claude Fable 5 96.4 · GPT-5.5 Pro 95.8 · Gemini 3.1 Pro 95.1.
Multi-turn tool use and natural language feedback. Top: Claude Fable 5 72.5 · GPT-5.5 Pro 71.2 · Gemini 3.1 Pro 69.4.
Berkeley Function Calling Leaderboard. Top: Claude Fable 5 87.6 · GPT-5.5 Pro 86.4 · Gemini 3.1 Pro 84.8.
Multi-environment evaluation of LLM agents. Top: Claude Fable 5 65.2 · GPT-5.5 Pro 63.8 · Gemini 3.1 Pro 61.7.
General AI Assistants benchmark — real-world tasks. Top: Claude Fable 5 74.6 · GPT-5.5 Pro 72.8 · Gemini 3.1 Pro 70.4.
Realistic web-task execution in a sandbox. Top: Claude Fable 5 58.4 · GPT-5.5 Pro 56.7 · Gemini 3.1 Pro 54.2.
Real desktop OS computer-use tasks. Top: Claude Fable 5 52.7 · GPT-5.5 Pro 50.4 · Gemini 3.1 Pro 48.6.
Real Upwork freelance software engineering tasks. Top: Claude Fable 5 46.8 · GPT-5.5 Pro 44.5 · Gemini 3.1 Pro 41.7.
American Invitational Math Exam — competition-grade problems. Top: Gemini 3.1 Pro 98.4 · GPT-5.5 Pro 97.0 · Claude Fable 5 96.7.
500 competition math problems across algebra, geometry, number theory. Top: Claude Fable 5 98.2 · Gemini 3.1 Pro 97.9 · GPT-5.5 Pro 97.4.
Prior-year AIME, widely cited in model release notes. Top: Gemini 3.1 Pro 98.1 · Claude Fable 5 97.5 · GPT-5.5 Pro 97.1.
Grade-school math word problems requiring chain-of-thought. Top: Claude Fable 5 98.6 · Gemini 3.1 Pro 98.3 · GPT-5.5 Pro 98.1.
Research-level mathematics from Epoch AI — extremely hard. Top: GPT-5.5 Pro 32.4 · Claude Fable 5 28.7 · Gemini 3.1 Pro 26.5.
Hendrycks 12,500 competition math problems. Top: Claude Fable 5 98.1 · GPT-5.5 Pro 97.8 · Gemini 3.1 Pro 97.2.
Bilingual Olympiad math and physics problems. Top: GPT-5.5 Pro 74.6 · Claude Fable 5 72.8 · Gemini 3.1 Pro 70.4.
Hierarchical math from primary to college level. Top: Claude Fable 5 89.4 · GPT-5.5 Pro 88.7 · Gemini 3.1 Pro 87.5.
Diverse grade-school arithmetic word problems. Top: Claude Fable 5 98.6 · GPT-5.5 Pro 98.4 · Gemini 3.1 Pro 98.0.
Math word problems with sensitivity perturbations. Top: Claude Fable 5 97.5 · GPT-5.5 Pro 97.1 · Gemini 3.1 Pro 96.4.
Multi-step arithmetic word problems. Top: Claude Fable 5 99.5 · GPT-5.5 Pro 99.4 · Gemini 3.1 Pro 99.2.
Theorem-application questions across STEM. Top: GPT-5.5 Pro 70.2 · Claude Fable 5 68.7 · Gemini 3.1 Pro 66.4.
Formalized Putnam competition problems in Lean/Coq/Isabelle. Top: GPT-5.5 Pro 18.4 · Claude Fable 5 16.9 · Gemini 3.1 Pro 14.2.
Hardest subset of the MATH benchmark. Top: Claude Fable 5 91.3 · GPT-5.5 Pro 90.7 · Gemini 3.1 Pro 88.6.
GSM8K with adversarially large numbers. Top: Claude Fable 5 88.7 · GPT-5.5 Pro 87.5 · Gemini 3.1 Pro 85.9.
Resolve real GitHub issues in real Python repos. Top: Claude Fable 5 79.4 · Claude Opus 4.8 74.5 · GPT-5.5 Pro 72.8.
Continuously updated coding contests — leak-resistant. Top: Gemini 3 Pro 91.7 · Gemini 3.1 Pro 90.8 · Claude Fable 5 89.3.
Scientific computing tasks: numerics, simulations, ML pipelines. Top: Claude Fable 5 60.2 · Gemini 3.1 Pro 58.9 · GPT-5.5 Pro 57.4.
Classic Python function synthesis with hardened tests. Top: Claude Fable 5 96.3 · GPT-5.5 Pro 95.7 · Gemini 3.1 Pro 95.1.
Mostly Basic Python Problems with hardened test cases. Top: Claude Fable 5 92.1 · GPT-5.5 Pro 91.3 · Gemini 3.1 Pro 90.4.
Multi-language code editing benchmark from the Aider project. Top: Claude Fable 5 84.7 · GPT-5.5 Pro 81.2 · Claude Opus 4.8 79.5.
OpenAI 164 Python coding problems (pass@1). Top: Claude Fable 5 99.4 · GPT-5.5 Pro 99.1 · Gemini 3.1 Pro 98.8.
Mostly Basic Python Problems (pass@1). Top: Claude Fable 5 95.8 · GPT-5.5 Pro 95.2 · Gemini 3.1 Pro 94.4.
Competitive programming problems from DeepMind. Top: GPT-5.5 Pro 58.7 · Claude Fable 5 56.9 · Gemini 3.1 Pro 54.3.
Coding interview problems across three difficulty tiers. Top: Claude Fable 5 62.8 · GPT-5.5 Pro 61.4 · Gemini 3.1 Pro 59.7.
Code understanding via input and output prediction. Top: Claude Fable 5 90.2 · GPT-5.5 Pro 89.4 · Gemini 3.1 Pro 87.6.
Practical coding tasks across 139 Python libraries. Top: Claude Fable 5 61.4 · GPT-5.5 Pro 60.2 · Gemini 3.1 Pro 57.8.
Repository-level code completion benchmark. Top: Claude Fable 5 72.6 · GPT-5.5 Pro 71.4 · Gemini 3.1 Pro 69.7.
300 curated real GitHub issue fixes. Top: Claude Fable 5 78.4 · GPT-5.5 Pro 76.9 · Gemini 3.1 Pro 74.2.
Class-level Python code generation. Top: Claude Fable 5 84.2 · GPT-5.5 Pro 83.4 · Gemini 3.1 Pro 81.6.
Data-science coding problems across 7 libraries. Top: Claude Fable 5 87.5 · GPT-5.5 Pro 86.4 · Gemini 3.1 Pro 84.6.
Harder, deduplicated successor to MMLU across 57 subjects. Top: Gemini 3.1 Pro 89.8 · Claude Fable 5 89.5 · GPT-5.5 Pro 88.2.
Short factual questions — hallucination stress test. Top: GPT-5.5 Pro 62.5 · Gemini 3.1 Pro 58.1 · Claude Fable 5 53.9.
Questions designed to elicit common human misconceptions. Top: Claude Fable 5 78.4 · GPT-5.5 Pro 76.2 · Gemini 3.1 Pro 74.8.
Massive multitask language understanding across 57 subjects. Top: Claude Fable 5 92.6 · GPT-5.5 Pro 92.1 · Gemini 3.1 Pro 91.5.
Open-domain trivia question answering. Top: Claude Fable 5 94.7 · GPT-5.5 Pro 94.2 · Gemini 3.1 Pro 93.5.
Real Google search queries answered from Wikipedia. Top: Claude Fable 5 78.3 · GPT-5.5 Pro 77.5 · Gemini 3.1 Pro 76.4.
Elementary-science QA requiring multi-hop reasoning. Top: Claude Fable 5 96.4 · GPT-5.5 Pro 95.8 · Gemini 3.1 Pro 95.1.
AI2 Reasoning Challenge — grade-school science. Top: Claude Fable 5 97.2 · GPT-5.5 Pro 96.8 · Gemini 3.1 Pro 96.3.
Commonsense sentence completion. Top: Claude Fable 5 96.1 · GPT-5.5 Pro 95.7 · Gemini 3.1 Pro 95.2.
Adversarial Winograd-style commonsense. Top: Claude Fable 5 94.5 · GPT-5.5 Pro 94.0 · Gemini 3.1 Pro 93.4.
Commonsense QA grounded in ConceptNet. Top: Claude Fable 5 90.7 · GPT-5.5 Pro 90.1 · Gemini 3.1 Pro 89.4.
Physical commonsense reasoning. Top: Claude Fable 5 92.5 · GPT-5.5 Pro 92.0 · Gemini 3.1 Pro 91.4.
Social commonsense reasoning about everyday interactions. Top: Claude Fable 5 87.4 · GPT-5.5 Pro 86.8 · Gemini 3.1 Pro 86.1.
Adherence to nuanced writing constraints and formats. Top: Claude Fable 5 86.2 · Gemini 3.1 Pro 84.7 · GPT-5.5 Pro 83.4.
Verifiable instruction-following constraints. Top: Claude Fable 5 92.4 · GPT-5.5 Pro 91.6 · Gemini 3.1 Pro 90.5.
Length-controlled win-rate vs GPT-4-Turbo. Top: Claude Fable 5 78.6 · GPT-5.5 Pro 76.2 · Gemini 3.1 Pro 73.4.
Fine-grained multi-turn dialogue capability. Top: Claude Fable 5 89.7 · GPT-5.5 Pro 88.4 · Gemini 3.1 Pro 87.0.
Fine-grained skill-set evaluation. Top: Claude Fable 5 87.2 · GPT-5.5 Pro 86.0 · Gemini 3.1 Pro 84.7.
Decomposed-requirement instruction following. Top: Claude Fable 5 90.5 · GPT-5.5 Pro 89.7 · Gemini 3.1 Pro 88.4.
Multi-level fine-grained constraint following. Top: Claude Fable 5 82.4 · GPT-5.5 Pro 81.2 · Gemini 3.1 Pro 79.8.
500 challenging real-user prompts from Chatbot Arena. Top: Claude Fable 5 94.2 · GPT-5.5 Pro 92.7 · Gemini 3.1 Pro 90.4.
MMLU-Pro translated into 14 languages. Top: Gemini 3.1 Pro 84.6 · Claude Fable 5 83.2 · GPT-5.5 Pro 82.0.
GSM8K translated into 10 typologically diverse languages. Top: Gemini 3.1 Pro 94.6 · Claude Fable 5 93.8 · GPT-5.5 Pro 92.5.
Causal commonsense reasoning across 11 languages. Top: Gemini 3.1 Pro 94.5 · Claude Fable 5 93.7 · GPT-5.5 Pro 92.8.
Natural language inference in 15 languages. Top: Gemini 3.1 Pro 89.4 · Claude Fable 5 88.5 · GPT-5.5 Pro 87.7.
Translation quality across 200 languages. Top: Gemini 3.1 Pro 62.8 · Claude Fable 5 61.4 · GPT-5.5 Pro 60.2.
Real human exam questions across 9 languages. Top: Gemini 3.1 Pro 76.8 · Claude Fable 5 75.4 · GPT-5.5 Pro 74.2.
Culturally-translated MMLU across 42 languages. Top: Gemini 3.1 Pro 85.7 · Claude Fable 5 84.6 · GPT-5.5 Pro 83.5.
Multilingual reading comprehension in 122 languages. Top: Gemini 3.1 Pro 91.4 · Claude Fable 5 90.5 · GPT-5.5 Pro 89.6.
College-level multimodal questions across 30 subjects. Top: Gemini 3.1 Pro 83.6 · Claude Fable 5 81.4 · GPT-5.5 Pro 80.7.
Visual mathematical reasoning across charts, diagrams, geometry. Top: Gemini 3.1 Pro 82.1 · Claude Fable 5 80.4 · GPT-5.5 Pro 79.2.
Question answering over real-world charts and plots. Top: Gemini 3.1 Pro 92.8 · Claude Fable 5 91.4 · GPT-5.5 Pro 90.7.
Bilingual multimodal capability evaluation. Top: Gemini 3.1 Pro 91.4 · Claude Fable 5 89.7 · GPT-5.5 Pro 88.5.
Generative multimodal comprehension across 12 dimensions. Top: Gemini 3.1 Pro 78.9 · Claude Fable 5 77.4 · GPT-5.5 Pro 76.2.
Real-world spatial understanding questions. Top: Gemini 3.1 Pro 78.2 · Claude Fable 5 76.5 · GPT-5.5 Pro 75.1.
Science-diagram question answering. Top: Gemini 3.1 Pro 95.1 · Claude Fable 5 94.3 · GPT-5.5 Pro 93.6.
Document question answering on scanned pages. Top: Gemini 3.1 Pro 96.4 · Claude Fable 5 95.7 · GPT-5.5 Pro 94.9.
Question answering over dense infographics. Top: Gemini 3.1 Pro 87.2 · Claude Fable 5 85.6 · GPT-5.5 Pro 84.1.
OCR-centric multimodal capability eval. Top: Gemini 3.1 Pro 92.8 · Claude Fable 5 91.4 · GPT-5.5 Pro 89.7.
Core visual perception puzzles trivial for humans. Top: Gemini 3.1 Pro 72.4 · Claude Fable 5 70.8 · GPT-5.5 Pro 69.5.
Integrated multimodal capability evaluation. Top: Gemini 3.1 Pro 84.6 · Claude Fable 5 83.2 · GPT-5.5 Pro 81.8.
GUI grounding accuracy across mobile, desktop, web. Top: Gemini 3.1 Pro 88.4 · Claude Fable 5 86.7 · GPT-5.5 Pro 85.2.