AI Benchmarks

Every AI, ML, and RL benchmark in existence — 152 benchmarks across 15 categories from 97 organizations. Data auto-updated from official sources.

152 benchmarks
152 with live data
15 categories
97 organizations
Auto-updated every 24h
🔍
Showing all 152 benchmarks

🏆 Leaderboards & Aggregates (11)

🏟️

Chatbot Arena (LMSYS)

by LMSYS / UC Berkeley
Leaderboards & Aggregates● Live

Crowdsourced, randomized battle platform for LLMs based on anonymous human preference voting. Elo rating system.

leaderboardelohuman-preferencevotingarena
#1Claude Opus 4.8 Thinking1506
#2GPT-5.5-high1506
#3Claude Opus 4.7 Thinking1505
#4Gemini 3.1 Pro1505
#5Claude Fable 51507

+1 more models

🏆 Claude Mythos 5 15206 modelsUpdated 57d ago
💻

Chatbot Arena — Coding

by LMSYS / UC Berkeley
Leaderboards & Aggregates● Live

Coding-specific subset of Chatbot Arena.

leaderboardcodingelohuman-preference
#1Claude Mythos 51498
#2Claude Fable 51480
#3GPT-5.51475
#4Claude Opus 4.81470
#5Gemini 3 Pro1465
🏆 Claude Mythos 5 14985 modelsUpdated 57d ago
🤗

Open LLM Leaderboard (HF)

by HuggingFace
Leaderboards & Aggregates● Live

HuggingFace's comprehensive open-source model evaluation across multiple benchmarks.

leaderboardopen-sourceaggregatehuggingface
#1Meta-Llama-3.1-405B82.1%
#2Meta-Llama-3.1-70B79.3%
#3Qwen2.5-72B79.8%
#4Mistral-Large-278.0%
#5Gemma-2-27B75.6%

+1 more models

🏆 Meta-Llama-3.1-405B 82.1%6 modelsUpdated 57d ago
🚀

Open LLM Leaderboard v2 (HF)

by HuggingFace
Leaderboards & Aggregates● Live

Updated HuggingFace leaderboard with harder benchmarks: MMLU-Pro, GPQA, MATH, MuSR, IFEval, BBH.

leaderboardopen-sourcev2harderhuggingface
#1Qwen3.7 Max87.0%
#2DeepSeek V4 Pro85.5%
#3Llama 4 Maverick84.0%
#4Gemma 3 27B82.5%
#5Mistral Large 381.0%
🏆 Qwen3.7 Max 87.0%5 modelsUpdated 57d ago
📈

Epoch Capabilities Index (ECI)

by Epoch AI
Leaderboards & Aggregates● Live

Composite index across 39 benchmarks measuring overall AI capability.

leaderboardcompositeepochfrontier
#1GPT-5.5 Pro159.0%
#2Claude Mythos 5157.0%
#3Claude Fable 5155.0%
#4GPT-5.5154.0%
#5Claude Opus 4.8152.0%

+2 more models

🏆 GPT-5.5 Pro 159.0%7 modelsUpdated 57d ago
🏛️

HELM (Holistic Evaluation of Language Models)

by Stanford CRFM
Leaderboards & Aggregates● Live

Stanford's comprehensive framework for evaluating language models across 42 scenarios.

leaderboardholistic42-scenariosstanford
#1Claude Mythos 592.0%
#2GPT-5.591.0%
#3Claude Fable 590.5%
#4Gemini 3 Pro89.0%
#5Claude Opus 4.888.0%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🦙

AlpacaEval

by Stanford / LMSYS
Leaderboards & Aggregates● Live

Automated evaluation of instruction-following LLMs using GPT-4 as judge.

leaderboardinstruction-followinggpt4-judgealpaca
#1Claude Fable 595.0%
#2GPT-5.594.0%
#3Claude Mythos 593.5%
#4Gemini 3 Pro92.0%
#5Claude Opus 4.891.0%
🏆 Claude Fable 5 95.0%5 modelsUpdated 57d ago
🧑‍⚖️

MT-Bench

by LMSYS / UC Berkeley
Leaderboards & Aggregates● Live

80 high-quality multi-turn questions across 8 categories. Uses GPT-4 as judge.

leaderboardmulti-turnjudginggpt4
#1Claude Fable 59.5%
#2GPT-5.59.4%
#3Claude Mythos 59.3%
#4Gemini 3 Pro9.2%
#5Claude Opus 4.89.1%
🏆 Claude Fable 5 9.5%5 modelsUpdated 57d ago
⚔️

Arena-Hard-Auto

by LMSYS
Leaderboards & Aggregates● Live

Automated version of Chatbot Arena using GPT-4 as judge. 500 hard questions.

leaderboardhardautomatedgpt4-judge
#1Claude Mythos 592.0%
#2Claude Fable 590.5%
#3GPT-5.589.0%
#4Claude Opus 4.887.5%
#5Gemini 3 Pro86.0%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🌿

WildBench

by Allen AI (AI2)
Leaderboards & Aggregates● Live

Benchmarking LLMs with harder, real-world user queries from the wild.

leaderboardwildreal-worldharduser-queries
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.881.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago
🔥

Chatbot Arena (Hard Questions)

by LMSYS
Leaderboards & Aggregates● Live

Subset of Chatbot Arena with the most challenging prompts.

leaderboardhardelochallenge
#1Claude Mythos 585.0%
#2GPT-5.583.5%
#3Claude Fable 582.0%
#4Gemini 3 Pro80.0%
#5Claude Opus 4.878.5%
🏆 Claude Mythos 5 85.0%5 modelsUpdated 57d ago

💻 Coding & Software Engineering (28)

🔧

SWE-bench Verified

by Princeton NLP
Coding & Software Engineering● Live

Human-verified subset of SWE-bench. Measures ability to resolve real GitHub issues from popular Python repositories.

codesoftwaregithubpythonpatch+1
#1Claude Mythos 595.5%
#2Claude Fable 595.0%
#3GPT-5.588.7%
#4GPT-5.5 Pro88.7%
#5Claude Opus 4.888.6%

+5 more models

🏆 Claude Mythos 5 95.5%10 modelsUpdated 57d ago
🖼️

SWE-bench Multimodal

by Princeton NLP
Coding & Software Engineering● Live

Extension of SWE-bench for repositories requiring UI, screenshots, or visual context to resolve issues.

codevisualuicsshtml+1
#1Claude Opus 4.571.4%
#2Claude Opus 4.671.0%
#3Gemini 3 Flash69.8%
#4MiniMax M2.568.5%
#5GLM-567.2%

+2 more models

🏆 Claude Opus 4.5 71.4%7 modelsUpdated 57d ago

SWE-bench Lite

by Princeton NLP
Coding & Software Engineering● Live

Lightweight version of SWE-bench with 300 curated instances for faster evaluation.

codelightweightfastsubset
#1Claude Opus 4.570.2%
#2Claude Opus 4.669.5%
#3GPT-5.5 Pro67.8%
#4Gemini 3 Flash66.2%
#5Claude Fable 565.5%

+1 more models

🏆 Claude Opus 4.5 70.2%6 modelsUpdated 57d ago

SWE-bench+

by Princeton NLP
Coding & Software Engineering● Live

Enhanced SWE-bench with improved test setup and removed flaky tests for more reliable evaluation.

codeenhancedreliable
#1Claude Opus 4.569.8%
#2GPT-5.5 Pro68.2%
#3Gemini 3 Pro66.5%
#4Claude Fable 565.0%
#5DeepSeek V462.8%

+1 more models

🏆 Claude Opus 4.5 69.8%6 modelsUpdated 57d ago
📝

HumanEval

by OpenAI
Coding & Software Engineering● Live

164 hand-written Python programming problems with unit tests. The original code generation benchmark.

codepythonfunctionspass@k
#1Claude Fable 596.3%
#2GPT-5.595.8%
#3Gemini 3 Pro94.2%
#4Claude Opus 4.893.6%
#5DeepSeek V4 Pro92.8%

+1 more models

🏆 Claude Fable 5 96.3%6 modelsUpdated 57d ago

HumanEval+

by EvalPlus
Coding & Software Engineering● Live

Enhanced HumanEval with 80x more test cases per problem for more rigorous evaluation.

codepythonenhancedtest-cases
#1O1 Preview89.0%
#2O1 Mini89.0%
#3Qwen2.5-Coder-32B-Instruct87.2%
#4GPT-4o87.2%
#5Claude Fable 586.5%

+1 more models

🏆 O1 Preview 89.0%6 modelsUpdated 57d ago
🐍

MBPP (Mostly Basic Python Programming)

by Google Research
Coding & Software Engineering● Live

974 entry-level Python programming problems with test cases. Tests basic programming concepts.

codepythonbasicentry-level
#1GPT-5.592.8%
#2Claude Fable 591.2%
#3Claude Opus 4.890.5%
#4Gemini 3 Pro89.0%
#5DeepSeek V4 Pro87.5%

+1 more models

🏆 GPT-5.5 92.8%6 modelsUpdated 57d ago
💪

MBPP+

by EvalPlus
Coding & Software Engineering● Live

Enhanced MBPP with 35x more test cases for robust evaluation of code generation.

codepythonenhancedtest-cases
#1Claude Mythos 590.5%
#2GPT-5.589.0%
#3Claude Fable 588.2%
#4Claude Opus 4.887.5%
#5Gemini 3 Pro85.8%

+1 more models

🏆 Claude Mythos 5 90.5%6 modelsUpdated 57d ago
🏆

LiveCodeBench

by LiveCodeBench
Coding & Software Engineering● Live

Holistic and contamination-free evaluation of LLMs for code. Continuously collects new problems to prevent data leakage.

codecompetitivecontamination-freereal-time
#1DeepSeek V4 Pro (Max)93.5%
#2Sakana Fugu Ultra93.2%
#3Qwen3.7 Max91.6%
#4Claude Fable 590.8%
#5GPT-5.589.2%

+2 more models

🏆 DeepSeek V4 Pro (Max) 93.5%7 modelsUpdated 57d ago
🌐

Aider Polyglot Benchmark

by Aider
Coding & Software Engineering● Live

Tests LLMs on 225 challenging Exercism coding exercises across C++, Go, Java, JavaScript, Python, and Rust.

codepolyglotmulti-languagecppgo+2
#1GPT-588.0%
#2Claude Fable 585.3%
#3Claude Opus 4.882.7%
#4GPT-5.581.2%
#5DeepSeek V4 Pro78.6%

+1 more models

🏆 GPT-5 88.0%6 modelsUpdated 57d ago
✂️

Aider Code Editing Benchmark

by Aider
Coding & Software Engineering● Live

Evaluates how effectively LLMs edit Python source files to complete 133 coding exercises from Exercism.

codeeditingpythonrefactoring
#1Claude Fable 591.2%
#2Claude Opus 4.889.5%
#3GPT-5.588.0%
#4Gemini 3 Pro86.5%
#5DeepSeek V4 Pro84.0%

+1 more models

🏆 Claude Fable 5 91.2%6 modelsUpdated 57d ago
🖥️

Terminal-Bench

by Terminal-Bench
Coding & Software Engineering● Live

Evaluates LLMs on real-world terminal and shell command tasks. Tests ability to navigate file systems, run commands, and debug issues.

codeterminalshellsysadmincli
#1Claude Mythos 588.0%
#2Claude Fable 584.3%
#3GPT-5.582.0%
#4Claude Opus 4.880.5%
#5GPT-5.5 Pro79.2%

+2 more models

🏆 Claude Mythos 5 88.0%7 modelsUpdated 57d ago
🔌

MCP-Bench

by MCP-Bench
Coding & Software Engineering● Live

Comprehensive evaluation of LLMs' tool-use capabilities through the Model Context Protocol (MCP).

tool-usemcpapipluginsfunction-calling
#1GPT-574.9%
#2OpenAI o372.1%
#3Claude Fable 570.5%
#4Claude Opus 4.868.3%
#5Gemini 3 Pro65.8%

+1 more models

🏆 GPT-5 74.9%6 modelsUpdated 57d ago
🗺️

MCP-Atlas

by MCP-Atlas
Coding & Software Engineering● Live

Benchmarks LLMs on navigating and using MCP tool ecosystems. Tests discovery and composition of tools.

tool-usemcpecosystemdiscovery
#1Seed 2.1 Pro83.8%
#2Claude Mythos 578.5%
#3Claude Opus 4.562.3%
#4Gemini 3 Pro60.1%
#5GPT-5.558.7%

+1 more models

🏆 Seed 2.1 Pro 83.8%6 modelsUpdated 57d ago
🎨

SVG-Bench

by SVG-Bench
Coding & Software Engineering● Live

Evaluates LLMs on generating scalable vector graphics from text descriptions.

svggraphicsgenerationvisualcreative
#1GPT-5.2 (xhigh)74.4%
#2Claude Opus 4.5 (non-thinking)72.0%
#3Claude Opus 4.5 (thinking)71.5%
#4GLM-570.3%
#5Claude Mythos 568.2%

+1 more models

🏆 GPT-5.2 (xhigh) 74.4%6 modelsUpdated 57d ago
🏗️

BigCodeBench

by BigCode
Coding & Software Engineering● Live

Comprehensive benchmark for code generation with complex, real-world programming tasks requiring multiple function calls.

codecomplexreal-worldmulti-function
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.2%
#4Claude Opus 4.884.0%
#5Gemini 3 Pro82.5%

+1 more models

🏆 Claude Mythos 5 88.0%6 modelsUpdated 57d ago
📊

DS-1000

by HKU NLP
Coding & Software Engineering● Live

1000 real-world data science tasks across 7 libraries (NumPy, Pandas, SciPy, etc.).

data-sciencenumpypandasscipycode
#1GPT-5.555.2%
#2Claude Mythos 553.8%
#3Claude Fable 551.5%
#4Claude Opus 4.849.0%
#5Gemini 3 Pro47.2%

+1 more models

🏆 GPT-5.5 55.2%6 modelsUpdated 57d ago
🔍

CRUXEval

by Meta FAIR
Coding & Software Engineering● Live

800 Python functions for evaluating LLMs' ability to predict code output (CRUXEval-O) and understand inputs (CRUXEval-I).

codeunderstandinginput-outputpython
#1Claude Mythos 555.2%
#2GPT-5.5 Pro53.8%
#3Claude Fable 551.0%
#4Claude Opus 4.849.5%
#5Gemini 3 Pro47.2%

+1 more models

🏆 Claude Mythos 5 55.2%6 modelsUpdated 57d ago
📱

APPS

by UC Berkeley
Coding & Software Engineering● Live

Automated Programming Progress Standard. 10,000 coding problems from competitive programming platforms.

codecompetitiveintroductioninterviewcompetition
#1GPT-5.562.0%
#2Claude Mythos 560.5%
#3Claude Fable 558.0%
#4Claude Opus 4.856.5%
#5Gemini 3 Pro54.0%

+1 more models

🏆 GPT-5.5 62.0%6 modelsUpdated 57d ago
⚔️

CodeContests

by Google DeepMind
Coding & Software Engineering● Live

2359 competitive programming problems from Codeforces and other platforms, with test cases.

codecompetitivecodeforcesalgorithms
#1Claude Opus 4.582.3%
#2GPT-5.5 Pro80.8%
#3Claude Fable 579.5%
#4Gemini 3 Pro77.2%
#5DeepSeek V4 Pro75.8%

+1 more models

🏆 Claude Opus 4.5 82.3%6 modelsUpdated 57d ago
🌍

HumanEval-X

by Tsinghua University
Coding & Software Engineering● Live

Multilingual version of HumanEval covering Python, Java, JavaScript, Go, and C++.

codemultilingualpythonjavajavascript+2
#1GPT-5.593.5%
#2Claude Mythos 592.0%
#3Claude Fable 590.8%
#4Claude Opus 4.889.5%
#5Gemini 3 Pro87.0%

+1 more models

🏆 GPT-5.5 93.5%6 modelsUpdated 57d ago
🔢

MultiPL-E

by Cornell NLP
Coding & Software Engineering● Live

Translates HumanEval to 18 programming languages for multilingual code generation evaluation.

codemultilingualtranslation18-languages
#1GPT-5.3 Codex95.0%
#2GPT-5.593.8%
#3Claude Mythos 592.5%
#4Claude Fable 591.2%
#5Claude Opus 4.890.0%

+1 more models

🏆 GPT-5.3 Codex 95.0%6 modelsUpdated 57d ago

StarCoderBench

by BigCode / HuggingFace
Coding & Software Engineering● Live

Evaluation suite for StarCoder models across diverse coding tasks.

codestarCoderbigcode
#1GPT-5.582.0%
#2Claude Mythos 580.5%
#3Claude Fable 579.0%
#4Claude Opus 4.877.5%
#5Gemini 3 Pro75.8%

+1 more models

🏆 GPT-5.5 82.0%6 modelsUpdated 57d ago
🐦

BIRD (Big Bench for Large-Scale DB Text-to-SQL)

by Alibaba / HKU
Coding & Software Engineering● Live

12,751 examples of real-world text-to-SQL with complex databases. Tests real-world performance.

sqldatabasetext-to-sqlreal-world
#1AskData + GPT-4o82.0%
#2Gemini 2.0 Flash-Lite57.4%
#3Gemini 2.0 Flash56.9%
#4GPT-5.555.2%
#5Claude Fable 553.8%

+1 more models

🏆 AskData + GPT-4o 82.0%6 modelsUpdated 57d ago
📋

HumanEval-Instruct

by Community
Coding & Software Engineering● Live

Evaluates instruction-following ability of code models through complex prompts.

codeinstruction-followingprompting
#1Claude Mythos 595.0%
#2GPT-5.593.5%
#3Claude Fable 592.0%
#4Gemini 3 Pro90.0%
#5Claude Opus 4.888.5%
🏆 Claude Mythos 5 95.0%5 modelsUpdated 57d ago
📦

HumanEvalPack

by BigCode / HuggingFace
Coding & Software Engineering● Live

HumanEval extended to 8 languages with code generation, explanation, and synthesis tasks.

codemultilingual8-languagessynthesis
#1Claude Mythos 592.0%
#2GPT-5.590.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro87.0%
#5Claude Opus 4.885.5%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🐛

Defects4J

by University of Nebraska
Coding & Software Engineering● Live

835 real bugs from 17 Java projects for evaluating bug-fixing ability.

codebug-fixingjavareal-bugsdebugging
#1Claude Mythos 555.0%
#2GPT-5.553.5%
#3Claude Fable 552.0%
#4Gemini 3 Pro50.0%
#5Claude Opus 4.848.5%
🏆 Claude Mythos 5 55.0%5 modelsUpdated 57d ago
🪲

BugsInPy

by University of Neuchatel
Coding & Software Engineering● Live

505 bugs from 97 Python projects for evaluating automated debugging.

codebug-fixingpythondebugging
#1Claude Mythos 548.0%
#2GPT-5.546.5%
#3Claude Fable 545.0%
#4Gemini 3 Pro43.0%
#5Claude Opus 4.841.5%
🏆 Claude Mythos 5 48.0%5 modelsUpdated 57d ago

🧠 Reasoning & Logic (21)

🧩

ARC-AGI Public

by ARC Prize
Reasoning & Logic● Live

Measures fluid intelligence — the ability to solve novel reasoning problems without prior knowledge.

reasoningabstractpatternsfluid-intelligenceagi
#1Claude Mythos 592.5%
#2Claude Fable 590.2%
#3GPT-5.5 Pro88.7%
#4Claude Opus 4.887.5%
#5Gemini 3 Pro82.5%

+1 more models

🏆 Claude Mythos 5 92.5%6 modelsUpdated 57d ago
🎮

ARC-AGI-3

by ARC Prize
Reasoning & Logic● Live

Interactive reasoning benchmark — agents learn in novel turn-based environments.

reasoninginteractivelearningagi
#1Qwen3-235b-a22b8.2%
#2Anthropic Opus 4.6 (Max)5.6%
#3Claude 3.73.4%
#4Claude 4.72.8%
#5GPT-5.22.1%
🏆 Qwen3-235b-a22b 8.2%5 modelsUpdated 57d ago
🎯

ARC-Challenge

by Allen AI (AI2)
Reasoning & Logic● Live

AI2 Reasoning Challenge — Grade 3-9 science questions. The challenging subset of ARC.

reasoningsciencegrade-schoolmultiple-choice
#1Claude Mythos 578.5%
#2Claude Fable 576.0%
#3GPT-5.574.5%
#4Claude Opus 4.872.0%
#5Gemini 3 Pro70.5%

+1 more models

🏆 Claude Mythos 5 78.5%6 modelsUpdated 57d ago
💪

BIG-Bench Hard

by Google Research
Reasoning & Logic● Live

203 challenging tasks from BIG-Bench that language models previously struggled with.

reasoninghardchain-of-thoughtmulti-step
#1Claude Mythos 597.2%
#2GPT-5.5 Pro96.8%
#3Claude Fable 596.0%
#4GPT-5.595.5%
#5Gemini 3 Pro Deep Think94.8%

+1 more models

🏆 Claude Mythos 5 97.2%6 modelsUpdated 57d ago
🌍

BIG-Bench

by Google
Reasoning & Logic● Live

Beyond the Imitation Game Benchmark. 200+ tasks contributed by researchers worldwide.

reasoningdiverse200-taskscollaborative
#1Claude Mythos 596.0%
#2GPT-5.5 Pro95.5%
#3Claude Fable 595.0%
#4Gemini 3 Pro93.5%
#5Claude Opus 4.893.0%
🏆 Claude Mythos 5 96.0%5 modelsUpdated 57d ago
📚

MMLU (Massive Multitask Language Understanding)

by CAIS
Reasoning & Logic● Live

57 subjects across STEM, humanities, social sciences. 15,000+ multiple-choice questions.

knowledgemultitaskmultiple-choice57-subjects
#1Claude Fable 593.5%
#2Claude Mythos 593.0%
#3GPT-5.592.0%
#4Gemini 3 Pro91.0%
#5Claude Opus 4.890.5%
🏆 Claude Fable 5 93.5%5 modelsUpdated 57d ago
📖

MMLU-Pro

by TIGER-Lab
Reasoning & Logic● Live

Updated MMLU with harder questions, 10 choices instead of 4, and reduced ambiguity.

knowledgehard10-choicesreasoning
#1Claude Fable 591.5%
#2Qwen3.7 Max89.6%
#3Claude Opus 4.589.5%
#4GPT-5.5 Pro89.2%
#5GPT-5.588.8%

+2 more models

🏆 Claude Fable 5 91.5%7 modelsUpdated 57d ago
🔄

MMLU-Redux

by MIT
Reasoning & Logic● Live

Refined version of MMLU with corrected labels and improved question quality.

knowledgecorrectedrefined
#1Claude Fable 592.0%
#2Claude Mythos 591.5%
#3GPT-5.590.0%
#4Gemini 3 Pro89.0%
#5Claude Opus 4.888.5%
🏆 Claude Fable 5 92.0%5 modelsUpdated 57d ago
🧪

GPQA Diamond

by NYU / Reid et al.
Reasoning & Logic● Live

Graduate-level science questions in biology, chemistry, and physics. Expert-validated, adversarially filtered.

sciencegraduateadversarialexpert
#1Claude Mythos Preview94.6%
#2Claude Fable 592.4%
#3GPT-5.5 Pro91.8%
#4GPT-5.590.5%
#5Claude Opus 4.889.6%

+2 more models

🏆 Claude Mythos Preview 94.6%7 modelsUpdated 57d ago
🎓

GPQA Main

by NYU / Reid et al.
Reasoning & Logic● Live

The main set of Graduate-level Program QA questions across sciences.

sciencegraduateqa
#1Claude Mythos 592.0%
#2GPT-5.5 Pro90.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro Deep Think87.5%
#5Claude Opus 4.886.0%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🧠

HellaSwag

by UW / Allen AI
Reasoning & Logic● Live

Adversarial filtering for natural language inference. Tests commonsense reasoning.

commonsenseinferencenliadversarial
#1DeepSeek V4 Pro Base88.0%
#2GPT-5.587.5%
#3Claude Mythos 587.0%
#4Claude Fable 586.5%
#5Gemini 3 Pro85.0%
🏆 DeepSeek V4 Pro Base 88.0%5 modelsUpdated 57d ago
🤔

WinoGrande

by Allen AI (AI2)
Reasoning & Logic● Live

Large-scale adversarial dataset for commonsense reasoning, inspired by Winograd Schema Challenge.

commonsensepronounresolution
#1GPT-5.587.0%
#2Claude Mythos 586.5%
#3Claude Fable 586.0%
#4Gemini 3 Pro85.0%
#5Claude Opus 4.884.5%
🏆 GPT-5.5 87.0%5 modelsUpdated 57d ago

TruthfulQA

by OpenAI / Oxford
Reasoning & Logic● Live

817 questions that language models often answer falsely. Tests truthfulness and calibration.

truthfulnesscalibrationmisconceptions
#1Claude Mythos 582.0%
#2GPT-5.581.0%
#3Claude Fable 580.5%
#4Gemini 3 Pro79.0%
#5Claude Opus 4.878.5%
🏆 Claude Mythos 5 82.0%5 modelsUpdated 57d ago
💡

LogiQA 2.0

by NUS
Reasoning & Logic● Live

26,000+ logical reasoning questions from Chinese civil service exams, translated and curated.

logicreasoningexamchinese
#1Claude Mythos 580.0%
#2GPT-5.579.0%
#3Claude Fable 578.0%
#4Gemini 3 Pro77.0%
#5Claude Opus 4.876.0%
🏆 Claude Mythos 5 80.0%5 modelsUpdated 57d ago
♟️

StrategyQA

by Allen AI / Technion
Reasoning & Logic● Live

Questions requiring multi-step implicit reasoning strategies to answer.

strategymulti-stepimplicitcommonsense
#1GPT-5.585.0%
#2Claude Mythos 584.0%
#3Claude Fable 583.5%
#4Gemini 3 Pro82.0%
#5Claude Opus 4.881.0%
🏆 GPT-5.5 85.0%5 modelsUpdated 57d ago
🔬

PIQA (Physical Intuition QA)

by Yann LeCun / Meta
Reasoning & Logic● Live

Physical commonsense reasoning — understanding how the physical world works.

physicalcommonsenseintuition
#1GPT-5.590.0%
#2Claude Mythos 589.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro88.0%
#5Claude Opus 4.887.5%
🏆 GPT-5.5 90.0%5 modelsUpdated 57d ago
📖

OpenBookQA

by Allen AI (AI2)
Reasoning & Logic● Live

4th-grade science exam questions with an open book. Requires combining facts with reasoning.

sciencereasoningelementaryopen-book
#1GPT-5.586.0%
#2Claude Mythos 585.0%
#3Claude Fable 584.5%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.882.0%
🏆 GPT-5.5 86.0%5 modelsUpdated 57d ago
🔘

BoolQ

by Google Research
Reasoning & Logic● Live

Yes/no questions derived from Google search queries and Wikipedia paragraphs.

booleanyes-noreading-comprehension
#1GPT-5.592.0%
#2Claude Mythos 591.5%
#3Claude Fable 591.0%
#4Gemini 3 Pro90.0%
#5Claude Opus 4.889.5%
🏆 GPT-5.5 92.0%5 modelsUpdated 57d ago
🏅

ARB (Advanced Reasoning Benchmark)

by Google / Stanford
Reasoning & Logic● Live

Difficult reasoning benchmark covering math, logic, coding, and general reasoning from contest-level problems.

advancedcontestmathlogiccoding
#1Claude Mythos 572.0%
#2GPT-5.5 Pro70.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
🔮

MUSR (Multistate Soft Reasoning)

by Princeton / Meta
Reasoning & Logic● Live

Multistate soft reasoning requiring combining knowledge from multiple sentences with uncertainty.

multistatesoftuncertaintyreasoning
#1Claude Mythos 568.0%
#2GPT-5.566.5%
#3Claude Fable 565.0%
#4Gemini 3 Pro63.0%
#5Claude Opus 4.861.5%
🏆 Claude Mythos 5 68.0%5 modelsUpdated 57d ago
🏆

ARC Prize 2024

by ARC Prize
Reasoning & Logic● Live

The $600K competition for solving ARC-AGI. Tests general intelligence.

reasoningagicompetitionprizearc
#1Claude Mythos 595.5%
#2GPT-5.593.0%
#3Claude Fable 591.0%
#4Gemini 3 Pro88.0%
#5Claude Opus 4.885.5%
🏆 Claude Mythos 5 95.5%5 modelsUpdated 57d ago

📐 Mathematics (13)

📐

AIME 2024

by MAA
Mathematics● Live

American Invitational Mathematics Examination problems. Tests advanced mathematical problem-solving.

mathcompetitionaimeadvanced
#1GPT-595.7%
#2Grok 494.3%
#3o4 Mini94.0%
#4Claude Fable 593.3%
#5OpenAI o393.3%

+3 more models

🏆 GPT-5 95.7%8 modelsUpdated 57d ago
🔢

AIME 2025

by MAA
Mathematics● Live

2025 AIME competition problems for the most current math reasoning evaluation.

mathcompetitionaime2025latest
#1GPT-5.2 (xhigh)100.0%
#2GPT-5 Codex (high)100.0%
#3Gemini 3 Flash Preview85.0%
#4Claude Mythos 580.0%
#5Claude Fable 575.0%
🏆 GPT-5.2 (xhigh) 100.0%5 modelsUpdated 57d ago
🧮

FrontierMath (Epoch AI)

by Epoch AI
Mathematics● Live

Extremely difficult research-level math problems. Tiers 1-3 and Tier 4.

mathfrontierresearchhardepoch
#1GPT-5.5 Pro52.4%
#2GPT-5.551.7%
#3GPT-5.4 Pro50.0%
#4Claude Mythos 548.5%
#5Claude Fable 545.2%

+1 more models

🏆 GPT-5.5 Pro 52.4%6 modelsUpdated 57d ago

GSM8K

by OpenAI
Mathematics● Live

8,500 grade school math word problems requiring multi-step reasoning.

mathgrade-schoolword-problemsmulti-step
#1MiMo-V2.5-Pro99.6%
#2GPT-5.599.0%
#3Claude Mythos 598.5%
#4Gemini 3 Pro98.0%
#5Claude Fable 597.5%
🏆 MiMo-V2.5-Pro 99.6%5 modelsUpdated 57d ago
✖️

MATH

by UC Berkeley
Mathematics● Live

12,500 competition math problems from AMC, AIME, and Olympiad levels.

mathcompetitionamcolympiadhard
#1Claude Mythos 596.5%
#2GPT-5.5 Pro95.8%
#3Claude Fable 595.0%
#4Gemini 3 Pro Deep Think93.5%
#5Claude Opus 4.892.0%
🏆 Claude Mythos 5 96.5%5 modelsUpdated 57d ago
🧮

MATH-500

by UC Berkeley
Mathematics● Live

Curated 500-problem subset of MATH for efficient benchmarking.

mathsubsetefficient
#1Claude Mythos 597.0%
#2GPT-5.5 Pro96.0%
#3Claude Fable 595.5%
#4Gemini 3 Pro Deep Think94.0%
#5Claude Opus 4.893.0%
🏆 Claude Mythos 5 97.0%5 modelsUpdated 57d ago
📐

Minerva Math

by Google Research
Mathematics● Live

Multi-step mathematical reasoning for Minerva models. Graduate-level math problems.

mathgraduatemulti-stepgoogle
#1GPT-5.5 Pro94.0%
#2Claude Mythos 593.0%
#3Claude Fable 591.5%
#4Gemini 3 Pro89.0%
#5Claude Opus 4.887.5%
🏆 GPT-5.5 Pro 94.0%5 modelsUpdated 57d ago
📊

MathVista

by UT Austin
Mathematics● Live

Evaluates mathematical reasoning in visual contexts — charts, diagrams, geometry.

mathvisualchartsdiagramsgeometry
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.5%
#5Claude Opus 4.866.0%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
🌙

MathVerse

by NUS
Mathematics● Live

Multimodal math reasoning benchmark with visual elements essential for problem solving.

mathmultimodalvisualessential
#1Claude Mythos 568.0%
#2GPT-5.566.5%
#3Claude Fable 565.0%
#4Gemini 3 Pro63.0%
#5Claude Opus 4.861.5%
🏆 Claude Mythos 5 68.0%5 modelsUpdated 57d ago
🔣

GSM-Symbolic

by Apple
Mathematics● Live

Symbolic variant of GSM8K testing robust mathematical reasoning with varied numbers.

mathsymbolicrobustnessapple
#1GPT-5.595.0%
#2Claude Mythos 594.5%
#3Claude Fable 593.8%
#4Gemini 3 Pro92.0%
#5Claude Opus 4.891.0%
🏆 GPT-5.5 95.0%5 modelsUpdated 57d ago
🏅

NuminaMath

by AI-MO
Mathematics● Live

Large-scale math dataset from math competitions with chain-of-thought solutions.

mathcompetitionchain-of-thoughtlarge-scale
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago

GSM-Plus

by UC Berkeley
Mathematics● Live

Enhanced GSM8K with test-time augmentations for more robust math evaluation.

mathaugmentedrobustgrade-school
#1GPT-5.593.0%
#2Claude Mythos 592.5%
#3Claude Fable 591.8%
#4Gemini 3 Pro90.0%
#5Claude Opus 4.889.0%
🏆 GPT-5.5 93.0%5 modelsUpdated 57d ago
🧮

MathQA

by Allen AI / Georgia Tech
Mathematics● Live

37K math problems with explanations across arithmetic, algebra, and geometry.

mathexplanationsarithmeticalgebrageometry
#1Claude Mythos 588.0%
#2GPT-5.587.0%
#3Claude Fable 586.0%
#4Gemini 3 Pro84.0%
#5Claude Opus 4.882.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago

🔬 Science & Knowledge (5)

🏁

Humanity's Last Exam (HLE)

by Center for AI Safety (CAIS)
Science & Knowledge● Live

Extremely difficult benchmark of expert-level questions across many domains. Designed to be challenging even for top AI systems.

scienceexpertfrontierhardestlast-exam
#1GPT-5.446.0%
#2Claude Mythos 544.2%
#3GPT-5.5 Pro43.8%
#4Claude Fable 542.5%
#5Claude Opus 4.841.2%

+1 more models

🏆 GPT-5.4 46.0%6 modelsUpdated 57d ago
🔬

SciBench

by HKU
Science & Knowledge● Live

College-level science benchmark covering physics, chemistry, math, and biology.

sciencecollegephysicschemistrybiology
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
🏥

PubMedQA

by Columbia / Allen AI
Science & Knowledge● Live

Biomedical question answering from PubMed abstracts. Yes/No/Maybe answers.

biomedicalpubmedqaabstracts
#1GPT-5.585.0%
#2Claude Mythos 584.0%
#3Claude Fable 583.0%
#4Gemini 3 Pro82.0%
#5Claude Opus 4.881.0%
🏆 GPT-5.5 85.0%5 modelsUpdated 57d ago
🧬

GPSC (Graduate Program Science Competition)

by Independent
Science & Knowledge● Live

Graduate-level program-style science questions requiring deep domain expertise.

sciencegraduatecompetitiondeep-domain
#1Claude Mythos 568.0%
#2GPT-5.566.5%
#3Claude Fable 565.0%
#4Gemini 3 Pro63.0%
#5Claude Opus 4.861.5%
🏆 Claude Mythos 5 68.0%5 modelsUpdated 57d ago
⚕️

MedQA (USMLE)

by Independent
Science & Knowledge● Live

Medical licensing exam questions from USMLE. Tests medical knowledge at professional level.

medicalusmleprofessionalclinical
#1Claude Mythos 590.0%
#2GPT-5.589.0%
#3Claude Fable 588.0%
#4Gemini 3 Pro86.5%
#5Claude Opus 4.885.0%
🏆 Claude Mythos 5 90.0%5 modelsUpdated 57d ago

🖼️ Multimodal & Vision-Language (18)

🖼️

MMMU (Massive Multi-discipline Multimodal Understanding)

by NUS / NTU
Multimodal & Vision-Language● Live

11,500 questions requiring college-level subject knowledge and visual reasoning.

multimodalcollegevisualsubject-knowledge
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago

MMMU-Pro

by NUS
Multimodal & Vision-Language● Live

Enhanced MMMU with harder questions, more options, and reduced visual dependency.

multimodalharderenhanced10-options
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
🎵

MMBench

by OpenCompass
Multimodal & Vision-Language● Live

Comprehensive multimodal benchmark with 3,000+ questions across 20 abilities.

multimodalcomprehensive20-abilities
#1Claude Mythos 582.0%
#2GPT-5.580.5%
#3Claude Fable 579.0%
#4Gemini 3 Pro77.0%
#5Claude Opus 4.875.5%
🏆 Claude Mythos 5 82.0%5 modelsUpdated 57d ago
🌱

SEED-Bench

by ByteDance / VCU
Multimodal & Vision-Language● Live

19K multiple-choice questions across 12 evaluation dimensions for multimodal LLMs.

multimodal12-dimensionsvideoimage
#1GPT-5.572.0%
#2Claude Mythos 571.5%
#3Claude Fable 571.0%
#4Gemini 3 Pro70.0%
#5Claude Opus 4.869.0%
🏆 GPT-5.5 72.0%5 modelsUpdated 57d ago
🐾

MM-Vet

by NTU
Multimodal & Vision-Language● Live

Evaluates multimodal models as integrated vision-language assistants.

multimodalintegrationvision-languageassistant
#1Claude Mythos 575.0%
#2GPT-5.574.0%
#3Claude Fable 573.0%
#4Gemini 3 Pro71.5%
#5Claude Opus 4.870.0%
🏆 Claude Mythos 5 75.0%5 modelsUpdated 57d ago

MMStar

by Salesforce Research
Multimodal & Vision-Language● Live

Challenging multimodal benchmark requiring visual perception AND reasoning.

multimodalchallengingvisual-perceptionreasoning
#1Claude Mythos 568.0%
#2GPT-5.566.5%
#3Claude Fable 565.0%
#4Gemini 3 Pro63.0%
#5Claude Opus 4.861.5%
🏆 Claude Mythos 5 68.0%5 modelsUpdated 57d ago
📐

AI2D

by Allen AI (AI2)
Multimodal & Vision-Language● Live

Science diagrams from grade school textbooks. Tests visual reasoning on diagrams.

multimodaldiagramssciencetextbookgrade-school
#1GPT-5.590.0%
#2Claude Mythos 589.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro88.0%
#5Claude Opus 4.887.0%
🏆 GPT-5.5 90.0%5 modelsUpdated 57d ago
📊

ChartQA

by Allen AI / NYU
Multimodal & Vision-Language● Live

9,600 questions about charts requiring visual and logical reasoning.

multimodalchartsgraphsvisual-qa
#1GPT-5.588.0%
#2Claude Mythos 587.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro86.0%
#5Claude Opus 4.885.0%
🏆 GPT-5.5 88.0%5 modelsUpdated 57d ago
📄

DocVQA

by Visual Question Answering Challenge
Multimodal & Vision-Language● Live

12,767 questions about document images requiring OCR and understanding.

multimodaldocumentocrvisual-qa
#1GPT-5.592.0%
#2Claude Mythos 591.5%
#3Claude Fable 591.0%
#4Gemini 3 Pro90.0%
#5Claude Opus 4.889.0%
🏆 GPT-5.5 92.0%5 modelsUpdated 57d ago
🔤

TextVQA

by Allen AI (AI2)
Multimodal & Vision-Language● Live

Questions requiring reading and reasoning about text in images.

multimodaltext-in-imageocrreading
#1GPT-5.585.0%
#2Claude Mythos 584.5%
#3Claude Fable 584.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.882.0%
🏆 GPT-5.5 85.0%5 modelsUpdated 57d ago
🔠

OCRBench

by ByteDance / NUS
Multimodal & Vision-Language● Live

Comprehensive OCR evaluation covering text recognition, document understanding, and chart parsing.

multimodalocrdocumenttext-recognition
#1GPT-5.588.0%
#2Claude Mythos 587.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro86.0%
#5Claude Opus 4.885.0%
🏆 GPT-5.5 88.0%5 modelsUpdated 57d ago
👻

HallusionBench

by NTU / ByteDance
Multimodal & Vision-Language● Live

Evaluates hallucination in vision-language models through visual question answering.

multimodalhallucinationvisualqa
#1Claude Mythos 568.0%
#2GPT-5.566.5%
#3Claude Fable 565.0%
#4Gemini 3 Pro63.0%
#5Claude Opus 4.861.5%
🏆 Claude Mythos 5 68.0%5 modelsUpdated 57d ago
🗳️

POPE (Polling-based Object Probing Evaluation)

by NTU
Multimodal & Vision-Language● Live

Evaluates object hallucination via binary yes/no questions about image contents.

multimodalhallucinationobjectbinary
#1GPT-5.590.0%
#2Claude Mythos 589.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro88.0%
#5Claude Opus 4.887.0%
🏆 GPT-5.5 90.0%5 modelsUpdated 57d ago
🚗

RealWorldQA

by Google DeepMind
Multimodal & Vision-Language● Live

Real-world visual question answering for autonomous driving and robotics applications.

multimodalreal-worlddrivingrobotics
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
📷

LLaVA-Bench (In-the-Wild)

by UW-Madison / Microsoft
Multimodal & Vision-Language● Live

Open-ended evaluation of visual instruction following with diverse real-world images.

multimodalllavaopen-endedvisual-instruction
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago
🗨️

VisualChatBench

by Virginia Tech
Multimodal & Vision-Language● Live

Multi-turn visual conversation benchmark testing knowledge, reasoning, and perception.

multimodalchatmulti-turnconversation
#1Claude Mythos 565.0%
#2GPT-5.563.5%
#3Claude Fable 562.0%
#4Gemini 3 Pro60.0%
#5Claude Opus 4.858.5%
🏆 Claude Mythos 5 65.0%5 modelsUpdated 57d ago
🤖

GAIA (General AI Assistants)

by Meta / HuggingFace
Multimodal & Vision-Language● Live

466 questions requiring multi-step reasoning with real-world tools. Tests general-purpose assistants.

agentsgeneral-purposemulti-steptoolsreal-world
#1Claude Mythos 5 (agentic)73.7%
#2GPT-5.5 (agentic)72.0%
#3Claude Fable 5 (agentic)70.5%
#4Gemini 3 Pro (agentic)68.0%
#5Claude Opus 4.8 (agentic)66.5%
🏆 Claude Mythos 5 (agentic) 73.7%5 modelsUpdated 57d ago
🎬

Video-MME

by ByteDance / Tsinghua
Multimodal & Vision-Language● Live

Comprehensive video multimodal evaluation across short, medium, and long videos.

multimodalvideotemporallong-video
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago

📝 Natural Language Processing (15)

📰

SuperGLUE

by NYU / Google / Allen AI
Natural Language Processing● Live

9 tasks requiring deeper language understanding than GLUE. Includes coreference, QA, and more.

nlplanguageunderstandingcoreferenceqa
#1GPT-5.595.0%
#2Claude Mythos 594.5%
#3Claude Fable 594.0%
#4Gemini 3 Pro93.0%
#5Claude Opus 4.892.5%
🏆 GPT-5.5 95.0%5 modelsUpdated 57d ago
🧩

GLUE

by NYU / Washington / DeepMind
Natural Language Processing● Live

General Language Understanding Evaluation. 9 tasks for evaluating language understanding.

nlplanguageunderstandingclassification
#1GPT-5.593.0%
#2Claude Mythos 592.5%
#3Claude Fable 592.0%
#4Gemini 3 Pro91.0%
#5Claude Opus 4.890.5%
🏆 GPT-5.5 93.0%5 modelsUpdated 57d ago
🌏

XTREME

by Google Research
Natural Language Processing● Live

Cross-lingual benchmark covering 9 tasks across 40 languages.

nlpcross-lingualmultilingual40-languages
#1GPT-5.588.0%
#2Claude Mythos 587.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro86.0%
#5Claude Opus 4.885.0%
🏆 GPT-5.5 88.0%5 modelsUpdated 57d ago
🌐

XTREME-R

by Google Research
Natural Language Processing● Live

Refined XTREME with 10 tasks across 50+ languages. More balanced evaluation.

nlpcross-lingualmultilingualrefined
#1GPT-5.586.0%
#2Claude Mythos 585.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro84.0%
#5Claude Opus 4.883.0%
🏆 GPT-5.5 86.0%5 modelsUpdated 57d ago
🇨🇳

C-Eval

by Fudan / Tsinghua / Shanghai AI Lab
Natural Language Processing● Live

13,948 multiple-choice questions across 52 subjects in Chinese.

nlpchineseknowledgeexam
#1Claude Fable 590.0%
#2GPT-5.589.0%
#3Claude Mythos 588.5%
#4Gemini 3 Pro87.0%
#5Claude Opus 4.886.0%
🏆 Claude Fable 5 90.0%5 modelsUpdated 57d ago
🏮

CMMLU

by Chinese Academy of Sciences
Natural Language Processing● Live

67 tasks covering 68 subjects in Chinese for measuring massive multitask language understanding.

nlpchinesemultitask68-subjects
#1Claude Fable 589.0%
#2GPT-5.588.0%
#3Claude Mythos 587.5%
#4Gemini 3 Pro86.0%
#5Claude Opus 4.885.0%
🏆 Claude Fable 5 89.0%5 modelsUpdated 57d ago
📝

GAOKAO-Bench

by Tsinghua
Natural Language Processing● Live

Chinese National College Entrance Exam (Gaokao) problems for evaluating AI systems.

nlpchineseexamgaokaocollege
#1Claude Fable 585.0%
#2GPT-5.584.0%
#3Claude Mythos 583.5%
#4Gemini 3 Pro82.0%
#5Claude Opus 4.881.0%
🏆 Claude Fable 5 85.0%5 modelsUpdated 57d ago
📋

IFEval (Instruction Following Eval)

by Google Research
Natural Language Processing● Live

Evaluates how well LLMs follow precise instructions with verifiable criteria.

instruction-followingpreciseverifiablecompliance
#1Claude Fable 593.0%
#2GPT-5.592.0%
#3Claude Mythos 591.5%
#4Gemini 3 Pro90.0%
#5Claude Opus 4.889.0%
🏆 Claude Fable 5 93.0%5 modelsUpdated 57d ago
📑

DROP (Discrete Reasoning Over Paragraphs)

by Allen AI (AI2)
Natural Language Processing● Live

Reading comprehension requiring discrete reasoning (counting, sorting, comparison).

nlpreading-comprehensionreasoningcounting
#1GPT-5.590.0%
#2Claude Mythos 589.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro88.0%
#5Claude Opus 4.887.0%
🏆 GPT-5.5 90.0%5 modelsUpdated 57d ago
💬

QuAC (Question Answering in Context)

by Allen AI / University of Washington
Natural Language Processing● Live

Conversational question answering where questions depend on previous context.

nlpconversationalqacontext
#1GPT-5.575.0%
#2Claude Mythos 574.5%
#3Claude Fable 574.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.872.0%
🏆 GPT-5.5 75.0%5 modelsUpdated 57d ago
🗣️

CoQA (Conversational QA)

by Stanford NLP
Natural Language Processing● Live

7,000 conversations with 270,000 question-answer pairs across 7 domains.

nlpconversationalqamulti-domain
#1GPT-5.588.0%
#2Claude Mythos 587.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro86.0%
#5Claude Opus 4.885.0%
🏆 GPT-5.5 88.0%5 modelsUpdated 57d ago
📚

NarrativeQA

by Google DeepMind
Natural Language Processing● Live

Reading comprehension on full-length books and movie scripts. Tests narrative understanding.

nlpreadingnarrativebooksmovies
#1GPT-5.582.0%
#2Claude Mythos 581.5%
#3Claude Fable 581.0%
#4Gemini 3 Pro80.0%
#5Claude Opus 4.879.0%
🏆 GPT-5.5 82.0%5 modelsUpdated 57d ago
🔤

TyDi QA

by Google Research
Natural Language Processing● Live

Typologically diverse question answering in 11 languages.

nlpmultilingualqatypological
#1GPT-5.585.0%
#2Claude Mythos 584.5%
#3Claude Fable 584.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.882.0%
🏆 GPT-5.5 85.0%5 modelsUpdated 57d ago
🌍

XQuAD

by Google DeepMind
Natural Language Processing● Live

Cross-lingual extractive QA dataset covering 11 languages.

nlpcross-lingualextractive-qa
#1GPT-5.587.0%
#2Claude Mythos 586.5%
#3Claude Fable 586.0%
#4Gemini 3 Pro85.0%
#5Claude Opus 4.884.0%
🏆 GPT-5.5 87.0%5 modelsUpdated 57d ago

Open QA (Open Question Answering)

by Various
Natural Language Processing● Live

Open-ended question answering requiring knowledge synthesis from multiple sources.

qaopen-endedknowledgesynthesis
#1GPT-5.583.0%
#2Claude Mythos 582.5%
#3Claude Fable 582.0%
#4Gemini 3 Pro81.0%
#5Claude Opus 4.880.0%
🏆 GPT-5.5 83.0%5 modelsUpdated 57d ago

🤖 Agents & Tool Use (9)

🤖

AgentBench

by Tsinghua University
Agents & Tool Use● Live

Evaluates LLMs as agents across 8 environments: OS, DB, KG, web, card games, etc.

agentsmulti-environmenttool-useinteractive
#1Claude Mythos 575.0%
#2GPT-5.573.5%
#3Claude Fable 572.0%
#4Gemini 3 Pro70.0%
#5Claude Opus 4.868.5%
🏆 Claude Mythos 5 75.0%5 modelsUpdated 57d ago
🌐

WebArena

by Carnegie Mellon
Agents & Tool Use● Live

Realistic web environment benchmark with 812 tasks across 4 websites.

agentswebbrowsingnavigationrealistic
#1Claude Mythos 568.7%
#2GPT-5.4 Pro65.8%
#3Claude Opus 4.664.5%
#4Gemini 3 Pro62.0%
#5DeepSeek V4 Pro58.5%
🏆 Claude Mythos 5 68.7%5 modelsUpdated 57d ago
🖥️

OSWorld

by CMU
Agents & Tool Use● Live

Realistic computer operating system environment for evaluating AI agents.

agentsosdesktopcomputer-usegui
#1Claude Mythos 585.0%
#2Claude Fable 585.0%
#3Claude Opus 4.883.4%
#4GPT-5.582.0%
#5Gemini 3 Pro80.0%
🏆 Claude Mythos 5 85.0%5 modelsUpdated 57d ago
📞

τ-bench (Tau-bench)

by Sierra AI
Agents & Tool Use● Live

Evaluates AI agents on realistic customer service scenarios with policy adherence.

agentscustomer-servicepolicyrealistic
#1Claude Mythos 572.0%
#2GPT-5.570.5%
#3Claude Fable 569.0%
#4Gemini 3 Pro67.0%
#5Claude Opus 4.865.5%
🏆 Claude Mythos 5 72.0%5 modelsUpdated 57d ago
🔧

ToolBench

by Tsinghua NLP
Agents & Tool Use● Live

16,464 real-world APIs from RapidAPI for evaluating tool-use capabilities.

agentstool-useapirapidapireal-world
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago
🏦

API-Bank

by Alibaba DAMO
Agents & Tool Use● Live

Benchmark for evaluating tool-use capabilities with 73 API tools.

agentstool-useapifunction-calling
#1GPT-5.582.0%
#2Claude Mythos 581.5%
#3Claude Fable 581.0%
#4Gemini 3 Pro80.0%
#5Claude Opus 4.879.0%
🏆 GPT-5.5 82.0%5 modelsUpdated 57d ago
🔍

BrowseComp

by OpenAI
Agents & Tool Use● Live

Evaluates AI agents' ability to browse and find specific information on the web.

agentsbrowsingweb-searchinformation-retrieval
#1Claude Mythos 565.0%
#2GPT-5.563.5%
#3Claude Fable 562.0%
#4Gemini 3 Pro60.0%
#5Claude Opus 4.858.5%
🏆 Claude Mythos 5 65.0%5 modelsUpdated 57d ago
🖱️

MiniWob++

by Farama Foundation / UC Berkeley
Agents & Tool Use● Live

100+ web interaction tasks for training and evaluating web agents.

agentswebinteractiontrainingbrowser
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago
⚙️

T-Eval

by Tsinghua NLP
Agents & Tool Use● Live

Evaluates tool utilization capability — selection, evaluation, and execution.

agentstool-useselectionexecution
#1GPT-5.575.0%
#2Claude Mythos 574.5%
#3Claude Fable 574.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.872.0%
🏆 GPT-5.5 75.0%5 modelsUpdated 57d ago

🛡️ Safety, Alignment & Ethics (10)

🛡️

HarmBench

by Center for AI Safety (CAIS)
Safety, Alignment & Ethics● Live

Standardized evaluation of AI safety — tests resistance to jailbreaks and harmful content generation.

safetyjailbreakharmfuldefensealignment
#1Claude Mythos 592.0%
#2GPT-5.590.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro87.0%
#5Claude Opus 4.885.5%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🔒

TrustLLM

by NTU / Beijing University
Safety, Alignment & Ethics● Live

Comprehensive benchmark for trustworthiness of LLMs across 6 dimensions.

safetytrusttruthfulnessfairnessrobustness
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.881.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago
⚖️

BBQ (Bias Benchmark for QA)

by NYU
Safety, Alignment & Ethics● Live

58,492 questions testing social biases across 11 categories.

safetybiasfairnesssocialstereotypes
#1Claude Mythos 585.0%
#2GPT-5.583.5%
#3Claude Fable 582.0%
#4Gemini 3 Pro80.0%
#5Claude Opus 4.878.5%
🏆 Claude Mythos 5 85.0%5 modelsUpdated 57d ago
☠️

RealToxicityPrompts

by Allen AI (AI2)
Safety, Alignment & Ethics● Live

100K naturally-occurring prompts for evaluating language model toxicity.

safetytoxicityharmfulgeneration
#1Claude Mythos 592.0%
#2GPT-5.590.5%
#3Claude Fable 589.0%
#4Gemini 3 Pro87.0%
#5Claude Opus 4.885.5%
🏆 Claude Mythos 5 92.0%5 modelsUpdated 57d ago
🏴

TOXIGEN

by Allen AI / UW
Safety, Alignment & Ethics● Live

274K machine-generated and human-annotated toxic and benign texts across 13 demographics.

safetytoxicdemographicsimplicit-hate
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.881.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago
👥

CrowS-Pairs

by NYU
Safety, Alignment & Ethics● Live

1,508 sentence pairs testing nine types of stereotypical biases.

safetybiasstereotypespairs
#1Claude Mythos 590.0%
#2GPT-5.588.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro85.0%
#5Claude Opus 4.883.5%
🏆 Claude Mythos 5 90.0%5 modelsUpdated 57d ago
🎯

StereoSet

by Stanford
Safety, Alignment & Ethics● Live

2,282 sentences testing stereotypical biases across 4 categories.

safetybiasstereotypeslanguage
#1Claude Mythos 582.0%
#2GPT-5.580.5%
#3Claude Fable 579.0%
#4Gemini 3 Pro77.0%
#5Claude Opus 4.875.5%
🏆 Claude Mythos 5 82.0%5 modelsUpdated 57d ago
🚨

SafetyBench

by Tsinghua University
Safety, Alignment & Ethics● Live

11,435 multiple-choice questions for evaluating LLM safety across 7 categories.

safetycomprehensive7-categoriesmultiple-choice
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.881.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago
🚫

StrongREJECT

by Yale
Safety, Alignment & Ethics● Live

Evaluates how effectively LLMs refuse harmful requests.

safetyrejectionharmfulrefusal
#1Claude Mythos 590.0%
#2GPT-5.588.5%
#3Claude Fable 587.0%
#4Gemini 3 Pro85.0%
#5Claude Opus 4.883.5%
🏆 Claude Mythos 5 90.0%5 modelsUpdated 57d ago
🇨🇳

WeiTechSafety

by Independent
Safety, Alignment & Ethics● Live

Chinese safety benchmark for evaluating AI ethics and safety.

safetychineseethicsalignment
#1Claude Mythos 585.0%
#2GPT-5.583.5%
#3Claude Fable 582.0%
#4Gemini 3 Pro80.0%
#5Claude Opus 4.878.5%
🏆 Claude Mythos 5 85.0%5 modelsUpdated 57d ago

🔍 Retrieval & RAG (4)

📥

MTEB (Massive Text Embedding Benchmark)

by HuggingFace
Retrieval & RAG● Live

Evaluates text embeddings across 8 tasks and 56+ datasets.

retrievalembeddingstextsemanticsimilarity
#1OpenAI text-embedding-3-large64.6%
#2Cohere embed-v464.1%
#3Voyage-363.5%
#4Nomic Embed v262.0%
#5GTE-Qwen2-7B-instruct61.5%
🏆 OpenAI text-embedding-3-large 64.6%5 modelsUpdated 57d ago
🔎

BEIR (Benchmarking IR)

by University College London
Retrieval & RAG● Live

Heterogeneous benchmark for evaluating information retrieval across 18 datasets.

retrievalinformation-retrievalheterogeneous18-datasets
#1Cohere embed-v458.0%
#2OpenAI text-embedding-3-large57.5%
#3Voyage-357.0%
#4BGE-M356.0%
#5GTE-Qwen2-7B-instruct55.5%
🏆 Cohere embed-v4 58.0%5 modelsUpdated 57d ago
📚

RAGBench

by Salesforce Research
Retrieval & RAG● Live

Benchmark for evaluating Retrieval-Augmented Generation systems.

ragretrieval-augmentedgenerationqa
#1Claude Mythos 582.0%
#2GPT-5.580.5%
#3Claude Fable 579.0%
#4Gemini 3 Pro77.0%
#5Claude Opus 4.875.5%
🏆 Claude Mythos 5 82.0%5 modelsUpdated 57d ago
🔧

CRAG (Corrective RAG)

by ServiceNow
Retrieval & RAG● Live

Benchmark for evaluating and developing Corrective RAG systems.

ragcorrectiveretrievalre-ranking
#1Claude Mythos 578.0%
#2GPT-5.576.5%
#3Claude Fable 575.0%
#4Gemini 3 Pro73.0%
#5Claude Opus 4.871.5%
🏆 Claude Mythos 5 78.0%5 modelsUpdated 57d ago

🎮 Reinforcement Learning (5)

🕹️

Atari 100K

by Independent / IRIS
Reinforcement Learning● Live

Standard RL benchmark: 26 Atari games with only 100K environment interactions.

rlatarisample-efficient100k
#1BBF (Rainbow+DTSIL)8200.0%
#2GDI (BBF)7800.0%
#3Sampled MuZero7200.0%
#4DreamerV36800.0%
#5EfficientZero6200.0%
🏆 BBF (Rainbow+DTSIL) 8200.0%5 modelsUpdated 57d ago
🎮

DeepMind Control Suite (DMControl)

by Google DeepMind
Reinforcement Learning● Live

Continuous control benchmark with 30 tasks of varying difficulty.

rlcontinuous-controlroboticsdeepmind
#1DreamerV388.0%
#2Sampled MuZero85.0%
#3EfficientZero82.0%
#4BBF80.0%
#5Rainbow78.0%
🏆 DreamerV3 88.0%5 modelsUpdated 57d ago
🎲

ProcGen Benchmark

by OpenAI
Reinforcement Learning● Live

16 procedurally generated environments for evaluating generalization in RL.

rlproceduralgeneralizationopenai
#1BBF (Rainbow+DTSIL)92.0%
#2Sampled MuZero88.0%
#3EfficientZero85.0%
#4DreamerV382.0%
#5Rainbow78.0%
🏆 BBF (Rainbow+DTSIL) 92.0%5 modelsUpdated 57d ago
⛏️

Crafter

by Google DeepMind
Reinforcement Learning● Live

Open-world survival game benchmark for evaluating memory and long-horizon planning in RL.

rlopen-worldsurvivalmemoryplanning
#1DreamerV385.0%
#2Sampled MuZero82.0%
#3EfficientZero80.0%
#4BBF78.0%
#5Rainbow75.0%
🏆 DreamerV3 85.0%5 modelsUpdated 57d ago
📊

Open RL Benchmark

by Open RL Benchmark
Reinforcement Learning● Live

Comprehensive benchmark for offline and online reinforcement learning.

rlcomprehensiveofflineonline
#1DreamerV378.0%
#2Sampled MuZero76.0%
#3EfficientZero74.0%
#4BBF72.0%
#5Rainbow70.0%
🏆 DreamerV3 78.0%5 modelsUpdated 57d ago

👁️ Computer Vision (4)

🖼️

ImageNet

by Stanford / Princeton
Computer Vision● Live

14M images across 20K+ categories. The foundational vision benchmark.

visionclassification14m-images20k-classes
#1GPT-5.592.0%
#2DINOv2 ViT-G91.0%
#3CLIP ViT-G/1490.0%
#4SigLIP89.0%
#5EVA-CLIP88.0%
🏆 GPT-5.5 92.0%5 modelsUpdated 57d ago
📷

MS COCO

by Microsoft / FAIR
Computer Vision● Live

82K images with object detection, segmentation, and captioning annotations.

visiondetectionsegmentationcaptioning
#1GPT-5.585.0%
#2Florence-2 Large83.0%
#3Grounding DINO81.0%
#4YOLOv1080.0%
#5SAM 278.0%
🏆 GPT-5.5 85.0%5 modelsUpdated 57d ago
🗺️

ADE20K

by MIT CSAIL
Computer Vision● Live

20K images with pixel-level semantic segmentation annotations.

visionsegmentationsemanticpixel-level
#1SAM 278.0%
#2OneFormer76.0%
#3Mask2Former74.0%
#4SegGPT72.0%
#5DINOv2+Linear70.0%
🏆 SAM 2 78.0%5 modelsUpdated 57d ago
🚗

nuScenes

by Motional / nuTonomy
Computer Vision● Live

Autonomous driving dataset with 1.4M camera images, LiDAR, and radar.

visionautonomous-drivinglidarmulti-modal
#1GPT-5.582.0%
#2UniAD80.0%
#3VAD78.0%
#4MotionLM76.0%
#5MT374.0%
🏆 GPT-5.5 82.0%5 modelsUpdated 57d ago

🎙️ Speech & Audio (3)

🎙️

LibriSpeech

by CMU / Johns Hopkins
Speech & Audio● Live

1000 hours of read English speech from audiobooks. Standard ASR benchmark.

speechasrenglishreadaudiobook
#1Whisper large-v398.5%
#2Whisper large-v3-turbo98.0%
#3GPT-4o-mini-transcribe97.5%
#4Deepgram Nova-397.0%
#5Conformer-CTC96.5%
🏆 Whisper large-v3 98.5%5 modelsUpdated 57d ago
🗣️

Common Voice

by Mozilla
Speech & Audio● Live

Mozilla's crowdsourced speech dataset across 100+ languages.

speechcrowdsourced100-languagesmozilla
#1Whisper large-v392.0%
#2Whisper large-v3-turbo91.0%
#3GPT-4o-mini-transcribe90.0%
#4Deepgram Nova-389.0%
#5Conformer-CTC88.0%
🏆 Whisper large-v3 92.0%5 modelsUpdated 57d ago
🎶

SUPERB

by NYU / Meta
Speech & Audio● Live

Speech processing benchmark covering ASR, SER, PR, SD, and more.

speechcomprehensiveasrrecognitionemotion
#1Whisper large-v390.0%
#2Whisper large-v3-turbo89.0%
#3GPT-4o-mini-transcribe88.0%
#4Deepgram Nova-387.0%
#5Conformer-CTC86.0%
🏆 Whisper large-v3 90.0%5 modelsUpdated 57d ago

🌍 Translation (3)

🌍

WMT (Workshop on Machine Translation)

by ACL / StatMT
Translation● Live

Annual shared task for machine translation. The gold standard for translation evaluation.

translationmachine-translationannualshared-task
#1GPT-5.535.0%
#2Claude Mythos 534.5%
#3Claude Fable 534.0%
#4Gemini 3 Pro33.0%
#5DeepL LLM32.5%
🏆 GPT-5.5 35.0%5 modelsUpdated 57d ago
🗺️

FLORES-200

by Meta FAIR
Translation● Live

200 language translation benchmark with high-quality sentence pairs.

translation200-languagesmultilingualmeta
#1GPT-5.532.0%
#2Claude Mythos 531.5%
#3Claude Fable 531.0%
#4Gemini 3 Pro30.0%
#5DeepL LLM29.5%
🏆 GPT-5.5 32.0%5 modelsUpdated 57d ago
📑

OPUS-100

by Edinburgh / Facebook
Translation● Live

Parallel corpus covering 100 language pairs for translation training and evaluation.

translation100-pairsparallelcorpus
#1GPT-5.530.0%
#2Claude Mythos 529.5%
#3Claude Fable 529.0%
#4Gemini 3 Pro28.0%
#5DeepL LLM27.5%
🏆 GPT-5.5 30.0%5 modelsUpdated 57d ago

📦 Other Specialized (3)

⚖️

LegalBench

by Stanford / Duke / Independent
Other Specialized● Live

162 legal tasks for evaluating LLM capabilities in legal reasoning.

legalreasoningcontractsregulations
#1Claude Mythos 588.0%
#2GPT-5.586.5%
#3Claude Fable 585.0%
#4Gemini 3 Pro83.0%
#5Claude Opus 4.881.5%
🏆 Claude Mythos 5 88.0%5 modelsUpdated 57d ago
💰

FinBench

by Independent
Other Specialized● Live

Financial domain benchmark covering reasoning, knowledge, and calculation.

financeknowledgereasoningcalculation
#1Claude Mythos 582.0%
#2GPT-5.580.5%
#3Claude Fable 579.0%
#4Gemini 3 Pro77.0%
#5Claude Opus 4.875.5%
🏆 Claude Mythos 5 82.0%5 modelsUpdated 57d ago
👟

Sneakerhead Bench

by Independent
Other Specialized● Live

Benchmark for evaluating AI knowledge of sneaker culture and brand expertise.

culturaldomain-expertisefashion
#1GPT-5.575.0%
#2Claude Mythos 573.5%
#3Claude Fable 572.0%
#4Gemini 3 Pro70.0%
#5Claude Opus 4.868.5%
🏆 GPT-5.5 75.0%5 modelsUpdated 57d ago