Benchmark Registry

Live TRACE scores aggregate independent and vendor-reported benchmark results. Methodology →

TRACE Aggregate Model Scores

Weighted average across all published benchmarks. Independent runs 3× weight; vendor runs 1×. Click a score to see source data.

ModelProviderTRACE ScoreBenchmarksIndependentVendorLast run
Qwen3-Coder-Next Alibaba (Qwen) 81 2 0 2 01/07/2026
GPT-5.6 Sol OpenAI 77 4 1 3 01/07/2026
Claude Fable 5 Anthropic 69.7 2 1 1 01/07/2026
Kimi K3 Moonshot AI 64.4 2 1 1 01/07/2026
4 models with insufficient data (fewer than 2 distinct benchmarks)

Claude Mythos 5, GLM-5.2, Gemini 3.5 Flash, Llama 4

Benchmark Sources

BenchmarkVersionOwnerDomainHealth
ARC-AGI-2
Abstract reasoning corpus designed to resist memorisation. Tests fluid intelligence.
2.0ARC Prize Foundationreasoning Healthy
Artificial Analysis Intelligence Index
Cross-model intelligence, latency, and pricing comparison.
2.0Artificial Analysisgeneral Healthy
GPQA Diamond
Graduate-level physics, chemistry, and biology questions — gold-standard hard benchmark.
1.0NYU / Anthropicscience Healthy
HumanEval+
Function synthesis benchmark with expanded test cases to catch false positives.
1.0OpenAI / EvalPluscoding Healthy
LMSYS Chatbot Arena
Human preference evaluation via blind pairwise comparisons.
2.0LMSYS / UC Berkeleygeneral Healthy
LiveBench
Contamination-resistant live benchmark updated with recent data.
2.0Abacus.AIgeneral Healthy
MLPerf Inference
Industry-standard hardware inference performance with reproducible results.
5.0MLCommonshardware Healthy
MMLU-Pro
Massive multitask language understanding with harder, reasoning-focused questions.
1.0University of Waterloogeneral Healthy
SWE-bench Verified
Evaluate coding agents on real-world GitHub issues with verified patches.
1.0Princeton NLP / OpenAIcoding Healthy
Stanford HELM
Holistic evaluation of language models across 40+ scenarios.
2.0Stanford CRFMgeneral Healthy