Benchmark Registry
Live TRACE scores aggregate independent and vendor-reported benchmark results. Methodology →
TRACE Aggregate Model Scores
Weighted average across all published benchmarks. Independent runs 3× weight; vendor runs 1×. Click a score to see source data.
4 models with insufficient data (fewer than 2 distinct benchmarks)
Claude Mythos 5, GLM-5.2, Gemini 3.5 Flash, Llama 4
Benchmark Sources
| Benchmark | Version | Owner | Domain | Health |
ARC-AGI-2 Abstract reasoning corpus designed to resist memorisation. Tests fluid intelligence. | 2.0 | ARC Prize Foundation | reasoning | Healthy |
Artificial Analysis Intelligence Index Cross-model intelligence, latency, and pricing comparison. | 2.0 | Artificial Analysis | general | Healthy |
GPQA Diamond Graduate-level physics, chemistry, and biology questions — gold-standard hard benchmark. | 1.0 | NYU / Anthropic | science | Healthy |
HumanEval+ Function synthesis benchmark with expanded test cases to catch false positives. | 1.0 | OpenAI / EvalPlus | coding | Healthy |
LMSYS Chatbot Arena Human preference evaluation via blind pairwise comparisons. | 2.0 | LMSYS / UC Berkeley | general | Healthy |
LiveBench Contamination-resistant live benchmark updated with recent data. | 2.0 | Abacus.AI | general | Healthy |
MLPerf Inference Industry-standard hardware inference performance with reproducible results. | 5.0 | MLCommons | hardware | Healthy |
MMLU-Pro Massive multitask language understanding with harder, reasoning-focused questions. | 1.0 | University of Waterloo | general | Healthy |
SWE-bench Verified Evaluate coding agents on real-world GitHub issues with verified patches. | 1.0 | Princeton NLP / OpenAI | coding | Healthy |
Stanford HELM Holistic evaluation of language models across 40+ scenarios. | 2.0 | Stanford CRFM | general | Healthy |