LMSYS Chatbot Arena
Healthy 2.0 · LMSYS / UC Berkeley · general
Human preference evaluation via blind pairwise comparisons.
- Reproducibility
- partially_reproducible
- Contamination concern
- low
- Saturation
- Not assessed
- Last reviewed
- 2026-07-20
Reviewed results
| Model | Score | Date | Provenance | Source |
| Kimi K3 | ~1300 | 2026-07-01 | Independent | Evidence |
| Claude Fable 5 | ~1330 | 2026-07-01 | Independent | Evidence |
| GPT-5.6 Sol | ~1350 | 2026-07-01 | Independent | Evidence |