GPT-5.6 Sol
Flagship OpenAI coding and agent model. Strongest on broad agentic coding benchmarks.
- active
- 2026-07-14
- unknown
- 128K
- text,code,image
- tool use, structured output, API
- 2026-07-20
Reviewed use cases
- agentic coding
- terminal use
- complex reasoning
Recorded limitations
- SWE-Bench Pro lags behind Claude Fable 5
Benchmark Performance
77
| Benchmark | Score | Provenance | Date | Source |
|---|---|---|---|---|
| SWE-bench Verified | 72.7% | Vendor | 01/07/2026 | Source |
| LMSYS Chatbot Arena | ~1350 | Independent | 01/07/2026 | Source |
| MMLU-Pro | 88.0% | Vendor | 01/07/2026 | Source |
| HumanEval+ | 95.0% | Vendor | 01/07/2026 | Source |