HumanEval+
Healthy 1.0 · OpenAI / EvalPlus · coding
Function synthesis benchmark with expanded test cases to catch false positives.
- Reproducibility
- reproducible
- Contamination concern
- medium
- Saturation
- Not assessed
- Last reviewed
- 2026-07-20
Reviewed results
| Model | Score | Date | Provenance | Source |
| Qwen3-Coder-Next | 94.0% | 2026-07-01 | Vendor-run | Evidence |
| GPT-5.6 Sol | 95.0% | 2026-07-01 | Vendor-run | Evidence |