HumanEval+

Healthy

1.0 · OpenAI / EvalPlus · coding

Function synthesis benchmark with expanded test cases to catch false positives.

Reproducibility
reproducible
Contamination concern
medium
Saturation
Not assessed
Last reviewed
2026-07-20

Reviewed results

ModelScoreDateProvenanceSource
Qwen3-Coder-Next94.0%2026-07-01Vendor-runEvidence
GPT-5.6 Sol95.0%2026-07-01Vendor-runEvidence