SWE-Bench Pro Is 40% Noise, OpenAI Data Shows
The benchmark everyone uses to prove their coding AI is best might be measuring noise instead of skill. OpenAI analyzed SWE-Bench Pro, a widely-used coding benchmark, and found significant reliability issues that call into question how we're measuring AI coding ability
Continue reading ›