OpenAI Audits SWE-Bench Pro — and Finds Nearly a Third of Tasks Broken
First OpenAI told the community to switch to SWE-Bench Pro. Now it's withdrawing that recommendation: up to 34% of tasks are defective. A wake-up call for anyone who takes coding benchmarks at face value.