Read news on SWE-Bench with our app.
Read more in the app
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
Anthropic overtakes OpenAI: Claude Opus 4 codes seven hours nonstop, sets record SWE-Bench score and reshapes enterprise AI
Some critical issues with the SWE-bench dataset