New DeepSWE Benchmark Upends LLM Coding Performance Rankings, Exposing Flaws in Industry Standards
A new coding benchmark, DeepSWE by Data Curve, challenges conventional LLM performance metrics, revealing significant disparities between models in real-world development tasks. Its findings suggest widespread issues with existing benchmarks, emphasizing the superior practical capabilities of leading OpenAI models.