/AIOpenAI withdraws endorsement of SWE-Bench Pro after audit finds 30% of its coding tasks are brokenThe benchmark's 70% noise ceiling prevents meaningful frontier model evaluation.FBHHTRGLFIJYAPTPKAEMPSRSWCV(T-T(MZMBACSKNTCombined views404.3K30 posts, first seen 48d ago3.3K likes commentsOpenAI withdraws endorsement of SWE-Bench Pro after audit finds 30% of its coding tasks are brokenThe benchmark's 70% noise ceiling prevents meaningful frontier model evaluation.FBHHTRGLFIJYAPTPKAEMPSRSWCV(T-T(MZMBACSKNTReactions23 postsFBFlorian Brand@xeophon6 weeks agoopenai looked at the data and it stared right back :( https://twitter.com/OpenAI/status/2074958149426241894likes: 59replies: 0bookmarks: 2reposts: 1TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoHow does this keep happening with AI benchmarks? Why don't these audits happen immediately? https://twitter.com/OpenAI/status/2074972179385720836likes: 50replies: 12bookmarks: 13reposts: 2TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoLike, no, these benchmarks need to be trustworthy from the start. Nothing to do with models improving. If they are not trustworthy, they are inherently bad benchmarks for coding! https://x.com/OpenAI/status/2074972185895342084?s=20likes: 22replies: 2bookmarks: 1reposts: 1TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoTerrible https://x.com/TaliaRinger/status/2075233130642919478?s=20likes: 5replies: 1bookmarks: 0reposts: 0TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoCredit: https://x.com/jpgabor/status/2075215748952203675?s=20likes: 6replies: 0bookmarks: 0reposts: 0GLgavin leech (Non-Reasoning)@gleech6 weeks agofinally https://twitter.com/OpenAI/status/2074972179385720836likes: 9replies: 1bookmarks: 1reposts: 0Show allCombined views404.3K30 posts, first seen 48d ago3.3K likes
FBFlorian Brand@xeophon6 weeks agoopenai looked at the data and it stared right back :( https://twitter.com/OpenAI/status/2074958149426241894likes: 59replies: 0bookmarks: 2reposts: 1
TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoHow does this keep happening with AI benchmarks? Why don't these audits happen immediately? https://twitter.com/OpenAI/status/2074972179385720836likes: 50replies: 12bookmarks: 13reposts: 2
TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoLike, no, these benchmarks need to be trustworthy from the start. Nothing to do with models improving. If they are not trustworthy, they are inherently bad benchmarks for coding! https://x.com/OpenAI/status/2074972185895342084?s=20likes: 22replies: 2bookmarks: 1reposts: 1
TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoTerrible https://x.com/TaliaRinger/status/2075233130642919478?s=20likes: 5replies: 1bookmarks: 0reposts: 0
TRTalia Ringer 🕊🪬@TaliaRinger6 weeks agoCredit: https://x.com/jpgabor/status/2075215748952203675?s=20likes: 6replies: 0bookmarks: 0reposts: 0
GLgavin leech (Non-Reasoning)@gleech6 weeks agofinally https://twitter.com/OpenAI/status/2074972179385720836likes: 9replies: 1bookmarks: 1reposts: 0