Sudo L7 benchmark tests coding agents on staff-level software engineering
A post introducing the benchmark says its 60 tasks are mostly in private production repositories at real companies.
TLDR
The sudo L7 announcement describes 60 staff-level software engineering tasks, mostly in private production repos at real companies and written by engineers who owned the systems. It says expert rubrics grade functional correctness, engineering craft, architectural judgment, thought partnership and unnecessary complexity. The post reports that the best coding agents succeed on only about 45% of tasks, arguing that writing code is only part of the job.
Combined views
516
2 Sources, first seen ago
