Supabase Launches Evals Benchmark for AI Coding Agents
Founders and investors react to Supabase's new benchmark testing AI coding agents on real tasks.
Supabase posted about Evals, a benchmark evaluating AI coding agents like Claude Code and Codex on real-world tasks involving its platform. Paul Graham replied that services agents cannot use will fail and predicted universal adoption. Dave Morin called it a positive trend toward evaluations that reflect actual use. The post drew these comments on X, with the benchmark positioned as practical rather than synthetic.
Introducing Supabase Evals. Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do.
Combined views
250K