Supabase Launches Evals Benchmark for AI Coding Agents
Founders and investors react to Supabase's new benchmark testing AI coding agents on real tasks.
TLDR
Supabase posted about Evals, a benchmark evaluating AI coding agents like Claude Code and Codex on real-world tasks involving its platform. Paul Graham replied that services agents cannot use will fail and predicted universal adoption. Dave Morin called it a positive trend toward evaluations that reflect actual use. The post drew these comments on X, with the benchmark positioned as practical rather than synthetic.
Supabase Launches Evals Benchmark for AI Coding Agents
Founders and investors react to Supabase's new benchmark testing AI coding agents on real tasks.