VeriHarness checks AI agent claims even when rollouts agree
A post about a Google paper says the verifier checks disputed claims against workspace evidence and probes unanimous answers for missed requirements.
TLDR
A post about a Google paper says VeriHarness uses the same base model to check disputed rollout claims against workspace evidence and probe agreed-upon claims for missed requirements. The post reports the best selection scores among tested baselines across five long-horizon benchmarks. It says evidence-backed revision added 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 with Claude Opus 4.8. The authors also released about 26,000 rollouts, according to the post.
Combined views
1 Source, first seen ago
