VERA paper describes training AI agents and their harnesses together
A user says VERA turns benchmark trajectories into more than 9,000 restartable sandboxes with rubric scoring.
TLDR
A user says Nvidia’s VERA paper describes updating model weights and agent harnesses together using verifiable environments. Harness edits must pass self-tests and add at least five points on the development set; model checkpoints are rejected if scores drop by more than 20%, the user says. They report that the 9B agent beat the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 on AutoMedBench. The environment corpus is open-sourced, they say.
Combined views
77
1 Source, first seen ago
