Verifying sampled shell commands can improve terminal-agent benchmark results
A post describing an NVIDIA paper says TerminalBench-Lite Pass@1 rose from 50.0% to 68.0% with a GPT-5.6 Sol verifier choosing among eight sampled actions.
TLDR
A post describing an NVIDIA paper says its Mid-Harness method verifies candidate shell commands before running one, without changing the generator or harness. With a GPT-5.6 Sol verifier choosing among eight sampled actions, TerminalBench-Lite Pass@1 rose from 50.0% to 68.0%. The post says extra samples added little with a weak verifier, while combining action and trajectory sampling reached higher success at lower estimated token cost than sampling full trajectories alone.
Combined views
387
1 Source, first seen ago
