Announcement
Open-source proxy turns coding-agent harnesses into reinforcement-learning environments
The projectโs post says its proxy records token IDs and log probabilities from vLLM while 10 harnesses run unmodified.
TLDR
The projectโs post says a capture proxy lets existing coding harnesses serve as reinforcement-learning environments without changes to the harnesses or training code. It reports 31% fewer tool calls on tasks the model already solved after adding a reward for using fewer calls, with reductions in every harness and about half as many calls under Codex. The post says the proxy, trainer, tasks, training code and seven trained models are open source.
Combined views
32.9K
3 Sources, first seen ago
561 likes73 comments453 saves132 reposts
