Announcement
Capture proxy enables reinforcement learning for open models across coding-agent harnesses
A post says the OpenEnv proxy records token IDs and log probabilities while harness runs remain unmodified.
TLDR
A post says a capture proxy added to OpenEnv sits between coding-agent harnesses and the model, recording token IDs and log probabilities without changing harness runs. Harbor supplies tasks and sandboxes, while TRL handles training with async GRPO. The post reports that after training across four harnesses, the average rose from 42% to 54%, and the Claude Code result rose from 33% to 49%.
Combined views
28.3K
6 Sources, first seen 4h ago
