Rohan Paul on Learned Context Policies for Agents
Tweet argues long-horizon agents need learned context policies rather than larger windows alone.
TLDR
Rohan Paul, a Bengaluru-based machine learning engineer, posted that long-running agents keep adding searches, tool outputs, and reasoning until the model spends more tokens on context than useful work. He stated that agents need to learn what deserves to stay in context, and that long-horizon agents work better when context management itself is a learned policy. The post included a screenshot of an arXiv paper title as attachment.
Combined views
5.3K
2 Sources, first seen 27d ago