Proxy Policy Steering adapts robot models without accessing their weights, researcher says
Instead of fine-tuning the base model, a researcher says PPS trains two lightweight proxy policies to steer it toward new tasks at inference time—when the model is running.
TLDR
A researcher describes Proxy Policy Steering (PPS) as a way to adapt vision-language-action models to new tasks without accessing their weights. The method reportedly trains two lightweight proxy policies and uses their velocity-space residual to steer the base model while leaving it unchanged. On September 11, 2026, the researcher also shared links to the paper, project website and code, and said the work was due to appear at CoRL 2026.
Combined views
286
1 Source, first seen 19d ago