Proxy Policy Steering adapts VLA models without weight access, a post says
Rather than fine-tuning the base model, PPS trains two lightweight proxy policies to steer its behavior while leaving it unchanged, according to the post.
TLDR
A post describes Proxy Policy Steering (PPS) as a way to specialize vision-language-action (VLA) models for new tasks at inference time—when the model is being run—without accessing their weights. According to the post, PPS trains two lightweight proxy policies and uses their velocity-space residual to steer the frozen base model instead of fine-tuning it directly.
Combined views
3.1K
1 Source, first seen 19d ago