• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Proxy Policy Steering adapts VLA models without weight access, a post says

    Rather than fine-tuning the base model, PPS trains two lightweight proxy policies to steer its behavior while leaving it unchanged, according to the post.

    KF
    1 Source, 19d ago, first seen 19d ago

    TLDR

    A post describes Proxy Policy Steering (PPS) as a way to specialize vision-language-action (VLA) models for new tasks at inference time—when the model is being run—without accessing their weights. According to the post, PPS trains two lightweight proxy policies and uses their velocity-space residual to steer the frozen base model instead of fine-tuning it directly.

    Combined views

    3.1K

    1 Source, first seen 19d ago

    Combined views

    3.1K

    1 Source, first seen 19d ago

    57 likes
    57 likes
    1 comments
    37 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    37 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @KuanFangProxy Policy Steering (PPS) specializes VLA models to new tasks at inference time without accessing their weights. Instead of direct fine-tuning, it trains two lightweight proxy policies and uses their velocity-space residual to steer the frozen base's behavior.

    1 Source

    @KuanFangProxy Policy Steering (PPS) specializes VLA models to new tasks at inference time without accessing their weights. Instead of direct fine-tuning, it trains two lightweight proxy policies and uses their velocity-space residual to steer the frozen base's behavior.