Announcement
QF3 promises one off-policy update for training humanoids and fine-tuning AI models
Its introducer calls QF3 “Fast Flow RL with Filtered Q-Gradients” and lists VLAs and image models as fine-tuning targets.
TLDR
QF3’s introducer says one simple off-policy update can train humanoids from scratch and fine-tune VLAs and image models. They call the method “Fast Flow RL with Filtered Q-Gradients.”
Combined views
5K
2 Sources, first seen ago
