Gaussian GRPO proposes mapping multimodal task rewards to a standard Gaussian
Its creators say reasoning- and perception-heavy tasks have different reward distributions and learning dynamics.
TLDR
The team behind Gaussian GRPO (G²RPO) says its method uses 1D optimal transport to map each task’s reward distribution to a standard Gaussian during multimodal LLM post-training. They say the resulting OpenVLThinker v2 performs strongly on knowledge, math, charts and document understanding, and outperforms closed-weight models more than 100 times larger on several benchmarks.
Combined views
369
1 Source, first seen ago