Composite rewards reportedly improve FLUX.2-dev and Ideogram 4 after post-training
Arena says its recipe pairs a human-preference model with rewards for prompt faithfulness, user constraints and avoiding reward-hacking.
TLDR
Arena says its approach combines a preference model trained on about 5.6 million pairwise human votes with faithfulness, constraint and anti-reward-hacking rewards. It reports that post-training raised FLUX.2-dev by 69 Elo points to 1202 and Ideogram 4 by 20 points to 1224 on its live text-to-image leaderboard. Arena said Ideogram 4 surpassed all publicly listed open models as of September 4, 2026. In offline tests judged by Gemini 3.5 Flash, it reports a 64.2% win rate against the base model after adding faithfulness and constraint rewards, rising to 66.0% with a weight-space ensemble.
Combined views
12K
2 Sources, first seen ago
