Code4Scene benchmarks agents building and editing 3D scenes in Unreal Engine
The benchmark’s creators say they tested 14 agent configurations on 95 public cases. Spatial composition was every agent’s weakest construction category.
TLDR
Code4Scene’s creators introduced a benchmark in which coding agents build Unreal Engine scenes from text or repair scenes using reference images. They say they reopened submitted scenes to assess task fulfillment, static physical validity and whether unrelated content stayed unchanged. Across 14 agent configurations and 95 public cases, spatial composition was every agent’s weakest construction category. They also report that 35.8% of edits that fully recovered the target introduced unintended changes elsewhere.
Combined views
1.7K
5 Sources, first seen 11h ago
