• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Code4Scene benchmarks agents building and editing 3D scenes in Unreal Engine

    The benchmark’s creators say they tested 14 agent configurations on 95 public cases. Spatial composition was every agent’s weakest construction category.

    LQ
    ZH
    XY
    5 Sources, ,

    TLDR

    Code4Scene’s creators introduced a benchmark in which coding agents build Unreal Engine scenes from text or repair scenes using reference images. They say they reopened submitted scenes to assess task fulfillment, static physical validity and whether unrelated content stayed unchanged. Across 14 agent configurations and 95 public cases, spatial composition was every agent’s weakest construction category. They also report that 35.8% of edits that fully recovered the target introduced unintended changes elsewhere.

    Combined views

    1.7K

    5 Sources, first seen 11h ago

    Combined views

    1.7K

    5 Sources, first seen 11h ago

    34 likes
    11h ago
    first seen 11h ago
    34 likes
    1 comments
    4 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    1 comments
    4 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @koe_ye40329Can coding agents reliably build a 3D scene, then change exactly what we ask? Introducing Code4Scene, our benchmark for constructing and editing 3D scenes in Unreal Engine through code. Agents build scenes from text or repair existing scenes from reference images. We reopen their submitted scenes to evaluate task fulfillment, static physical validity, and whether unrelated scene content stays unchanged. We evaluate 14 agent configurations on 95 public cases. Spatial composition is the weakest construction category for every agent. And 35.8% of edits that fully recover the target still introduce unintended changes elsewhere. Here’s what we test, how we evaluate it, and what today’s coding agents still struggle with 🧵
    @LianhuiqRT @koe_ye40329: Can coding agents reliably build a 3D scene, then change exactly what we ask? Introducing Code4Scene, our benchmark for c…
    @simworld_ai🌍 AI can build 3D worlds. But can it build them correctly and edit without breaking something else? Introducing SimWorld-Code4Scene: 190 Unreal Engine cases testing spatial reasoning through text-to-scene construction and precise control through image-guided editing. We reopen and evaluate the actual generated 3D scenes, not just screenshots, for instruction following, physical validity, and scene integrity. Astra leads overall, but important gaps remain: spatial composition is the weakest construction category across all 14 tested configurations. And 35.8% of edits that fully recover their targets still introduce unintended changes elsewhere. Paper and code: ⤵️
    @ZhitingHuRT @simworld_ai: 🌍 AI can build 3D worlds. But can it build them correctly and edit without breaking something else? Introducing SimWorld-…

    5 Sources

    @koe_ye40329Can coding agents reliably build a 3D scene, then change exactly what we ask? Introducing Code4Scene, our benchmark for constructing and editing 3D scenes in Unreal Engine through code. Agents build scenes from text or repair existing scenes from reference images. We reopen their submitted scenes to evaluate task fulfillment, static physical validity, and whether unrelated scene content stays unchanged. We evaluate 14 agent configurations on 95 public cases. Spatial composition is the weakest construction category for every agent. And 35.8% of edits that fully recover the target still introduce unintended changes elsewhere. Here’s what we test, how we evaluate it, and what today’s coding agents still struggle with 🧵
    @LianhuiqRT @koe_ye40329: Can coding agents reliably build a 3D scene, then change exactly what we ask? Introducing Code4Scene, our benchmark for c…
    @simworld_ai🌍 AI can build 3D worlds. But can it build them correctly and edit without breaking something else? Introducing SimWorld-Code4Scene: 190 Unreal Engine cases testing spatial reasoning through text-to-scene construction and precise control through image-guided editing. We reopen and evaluate the actual generated 3D scenes, not just screenshots, for instruction following, physical validity, and scene integrity. Astra leads overall, but important gaps remain: spatial composition is the weakest construction category across all 14 tested configurations. And 35.8% of edits that fully recover their targets still introduce unintended changes elsewhere. Paper and code: ⤵️
    @ZhitingHuRT @simworld_ai: 🌍 AI can build 3D worlds. But can it build them correctly and edit without breaking something else? Introducing SimWorld-…