• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    SPACE Framework Chunks Actions for Long-Horizon LLM Agents

    Rutgers paper proposes skill-guided chunking to reduce repeated LLM decisions on extended tasks.

    EL
    RP
    DA
    4 Sources, 27d ago, first seen 27d ago

    TLDR

    DAIR.AI highlighted a paper by Yanting Yang, Can Jin, Dimitris Metaxas and colleagues at Rutgers. Titled Act More, Decide Less, it introduces SPACE, which lets agents emit variable-length action chunks distilled from skills. The approach targets ReAct-style protocols that issue one primitive action per LLM round. On long-horizon interactive tasks this leads agents to re-decide routine sequences repeatedly. The arXiv preprint describes how the method supports more efficient planning without losing adaptability.

    Combined views

    15.6K

    4 Sources, first seen 27d ago

    Combined views

    15.6K

    4 Sources, first seen 27d ago

    175 likes
    175 likes
    15 comments
    178 saves
    40 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    178 saves
    40 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @dair_aiBrilliant paper on long-horizon agents. They cut 78.9% of an agent's LLM calls while raising its success rate. Here is how: It turns out that ReAct issues one primitive action per model round. That allows frequent replanning, and on long-horizon tasks it spends most of the episode re-deciding routine sequences that were never in doubt. Training an agent to emit variable-length action chunks with standard RL fails because the policy never learns where a chunk should end, so it either falls back to single actions or commits to sequences that run far too long. SPACE derives the supervision from data it already has. It induces two-level programmatic skills from successful trajectories and uses the subskill boundaries as direct chunk-boundary labels, then distills the temporal structure into a primitive-chunk policy with hybrid on-policy and off-policy optimization and chunk-aware credit assignment. On ALFWorld and ScienceWorld it improves success rates by 7.0 to 31.3% over the strongest baseline in each setting while reducing average LLM decision rounds by up to 78.9%. Paper: https://arxiv.org/abs/2609.02042 Chat with Paper: https://academy.dair.ai/papers/act-more-decide-less-skill-guided-adaptive-action-chunking-for-long-horizon-llm-2609.02042
    @omarsar0RT @dair_ai: Brilliant paper on long-horizon agents. They cut 78.9% of an agent's LLM calls while raising its success rate. Here is how:…
    @rohanpaul_aiNew Amazon Microsoft paper shows long-horizon agents should not need an LLM decision after every tiny action; the hard part is knowing which actions can safely run together. SPACE learns those boundaries from successful trajectories. It converts trajectories into programmatic skills, treats subskill boundaries as labels for meaningful chunks, then distills them into a policy that emits variable-length primitive actions with no skill library at test time. On ScienceWorld, success rose from 35.9% to 67.2%, while average LLM rounds fell from 10.2 to 5.2. do not make agents reconsider every tiny step, and do not blindly batch actions either. Train them to learn when to keep acting and when to look again.

    4 Sources

    @dair_aiBrilliant paper on long-horizon agents. They cut 78.9% of an agent's LLM calls while raising its success rate. Here is how: It turns out that ReAct issues one primitive action per model round. That allows frequent replanning, and on long-horizon tasks it spends most of the episode re-deciding routine sequences that were never in doubt. Training an agent to emit variable-length action chunks with standard RL fails because the policy never learns where a chunk should end, so it either falls back to single actions or commits to sequences that run far too long. SPACE derives the supervision from data it already has. It induces two-level programmatic skills from successful trajectories and uses the subskill boundaries as direct chunk-boundary labels, then distills the temporal structure into a primitive-chunk policy with hybrid on-policy and off-policy optimization and chunk-aware credit assignment. On ALFWorld and ScienceWorld it improves success rates by 7.0 to 31.3% over the strongest baseline in each setting while reducing average LLM decision rounds by up to 78.9%. Paper: https://arxiv.org/abs/2609.02042 Chat with Paper: https://academy.dair.ai/papers/act-more-decide-less-skill-guided-adaptive-action-chunking-for-long-horizon-llm-2609.02042
    @omarsar0RT @dair_ai: Brilliant paper on long-horizon agents. They cut 78.9% of an agent's LLM calls while raising its success rate. Here is how:…
    @rohanpaul_aiNew Amazon Microsoft paper shows long-horizon agents should not need an LLM decision after every tiny action; the hard part is knowing which actions can safely run together. SPACE learns those boundaries from successful trajectories. It converts trajectories into programmatic skills, treats subskill boundaries as labels for meaningful chunks, then distills them into a policy that emits variable-length primitive actions with no skill library at test time. On ScienceWorld, success rose from 35.9% to 67.2%, while average LLM rounds fell from 10.2 to 5.2. do not make agents reconsider every tiny step, and do not blindly batch actions either. Train them to learn when to keep acting and when to look again.