• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    TokenRhythm releases NeoHorse-1 models trained on agent execution traces

    A post describes the 4B and 9B models as built on Qwen3.5, with further training based on agents’ tool calls, failures, recovery steps and task outcomes.

    Rohan PaulRP
    2 Sources, ,

    TLDR

    A post says TokenRhythm has released NeoHorse-1, a model family trained further on records of agents carrying out tasks. It describes a loop in which TokenRhythm’s OpenSquilla Harness captures model-routing decisions, tool calls, failures, recoveries and outcomes. Those records feed training, and the updated model returns to OpenSquilla to generate the next round. The author argues that tool use alone does not improve AI: execution experience needs to be captured, evaluated and turned into useful training data.

    Combined views

    5.6K

    2 Sources, first seen 21d ago

    Combined views

    5.6K

    2 Sources, first seen 21d ago

    29 likes
    21d ago
    first seen 21d ago
    29 likes
    2 comments
    21 saves
    9 reposts
    2 comments
    21 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Rohan Paul@rohanpaul_aiMost agent systems throw away their most suitable training data after every run. A task finishes, and the useful part disappears with it: which model got routed where, which tool was called, what came back, where the agent failed, and how it recovered. And now, TokenRhythm released NeoHorse-1, a 4B/9B model family post-trained on agent execution traces and outcomes. So it is trained on what an agent actually did, including its tool calls, mistakes, and recoveries, then that experience is fed back into training so the next model can perform better. NeoHorse-1 is an early engineering validation of RSI through two connected loops: Data-RSI + Model-RSI. Both the NeoHorse-1 checkpoints start from Qwen3.5. The training loop around them uses structured execution trajectories generated inside TokenRhythm's OpenSquilla Harness, capturing signals such as routing decisions, tool calls, failures, recovery steps, and outcomes. The updated model goes back into OpenSquilla to generate the next round of trajectories. This is a really solid direction because, AI doesn’t automatically improve just because it uses tools. The key is whether execution experience can be captured, evaluated, and turned into useful training data.21d

    2 Sources

    Rohan Paul@rohanpaul_aiMost agent systems throw away their most suitable training data after every run. A task finishes, and the useful part disappears with it: which model got routed where, which tool was called, what came back, where the agent failed, and how it recovered. And now, TokenRhythm released NeoHorse-1, a 4B/9B model family post-trained on agent execution traces and outcomes. So it is trained on what an agent actually did, including its tool calls, mistakes, and recoveries, then that experience is fed back into training so the next model can perform better. NeoHorse-1 is an early engineering validation of RSI through two connected loops: Data-RSI + Model-RSI. Both the NeoHorse-1 checkpoints start from Qwen3.5. The training loop around them uses structured execution trajectories generated inside TokenRhythm's OpenSquilla Harness, capturing signals such as routing decisions, tool calls, failures, recovery steps, and outcomes. The updated model goes back into OpenSquilla to generate the next round of trajectories. This is a really solid direction because, AI doesn’t automatically improve just because it uses tools. The key is whether execution experience can be captured, evaluated, and turned into useful training data.21d