• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Microsoft Paper Tests Reusing Rules from Past Agent Runs

    Tweet by Rohan Paul summarizes a Microsoft paper on cheaper agent reasoning.

    RP
    1 Source, 24d ago, first seen 24d ago

    TLDR

    Rohan Paul shared details from a new Microsoft paper on AI agent reasoning. The paper suggests collecting 35 to 50 past trajectories to learn a compact set of rules. These rules could substitute for expensive test-time reasoning in subsequent tasks. The post poses the question of paying the reasoning cost once and reusing the model's learned insights across multiple future tasks. It describes testing this cheaper alternative for agent runs. The tweet includes a static screenshot of the research paper as its attachment.

    Combined views

    5.6K

    1 Source, first seen 24d ago

    Combined views

    5.6K

    1 Source, first seen 24d ago

    37 likes
    37 likes
    9 comments
    34 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    34 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rohanpaul_aiWhat if you could pay the reasoning cost once, then reuse what the model learned across future tasks? New Microsoft paper finds that some expensive test-time reasoning can be replaced with a small set of rules learned from previous agent runs. The paper tests a cheaper alternative: collect 35–50 past trajectories, have a coding agent extract recurring failure patterns, then turn those patterns into a small markdown skill added to the non-reasoning model’s system prompt. For GPT-5.4-mini, those skills recovered 55%–100%+ of the gap between non-reasoning and reasoning modes across 4 agent benchmarks, while using 2.9–4.5× fewer output tokens than reasoning. On ALFWorld and τ²-retail, the skilled non-reasoning model actually beat the reasoning mode. The useful part is that the distiller did not need reasoning traces: skills built only from cheap non-reasoning rollouts were competitive across all 4 domains. The limit is equally useful: reasoning still won on telecom and SpreadsheetBench, where each task contains more instance-specific dependencies that a fixed skill cannot capture. So the practical split is: distill repeated procedures once, then reserve expensive test-time reasoning for the tasks that genuinely need fresh search.

    1 Source

    @rohanpaul_aiWhat if you could pay the reasoning cost once, then reuse what the model learned across future tasks? New Microsoft paper finds that some expensive test-time reasoning can be replaced with a small set of rules learned from previous agent runs. The paper tests a cheaper alternative: collect 35–50 past trajectories, have a coding agent extract recurring failure patterns, then turn those patterns into a small markdown skill added to the non-reasoning model’s system prompt. For GPT-5.4-mini, those skills recovered 55%–100%+ of the gap between non-reasoning and reasoning modes across 4 agent benchmarks, while using 2.9–4.5× fewer output tokens than reasoning. On ALFWorld and τ²-retail, the skilled non-reasoning model actually beat the reasoning mode. The useful part is that the distiller did not need reasoning traces: skills built only from cheap non-reasoning rollouts were competitive across all 4 domains. The limit is equally useful: reasoning still won on telecom and SpreadsheetBench, where each task contains more instance-specific dependencies that a fixed skill cannot capture. So the practical split is: distill repeated procedures once, then reserve expensive test-time reasoning for the tasks that genuinely need fresh search.