• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Personal agents may miss more rules when memory gets too long

    A post reports that Claude Haiku 4.5’s rule violations fell from 77% without memory to 20% with 10 lines, then rose with longer memories.

    RP
    1 Source, 42m ago, first seen 42m ago

    TLDR

    A post argues that personal agents should keep memory short and use code for things they must track. It reports Claude Haiku 4.5’s rule violations fell from 77% without memory to 20% with 10 lines, then rose to about 25% with longer memories. In a spending-total test, a written rule failed 44% of the time, while code tracking the total failed 0%. The best methods for rewriting memory from complaints ended near 48% violations, versus 7.1% when the agent was told every preference.

    Combined views

    1.6K

    1 Source, first seen 42m ago

    Combined views

    1.6K

    1 Source, first seen 42m ago

    10 likes
    10 likes
    10 comments
    9 saves
    1 reposts
    10 comments
    9 saves
    1 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rohanpaul_aiMore memory does not keep helping personal agents, because relevant notes help up to a point and then extra lines start making the agent miss rules. A personal agent can't learn every user preference through memory notes, and more notes eventually make it worse, so use code for anything it must count or track and keep memory short. Written rules work for style, like signing texts with the user's first name. For a running spending total, a stated rule still failed 44% of the time, while code that kept the total failed 0%. Memory size has a sweet spot. With Claude Haiku 4.5, violations dropped from 77% with no memory to 20% at 10 lines, then rose to about 25% with longer memories. Agents that rewrite their own memory from user complaints improve early, then stall. The best methods ended near 48% violations, against 7.1% when the agent was simply told every preference.42m