• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Fuse's simulated test of how AI assistants infer hidden motives

    DAIR.AI describes Fuse, a Google Research framework that tests assistants using simulated social interactions. Its summary says biased user framing shifts assistants' answers, while longer conversations with room for clarifying questions did not reliably help.

    DA
    2 Sources, 14d ago, first seen 14d ago

    TLDR

    DAIR.AI's summary of a Google Research paper describes Fuse: a target agent with a hidden motive interacts with other agents, including one playing the user. The user agent recounts what happened to an assistant, which must infer the motive. The simulation supplies a known motive against which to check the answer.

    According to DAIR.AI, the team validated the simulations with 24,000 human annotations and tested 12 language models. Hearing events through the user made the task harder, biased framing shifted answers, and models sometimes needed more detail than humans. Longer conversations with room for clarifying questions did not reliably help. DAIR.AI says the team released the framework and 21,000 examples.

    Combined views

    10.9K

    2 Sources, first seen 14d ago

    Combined views

    10.9K

    2 Sources, first seen 14d ago

    108 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    108 likes
    8 comments
    122 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 comments
    122 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @dair_aiBanger paper from Google Research. This one is on how LLM assistants reason about the people in a user's life. People ask assistants for social advice constantly, and the assistant only hears the user's version of events. Measuring whether it reads the situation correctly is hard, because other people's intentions have no ground truth. Fuse builds that ground truth with simulation. A target agent with a hidden motive interacts with other agents, including one playing the user. The user agent then describes what happened to the assistant, which has to infer the motive. The team validated the simulations with 24k human annotations and tested 12 LLMs. Hearing events through the user makes the task harder. Biased framing from the user shifts the assistant's answer. Models sometimes need more detail than humans do, and longer conversations with room for clarifying questions did not reliably help. They release the framework and 21k examples. Paper: https://academy.dair.ai/papers/verifiable-social-reasoning-for-llm-assistants-2609.17496

    2 Sources

    @dair_aiBanger paper from Google Research. This one is on how LLM assistants reason about the people in a user's life. People ask assistants for social advice constantly, and the assistant only hears the user's version of events. Measuring whether it reads the situation correctly is hard, because other people's intentions have no ground truth. Fuse builds that ground truth with simulation. A target agent with a hidden motive interacts with other agents, including one playing the user. The user agent then describes what happened to the assistant, which has to infer the motive. The team validated the simulations with 24k human annotations and tested 12 LLMs. Hearing events through the user makes the task harder. Biased framing from the user shifts the assistant's answer. Models sometimes need more detail than humans do, and longer conversations with room for clarifying questions did not reliably help. They release the framework and 21k examples. Paper: https://academy.dair.ai/papers/verifiable-social-reasoning-for-llm-assistants-2609.17496