Fuse's simulated test of how AI assistants infer hidden motives
DAIR.AI describes Fuse, a Google Research framework that tests assistants using simulated social interactions. Its summary says biased user framing shifts assistants' answers, while longer conversations with room for clarifying questions did not reliably help.
TLDR
DAIR.AI's summary of a Google Research paper describes Fuse: a target agent with a hidden motive interacts with other agents, including one playing the user. The user agent recounts what happened to an assistant, which must infer the motive. The simulation supplies a known motive against which to check the answer.
According to DAIR.AI, the team validated the simulations with 24,000 human annotations and tested 12 language models. Hearing events through the user made the task harder, biased framing shifted answers, and models sometimes needed more detail than humans. Longer conversations with room for clarifying questions did not reliably help. DAIR.AI says the team released the framework and 21,000 examples.
Combined views
10.9K
2 Sources, first seen 14d ago