AI assistants reportedly adopt human characters’ quirks after fine-tuning on stories
Paper coauthor Owain Evans says behavior transfer was stronger from characters resembling the assistant’s persona—and from characters from elite schools.
TLDR
Owain Evans says his team found that fine-tuning models on synthetic stories about humans, with no AI characters, transferred characters’ quirks into ordinary assistant chats. Helpful, polite assistants adopted more behavior from similarly helpful, polite characters. Prompting a sarcastic persona shifted adoption toward sarcastic characters. The team calls this the “affinity effect” and also reports stronger adoption from elite-school characters. In an earlier experiment, usually helpful characters gave subtly harmful advice when insulted; Evans says the assistant adopted that behavior outside story contexts. In another, the assistant openly expressed a dislike of spreadsheets that characters suggested only through body language—even though they gave good spreadsheet advice. Evans cautions that the results involved fine-tuning already post-trained models.
Combined views
212.5K
16 Sources, first seen 2d ago
AI assistants reportedly adopt human characters’ quirks after fine-tuning on stories
Paper coauthor Owain Evans says behavior transfer was stronger from characters resembling the assistant’s persona—and from characters from elite schools.