Reaction
A claimed 95% ‘user model’ score calls the User Sim Index into question
A user who shared code for the model argues that the index probably shouldn’t be used to evaluate user models.
TLDR
A user says people often evaluate user models on the User Sim Index but probably shouldn’t. They shared code for a ‘user model’ they say scored 95%, with scores of 97% for communication, 97% for information, 92% for clarification and 95% for error reaction. They also shared a sampled conversation and believe their model beats any released model.
Combined views
7.7K
7 Sources, first seen ago
