Three proposed dimensions of language models’ behavioral self-knowledge
A post highlights a paper proposing a 10-task benchmark and says it shows that models can improve their self-knowledge with training.
TLDR
A post discussing a 10-task benchmark proposes organizing language models’ behavioral self-knowledge along three dimensions: whether it concerns the current scenario or a hypothetical changed one; whether the model predicts how its behavior would differ or only whether it would differ; and whether the prediction covers one input or a distribution of inputs. The author says the paper shows that models can improve with training.
Three proposed dimensions of language models’ behavioral self-knowledge
A post highlights a paper proposing a 10-task benchmark and says it shows that models can improve their self-knowledge with training.