LLM answer accuracy versus information-gathering quality
The post argues that large language models frequently underestimate how much information they need, and calls for testing when they stop rather than relying on answer accuracy alone.
TLDR
A post linking to “Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking” argues that a correct answer can mask weak evidence gathering. A model may use prior knowledge or guess correctly without collecting enough evidence, it says—so evaluations should test when models stop seeking information, not just whether their answers are right.
Combined views
5.8K
2 Sources, first seen 18d ago