A proposed algorithm to trace citation errors in AI deep research reports
An EMNLP 2026 paper's author says many sentences in deep research reports aren't supported by their citations, reporting recall of 58.7% in NVIDIA AI-Q and 7.1% in TrajectoryKit.
TLDR
An author of an EMNLP 2026 main-conference paper reports that many sentences in deep research reports aren't supported by their citations: recall is 58.7% in NVIDIA AI-Q and 7.1% in TrajectoryKit. The author says the paper proposes an algorithm to trace each error to the agent that introduced it and identify the error type.
A proposed algorithm to trace citation errors in AI deep research reports
An EMNLP 2026 paper's author says many sentences in deep research reports aren't supported by their citations, reporting recall of 58.7% in NVIDIA AI-Q and 7.1% in TrajectoryKit.
TLDR
An author of an EMNLP 2026 main-conference paper reports that many sentences in deep research reports aren't supported by their citations: recall is 58.7% in NVIDIA AI-Q and 7.1% in TrajectoryKit. The author says the paper proposes an algorithm to trace each error to the agent that introduced it and identify the error type.
