Publicly shared AI reasoning traces may expose private data and enable hidden prompt injections
A paper quoted in a post says 315,320 reasoning blocks from public repositories yielded 367 PII artifacts and 182 credentials.
TLDR
A user shared a paper about extracting reasoning traces from proprietary LLMs. The quoted passage says decoding 315,320 reasoning blocks scraped from public repositories recovered 367 personally identifiable information artifacts and 182 credentials. It warns that traces can reveal hazardous information even when a model’s visible answer rejects a malicious request, and that attackers can hide prompt-injection payloads in encrypted blocks.
Combined views
738
1 Source, first seen ago