Local AI models can leak answers through CPU cache, a post warns
The post says an attacker must already be on the same machine, share relevant CPU resources and profile the same long-lived model process to reconstruct its responses.
TLDR
A post describing a paper says converting generated tokens into readable text leaves repeatable CPU-cache patterns. Another local process can learn those patterns and reconstruct later answers without reading the model’s memory, according to the account. The post cites full-response attack success of about 56%–93% on text tasks, reaching 95.87% in one code setting and 30.12% in an end-to-end OpenClaw attack. It stresses that attackers need access to the same machine, shared CPU resources and profiling of the same long-lived model process. For sensitive deployments, the post says the paper points to stronger CPU isolation, shorter-lived processes and disabling simultaneous multithreading (SMT) where the security tradeoff justifies it.
Combined views
1 Source, first seen 19d ago