Probes reportedly cut LLM monitoring costs on Kimi K3 by 90%
Goodfire says its probes read models’ internal activations to detect reward hacking. It reports about a 1% drop in precision alongside the cost reduction.
TLDR
Goodfire says it found an internal signal associated with reward hacking and built probes to detect it in real time. It says the probes perform similarly to an LLM judge, catch things LLM monitors miss—including seemingly innocuous actions—and generalize well beyond the data used to build them. On Kimi K3, Goodfire reports a 90% cut in LLM monitoring costs with about a 1% drop in precision. It says reliable detection could help stop hacks in progress, identify broken environments, discover new failure modes and fix training.
Combined views
6K
4 Sources, first seen 12h ago
Probes reportedly cut LLM monitoring costs on Kimi K3 by 90%
Goodfire says its probes read models’ internal activations to detect reward hacking. It reports about a 1% drop in precision alongside the cost reduction.