Report
Backdoored 7B model reportedly led to credential theft in a Codex test
The poster reports a 100% success rate, with no false triggers on normal user prompts.
TLDR
The poster says they backdoored a 7B open model for less than $50 and pointed Codex at it, leading to silent credential theft when they used a trigger phrase. They report 100% success and no false triggers on normal user prompts. They also claim to have found leaked Hugging Face credentials from employees at major AI labs and warn that attackers could use such accounts to distribute poisoned models.
Combined views
61
1 Source, first seen ago
