Announcement
At 524k width, BeFOND is claimed to match Gemma Scope’s single-feature probing accuracy with over 1,000 times fewer samples
A co-lead describes BeFOND as an encoder-free iterative sparse coding model for interpreting language models.
TLDR
A BeFOND co-lead says the model is designed to help interpret language models. At a width of 524k, they claim it matches Gemma Scope’s single-feature probing accuracy with over 1,000 times fewer samples—roughly 6 million versus 8 billion SAE-training tokens. They also say BeFOND keeps improving as SAE width increases while alternatives plateau, and describe its inference and learning rules as closed-form.
Combined views
1.8K
3 Sources, first seen ago
