OCR-related attention heads reportedly verbalize more than just text
A researcher sharing a new preprint says they used the heads to create a simple logit lens for image tokens after finding they were responsible for OCR in vision-language models.
TLDR
A researcher sharing a new preprint says they found attention heads responsible for OCR in vision-language models. The heads could verbalize more than just text, they say, so they used them to create a simple logit lens for image tokens.
Combined views
3K
2 Sources, first seen 6h ago
46 likes
