How tokenizer choices could affect tools for interpreting language models
A post considers how logit lens and J-lens might work with character-level or multiword tokens, arguing that even simple interpretability tools depend on data and low-level design choices.
TLDR
A user wonders whether tools such as logit lens and J-lens would be harder to use if models split text into much smaller units, down to single characters, or larger, multiword units. They also note that when other modalities enter language models, vocabulary embeddings and individual tokens in a sequence are rarely meaningful.
Combined views
1.4K
2 Sources, first seen 10h ago