Users praise the visualization of Gemma 4 predicting tokens from raw image patches as beautiful and mind-blowing, while a few call the format unsettling or dismiss it as hallucination.
Based on 16 visible X reactions from 35 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@emollick The clarity of the data flow here really makes the underlying logic pop—great job.
@matthen2 Not your intention but something about this particular format is quite unsettling.
@matthen2 oh this is sick
@matthen2 this is so cool
@matthen2 This is very cool
@matthen2 raw hallucination
what is a multimodal LLM thinking as it watches a video? Gemma 4 12B reads raw image patches, as if they were tokens. It was never trained to predict anything at these 'tokens' - but this video shows what it would predict if you did sample from its next token prediction head
at each patch I am showing a word cloud of the top 3 token distribution. But this is after dividing out the global distribution of tokens over the whole image. This helps to show what each individual square predicts that the rest of the image doesn't
this is cleaner to do with Gemma 4 12B's encoder-free architecture. Most VLMs would encode through a dedicated vision network, so there wouldn't be such a clean mapping from patch -> token
and the distribution is smoothed through the video. word clouds are morphed from frame to frame, with a little physics sim, and resetting at scene cuts
in this frame you can see it is reading the text "National Galleries of Scotland"
This is a wonderful visualization.
Users praise the visualization of Gemma 4 predicting tokens from raw image patches as beautiful and mind-blowing, while a few call the format unsettling or dismiss it as hallucination.
Based on 16 visible X reactions from 35 accounts; directional sample.
Ask a question below.
Published answers will appear here.