Positive users are excited about hierarchical upscaling with Gemma LLMs to generate cooler objects and textures like in WOMBO, while negative users dislike the FFT step in the optimization process.
Based on 3 visible X reactions from 11 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@matthen2 Used to do this back in the days. So many tricks to improve quality of img. Not a fan of the FFT step. Just signing the grads gets you far. Adding a small gaussian blur or offsetting by 1 pixel back and forth. Also adding a 3x3 transform on RGB channels. Never found golden rule
@matthen2 hierarchically please:( I need to know if things look even more cool that way... I.e. start at 128p, dream, upscale, dream, upscale...
@matthen2 WOMBO was awesome
Deep dreams on modern LLMs are so cool (optimizing an image to maximize P(target caption)) Gemma 12B (left) has no vision encoder — it reads pixels like token embeddings — and stamps recognisable objects around the canvas. E4B (right) has one, and drifts to texture instead.
these runs are doing gradient descent to optimize the probability of the caption given an input image, represented as an FFT. Higher frequencies are suppressed in the FFT, and the RGB colours are decorrelated and normalized to a target standard deviation, then tanh
of course 12B is a bigger model, so it is not a fully fair comparison. But it's an interesting glimpse into how these models see.
final frames
Positive users are excited about hierarchical upscaling with Gemma LLMs to generate cooler objects and textures like in WOMBO, while negative users dislike the FFT step in the optimization process.
Based on 3 visible X reactions from 11 accounts; directional sample.
Ask a question below.
Published answers will appear here.