Reactions from ranked influencers
3 postsProud of my just-graduated PhD student @Smearle_RH who just won a well-deserved best paper award at @GeccoConf for his work on investigating the components of open-endedness through having VLMs play PicBreeder!
Proud to have received a Best Paper Award for our AI Picbreeder work at GECCO 2026 @GeccoConf in the Complex Systems track. Blog post here: https://pub.sakana.ai/picbreeder-vlm/. (Along with an Outstanding Reviewer award in the Evolutionary Machine Learning track!🧎🏻♂️🙏🏻) https://twitter.com/smearle_rh/status/2075655209229943095
“In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models” won the Best Paper Award in the Complex Systems track at #GECCO2026. Congratulations to the co-authors, and thanks to our collaborators!! Blog: http://pub.sakana.ai/picbreeder-vlm 🦎
Proud to have received a Best Paper Award for our AI Picbreeder work at GECCO 2026 @GeccoConf in the Complex Systems track. Blog post here: https://pub.sakana.ai/picbreeder-vlm/. (Along with an Outstanding Reviewer award in the Evolutionary Machine Learning track!🧎🏻♂️🙏🏻)
Our new work, The AI Picbreeder Experiment, explores the use of frontier models as drivers of synthetic open-endedness. If we're serious about putting these things in the driver's seat of a new and automatic science, then we need to know what they're really made of in terms of the ability to create and discover through intuition. Giving shape to the formless, making decisions based on vibes, having "taste"—whatever you want to call it—Can they do it? Do they have the sauce? Picbreeder, a website where human users collaborated to spontaneously evolve images, is just the sauce-bearing test we need. Here, images were represented as neural networks that could be bred and mutated, with humans playing the role of natural selectors. By design, this interface prohibits the creative baggage of premeditation, of having goals in advance, and demands the artist patiently follow the flow of the work and seize upon serendipitous opportunities when they arise. It's more like catching fish from a stream than drawing a picture. And yet, distributing their work across many sessions, and branching and remixing each other's creations, humans were ultimately able to bend these neural networks into all manner of interesting, evocative, and striking images. So, can large vision language models do the same? On the blog, we've built an interactive archive viewer that allows visitors to walk through galleries of Picbreeder images created by both humans and AI, and judge for themselves. Call us old fashioned, but we're pretty sure the human output has something special that the AI can't quite yet replicate. We design a number of evaluation metrics to get at this quality. We ask: "How visually different are the images in the archive? How much do they look like real things? How different are the things they look like?" The numbers show the humans coming out on top. And looking at the AI-generated archives and lineages, we find traces of an anxious attachment to plans and objectives. Often, even when the AI makes an apparent creative leap—e.g. transforming an image of a hood ornament into a side view of a car—it really stays stuck in place in some broader semantic/thematic space. And that's to say nothing of the handful of archives littered almost entirely with top-down views of soda can pull tabs, or high frequency circular patterns that appear chaotic and uninteresting to us, but apparently scratch some perceptual itch in the agents. And yet we're optimistic. Though the AI's output is less refined, its movement through the stream of images less graceful and vivacious than our own, what we have here is a plausible model organism of open-endedness. The agents indeed (re)discover distributions of novel and interesting images when left to their own devices. They display a keen eye (even sometimes discovering optical illusions that might slip by a casual glance from a human), and explore persistently under considerable creative constraints. This allows us to model factors that are consequential to such open-ended exploration; i.e. injecting noise into the agents' decision making process, playing with their memory, and seeding them with subtly distinct personalities—all of which can be beneficial in the right doses. And there's something to be said for searching without objectives. Prior work shows that if we optimize Picbreeder's pattern-producing neural networks to resemble a particular image (say, a skull), these representations will be fractured (meddling with their internal weights will immediately explode the skull beyond recognition), while the same neural image found by humans via open-ended exploration is robust to such perturbations, and even shows meaningful variations across them (e.g. the jaw opening and closing). Our VLM agents also stumbled upon images of skulls. Their representations are not as neatly semantically factorized as those discovered by humans, but neither are they nearly as fractured as those discovered by optimization. This suggests that if we want to have AI build the next generation of AI, then it will be crucial to let them attack this problem through aimless wandering. Without this freedom, future models will be brittle and myopic; with it, they will have developed a more thorough model of the world, and an improved capacity for the kind of creative insight that is so quietly fundamental to the most meaningful of human endeavors.
Combined views
10.3K
3 posts, first seen 1d ago