Poolside releases Laguna S 2.1 agentic coding model
Poolside says its new open-weight coding model can punch above its size class and still run on a single NVIDIA DGX Spark.
Entities: Laguna S 2.1, Poolside, Eiso Kant
In a launch post on X, Poolside co-founder Eiso Kant said the company is releasing Laguna S 2.1, an open-weight agentic coding model with 118B total parameters, 8B active per token, and support for running on a single NVIDIA DGX Spark. A separate Poolside post on X calls it the company's most capable model to date, while the company's models page lists Laguna S 2.1 at 118B parameters, 8B active and 1M context.
Combined views
707.4K
78 posts, first seen 21h ago
Reactions from ranked influencers
78 postsYou should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm
Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class. On Terminal-Bench 2.1 it scores 70.2, sitting beside models 5–25x its size and ahead of several of them. And on DeepSWE from @datacurve, the hardest long-horizon benchmark we ran, Laguna S 2.1 scores 40.4, outperforming some open models with more than 1T parameters. For every score we publish today, we're releasing the full trajectory of every trial in the final evaluation set at http://trajectories.poolside.ai
Big day for American open-source AI. For the launch of Laguna S, I sat down with @eisokant to discuss its architecture, the economics of open weights, and the question of who gets to build intelligence. Timestamps: 0:00 Intro 1:50 Why Poolside started opening its models: the oligopoly on intelligence 4:28 Getting nerd-sniped by Karpathy, building LLMs before anyone cared 11:20 Laguna S: 118B parameters, 8B active, built in 8 weeks 13:05 The future of software engineering: behaviors over IQ 14:25 Sliding window attention, 1M context, the model factory 15:50 Being an American open-source lab 20:47 The economics of open weights 24:22 Who gets to build intelligence? The 12–18 month window 30:28 Erdős 397 in 30 minutes 31:48 Extracting transcripts with a debugger Congrats to the Poolside team on the launch!
Turns out, the next Laguna also fits on a @NVIDIAAI DGX Spark 👀
Still blows my mind that this model fits on a single DGX Spark 🤯
🎉 Day-0 support for Laguna S 2.1 from @poolsideai is now live on SGLang! 118B total params, 1M context, with thinking & no-thinking modes. ✅ Long-horizon persistence: keeps planning, testing, and self-correcting up to ~24h runs with little intervention ✅ Built for agentic coding: 70.2% Terminal-Bench 2.1, 78.5% SWE-bench Multilingual ✅ Sparse MoE efficiency: only 8B active params per token, making long agent runs and RL loops fast and cheap ✅ Runs locally: official NVFP4 quantization fits on a single NVIDIA DGX Spark Cookbook: https://docs.sglang.io/cookbook/autoregressive/Poolside/Laguna-S-2.1 Run it now with SGLang!
@jasoncwarner 52 days is obscene. I've often said anything over 70 on terminal bench 2.1 is useful for most every coding task out there. 40 on deepswe is a huge bonus did you train it for context compaction?
OP
Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, with context up to 1M tokens https://poolside.ai/blog/introducing-laguna-s-2-1 Laguna S 2.1 sits at the top of its weight class and competes with open models many times larger, while remaining small enough to run on a single NVIDIA DGX Spark It is available today under OpenMDW-1.1 via Hugging Face, OpenRouter, the Poolside API, and pool This is a remarkable model by any measure. As much as we can gather, it's the best open weight model in the West, regardless of size But the model itself is only part of the story There is a prevailing narrative that building capable models requires ever more capital, compute, and people. We believe the more important question is how efficiently you can turn those resources into intelligence At Poolside, we approach model building as an industrialized process spanning data, pre-training, reinforcement learning, evaluation, and inference. We call that system the Model Factory. We have written about our approach to model building extensively in a 6 part blog series: https://poolside.ai/blog/introducing-the-model-factory Laguna S 2.1 went from the Model Factory kicking it off to release in 52 days. The model is the output. The ability to keep building better models, faster and more efficiently each time is the actual innovation. The Model Factory is Poolside's compounding asset And so here we are, 52 days after kicking this off, releasing a 118B/8B MOE that tops 70% on TB 2.1, 78% on SWE-Bench Multi, 59% on SWE-Bench Pro, and 40% on DeepSWE. And it's fully open, from an American company I grew up in the Linux and Python communities, then spent much of my career at Canonical (Ubuntu), Heroku and GitHub. Those communities shaped my belief that important technology becomes more useful when people can understand it, challenge it, and build upon it. And even more important than being useful is being trusted. That is what open weights mean to me. They are not a marketing or distribution exercise. They give people control: the ability to inspect the work, reproduce the claims, modify the model, and run it inside their own environment I believe the West needs a credible open path to frontier intelligence. I believe it should be from an American company. We intend to be that company Laguna S 2.1 is another step in that direction and yet more evidence of Poolside's long held, often times contrarian beliefs and views on how intelligence will be built. One last note. We are doing something very different that we hope becomes industry norm going forward. We recognize that releasing a 118B/8B MOE that performs as well as Laguna S does would be met with some degree of skepticism. Models of this size are not supposed to outperform models 4-25x larger. We double and triple checked our benchmark trajectories to be sure. But we wanted to go another step and release those trajectories for you to see and help us quadruple check them. https://trajectories.poolside.ai/ If you find something we missed, we genuinely want to hear about Have fun building whatever thing you can think of with the most persistent little model that could
Combined views
707.4K
78 posts, first seen 21h ago