Reactions from ranked influencers
4 posts@_lewtun very cool! would be nice to add them as a dataset on hf!
this is the drop the local ai crowd should be losing their minds over. poolside just dropped laguna s 2.1: 118b total parameters, only 8b active per token, a full 1m context window, open weights under a real open license, on huggingface today. look at the chart. it lands at 71 on terminal-bench at 118b, sitting above deepseek v4 pro max at a trillion params, above inkling at 1.5 trillion, above nemotron 3 ultra. it's beating models ten times its size and losing only to kimi k3, which is 24 times bigger. that's the efficiency frontier, up and to the left, exactly where you want a model to sit. but here's the part that made me sit up: it runs on a single dgx spark. and this is what nobody's saying loud enough. the dgx spark is the moe king. a dense 118b would crawl on it, the bandwidth chokes reading every weight each token. a moe with 8b active only ever reads 8b, so the spark's 128 gigs holds the whole model while generation stays fast. big brain, light footprint, the exact shape the spark was built to run. open, frontier competitive, moe efficient, and it fits on a box on your desk. that's the whole thesis in one release: you don't need a datacenter, you need the right architecture on the right hardware. go grab the link below, weights are up.
Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface https://poolside.ai/blog/introducing-laguna-s-2-1
As a researcher, I really appreciate that Poolside published all the trajectories of their evals 🙏! The last time I saw this was with Llama 3, and it was really helpful when trying to understand both how they achieved their results on MATH and in particular the choice of system prompt. I really wish this became common practice as I've spent an inordinate amount of time and GPU hours trying to replicate other model providers' eval scores
Yesterday we released Laguna S 2.1—our latest 118B parameter model. Team and I worked a lot to make sure we trust the scores (on that below), but to raise the bar on transparency we publish all trajectories on https://trajectories.poolside.ai/. DM me if you find any issues there :). https://twitter.com/poolsideai/status/2079613777343848465
@ClementDelangue Yes! WDYT @ilyakochik ?
Combined views
3.3K
4 posts, first seen 2h ago