• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Mostik AI Claims ARC-AGI Leaderboard First Place

    Robert Scoble retweets announcement about a team of 12 PhDs building Mostik AI.

    SZ
    NN
    YD
    15 Sources, 28d ago, first seen 28d ago

    TLDR

    Robert Scoble, a tech blogger, author, and commentator focused on AI, robots, spatial computing, and emerging tech who previously worked at Microsoft and Rackspace, retweeted a post by @aimalysheva. The post introduces @mostik_ai and asks what happens when 12 PhDs share one room for four months. It states the result was first place on the ARC-AGI leaderboard. The evidence packet contains only this retweet and offers no independent confirmations or further details from other sources.

    Combined views

    1.4M

    15 Sources, first seen 28d ago

    Combined views

    1.4M

    15 Sources, first seen 28d ago

    5.9K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5.9K likes
    412 comments
    3.1K saves
    631 reposts
    412 comments
    3.1K saves
    631 reposts

    Sentiment

    Positive32.3%67.7%Negative

    Summary

    Sentiment

    Positive32.3%67.7%Negative

    Positive replies praised the 753B-plus-4B latent bridge for efficiency gains and edge deployment, while negative replies dismissed it as research slop or no better than draft models.

    Based on 67 sentiment-bearing replies from 62 accounts across 5 conversations.

    Summary

    Positive replies praised the 753B-plus-4B latent bridge for efficiency gains and edge deployment, while negative replies dismissed it as research slop or no better than draft models.

    Based on 67 sentiment-bearing replies from 62 accounts across 5 conversations.

    15 Sources

    @aimalyshevameet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: http://wired.com/story/russian-startup-mostik-ai-models-communication/ full writeup, the setup, and all the numbers: http://mostik.ai/read-more
    @ScobleizerRT @aimalysheva: meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, w…
    @yurisThis is one of the coolest breakthroughs I've seen with LLMs. A 753B model "thinks" about the answer, and a 4B model "writes" the answer. The communication between the models happens in latent space, resulting in performance nearly as good as the big model, but 20x faster.⬇️
    @denisyarats👀
    @kimmonismusA 753B model reads the problem: A 4B model writes the answer. Mostik reports that its latent bridge lets the 4B model close half the performance gap to the 753B model The bridged pair uses 2.5x less compute than a score-matched mid-sized model A fascinating alternative to the text-based handoffs used by most multi-model systems today. Incredible!
    @DanielleFongRT @yuris: This is one of the coolest breakthroughs I've seen with LLMs. A 753B model "thinks" about the answer, and a 4B model "writes"…
    @yuntiandengWow, this brought back a project Woojeong Kim, @srush_nlp, and I worked on in 2024. A strong model read a task description, turned it into "ability vectors", and passed them to a weak model. Sasha even came up with a great name for it: Strong-to-Weak Ability Transfer (SWAT) 😆 We never published it, but the idea eventually developed into ProgramAsWeights, which compiles task descriptions into small reusable neural programs that run locally. Very cool to see @aimalysheva and the Mostik team make a related idea work so well at this scale. https://programasweights.com
    @suchenzangwhat the helly is this... not spec dec, not disagg, but some lobotomized franken-glueing where the glued models are frozen? why? if you have access to latents, you probably also have access to model weights. freezing weights seem completely unnecessary, not to mention what kind of cursed deployment hell would need this kind of setup...
    @srchvrsRT @aimalysheva: meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, w…
    @m2saxon@suchenzang But they benchmaxxed!

    15 Sources

    @aimalyshevameet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: http://wired.com/story/russian-startup-mostik-ai-models-communication/ full writeup, the setup, and all the numbers: http://mostik.ai/read-more
    @ScobleizerRT @aimalysheva: meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, w…
    @yurisThis is one of the coolest breakthroughs I've seen with LLMs. A 753B model "thinks" about the answer, and a 4B model "writes" the answer. The communication between the models happens in latent space, resulting in performance nearly as good as the big model, but 20x faster.⬇️
    @denisyarats👀
    @kimmonismusA 753B model reads the problem: A 4B model writes the answer. Mostik reports that its latent bridge lets the 4B model close half the performance gap to the 753B model The bridged pair uses 2.5x less compute than a score-matched mid-sized model A fascinating alternative to the text-based handoffs used by most multi-model systems today. Incredible!
    @DanielleFongRT @yuris: This is one of the coolest breakthroughs I've seen with LLMs. A 753B model "thinks" about the answer, and a 4B model "writes"…
    @yuntiandengWow, this brought back a project Woojeong Kim, @srush_nlp, and I worked on in 2024. A strong model read a task description, turned it into "ability vectors", and passed them to a weak model. Sasha even came up with a great name for it: Strong-to-Weak Ability Transfer (SWAT) 😆 We never published it, but the idea eventually developed into ProgramAsWeights, which compiles task descriptions into small reusable neural programs that run locally. Very cool to see @aimalysheva and the Mostik team make a related idea work so well at this scale. https://programasweights.com
    @suchenzangwhat the helly is this... not spec dec, not disagg, but some lobotomized franken-glueing where the glued models are frozen? why? if you have access to latents, you probably also have access to model weights. freezing weights seem completely unnecessary, not to mention what kind of cursed deployment hell would need this kind of setup...
    @srchvrsRT @aimalysheva: meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, w…
    @m2saxon@suchenzang But they benchmaxxed!