• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Jeff Huber Shares Anthropic Mythos 5.1 Benchmark Claim

    Tweet from investor Jeff Huber highlights Anthropic model results on protein benchmarks.

    JH
    2 Sources, 29d ago, first seen 29d ago

    TLDR

    Jeff Huber posted that Anthropic shipped Mythos 5.1, which he said leads the ProteinGym benchmark at 49.3% rank correlation against real lab measurements. He also noted a prior publication from two weeks earlier in which Claude achieved a 27% hit rate on autonomous protein binder design campaigns, compared with the 10-15% rate he described as typical. Huber called both results impressive while suggesting they point to further implications. The post includes an attached infographic from Anthropic. No independent confirmation of the claims appears in the packet.

    Combined views

    7.5K

    2 Sources, first seen 29d ago

    Combined views

    7.5K

    2 Sources, first seen 29d ago

    59 likes
    59 likes
    10 comments
    33 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    33 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @jhuberAnthropic shipped Mythos 5.1 today, leading the ProteinGym benchmark at 49.3% rank correlation against real lab measurements. Two weeks ago they published Claude running autonomous protein binder design campaigns at a 27% hit rate versus the 10-15% that's typical today. Both results are impressive, but both point to a problem most people aren't (yet) seeing. Start with the shape of that ProteinGym chart: Mythos 5.1 - 49.3% Opus 5 - 47.7% Mythos 5 - 45.8% Gemini 3.1 Pro - 37.0% Sonnet 5 - 36.6% GPT-5.6 - 35.5% Six frontier general-purpose models inside a 14-point band. None of them dedicated protein models. All trained on roughly the same public corpus. That's not anyone achieving a moat; it's a a floor rising. What ProteinGym actually measures: ~217 deep mutational scanning assays covering millions of variants, where every one of them is an experiment somebody already ran and published. Predicting results that exist is a different problem from generating results that don't. The binder campaign has the same shape, and Anthropic says so in its own technical report. Co-folding confidence turned out to be a useful filter but not a guarantee, and "experimental screening remains the only way to learn" which targets a campaign actually succeeded on. Adaptyv, who ran the wet lab, called it an open-loop experiment and said the next step is closing the loop. You can't close that loop at 30 designs per target with a multi-week CRO turnaround. You get one shot and no statistical power to learn anything from it. In contrast – last March, @ManifoldBio and @NVIDIA tested 1,000,000 designs against 127 targets in a single multiplexed experiment, measuring over 100 million protein-protein interactions. Roughly 750x the design count of the Anthropic campaign, in one run. (disclosure: I'm on Manifold's board, and an investor.) The model they were validating in that study was Proteina-Complexa, which is one of the ten open-source design models Claude used in the Anthropic campaign. Design is converging on a shared toolkit, and ... commoditizing. Measurement is not. Which brings me to the ladder below. Every public benchmark, and both of Anthropic's recent headline results, live near the bottom of it. Expression. Binding. In vitro affinity. There is no ProteinGym for biodistribution and for what actually translates into human biology (and into therapeutics that work). Not because nobody wants one, but because the data doesn't exist at benchmark scale. Someone has to generate it, at billion-scale. Whoever does has the data that matters to train the models that will actually make a difference. My bet is that the first to million-scale will be the first to billion-scale.

    2 Sources

    @jhuberAnthropic shipped Mythos 5.1 today, leading the ProteinGym benchmark at 49.3% rank correlation against real lab measurements. Two weeks ago they published Claude running autonomous protein binder design campaigns at a 27% hit rate versus the 10-15% that's typical today. Both results are impressive, but both point to a problem most people aren't (yet) seeing. Start with the shape of that ProteinGym chart: Mythos 5.1 - 49.3% Opus 5 - 47.7% Mythos 5 - 45.8% Gemini 3.1 Pro - 37.0% Sonnet 5 - 36.6% GPT-5.6 - 35.5% Six frontier general-purpose models inside a 14-point band. None of them dedicated protein models. All trained on roughly the same public corpus. That's not anyone achieving a moat; it's a a floor rising. What ProteinGym actually measures: ~217 deep mutational scanning assays covering millions of variants, where every one of them is an experiment somebody already ran and published. Predicting results that exist is a different problem from generating results that don't. The binder campaign has the same shape, and Anthropic says so in its own technical report. Co-folding confidence turned out to be a useful filter but not a guarantee, and "experimental screening remains the only way to learn" which targets a campaign actually succeeded on. Adaptyv, who ran the wet lab, called it an open-loop experiment and said the next step is closing the loop. You can't close that loop at 30 designs per target with a multi-week CRO turnaround. You get one shot and no statistical power to learn anything from it. In contrast – last March, @ManifoldBio and @NVIDIA tested 1,000,000 designs against 127 targets in a single multiplexed experiment, measuring over 100 million protein-protein interactions. Roughly 750x the design count of the Anthropic campaign, in one run. (disclosure: I'm on Manifold's board, and an investor.) The model they were validating in that study was Proteina-Complexa, which is one of the ten open-source design models Claude used in the Anthropic campaign. Design is converging on a shared toolkit, and ... commoditizing. Measurement is not. Which brings me to the ladder below. Every public benchmark, and both of Anthropic's recent headline results, live near the bottom of it. Expression. Binding. In vitro affinity. There is no ProteinGym for biodistribution and for what actually translates into human biology (and into therapeutics that work). Not because nobody wants one, but because the data doesn't exist at benchmark scale. Someone has to generate it, at billion-scale. Whoever does has the data that matters to train the models that will actually make a difference. My bet is that the first to million-scale will be the first to billion-scale.