• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    METR Evaluates Mythos 5.1 With Conceptual Reasoning Benchmark

    Daniel Filan shares screenshots from a METR preliminary assessment report on X.

    AC
    LA
    CP
    5 Sources, 29d ago, first seen 29d ago

    TLDR

    Daniel Filan, Technical Staff at METR, posted on X that the lab used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1. The tweet includes three screenshots of text excerpts from the preliminary assessment report. Filan states the report supplies more METR text on Mythos 5.1 than he has seen in prior model cards. The post links to the attached images of the report excerpts.

    Combined views

    25.1K

    5 Sources, first seen 29d ago

    Combined views

    25.1K

    5 Sources, first seen 29d ago

    346 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    346 likes
    9 comments
    40 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    40 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 Sources

    @scaling01Mythos 5.1 still has shortcomings compared to human researchers many of which are related to "behavioral or alginment-adjacent issues" "the main issues we observe are around epistemic quality and instruction following"
    @dfrsrchtwtsApparently METR used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1. We also get more text from METR on Mythos 5.1 than I've seen in past model cards.
    @AndrewCurran_METR: 'we believe that Mythos 5.1 is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks. At the same time, we believe that Mythos 5.1 is still likely to noticeably accelerate researchers and automate limited aspects of R&D. For instance, we expect that Mythos 5.1 likely provides a higher productivity uplift than Mythos Preview.'
    @ChrisPainterYupRT @dfrsrchtwts: Apparently METR used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1. We also get more text from MET…

    5 Sources

    @scaling01Mythos 5.1 still has shortcomings compared to human researchers many of which are related to "behavioral or alginment-adjacent issues" "the main issues we observe are around epistemic quality and instruction following"
    @dfrsrchtwtsApparently METR used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1. We also get more text from METR on Mythos 5.1 than I've seen in past model cards.
    @AndrewCurran_METR: 'we believe that Mythos 5.1 is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks. At the same time, we believe that Mythos 5.1 is still likely to noticeably accelerate researchers and automate limited aspects of R&D. For instance, we expect that Mythos 5.1 likely provides a higher productivity uplift than Mythos Preview.'
    @ChrisPainterYupRT @dfrsrchtwts: Apparently METR used Redwood/Acorn's conceptual reasoning benchmark to evaluate Mythos 5.1. We also get more text from MET…