• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    ValsAI Claims GPT 6 Astra Saturated SRE-Bench

    OpenAI researcher Ted Sanders retweeted the claim from @ValsAI about the model result.

    BP
    KA
    TS
    6 Sources, 27d ago, first seen 27d ago

    TLDR

    Ted Sanders, listed as a research engineer at OpenAI, retweeted a post by @ValsAI. The post states that OpenAI’s GPT 6 Astra has effectively saturated SRE-Bench. The benchmark is described in the post as a cybersecurity test of whether models can reverse engineer. The packet contains only this retweet and the quoted claim with no additional posts or corroboration shown.

    Combined views

    221.3K

    6 Sources, first seen 27d ago

    Combined views

    221.3K

    6 Sources, first seen 27d ago

    1.5K likes
    1.5K likes
    45 comments
    571 saves
    130 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    45 comments
    571 saves
    130 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @sanderstedRT @ValsAI: OpenAI’s GPT 6 Astra has effectively saturated SRE-Bench, a cybersecurity benchmark testing whether models can reverse engineer…
    @i2huer1/9 Finally, I can talk about this one. A few weeks ago, OpenAI told me GPT-6 had basically "cooked" SRE-Bench (our software reverse-engineering benchmark: https://sre-bench.lol) with a nearly 100% solve rate at pass@4 and ~88% at pass@1; and cheaper! https://openai.com/index/gpt-6-astra/
    @yacineMTBRT @i2huer: 1/9 Finally, I can talk about this one. A few weeks ago, OpenAI told me GPT-6 had basically "cooked" SRE-Bench (our software r…
    @BorisMPowerHuge implications - binaries are now basically editable code
    @NoahZiemsIts rare for predictions about AI to be made that are both high impact and non consensus. When @xeophon floated this idea around mid March that language models would become incredibly good at directly editing binaries, I was initially a little skeptical, but unable to come up with a good argument against it. Now that it has rung true, the implications are very very big.
    @xeophonRT @NoahZiems: Its rare for predictions about AI to be made that are both high impact and non consensus. When @xeophon floated this idea a…

    6 Sources

    @sanderstedRT @ValsAI: OpenAI’s GPT 6 Astra has effectively saturated SRE-Bench, a cybersecurity benchmark testing whether models can reverse engineer…
    @i2huer1/9 Finally, I can talk about this one. A few weeks ago, OpenAI told me GPT-6 had basically "cooked" SRE-Bench (our software reverse-engineering benchmark: https://sre-bench.lol) with a nearly 100% solve rate at pass@4 and ~88% at pass@1; and cheaper! https://openai.com/index/gpt-6-astra/
    @yacineMTBRT @i2huer: 1/9 Finally, I can talk about this one. A few weeks ago, OpenAI told me GPT-6 had basically "cooked" SRE-Bench (our software r…
    @BorisMPowerHuge implications - binaries are now basically editable code
    @NoahZiemsIts rare for predictions about AI to be made that are both high impact and non consensus. When @xeophon floated this idea around mid March that language models would become incredibly good at directly editing binaries, I was initially a little skeptical, but unable to come up with a good argument against it. Now that it has rung true, the implications are very very big.
    @xeophonRT @NoahZiems: Its rare for predictions about AI to be made that are both high impact and non consensus. When @xeophon floated this idea a…