• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Muse Spark 1.3 Trails on Cybersecurity CVE Benchmark

    AI security researcher shares benchmark scores for Muse Spark 1.3 against other models.

    T(
    P(
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    AI Pentest Lead pilvar222 posted results from running Muse Spark 1.3 on a cybersecurity benchmark. At pass@1 it rediscovered an average of 19 out of 32 CVEs, while Grok 4.6 scored 23.3 out of 32. Pooling three runs raised Muse to 24 out of 32, with DeepSeek V4 Pro reaching 28 out of 32. The post noted pricing remains competitive. An attached animated chart from aikido/research plots cost versus recall across one to three runs for several models. The post was retweeted by Teortaxes, a pseudonymous commentator focused on DeepSeek.

    Combined views

    17.6K

    2 Sources, first seen 27d ago

    Combined views

    17.6K

    2 Sources, first seen 27d ago

    115 likes
    115 likes
    4 comments
    20 saves
    12 reposts
    4 comments
    20 saves
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @pilvar222We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32 - When pooling the results of 3 runs (pass@3), Muse scores 24/32. DeepSeek V4 Pro gets 28/32 - The pricing is good compared to many frontier models, but compared to some of the Chinese ones, its use is hard to justify - We had a cached input token rate of 45%, much lower than we usually get. The model is new, so we decided to include the results as if we had had 90% Overall, I would say that @AIatMeta is slowly catching up. The model seems good at coding, but it's definitely not there yet for cybersecurity. 🧵 1/3
    @teortaxesTexRT @pilvar222: We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers a…

    2 Sources

    @pilvar222We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32 - When pooling the results of 3 runs (pass@3), Muse scores 24/32. DeepSeek V4 Pro gets 28/32 - The pricing is good compared to many frontier models, but compared to some of the Chinese ones, its use is hard to justify - We had a cached input token rate of 45%, much lower than we usually get. The model is new, so we decided to include the results as if we had had 90% Overall, I would say that @AIatMeta is slowly catching up. The model seems good at coding, but it's definitely not there yet for cybersecurity. 🧵 1/3
    @teortaxesTexRT @pilvar222: We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers a…