Muse Spark 1.3 Trails on Cybersecurity CVE Benchmark
AI security researcher shares benchmark scores for Muse Spark 1.3 against other models.
TLDR
AI Pentest Lead pilvar222 posted results from running Muse Spark 1.3 on a cybersecurity benchmark. At pass@1 it rediscovered an average of 19 out of 32 CVEs, while Grok 4.6 scored 23.3 out of 32. Pooling three runs raised Muse to 24 out of 32, with DeepSeek V4 Pro reaching 28 out of 32. The post noted pricing remains competitive. An attached animated chart from aikido/research plots cost versus recall across one to three runs for several models. The post was retweeted by Teortaxes, a pseudonymous commentator focused on DeepSeek.
Combined views
17.6K
2 Sources, first seen 27d ago