Muse Spark 1.2 Scores 60.3 Percent on WeirdML Benchmark
Researcher reports model performance comparable to GPT-5 variants on expanded test.
Håvard Ihle posted that Muse Spark 1.2 (xhigh) scored 60.3 percent on WeirdML. The post says this result is comparable to GPT-5 or GPT-5.4 Mini while remaining far from the frontier. WeirdML v2 now covers 19 tasks and tracks API costs across models. Teortaxes shared the update and noted they had not benchmarked the model yet. Yacine retweeted the claim from Ihle without additional comment.
Muse Spark 1.2 (xhigh) scores 60.3% on WeirdML, comparable to GPT-5 or GPT 5.4 Mini. It's pretty good, but far from the frontier.
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two…
