Kimi K3 trails leading US models on ExploitBench cyber benchmark
A joint US-UK study places Kimi nine months behind.
Entities: Kimi K3, AI Security Institute (AISI), U.S. Department of Commerce, David Sacks, Howard Lutnick
A joint US-UK study places Kimi nine months behind.
Entities: Kimi K3, AI Security Institute (AISI), U.S. Department of Commerce, David Sacks, Howard Lutnick
Positive users endorse US models leading Kimi K3 in cyber tests and defend related reports, while negative users insult the benchmarks as biased or call the claims bullshit.
Based on 132 visible X reactions from 271 accounts.
Ask a question below.
Published answers will appear here.
Together with the US Center for AI Standards and Innovation (@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
Really good job Kimi team! Calling Kimi K3 weak at cyber is the wrong takeaway. It is most likely a calculated competence tax. Moonshot AI knows they are releasing open weights, so they probably scrubbed actionable cyber and CBRN data from the training pipeline and baked in strict safety guardrails for such cases. They are acting like adults making sure the open model is powerful for general use but not a weaponized anarchy engine. I would love to read an article from their team on how they actually pulled off making a frontier model safe for an open-weight release. P.S. Looks like Qwen-3.8-Max is following the exact same blueprint.
@DavidSacks @howardlutnick History suggests closed products can lead temporarily, but open ecosystems often win in the long run. Linux, Kubernetes, PyTorch, and LLVM became dominant because everyone could improve them. AI may follow a similar path.
@AravSrinivas Cope! BTW, how’s perplexity innovating today? Is it still around?
@DavidSacks @howardlutnick I could not agree more!!
@AISecurityInst @NIST Thanks for the update.
@AISecurityInst @NIST Graph slop. Delete your account
Secretary @howardlutnick is right. The Kimi Panic needs to stop. — American frontier models are still ahead. When you factor in what’s in the lab, the gap is even larger. As long as we keep releasing, we will stay ahead. Let our horses run. — As Ben Thompson showed, Kimi’s apparent cost advantage largely disappears once you account for higher token usage and the real cost of running a model this size. Open weights still require expensive infrastructure. — Anthropic and OpenAI are growing revenue at rates that Silicon Valley has never seen before at this scale. This remains the clearest test of who is winning the market. President Trump’s light-touch regulatory approach is working. We should remain confident in American innovation. As long as we don’t sabotage ourselves with unnecessary rules, the U.S. will continue to win.
Positive users endorse US models leading Kimi K3 in cyber tests and defend related reports, while negative users insult the benchmarks as biased or call the claims bullshit.
Based on 132 visible X reactions from 271 accounts.
Ask a question below.
Published answers will appear here.
"Kimi K3 performs significantly below the most recent frontier cyber-capable models" On UK AISI's cyber range "The Last Ones" Kimi-K3 reaches on average step 17 out of 32. This seems to be around the level of Opus 4.6, a 6 month old model.
I don't think these datapoints support the linear fit bros here's my eyeball fit of course, Blue Team is not linear either please don't be so sloppy, otherwise people may be spooked
more cheeky fit and near term models (I half believe this will be proven true, if my understanding of the dynamics in compute/RL stack maturity/data engine is correct) @stalkermustang @scaling01 @zephyr_z9 @IbrahimDagher20
The issue, of course, is that Mythos Preview was trained with cyberoffense in mind Kimi is a strong model. If China wants to make a cyber-strong model, they'll do it
CAISI is pronounced Casey.