Secretary @howardlutnick is right. The Kimi Panic needs to stop. — American frontier models are still ahead. When you factor in what’s in the lab, the gap is even larger. As long as we keep releasing, we will stay ahead. Let our horses run. — As Ben Thompson showed, Kimi’s…
British and American government safety institutes jointly published a report comparing Kimi K3 versus top frontier US models. Kimi K3 stopped at step 17 on average, while the strongest American models reached 28.5. Kimi K3 remains substantially behind frontier U.S. models in…
"Kimi K3 performs significantly below the most recent frontier cyber-capable models" On UK AISI's cyber range "The Last Ones" Kimi-K3 reaches on average step 17 out of 32. This seems to be around the level of Opus 4.6, a 6 month old model.…
I don't think these datapoints support the linear fit bros here's my eyeball fit of course, Blue Team is not linear either please don't be so sloppy, otherwise people may be spooked https://twitter.com/CommerceGov/status/2080341953086886387
This puts K3 on the level of Mythos Preview on its best attempt and between Opus 4.6 and Mythos Preview on average https://twitter.com/aisecurityinst/status/2080343068285280753
The issue, of course, is that Mythos Preview was trained with cyberoffense in mind Kimi is a strong model. If China wants to make a cyber-strong model, they'll do it https://twitter.com/xeophon/status/2080350110400049393
@teortaxesTex @stalkermustang @scaling01 @zephyr_z9 @IbrahimDagher20 if you do slop, go all out also idk whether i like elo for this (esp if based on avg)