UK US Safety Institutes Find Kimi K3 Trails Frontier Models
Joint safety institutes assess Chinese model against American leaders on cyber benchmarks.
Combined views
1.1M
31 Sources, first seen ago
Sources
LA Lisan al Gaib@scaling01
"Kimi K3 performs significantly below the most recent frontier cyber-capable models" On UK AISI's cyber range "The Last Ones" Kimi-K3 reaches on average step 17 out of 32. This seems to be around the level of Opus 4.6, a 6 month old model.…
- likes: 350
- replies: 33
- bookmarks: 81
- reposts: 19
DS David Sacks@DavidSacks
Secretary @howardlutnick is right. The Kimi Panic needs to stop. — American frontier models are still ahead. When you factor in what’s in the lab, the gap is even larger. As long as we keep releasing, we will stay ahead. Let our horses run. — As Ben Thompson showed, Kimi’s…
- likes: 4.8K
- replies: 300
- bookmarks: 618
- reposts: 406
XL xlr8harder@xlr8harder
Okay but now do it with the classifiers blocking you from using the actual frontier models. This is an irrelevant point if no one can use them. https://twitter.com/CommerceGov/status/2080341953086886387
- likes: 59
- replies: 2
- bookmarks: 5
- reposts: 3
FB Florian Brand@xeophon
Overall, this seems to put K3 in cyber in the 3-9mo range. However, aggregate and avg scores should be interpreted with caution, as for safety evals, the worst outcome is what you should look at/prepare for. https://twitter.com/xeophon/status/2080350110400049393
- likes: 25
- replies: 1
- bookmarks: 4
- reposts: 0
MC Matt Clifford@matthewclifford
Great work from AISI. Worth a read. https://twitter.com/aisecurityinst/status/2080343066389479706
- likes: 21
- replies: 3
- bookmarks: 8
- reposts: 2
ND Nick Dobos@NickADobos
US government currently posting Ai evals shitting on kimi as cyberwar military propaganda wild time to be alive https://twitter.com/commercegov/status/2080341953086886387
- likes: 27
- replies: 1
- bookmarks: 5
- reposts: 0
Combined views
1.1M
31 Sources, first seen ago