UK AISI: Open-weight hacking gap shrinks to 4–7 months
GLM-5.2 matched the performance of Anthropic's Claude Opus 4.5.


Combined views
102.1K
21 Sources, first seen ago
Sources
LA Lisan al Gaib@scaling01
Mythos Preview outperforms both GPT-5.6-Sol and Mythos 5 in the first ~30ish million tokens Mythos Preview's best attempt is also slightly faster than GPT-5.6-Sol's best attempt https://twitter.com/scaling01/status/2078105835364893159
- likes: 147
- replies: 7
- bookmarks: 25
- reposts: 5
SH Shakeel@ShakeelHashim
Quite brave to publish this today! https://twitter.com/aisecurityinst/status/2078103148665667648
- likes: 30
- replies: 2
- bookmarks: 4
- reposts: 1
SA Samuel Albanie 🇬🇧@SamuelAlbanie
- likes: 4
- replies: 0
- bookmarks: 0
- reposts: 0
IH Ian Hogarth@soundboy
Important analysis from @AISecurityInst on the open/closed weight gap in frontier cyber capabilities. https://twitter.com/AISecurityInst/status/2078103148665667648
- likes: 12
- replies: 0
- bookmarks: 4
- reposts: 3
RN Ramez Naam@ramez
1. OpenAI's 5.6 Sol beats Anthropic's still not fully available Mythos in this hacking evaluation. 2. The best open model, Kimi K3, will probably be quite similar to the US proprietary leaders. Everyone now has frontier hacking capabilities.…
- likes: 24
- replies: 2
- bookmarks: 3
- reposts: 3
SH Shakeel@ShakeelHashim
- likes: 22
- replies: 1
- bookmarks: 2
- reposts: 2
Combined views
102.1K
21 Sources, first seen ago