Reaction
GLM-5.3 Flash is claimed capable of exceeding Mythos Preview’s ExploitBench scores
The user argues that scaling test-time compute ultimately destroys categorical capability tiers between AI models.
TLDR
A user claims GLM-5.3 Flash can exceed Mythos Preview’s scores on ExploitBench. They argue that scaling test-time compute ultimately destroys categorical capability tiers and say models are now “AGI enough” to keep making progress.
Combined views
5K
1 Source, first seen ago