Training environments may matter more than active parameter count in AI cyber-capability evaluations
A user argues that a lab could train even a model with 10B active parameters to score well if its environments cover the domain.
TLDR
A user argues that active parameter count matters little when evaluating cyber capabilities and related RLVR tasks. In their view, a lab with environments that reasonably cover the domain could train even a model with 10B active parameters to score well; without them, 104B would not help. They speculate that scaling network size and inference budget, as they suggest Anthropic may be doing with Mythos-Preview, could be another way to address narrow generalization.
Combined views
1.3K
1 Source, first seen ago