• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Training environments may matter more than active parameter count in AI cyber-capability evaluations

    A user argues that a lab could train even a model with 10B active parameters to score well if its environments cover the domain.

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    1 Source, 1h ago, first seen 1h ago

    TLDR

    A user argues that active parameter count matters little when evaluating cyber capabilities and related RLVR tasks. In their view, a lab with environments that reasonably cover the domain could train even a model with 10B active parameters to score well; without them, 104B would not help. They speculate that scaling network size and inference budget, as they suggest Anthropic may be doing with Mythos-Preview, could be another way to address narrow generalization.

    Combined views

    1.3K

    1 Source, first seen 1h ago

    Combined views

    1.3K

    1 Source, first seen 1h ago

    15 likes
    15 likes
    2 comments
    3 saves
    2 comments
    3 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexHot take: I think active param basically doesn't matter for evaluating cyber capabilities and related RLVR domains. If a lab has environments reasonably covering the domain, they can RL even a 10B active to score well. If no, 104B won't save them. Narrow generalization. Though I suspect Anthropic said "if your generalization is too narrow, you're not scaling the network and inference budget far enough!" with Mythos-Preview. It is an obvious solution when you just don't have environments yet and can speed ahead with compute. It saves a few months.1h

    1 Source

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexHot take: I think active param basically doesn't matter for evaluating cyber capabilities and related RLVR domains. If a lab has environments reasonably covering the domain, they can RL even a 10B active to score well. If no, 104B won't save them. Narrow generalization. Though I suspect Anthropic said "if your generalization is too narrow, you're not scaling the network and inference budget far enough!" with Mythos-Preview. It is an obvious solution when you just don't have environments yet and can speed ahead with compute. It saves a few months.1h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet