• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    LLM token-loss gaps reportedly stay roughly constant across optimizers but not changes to training data

    A researcher sharing a paper calls this pattern relative generalization invariance: two models’ token-wise losses differ by a constant gap. They say it largely holds across optimizers and moderate architecture changes, but changing the training data stream breaks it.

    RA
    1 Source, 6h ago, first seen 6h ago

    TLDR

    A researcher sharing a paper says two models satisfy relative generalization invariance when their token-wise losses differ by a constant gap. They report that this pattern largely holds across optimizers and moderate architecture changes, producing an approximately uniform shift in token-wise loss. Changing the training data stream breaks the pattern, they say.

    Combined views

    54

    1 Source, first seen 6h ago

    Combined views

    54

    1 Source, first seen 6h ago

    14 reposts
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @_arohan_RT @FengzhuoZhang: How do optimizers, architectures, and data shape LLM generalization differently? 🚀 We study this through Relative Gener…

    1 Source

    @_arohan_RT @FengzhuoZhang: How do optimizers, architectures, and data shape LLM generalization differently? 🚀 We study this through Relative Gener…