• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Model run reportedly had no spikes amid four changes

    A reply says attention-gate weight decay was added over concerns the model wasn't using early layers enough, while PDL was disabled in QuACK grouped GEMMs because of hangs.

    Yoav ArtziYA
    David HallDH
    Manish Maheshwari MM
    4 Sources, 20d ago,

    TLDR

    Asked whether different colors indicated that the run had spiked and resumed, a participant replied: “Nope, no spikes!” They listed four changes: at about 58k, adding attention-gate weight decay because of concerns the model wasn't using early layers enough; at about 82k, adding a faster EP kernel; at about 108k, changing the data mix a bit; and at about 109k, disabling PDL in QuACK grouped GEMMs because of hangs.

    Combined views

    971

    4 Sources, first seen 20d ago

    Combined views

    971

    4 Sources, first seen 20d ago

    16 likes
    first seen 20d ago
    16 likes
    4 comments
    2 saves
    2 reposts
    4 comments
    2 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    David Hall@dlwh@yoavartzi @percyliang Nope, no spikes! Changes: - At ~58k, we added attn gate weight decay b/c we were worried the model wasn't using early layers enough - At ~82k, we added a faster EP kernel - At ~108k, we changed the data mix a bit - At ~109k, we disabled PDL in the QuACK grouped GEMMs b/c hangs20d
    Manish Maheshwari @Manish_m09@dlwh @yoavartzi @percyliang In W&B, if we added the ability to manually annotate a line plot at a specific step, e.g., "58k: added attention gate decay", would that be useful? The annotation would stay attached to that point in the run, so you could later see what changed and why directly on the chart.20d
    Yoav Artzi@yoavartziRT @Manish_m09: @dlwh @yoavartzi @percyliang In W&B, if we added the ability to manually annotate a line plot at a specific step, e.g., "58…20d

    4 Sources

    David Hall@dlwh@yoavartzi @percyliang Nope, no spikes! Changes: - At ~58k, we added attn gate weight decay b/c we were worried the model wasn't using early layers enough - At ~82k, we added a faster EP kernel - At ~108k, we changed the data mix a bit - At ~109k, we disabled PDL in the QuACK grouped GEMMs b/c hangs20d
    Manish Maheshwari @Manish_m09@dlwh @yoavartzi @percyliang In W&B, if we added the ability to manually annotate a line plot at a specific step, e.g., "58k: added attention gate decay", would that be useful? The annotation would stay attached to that point in the run, so you could later see what changed and why directly on the chart.20d
    Yoav Artzi@yoavartziRT @Manish_m09: @dlwh @yoavartzi @percyliang In W&B, if we added the ability to manually annotate a line plot at a specific step, e.g., "58…20d