• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Summing instead of averaging training loss reportedly cuts accuracy from 58.1% to 51.2%

    A user says the averaging convention was doing some of the tuning: adding the losses for two outputs doubled the update at an unchanged learning rate.

    AK
    1 Source, 14d ago, first seen 14d ago

    TLDR

    A user reports that switching to binary cross-entropy loss on only the current pair’s two outputs, while keeping stochastic gradient descent (SGD), raised final ten-class accuracy from 19.4% to 58.1%. In a follow-up, they say summing rather than averaging that loss cut accuracy to 51.2%. Their explanation: at a fixed SGD learning rate, summing the two outputs’ losses doubles the update, so averaging had been doing some of the tuning.

    Combined views

    111

    1 Source, first seen 14d ago

    Combined views

    111

    1 Source, first seen 14d ago

    4 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 likes
    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @BlackHCA nice gotcha: mean BCE averages over the pair’s two outputs. Sum them instead (loss × 2), and accuracy falls from 58.1% to 51.2%. At fixed SGD learning rate, that doubles the update. The averaging convention was doing some of the tuning for us.

    1 Source

    @BlackHCA nice gotcha: mean BCE averages over the pair’s two outputs. Sum them instead (loss × 2), and accuracy falls from 58.1% to 51.2%. At fixed SGD learning rate, that doubles the update. The averaging convention was doing some of the tuning for us.