Summing instead of averaging training loss reportedly cuts accuracy from 58.1% to 51.2%
A user says the averaging convention was doing some of the tuning: adding the losses for two outputs doubled the update at an unchanged learning rate.
TLDR
A user reports that switching to binary cross-entropy loss on only the current pair’s two outputs, while keeping stochastic gradient descent (SGD), raised final ten-class accuracy from 19.4% to 58.1%. In a follow-up, they say summing rather than averaging that loss cut accuracy to 51.2%. Their explanation: at a fixed SGD learning rate, summing the two outputs’ losses doubles the update, so averaging had been doing some of the tuning.
Combined views
111
1 Source, first seen 14d ago