Digit-learning model reportedly scores 19.4% across ten digits, 97.5% within pairs
A user’s Split MNIST experiment suggests poor accuracy across all ten digits can hide a model’s retained ability to distinguish digits when given the correct pair of choices.
TLDR
A user describes training a small neural network on handwritten-digit pairs in sequence: 0/1, 2/3, 4/5, 6/7 and 8/9. Each stage used only that pair’s real training images, followed by a final test on all ten digits.
With cross-entropy over all ten outputs and SGD, the user reports 19.4% final accuracy—but 97.5% when answers were restricted to the correct digit pair. With Adam, the corresponding results were 19.6% and 71.3%. Results averaged three seeds, with the setup starting at two epochs per task.
The user’s takeaway: poor ten-class accuracy can hide retained within-pair discrimination.
Combined views
177
1 Source, first seen 14d ago