Vision transformer shrinks 54.5-fold in size, a post reports
A post says the compressed model matched a full-precision result of 95.13% on a cross-village chilli test, but argues that the benefit of distillation remains unshown.
TLDR
A post reports that a vision transformer shrank from 327.42MB to 6.01MB while matching the full-precision (FP32) result of 95.13% on a cross-village chilli test. A same-size 8-bit (INT8) student model trained directly reached 94.87%, according to the post. The author calls the deployment win clear, but says the gain from distillation—training a student model to learn from another model—remains unshown.
Combined views
13.7K
2 Sources, first seen 24d ago