Report
Samsung's LittleBit reportedly shrinks a 13B-parameter language model below 1 GB
A user says Samsung open-sourced the method, which uses latent factorization to compress model weights.
TLDR
LittleBit can shrink a 13-billion-parameter language model to under 1 GB, according to a user who says Samsung open-sourced the method. The post says some configurations reach 0.1 bits per weight and claims an 11.6x inference speedup relative to standard FP16 models.
Combined views
3.2K
2 Sources, first seen ago
