Intel reportedly compresses a 1.58-bit LLM to 1.485 bits without changing its weights
The New Stack reports GPU decoding speed gains of up to 27% for Intel's BITCOS format, which compresses ternary model weights by exploiting zero-heavy distributions.
TLDR
The New Stack reports that Intel's BITCOS format compresses a 1.58-bit language model to 1.485 bits without changing any weights. The format takes advantage of weight distributions containing many zero values. The outlet also reports decoding speed gains of up to 27% on GPUs.
Intel reportedly compresses a 1.58-bit LLM to 1.485 bits without changing its weights
The New Stack reports GPU decoding speed gains of up to 27% for Intel's BITCOS format, which compresses ternary model weights by exploiting zero-heavy distributions.
TLDR
The New Stack reports that Intel's BITCOS format compresses a 1.58-bit language model to 1.485 bits without changing any weights. The format takes advantage of weight distributions containing many zero values. The outlet also reports decoding speed gains of up to 27% on GPUs.