Intel's BITCOS reportedly compresses a 1.58-bit LLM to 1.485 bits
The New Stack reports that the format exploits zero-heavy weight distributions, boosting decoding speed by up to 27% on GPUs.
TLDR
The New Stack reports that Intel's BITCOS compresses ternary model weights below 1.58 bits by taking advantage of zero-heavy distributions. The outlet describes a reduction to 1.485 bits without changing a single weight and reports decoding speed gains of up to 27% on GPUs.
Combined views
413
1 Source, first seen 7h ago
Intel's BITCOS reportedly compresses a 1.58-bit LLM to 1.485 bits
The New Stack reports that the format exploits zero-heavy weight distributions, boosting decoding speed by up to 27% on GPUs.
TLDR
The New Stack reports that Intel's BITCOS compresses ternary model weights below 1.58 bits by taking advantage of zero-heavy distributions. The outlet describes a reduction to 1.485 bits without changing a single weight and reports decoding speed gains of up to 27% on GPUs.