Get the latest tech news
Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight
Intel's BITCOS format compresses ternary model weights below 1.58 bits by exploiting zero-heavy distributions, boosting decoding speed up to 27% on GPUs.
None
Or read this on r/technology

