Sub-1-Bit LLM Compression: A New Approach via Latent Factorization
The LittleBit Project enhances language model efficiency through sub-1-bit compression.
The LittleBit Project compresses large language models into the sub-1-bit range by factorizing dense weight matrices into low-rank latent factors. This enables extreme compression down to 0.1 bits per weight while preserving the original model architecture. LittleBit-2 improves upon this by addressing latent geometry misalignment during initialization, enhancing performance.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work