Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
• 17
None defined yet.
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss