Hugging Face Blog·· 2026-08-25
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
AI summary
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Selection record
Threshold 60Official, first-handFirst 62Second 62
AdmittedSum of both 124 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:量化感知修复压缩4-bit大模型技术
- Why it was chosen
- Provides a concrete recipe for distilling from the original teacher after compression and quantisation, with nine benchmark comparisons to assess accuracy trade-offs in 4-bit deployment.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Hugging Face Blog · huggingface.co