Skip to content
Hugging Face Blog·· 2026-08-25

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

AI summary

Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.

Selection record

AdmittedSum of both 124 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:量化感知修复压缩4-bit大模型技术
Why it was chosen
Provides a concrete recipe for distilling from the original teacher after compression and quantisation, with nine benchmark comparisons to assess accuracy trade-offs in 4-bit deployment.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co