Hugging Face Blog·· 2026-08-21
Up to 3.2x Faster Inference with LFM2.5-DSpark
Up to 3.2x Faster Inference with LFM2.5-DSpark
AI summary
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Selection record
Threshold 60Official, first-handFirst 62Second 62
AdmittedSum of both 124 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:发布AI模型推理加速技术
- Why it was chosen
- Measured throughput and function-call latency for three LFM2.5 draft models show the scope for accelerating on-device inference.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Hugging Face Blog · huggingface.co