Skip to content
Hugging Face Blog·· 2026-08-21

Up to 3.2x Faster Inference with LFM2.5-DSpark

Up to 3.2x Faster Inference with LFM2.5-DSpark

AI summary

Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.

Selection record

AdmittedSum of both 124 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:发布AI模型推理加速技术
Why it was chosen
Measured throughput and function-call latency for three LFM2.5 draft models show the scope for accelerating on-device inference.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co