Liquid AI Releases LFM2.5-DSpark Draft Models for Up to 3.2x Faster Inference

Liquid AI has released LFM2.5-DSpark, a set of speculative decoding draft models that accelerate inference on their LFM2.5 series by up to 3.18x without altering model outputs. The approach uses a lightweight draft model to predict multiple tokens ahead, which the main model then verifies in parallel, dramatically reducing wall-clock latency at no cost to output quality. This is a meaningful advance for developers deploying LFM2.5 in latency-sensitive applications such as real-time chat, coding assistants, or streaming inference pipelines. The models are available on Hugging Face, making them immediately accessible for integration into existing inference stacks. Developers already using LFM2.5 can adopt DSpark draft models with minimal changes to their serving infrastructure to achieve substantial throughput gains.
Read original source ↗Part of the 2026-08-21 briefing→