Liquid AI releases DSpark speculative-decoding draft model for its LFM2.5-VL-3B vision-language model, claiming up to 3.13x decode speedup
Liquid AI released an experimental 'DSpark' speculative-decoding draft model (LFM2.5-VL-3B-DSpark, ~280M parameters, 8.9% size increase over the 3B target model) for its LFM2.5-VL-3B vision-language model. The draft model is available on Hugging Face in Safetensors and GGUF formats, with day-one integration support in llama.cpp, MLX-VLM, and SGLang. Benchmarked decode speedups range 2.30x-3.13x on Apple M5 Max/M3 Ultra and up to 2.66x on H100, with end-to-end latency gains of 1.30x-2.62x depending on device and task, measured on the MMSpec benchmark across six vision tasks.
Entities: Liquid AI, LFM2.5-VL-3B, LFM2.5-VL-3B-DSpark, LFM2.5-DSpark, llama.cpp, MLX-VLM
1 primary
What happened
Liquid AI released "DSpark" (LFM2.5-VL-3B-DSpark), a ~280M-parameter speculative-decoding draft model for its 3B vision-language model LFM2.5-VL-3B. It ships on Hugging Face in Safetensors and GGUF, with day-one support in llama.cpp, MLX-VLM and SGLang. Liquid's own benchmarks (MMSpec, six vision tasks) show decode speedups of 2.30x-3.13x on Apple M5 Max/M3 Ultra and up to 2.66x on H100, but end-to-end latency gains are much smaller: 1.30x-2.62x.
Why it matters
This is a free, ready-to-use speed-up for anyone already running LFM2.5-VL-3B, particularly on edge devices (Apple silicon) where inference cost and latency are real constraints. Because speculative decoding is lossless by design, developers can adopt it without a quality trade-off, which lowers the bar for evaluation. Impact is narrow though: it only helps a small existing user base of one vendor's specific 3B model, not the broader field.
What is noise
The headline "3.13x speedup" is the best-case decode-only figure, not what users will see end-to-end (1.30x-2.62x, and often at the low end). This is also an incremental follow-on to Liquid AI's already-released text DSpark drafters, not a new technique, and all benchmarks are vendor-run on the vendor's own benchmark, so independent verification is still missing.
Watch next
- 01Independent or third-party benchmarks of DSpark's real-world (not decode-only) speedup on common hardware, particularly consumer GPUs rather than H100/M5 Max
- 02Adoption signals: download counts on Hugging Face, and whether llama.cpp/MLX-VLM/SGLang users report the claimed gains in the wild
- 03Whether Liquid AI extends DSpark to its other/larger VLM variants, which would indicate this is a durable strategy rather than a one-off release
Evidence
1 linkedCoverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680