Signum
Feed
Useful signal24 Sept 2026high confidence

Liquid AI releases DSpark speculative-decoding draft model for its LFM2.5-VL-3B vision-language model, claiming up to 3.13x decode speedup

Liquid AI released an experimental 'DSpark' speculative-decoding draft model (LFM2.5-VL-3B-DSpark, ~280M parameters, 8.9% size increase over the 3B target model) for its LFM2.5-VL-3B vision-language model. The draft model is available on Hugging Face in Safetensors and GGUF formats, with day-one integration support in llama.cpp, MLX-VLM, and SGLang. Benchmarked decode speedups range 2.30x-3.13x on Apple M5 Max/M3 Ultra and up to 2.66x on H100, with end-to-end latency gains of 1.30x-2.62x depending on device and task, measured on the MMSpec benchmark across six vision tasks.

CapabilityInfrastructureAccess

Entities: Liquid AI, LFM2.5-VL-3B, LFM2.5-VL-3B-DSpark, LFM2.5-DSpark, llama.cpp, MLX-VLM

69Useful signal
1 source
1 primary
Was this useful?
01

What happened

Liquid AI released "DSpark" (LFM2.5-VL-3B-DSpark), a ~280M-parameter speculative-decoding draft model for its 3B vision-language model LFM2.5-VL-3B. It ships on Hugging Face in Safetensors and GGUF, with day-one support in llama.cpp, MLX-VLM and SGLang. Liquid's own benchmarks (MMSpec, six vision tasks) show decode speedups of 2.30x-3.13x on Apple M5 Max/M3 Ultra and up to 2.66x on H100, but end-to-end latency gains are much smaller: 1.30x-2.62x.

02

Why it matters

This is a free, ready-to-use speed-up for anyone already running LFM2.5-VL-3B, particularly on edge devices (Apple silicon) where inference cost and latency are real constraints. Because speculative decoding is lossless by design, developers can adopt it without a quality trade-off, which lowers the bar for evaluation. Impact is narrow though: it only helps a small existing user base of one vendor's specific 3B model, not the broader field.

03

What is noise

The headline "3.13x speedup" is the best-case decode-only figure, not what users will see end-to-end (1.30x-2.62x, and often at the low end). This is also an incremental follow-on to Liquid AI's already-released text DSpark drafters, not a new technique, and all benchmarks are vendor-run on the vendor's own benchmark, so independent verification is still missing.

04

Watch next

  1. 01Independent or third-party benchmarks of DSpark's real-world (not decode-only) speedup on common hardware, particularly consumer GPUs rather than H100/M5 Max
  2. 02Adoption signals: download counts on Hugging Face, and whether llama.cpp/MLX-VLM/SGLang users report the claimed gains in the wild
  3. 03Whether Liquid AI extends DSpark to its other/larger VLM variants, which would indicate this is a durable strategy rather than a one-off release

Evidence

1 linked

Coverage

1 story

More capability signals

Full feed →