What happened
- On September 24, Liquid AI published LFM2.5-VL-3B-DSpark on the Hugging Face blog, a version of its 3-billion-parameter vision-language model accelerated with speculative decoding.
- The technique uses an auxiliary model (a draft) that proposes several tokens at once and lets the large model just verify them. The draft has 280 million parameters, an 8.9% overhead on the main model.
- The figures published by the company: on a laptop with an M5 Max chip, decoding speeds up between 2.30 and 3.13 times, with an overall gain of up to 2.62 times. On an M3 Ultra with llama.cpp, between 1.57 and 2.14 times. On an H100 GPU, 2.66 times for decoding and between 1.64 and 2.27 times end to end.
- The weights are open on Hugging Face, in Safetensors and GGUF formats, with day-one support in llama.cpp, MLX-VLM and SGLang.
Why it matters
- The use case this enables in the region isn’t the lab: it’s the field. A model that reads images at that speed on a laptop can be used to review product listings, read receipts or classify inventory photos without uploading anything to an external server.
- That goes straight to the regulatory problem. Under Law 21.719, data that never leaves the device avoids a good part of the international transfer obligations. The speedup isn’t a performance improvement: it’s a condition for local processing to be viable.
- The consequence the announcement doesn’t mention: the gain is measured on expensive hardware. An M5 Max or an H100 isn’t the average installed base of a Chilean small business, and the number that matters (how much it speeds up on a mid-range machine) isn’t published.
The number
8.9% parameter overhead is what the draft costs, and in the best measured case it more than triples decoding speed.
Context
It’s the second publication in three days on the same bottleneck. On September 23, an MBZUAI group presented Flash-dLLM, which speeds up diffusion inference without retraining the model. Both go after the cost per token, not capability.
What’s next
- No timelines announced for versions in other sizes or for measurements on mid-range hardware.
- The weights are already available; there’s no committed date for a license different from the current one.
Bottom line
The open-weights competition is no longer measured in benchmark scores. Xiaomi showed that by publishing MiMo with open weights and landing above closed models on two tests. What’s being fought over now is where the model runs, not how well it answers.
Sources
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


