Open models

Liquid AI triples the speed of its vision model with a small draft model


A 280-million-parameter auxiliary model speeds up decoding of a 3-billion-parameter one by up to 3.13 times, and it runs on a laptop.

September 25, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

8.9% parameter overhead is what the draft costs, and in the best measured case it more than triples decoding speed.

Context

It’s the second publication in three days on the same bottleneck. On September 23, an MBZUAI group presented Flash-dLLM, which speeds up diffusion inference without retraining the model. Both go after the cost per token, not capability.

What’s next

Bottom line

The open-weights competition is no longer measured in benchmark scores. Xiaomi showed that by publishing MiMo with open weights and landing above closed models on two tests. What’s being fought over now is where the model runs, not how well it answers.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes