Open models

A general-purpose model matches a specialized classifier at 62 euros per million


The test across 28 datasets leaves a median gap of 0.7 points, with no statistical significance, and a cost almost four times higher.

September 27, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

0.7 percentage points of median difference between the general-purpose model and the specialized classifier.

Context

Lightweight models from Chinese providers have spent months competing on inference price more than on frontier capability. This test moves the discussion a step further: it doesn’t compare two language models with each other, but a language model against the category of tools it supposedly wasn’t meant to touch.

What’s next

Bottom line

The usual argument for keeping your own classifier was accuracy. This measurement reduces it to cost per decision and latency depending on where the server is, which are two engineering problems with known answers.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes