What happened
- On September 24, Privatemode published a test in which it turns GLM-5.3-Flash, a general-purpose language model, into a system of typed decisions with a probability for each option, in a single pass.
- The method requires no fine-tuning of the model. It numbers the options, prefixes the response format and reads the probability distribution the model assigns to each number.
- Across 28 text datasets, the result ties with Jev, a specialized classifier: each wins on 10 datasets, and the median gap of 0.7 percentage points in Jev’s favor doesn’t reach statistical significance (p = 0.64).
- Costs don’t tie: 62 euros per million decisions versus 16 euros for the specialized one. Latency depends on where the request comes from: 180 ms versus 264 ms from Germany, 299 ms versus 164 ms from the United States.
- The general-purpose model adds something the specialized one doesn’t do: it processes images, with 70.2% accuracy classifying scanned documents into 16 categories across 1,600 samples.
Why it matters
- For teams that currently maintain their own classifier, the question stops being about accuracy and becomes about operations: how much the trained model plus the retraining cycle costs, versus four times the inference cost and zero model maintenance.
- The latency difference by geography is the figure the post doesn’t underline, and it decides the case in the region. Measured from the United States, the advantage flips; measured from South America, nobody has measured it, and that’s where local traffic runs.
- Image processing changes the scope: classifying scanned invoices or forms with the same component that classifies text removes an entire integration from the diagram.
The number
0.7 percentage points of median difference between the general-purpose model and the specialized classifier.
Context
Lightweight models from Chinese providers have spent months competing on inference price more than on frontier capability. This test moves the discussion a step further: it doesn’t compare two language models with each other, but a language model against the category of tools it supposedly wasn’t meant to touch.
What’s next
- Privatemode released the Python library, an interactive test environment and the reproducible methodology across 29 datasets. No timelines announced for a second round of measurements.
Bottom line
The usual argument for keeping your own classifier was accuracy. This measurement reduces it to cost per decision and latency depending on where the server is, which are two engineering problems with known answers.
Sources
- Privatemode, “System One from GLM Flash”
- Our previous coverage: DeepSeek Flash’s time-of-day pricing, Xiaomi’s open MiMo weights and Mistral’s free plan and training on user data.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


