Notes

DeepSeek V4.1 Flash cuts prices, and its peak rate falls at night in Chile


It launched on September 10 with open weights and a 1-million-token context. Two days earlier, three U.S. agencies accused DeepSeek of distillation.

September 11, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

$0.003. Price per million cached input tokens off-peak. It’s what makes rereading the same long context cheap, the typical usage pattern of a coding agent.

Context

The price war runs in different categories. Astra launched at $50 per million output tokens, about 83 times what V4.1 Flash charges, although the two models don’t compete for the same tasks. Google already made Gemini’s pricing more flexible to retain companies that look at the bill before the leaderboard.

What’s next

Bottom line

V4.1 Flash’s weights are downloaded from Hugging Face, the repository Nvidia announced it would buy on September 3. The Chinese model flagged by three U.S. agencies is distributed from a platform owned by the largest chip company in the United States.

Sources


Edited by Rodrigo Cornejo. How we select and verify the facts, in who writes.

Related notes

← All notes