What happened
- DeepSeek launched V4.1 Flash on September 10, a 552-billion-parameter model that activates about 8 billion for each token, the unit of text it processes. It reads text and images, supports a 1-million-token context, and its weights are published under an MIT license.
- Off-peak price: $0.15 per million input tokens, $0.60 per million output tokens and $0.003 if the input is already cached. At peak hours, double.
- Peak hours run from 01:00 to 04:00 and from 06:00 to 10:00 UTC, Monday through Friday. DeepSeek has billed under this scheme since August 16.
- From September 14, requests to V4 Pro will be served by V4.1 Flash, at its price, until a V4.1 Pro exists. The company says Flash already beats Pro on performance, cost and speed.
Why it matters
- In Chile, which moved to UTC-3 on September 6, peak hours fall between 22:00 and 01:00 and between 03:00 and 07:00. The entire working day, from Mexico City to Santiago, sits in the low rate. A process scheduled in the small hours to save money pays double.
- DeepSeek tops the list of companies that the NSA, the FBI and CISA, three U.S. security agencies, accused on September 8 of extracting capabilities from American models. For a procurement team, the lowest price on the market comes with that warning attached.
- Open weights make it possible to run the model on your own servers, without sending data to the company’s API. A team handling Chileans’ personal data can use it without that information leaving its infrastructure, just as the data protection law the government proposes to postpone defines how such data is transferred abroad.
The number
$0.003. Price per million cached input tokens off-peak. It’s what makes rereading the same long context cheap, the typical usage pattern of a coding agent.
Context
The price war runs in different categories. Astra launched at $50 per million output tokens, about 83 times what V4.1 Flash charges, although the two models don’t compete for the same tasks. Google already made Gemini’s pricing more flexible to retain companies that look at the bill before the leaderboard.
What’s next
- September 14, 2026, 12:00 Beijing time (01:00 in Chile): V4 Pro starts being served by V4.1 Flash.
- V4.1 Pro: no date announced.
Bottom line
V4.1 Flash’s weights are downloaded from Hugging Face, the repository Nvidia announced it would buy on September 3. The Chinese model flagged by three U.S. agencies is distributed from a platform owned by the largest chip company in the United States.
Sources
- DeepSeek-V4.1-Flash debuts with $0.003/1M off-peak cached-input rate — VentureBeat, September 10, 2026
- DeepSeek formally launches V4.1 Flash, routes V4 Pro requests to Flash — TechNode, September 10, 2026
- China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies — NSA, CISA and FBI, September 2026
Edited by Rodrigo Cornejo. How we select and verify the facts, in who writes.



%2016.40.34.D86RbLdS_Z2hOnWV.webp)
