DeepSeek has officially transitioned its API ecosystem to the new V4.1 Flash model, marking a strategic shift where the "lightweight" tier now outperforms the previous flagship. Released on September 10, 2026, V4.1 Flash is not merely an incremental update but a complete architectural replacement for the previous V4 Flash and V4 Flash Vision Exp models.
The most disruptive aspect of this release is the planned retirement of V4 Pro. Starting September 14, 2026, at 12:00 Beijing Time, all requests directed to the Pro model will be automatically routed to V4.1 Flash. This move is backed by internal and external testing showing that the new Flash model comprehensively surpasses V4 Pro across all key metrics, including raw performance, inference speed, and task completion time.
Asymmetric Architecture and Multimodality

Scraped da https://x.com/deepseek_ai/status/2097930608790167907 As previously detailed in the beta release, V4.1 Flash utilizes a brand-new Causal-Encoder-Decoder architecture based on a Mixture-of-Experts (MoE) framework with 552 billion parameters. The core innovation lies in its asymmetric design: only 8B parameters are activated during the input phase, while 16B parameters are activated for output. This imbalance allows the model to maintain a high capability ceiling while drastically reducing the computational cost of processing long prompts.
Furthermore, V4.1 Flash integrates native visual understanding directly into the base architecture. Unlike previous iterations where vision was often a separate module or an experimental addition, this native integration allows for higher throughput and more seamless multimodal reasoning.
New Dynamic Pricing Structure
The launch introduces a tiered pricing model that fluctuates based on demand, reflecting the company's move away from the "ultra-cheap" era toward a more sustainable infrastructure model. The pricing is split into off-peak and peak hours (1:00–4:00 AM and 6:00–10:00 AM UTC, Monday to Friday).
| Token Type | Off-peak Price (per 1M) | Peak Price (per 1M) |
|---|---|---|
| Input (Cache Hit) | $0.003 | $0.006 |
| Input (Cache Miss) | $0.15 | $0.3 |
| Output Tokens | $0.6 | $1.2 |
Compared to the previous V4 Flash, these rates represent a reduction of approximately 11% to 57% depending on the token type, effectively lowering the barrier for high-volume agentic workflows.
Strategic Implications: The Death of the Mid-Tier
The decision to route Pro requests to a Flash model signals a paradigm shift in AI scaling. For months, the industry has followed a strict hierarchy: Flash for speed, Pro/Ultra for intelligence. By proving that an optimized MoE architecture can deliver "Pro" level reasoning at "Flash" speeds and costs, DeepSeek is challenging the necessity of massive, dense models for most enterprise tasks.
This evolution follows a period of intense competition where DeepSeek's previous versions, such as the V4-Flash 0731, had already begun to challenge the cost-performance ratio of models like GPT-5.6 Luna. The current transition suggests that DeepSeek is prioritizing architectural efficiency over raw parameter count, aiming for a scalable framework that can be expanded to even larger models without the linear increase in latency.

No comments yet. Be the first!