DeepSeek’s September 10, 2026 product post for V4.1-Flash leads with KV-cache compression and API price cuts, not just another parameter headline. The company also says V4-Pro traffic will route to Flash at Flash prices from Sept 14 UTC until V4.1-Pro ships—schedule language, not independent proof that Flash beats Pro everywhere.
Deepseek V4.1 Flash (Fully Tested): 200 TPS & Beats Astra!? (+New Architecture Overview) · AICodeKing
Quick Take
- DeepSeek launched V4.1-Flash on the API with a company focus on smaller KV cache and lower agent costs.
- Architecture claims (company): 552B MoE with 8B active on input and 16B on output via a Causal Encoder–Decoder design, plus native visual understanding.
- From 04:00 UTC Sept 14, 2026, DeepSeek says
deepseek-v4-prorequests route to V4.1-Flash at Flash pricing until V4.1-Pro arrives.
What DeepSeek actually announced
On September 10, 2026, DeepSeek published Introducing DeepSeek-V4.1-Flash. The post positions Flash as the smallest model in a new architecture family, live on the API when callers set the model to deepseek-flash. Prior V4-Flash and vision-exp aliases temporarily route to V4.1-Flash for compatibility, per the company.
The economic lead is cache size. DeepSeek says V4.1-Flash’s KV cache needs roughly one-quarter the HBM and one-eighth the SSD storage versus the previous generation, arguing that cache-hit charges often dominate agent bills. Separately, it says new API prices took effect at 04:00 UTC on Sept 10, with off-peak rates at 50% of peak.
The Sept 14 Pro routing schedule
DeepSeek also states that, starting 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches. That is a company migration schedule—useful for operators planning keys and budgets—not independent proof that Flash is superior to Pro on every task.
The same post claims “tests by multiple parties” put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, and that DeepSeek is phasing out V4-Pro. Attribute those comparisons. Where Hugging Face cards or third-party latency/price pages list the model, treat them as listings, not a settled bake-off.
How to cover without hype
Parameter totals still grab headlines; DeepSeek’s own post spends more energy on active parameters and cache compression. For practical readers, the actionable facts are API alias behavior, peak/off-peak pricing, and the dated Pro→Flash route. Capability bragging stays inside quotation marks until independent evals with disclosed prompts catch up.
Operator checklist
If you already call deepseek-v4-pro, read the Sept 14 UTC cutover as a billing and behavior change, not a quiet alias. Confirm whether your SDK pins a model string that will suddenly resolve to Flash, and whether your eval harness still compares against a true Pro endpoint afterward. If you care about multimodal inputs, note that DeepSeek markets native visual understanding on Flash—verify your client path actually sends images the way the new API expects.
Open-source deployment remains a separate track: DeepSeek says it will work with the community on inference support and invites large GPU+storage deployments to talk. That is not the same as a turnkey self-host guarantee on day one.
Bottom Line
V4.1-Flash is a cache-and-price story with a hard calendar note for Pro callers on Sept 14 UTC. Cover DeepSeek’s schedule and cost claims as company statements. Do not upgrade them into a universal verdict that Flash replaces Pro on quality.