Global foundation-model progress briefing
English Edition中文
Enter keywords to search ingested stories.

Today / Sunday, September 27, 2026

limbo logolimbo

Data updated

Sep 27, 01:01 AM

Live sources

17

Ingestion status

Live ingest

CompaniesDeepSeek Official

DeepSeek-V4.1-Flash 发布,文本与 Agent 性能全面提升、兼具原生多模态视觉理解能力,能力更强、速度更快、且成本更低,欢迎测试和反馈

Summary

No summary yet; AI summarization will be added later.

Original Article

Captured source content or English translation, normalized into this reading format.

Read Source

DeepSeek V4.1 Flash: Stronger, Faster, More Accessible

Today, we are officially releasing the DeepSeek V4.1 Flash model. It is the smallest model in our entirely new model architecture series, featuring native multimodal visual understanding capabilities. The design goals of the new model architecture are: a higher capability ceiling, faster inference, greater throughput, and scalability to larger-parameter models.

DeepSeek V4.1 Flash is a 552B-parameter MoE model that adopts an entirely new Causal-Encoder-Decoder architecture, with asymmetric input and output: only 8B input activation and 16B output activation, making its cost significantly lower than known models of the same size. At the same time, V4.1 Flash also adopts a new pretraining approach and has undergone larger-scale reinforcement learning post-training. In benchmark tests, it successfully surpasses the intelligence level of a range of flagship models, including DeepSeek V4 Pro.

Performance comparison of DeepSeek-V4.1-Flash and mainstream frontier models on Agentic Benchmark

The new-generation model greatly reduces the size of the KV Cache. Compared with the previous-generation model, the demand for HBM is reduced to 1/4, and the demand for SSD is reduced to 1/8. In Agent usage scenarios, the cost of cache hits often accounts for a relatively high proportion, and the compression of the KV Cache greatly reduces the usage cost of Agent-type tasks.

As shown in the figure, DeepSeek's continued progress in reducing context storage. Relative to the first-generation model, the KV Cache has been reduced by 437 times.

DeepSeek V4.1 Flash is now also available on the DeepSeek API, with native multimodal support. Simply change the model name to deepseek-flash to call the latest V4.1 Flash model. The old-version models V4 Flash and V4 Flash Vision Exp are now offline. For compatibility considerations, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will be temporarily routed to V4.1 Flash.

At the same time, after testing by multiple parties, V4.1 Flash has comprehensively surpassed DeepSeek V4 Pro across various metrics including performance, cost, speed, and total time. Therefore, we plan to gradually take the V4 Pro model offline. After 12:00 Beijing Time on September 14, 2026, and until the future launch of V4.1 Pro, user requests accessing deepseek-v4-pro will all be routed to V4.1 Flash and billed at the V4.1 Flash unit price.

WorkBuddy (including CodeBuddy) and OpenCode, as official partners, have now fully integrated DeepSeek V4.1 Flash. Welcome to use them!

Thanks to the innovation in model architecture, DeepSeek V4.1 Flash can serve more users at a lower cost, so we have accordingly lowered the pricing of V4.1 Flash. At the same time, in order to allocate resources more reasonably, we still adopt peak/off-peak pricing, with the off-peak price being half of the peak-hour price, encouraging users to adjust task timing according to their actual usage. The new prices take effect starting at 12:00 on September 10, 2026.

We will fully support the open-source community in adapting inference for the new model, and will try various ways to expand the scope of deployment. If you have large-scale deployment needs and possess the corresponding resources (2k GPU cards, a storage cluster), welcome to contact us.

Region

China

Heat Score

89

Category

Companies

Language

zh