Overview
- The company publicly released V4.1‑Flash on Thursday, Sept. 10, 2026, offering API access and publishing model weights under an MIT license on Hugging Face for download.
- DeepSeek describes the model as a 552 billion‑parameter Mixture‑of‑Experts with an asymmetric Causal‑Encoder‑Decoder that fires only a small subset of parameters per request to cut compute.
- The firm and community tests report throughput above 400 tokens per second and a 1 million‑token context window supported by a reduced KV cache of about 890 bytes per token.
- DeepSeek said it will route V4‑Pro traffic to the cheaper Flash model starting Sept. 14 and has signaled hiring and capacity builds tied to an expected STAR Market IPO.
- Independent third‑party verification and long‑term stability data remain limited, raising questions about operational reliability, benchmark reproducibility, and regulatory or data‑sovereignty risks for users who deploy the open weights.