Particle.news
Download on the App Store

DeepSeek Launches V4.1‑Flash, a Native Multimodal MoE Model Built for Speed and Low Cost

DeepSeek hopes to force down AI inference costs with a sparse Mixture‑of‑Experts design as it moves toward a STAR Market IPO.

Overview

  • The company publicly released V4.1‑Flash on Thursday, Sept. 10, 2026, offering API access and publishing model weights under an MIT license on Hugging Face for download.
  • DeepSeek describes the model as a 552 billion‑parameter Mixture‑of‑Experts with an asymmetric Causal‑Encoder‑Decoder that fires only a small subset of parameters per request to cut compute.
  • The firm and community tests report throughput above 400 tokens per second and a 1 million‑token context window supported by a reduced KV cache of about 890 bytes per token.
  • DeepSeek said it will route V4‑Pro traffic to the cheaper Flash model starting Sept. 14 and has signaled hiring and capacity builds tied to an expected STAR Market IPO.
  • Independent third‑party verification and long‑term stability data remain limited, raising questions about operational reliability, benchmark reproducibility, and regulatory or data‑sovereignty risks for users who deploy the open weights.