modelsSEP 10 12:40 UTC
DeepSeek releases V4.1-Flash multimodal model with reduced KV cache memory
DeepSeek has introduced V4.1-Flash, a multimodal model with 552 billion parameters, though only 16 billion are active per token. The release cuts KV cache memory requirements to a quarter of the previous generation. On the DeepSWE coding benchmark, it slightly outperforms Opus 5 and GPT-5.6 Sol.