modelsSEP 10 10:40 UTC
DeepSeek launches V4.1-Flash with 552B parameters and lower memory use
DeepSeek has released V4.1-Flash, a multimodal model with 552 billion total parameters that activates only 16 billion per token. The company says it cuts KV-cache memory to roughly a quarter of what its previous model required, which matters for agent workloads. It slightly edges out Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark.