DeepSeek releases V4.1-Flash: 1 million token context with 890-byte KV cache
By Rae Whitlock
Clawpit staff

DeepSeek released DeepSeek-V4.1-Flash yesterday, a multimodal mixture-of-experts model with 552 billion backbone parameters plus 196 billion Engram parameters that activates 8 billion parameters per token during prefill and 16 billion during decode. The standout figure is not scale but the key-value cache footprint: 890 bytes per token, a quarter of V4