DeepSeek V4.1 Flash (Reasoning, Max Effort)
描述
DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the released checkpoint. Its causal encoder-decoder architecture activates 8B parameters per token during prefill and 16B during decode. Trained on 45T multimodal tokens, it supports a 1M-token context, up to 384K output tokens on the DeepSeek API, and continuously adjustable reasoning effort from 1 to 100. CSA2 attention and FP4 KV caching reduce global KV cache storage to 890 bytes per token. The API model name is deepseek-flash. Catalog prices are peak rates per million tokens: $0.30 input, $0.006 cached input, and $1.20 output. Off-peak rates are $0.15, $0.003, and $0.60 respectively, effective September 10, 2026 at 04:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; all other times are off-peak (the launch pricing announcement also lists public holidays as off-peak).
能力雷達圖
排行榜排名
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))Agents
Code
Math
Reasoning
Vision
AA 評測指數
(Artificial Analysis)LLM Stats 分類評分
(LLM Stats (zeroeval))定價
速度
供應商價格排行
供應商價格排行
13 個供應商
比較該模型在不同 API 供應商之間的定價。