DeepSeek V4.1 Flash (Reasoning, Max Effort)
描述
DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the released checkpoint. Its causal encoder-decoder architecture activates 8B parameters per token during prefill and 16B during decode. Trained on 45T multimodal tokens, it supports a 1M-token context, up to 384K output tokens on the DeepSeek API, and continuously adjustable reasoning effort from 1 to 100. CSA2 attention and FP4 KV caching reduce global KV cache storage to 890 bytes per token. The API model name is deepseek-flash. Catalog prices are peak rates per million tokens: $0.30 input, $0.006 cached input, and $1.20 output. Off-peak rates are $0.15, $0.003, and $0.60 respectively, effective September 10, 2026 at 04:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; all other times are off-peak (the launch pricing announcement also lists public holidays as off-peak).
能力雷达图
排行榜排名
基准测试分数 (LLM Stats)
(LLM Stats (zeroeval))Agents
Code
Math
Reasoning
Vision
AA 评测指数
(Artificial Analysis)LLM Stats 分类评分
(LLM Stats (zeroeval))定价
速度
供应商价格排行
供应商价格排行
13 个供应商
比较该模型在不同 API 供应商之间的定价。