DeepSeek-V4-Flash-0423
描述
DeepSeek-V4-Flash-0423 is the preview release of DeepSeek-V4-Flash, a 284B-parameter MoE model with 13B activated parameters and a 1M-token context window, evaluated here at the default high reasoning effort. It shares the V4 series' hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) for dramatically improved long-context efficiency, Manifold-Constrained Hyper-Connections (mHC) for stable signal propagation, and the Muon optimizer for faster convergence. Pre-trained on more than 32T tokens and post-trained with a two-stage paradigm of domain-specific expert cultivation followed by on-policy distillation, V4-Flash offers reasoning capabilities that closely approach V4-Pro with faster responses and highly cost-effective pricing.
能力雷達圖
缺少專門科學評測時,Science 由 LLM Stats 科學得分或推理能力估算。
排行榜排名
| 領域 | #排名 | 分數 | 來源 |
|---|---|---|---|
| 智慧體能力模型榜 | 117 | 35.0 | LS |
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))Agents
Biology
Code
Factuality
Finance
General
Math
AA 評測指數
(Artificial Analysis)暫無 AA 評測資料
LLM Stats 分類評分
(LLM Stats (zeroeval))定價
速度
暫無速度資料
供應商價格排行
供應商價格排行
3 個供應商
比較該模型在不同 API 供應商之間的定價。