DeepSeek-V4-Flash-0423
설명
DeepSeek-V4-Flash-0423 is the preview release of DeepSeek-V4-Flash, a 284B-parameter MoE model with 13B activated parameters and a 1M-token context window, evaluated here at the default high reasoning effort. It shares the V4 series' hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) for dramatically improved long-context efficiency, Manifold-Constrained Hyper-Connections (mHC) for stable signal propagation, and the Muon optimizer for faster convergence. Pre-trained on more than 32T tokens and post-trained with a two-stage paradigm of domain-specific expert cultivation followed by on-policy distillation, V4-Flash offers reasoning capabilities that closely approach V4-Pro with faster responses and highly cost-effective pricing.
능력 레이더
전용 과학 벤치마크가 없을 때 Science는 LLM Stats 과학 점수 또는 추론 능력에서 추정합니다.
랭킹
| 도메인 | #순위 | 점수 | 소스 |
|---|---|---|---|
| 에이전트형 역량 | 117 | 35.0 | LS |
벤치마크 점수 (LLM Stats)
(LLM Stats (zeroeval))Agents
Biology
Code
Factuality
Finance
General
Math
AA 평가 지수
(Artificial Analysis)AA 평가 데이터가 없습니다
LLM Stats 카테고리 점수
(LLM Stats (zeroeval))가격
속도
속도 데이터가 없습니다
공급자 가격 순위
공급자 가격 순위
3개 공급자
이 모델의 다양한 API 공급자 간 가격 비교.