Перейти к основному содержанию

DeepSeek V4.1 Flash (Reasoning, Max Effort)

DeepSeekDeepSeekОткрытые весаMIT · Коммерческое использование

Описание

DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the released checkpoint. Its causal encoder-decoder architecture activates 8B parameters per token during prefill and 16B during decode. Trained on 45T multimodal tokens, it supports a 1M-token context, up to 384K output tokens on the DeepSeek API, and continuously adjustable reasoning effort from 1 to 100. CSA2 attention and FP4 KV caching reduce global KV cache storage to 890 bytes per token. The API model name is deepseek-flash. Catalog prices are peak rates per million tokens: $0.30 input, $0.006 cached input, and $1.20 output. Off-peak rates are $0.15, $0.003, and $0.60 respectively, effective September 10, 2026 at 04:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; all other times are off-peak (the launch pricing announcement also lists public holidays as off-peak).

Дата выхода
2026-09-10
Параметры
763.2B
Длина контекста
1.0M
Модальности
image, text

Радар способностей

39
general
52
coding
100
reasoning
47
science
50
agents
70
multimodal

Рейтинги

Оценки бенчмарков (LLM Stats)

(LLM Stats (zeroeval))

Agents

SEC-bench Pro62.8%Сам.
AutomationBench54.8%Сам.
Agents' Last Exam31.8%Сам.
Program Bench20.3%Сам.
ExploitGym15.3%Сам.

Code

CyberGym88.1%Сам.
DeepSWE 1.174.2%Сам.
NL2Repo64.0%Сам.

Math

CodeForces3471.00 / 3000Сам.
MathArena Apex65.6%Сам.

Reasoning

GPQANYU + Cohere + Anthropic (2023)90.9%Сам.
Terminal-Bench 2.190.6%Сам.
Humanity's Last Exam (with tools)63.9%Сам.
Humanity's Last Exam (no tools, text-only)39.1%Сам.
Humanity's Last Exam36.8%Сам.
Terminal-Bench 4.031.2%Сам.
Terminal-Bench 3.030.0%Сам.

Vision

BabyVision89.6%Сам.
Chartography78.9%Сам.
ZEROBench0.49 / 100Сам.

Индексы оценки AA

(Artificial Analysis)
Lcr(Artificial Analysis)
84.0
Scicode(UIUC + Argonne National Lab (2024))
51.9
Intelligence Index(Artificial Analysis)
39.5
Hle(Center for AI Safety + Scale AI (2025))
39.2

Оценки категорий LLM Stats

(LLM Stats (zeroeval))
Math
100
Reasoning
100
General
100
Physics
90
Biology
90
Chemistry
90
Multimodal
70
Safety
60
Vision
60
Agents
50
Code
50
Tool Calling
50

Цены

Цена ввода$0.3 / 1M токенов
Цена вывода$1.2 / 1M токенов
Смешанная цена (3:1)$0.525 / 1M токенов
Цена чтения кэша$0.003 / 1M токенов

Скорость

Токенов/сек267.1
Задержка первого токена0.90s
Время до первого ответа8.38s

Рейтинг цен провайдеров

Рейтинг цен провайдеров

13 провайдеров

Самый дешевый: FireworksСамый дорогой: Venice AI
ПровайдерВводВывод
1FireworksСамый дешевый
$0
$0
2DeepSeek
$0
$0
3Novita
$0
$0
4DeepInfra
$0
$0
5NanoGPT
$0.15
$0.6
6OpenRouter
$0.15
$0.6
7Merge Gateway
$0.15
$0.6
8LLM Gateway
$0.15
$0.6
9CrossModel
$0.27
$1.08
10Kilo Gateway
$0.3
$1.2
11Vercel AI Gateway
$0.3
$1.2
12Ofox
$0.3
$1.2
13Venice AI
$0.375
$1.5

Сравнение цен разных API-провайдеров для этой модели.

Внешние ссылки