Перейти к основному содержанию

Claude Opus 4.8

AnthropicClaudeProprietary

Описание

Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

Дата выхода
2026-05-28
Параметры
Длина контекста
1.0M
Модальности
image, pdf, text

Радар способностей

61
general
70
coding
70
reasoning
77
scienceоцен.
70
agents
70
multimodal

Science использует прокси на основе рассуждений, когда специализированные научные бенчмарки недоступны.

Рейтинги

Оценки бенчмарков (LLM Stats)

(LLM Stats (zeroeval))

Agents

GDPval-AA1638.00 / 3000
DeepSearchQA93.1%Сам.
BrowseCompOpenAI (2025)84.3%Сам.
OSWorld-Verified83.4%Сам.
MCP Atlas82.2%Сам.
CyberGym78.8%Сам.
FrontierSWE75.0%
Terminal-Bench 2.0Stanford × Laude Institute (2026)74.6%Сам.
SWE-Bench ProPrinceton NLP (2024)69.2%Сам.
OfficeQA Pro66.2%Сам.
Toolathlon59.9%Сам.
DeepSWE 1.159.0%
Finance Agent v253.9%
Finance Agent53.9%Сам.
FrontierCode 1.146.5%
SWE-Bench Multimodal38.4%Сам.

Biology

GPQANYU + Cohere + Anthropic (2023)93.6%Сам.

Code

SWE-Bench Verified88.6%Сам.
SWE-bench Multilingual84.4%Сам.

General

Include87.6%Сам.
LiveBench77.2%

Grounding

ScreenSpot Pro87.9%Сам.

Healthcare

HealthBench Professional55.8%Сам.

Long Context

Graphwalks parents >128k83.3%Сам.
Graphwalks BFS >128k68.1%Сам.

Math

Humanity's Last Exam57.9%Сам.

Multimodal

CharXiv-R89.9%Сам.

Индексы оценки AA

(Artificial Analysis)

Нет данных AA оценки

Оценки категорий LLM Stats

(LLM Stats (zeroeval))
Legal
100
Finance
100
Agents
100
Reasoning
83
General
61
Physics
90
Search
90
Frontend Development
90
Grounding
90
Biology
90
Chemistry
90
Long Context
80
Safety
80
Spatial Reasoning
80
Math
70
Multimodal
70
Code
70
Tool Calling
70
Vision
70
Healthcare
60

Цены

Цена ввода$0.00001 / 1M токенов
Цена вывода$0.00003 / 1M токенов
Смешанная цена (3:1)$0.00001 / 1M токенов
Цена чтения кэша$0.5 / 1M токенов
Цена записи кэша$6.25 / 1M токенов

Скорость

Нет данных о скорости

Рейтинг цен провайдеров

Рейтинг цен провайдеров

16 провайдеров

Самый дешевый: AnthropicСамый дорогой: Venice AI
ПровайдерВводВывод
1AnthropicОсновной
$0.00001
$0.00003
2UnoRouter
$0.425
$2.125
3Xpersona
$1.5
$9.25
4Abacus
$5
$25
5OpenCode Zen
$5
$25
6AIHubMix
$5
$25
7Azure Cognitive Services
$5
$25
8LLM Gateway
$5
$25
9Azure
$5
$25
10routing.run
$5
$25
11FreeModel
$5
$25
12Neon
$5
$25
13Pioneer
$5
$25
14DaoXE
$5
$25
15Modelis
$5
$25
16Venice AI
$6
$30

Сравнение цен разных API-провайдеров для этой модели.

Внешние ссылки