メインコンテンツへスキップ

Claude Opus 4.8

AnthropicClaudeProprietary

説明

Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

リリース日
2026-05-28
パラメータ
コンテキスト長
1.0M
モダリティ
image, pdf, text

能力レーダー

61
general
70
coding
70
reasoning
77
science推定
70
agents
70
multimodal

専門的な科学ベンチマークが利用できない場合、Scienceは推論プロキシを使用して推定します。

ランキング

ドメイン#順位スコアソース
エージェント能力7
72.0
LS
数学的推論7
94.0
LB
マルチモーダルランキング13
63.0
LS
推論7
89.0
LB

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Agents

GDPval-AA1638.00 / 3000
DeepSearchQA93.1%自己申告
BrowseCompOpenAI (2025)84.3%自己申告
OSWorld-Verified83.4%自己申告
MCP Atlas82.2%自己申告
CyberGym78.8%自己申告
FrontierSWE75.0%
Terminal-Bench 2.0Stanford × Laude Institute (2026)74.6%自己申告
SWE-Bench ProPrinceton NLP (2024)69.2%自己申告
OfficeQA Pro66.2%自己申告
Toolathlon59.9%自己申告
DeepSWE 1.159.0%
Finance Agent v253.9%
Finance Agent53.9%自己申告
FrontierCode 1.146.5%
SWE-Bench Multimodal38.4%自己申告

Biology

GPQANYU + Cohere + Anthropic (2023)93.6%自己申告

Code

SWE-Bench Verified88.6%自己申告
SWE-bench Multilingual84.4%自己申告

General

Include87.6%自己申告
LiveBench77.2%

Grounding

ScreenSpot Pro87.9%自己申告

Healthcare

HealthBench Professional55.8%自己申告

Long Context

Graphwalks parents >128k83.3%自己申告
Graphwalks BFS >128k68.1%自己申告

Math

Humanity's Last Exam57.9%自己申告

Multimodal

CharXiv-R89.9%自己申告

AA評価指数

(Artificial Analysis)

AA評価データがありません

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Legal
100
Finance
100
Agents
100
Reasoning
83
General
61
Physics
90
Search
90
Frontend Development
90
Grounding
90
Biology
90
Chemistry
90
Long Context
80
Safety
80
Spatial Reasoning
80
Math
70
Multimodal
70
Code
70
Tool Calling
70
Vision
70
Healthcare
60

価格設定

入力価格$0.00001 / 1Mトークン
出力価格$0.00003 / 1Mトークン
混合価格(3:1)$0.00001 / 1Mトークン
キャッシュ読み取り価格$0.5 / 1Mトークン
キャッシュ書き込み価格$6.25 / 1Mトークン

速度

速度データがありません

プロバイダー価格ランキング

プロバイダー価格ランキング

16 プロバイダー

最安: Anthropic最高: Venice AI
プロバイダー入力出力
1Anthropicプライマリ
$0.00001
$0.00003
2UnoRouter
$0.425
$2.125
3Xpersona
$1.5
$9.25
4Abacus
$5
$25
5OpenCode Zen
$5
$25
6AIHubMix
$5
$25
7Azure Cognitive Services
$5
$25
8LLM Gateway
$5
$25
9Azure
$5
$25
10routing.run
$5
$25
11FreeModel
$5
$25
12Neon
$5
$25
13Pioneer
$5
$25
14DaoXE
$5
$25
15Modelis
$5
$25
16Venice AI
$6
$30

このモデルの異なるAPIプロバイダー間の価格を比較。

外部リンク