跳转到主要内容

Claude Opus 4.8

AnthropicClaudeProprietary

描述

Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

发布日期
2026-05-28
参数规模
上下文长度
1.0M
支持模态
image, pdf, text

能力雷达图

61
general
70
coding
70
reasoning
77
science估算
70
agents
70
multimodal

Science 在缺少专门科学评测时使用推理能力代理估算。

排行榜排名

领域#排名分数来源
智能体能力模型榜7
72.0
LS
数学推理7
94.0
LB
多模态榜13
63.0
LS
推理能力7
89.0
LB

基准测试分数 (LLM Stats)

(LLM Stats (zeroeval))

Agents

GDPval-AA1638.00 / 3000
DeepSearchQA93.1%自报
BrowseCompOpenAI (2025)84.3%自报
OSWorld-Verified83.4%自报
MCP Atlas82.2%自报
CyberGym78.8%自报
FrontierSWE75.0%
Terminal-Bench 2.0Stanford × Laude Institute (2026)74.6%自报
SWE-Bench ProPrinceton NLP (2024)69.2%自报
OfficeQA Pro66.2%自报
Toolathlon59.9%自报
DeepSWE 1.159.0%
Finance Agent v253.9%
Finance Agent53.9%自报
FrontierCode 1.146.5%
SWE-Bench Multimodal38.4%自报

Biology

GPQANYU + Cohere + Anthropic (2023)93.6%自报

Code

SWE-Bench Verified88.6%自报
SWE-bench Multilingual84.4%自报

General

Include87.6%自报
LiveBench77.2%

Grounding

ScreenSpot Pro87.9%自报

Healthcare

HealthBench Professional55.8%自报

Long Context

Graphwalks parents >128k83.3%自报
Graphwalks BFS >128k68.1%自报

Math

Humanity's Last Exam57.9%自报

Multimodal

CharXiv-R89.9%自报

AA 评测指数

(Artificial Analysis)

暂无 AA 评测数据

LLM Stats 分类评分

(LLM Stats (zeroeval))
Legal
100
Finance
100
Agents
100
Reasoning
83
General
61
Physics
90
Search
90
Frontend Development
90
Grounding
90
Biology
90
Chemistry
90
Long Context
80
Safety
80
Spatial Reasoning
80
Math
70
Multimodal
70
Code
70
Tool Calling
70
Vision
70
Healthcare
60

定价

输入价格$0.00001 / 1M tokens
输出价格$0.00003 / 1M tokens
混合价格(3:1)$0.00001 / 1M tokens
缓存读取价格$0.5 / 1M tokens
缓存写入价格$6.25 / 1M tokens

速度

暂无速度数据

供应商价格排行

供应商价格排行

16 个供应商

最便宜: Anthropic最贵: Venice AI
供应商输入输出
1Anthropic主要
$0.00001
$0.00003
2UnoRouter
$0.425
$2.125
3Xpersona
$1.5
$9.25
4Abacus
$5
$25
5OpenCode Zen
$5
$25
6AIHubMix
$5
$25
7Azure Cognitive Services
$5
$25
8LLM Gateway
$5
$25
9Azure
$5
$25
10routing.run
$5
$25
11FreeModel
$5
$25
12Neon
$5
$25
13Pioneer
$5
$25
14DaoXE
$5
$25
15Modelis
$5
$25
16Venice AI
$6
$30

比较该模型在不同 API 供应商之间的定价。

外部链接