Claude Mythos Preview
描述
Claude Mythos Preview is an unreleased general-purpose frontier model from Anthropic, a new tier above Opus (internal codename 'Capybara'). It identified thousands of zero-day vulnerabilities across every major operating system and web browser as part of Project Glasswing, a cross-industry cybersecurity initiative with 12 partners including AWS, Apple, Microsoft, and Google. State-of-the-art on SWE-bench Verified (93.9%), GPQA Diamond (94.6%), USAMO (97.6%), Terminal-Bench 2.0 (82.0%), CyberGym (83.1%), and Cybench (100% pass@1, saturated). Represents a 4.3x increase over the previous trendline for model performance. Deployed under ASL-3 Standard. Best-aligned Claude model to date per Anthropic's risk report, with the first-ever 24-hour internal alignment review before deployment. Not planned for general availability. Pricing for participants: $25/$125 per million tokens (input/output). 244-page system card.
能力雷達圖
缺少專門科學評測時,Science 由 LLM Stats 科學得分或推理能力估算。
排行榜排名
| 領域 | #排名 | 分數 | 來源 |
|---|---|---|---|
| 多模態榜 | 61 | 55.0 | LS |
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))Code
Language
Math
Multimodal
Reasoning
Safety
AA 評測指數
(Artificial Analysis)暫無 AA 評測資料
LLM Stats 分類評分
(LLM Stats (zeroeval))定價
暫無定價資料
速度
暫無速度資料
供應商價格排行
暫無提供商資料