Saltar al contenido principal

Claude Mythos Preview

AnthropicClaudeProprietary

Descripción

Claude Mythos Preview is an unreleased general-purpose frontier model from Anthropic, a new tier above Opus (internal codename 'Capybara'). It identified thousands of zero-day vulnerabilities across every major operating system and web browser as part of Project Glasswing, a cross-industry cybersecurity initiative with 12 partners including AWS, Apple, Microsoft, and Google. State-of-the-art on SWE-bench Verified (93.9%), GPQA Diamond (94.6%), USAMO (97.6%), Terminal-Bench 2.0 (82.0%), CyberGym (83.1%), and Cybench (100% pass@1, saturated). Represents a 4.3x increase over the previous trendline for model performance. Deployed under ASL-3 Standard. Best-aligned Claude model to date per Anthropic's risk report, with the first-ever 24-hour internal alignment review before deployment. Not planned for general availability. Pricing for participants: $25/$125 per million tokens (input/output). 244-page system card.

Fecha de lanzamiento
2026-05-07
Parámetros
—
Longitud del contexto
—
Modalidades
image, text

Radar de capacidades

80
general
80
coding
80
reasoning
77
scienceest.
80
agents
80
multimodal

Science se estima a partir de las puntuaciones científicas de LLM Stats o del razonamiento cuando no hay benchmarks científicos dedicados.

Rankings

Dominio#PosiciónPuntuaciónFuente
Ranking multimodal61
55.0
LS

Puntuaciones de benchmarks (LLM Stats)

(LLM Stats (zeroeval))

Code

CyberGym83.1%Aut.

Language

MMMLU92.7%Aut.

Math

USAMO2597.6%Aut.

Multimodal

OSWorld-Verified79.6%Aut.

Reasoning

GPQANYU + Cohere + Anthropic (2023)94.6%Aut.
SWE-Bench Verified93.9%Aut.
CharXiv-R93.2%Aut.
SWE-bench Multilingual87.3%Aut.
BrowseCompOpenAI (2025)86.9%Aut.
Terminal-Bench 2.0Stanford × Laude Institute (2026)82.0%Aut.
Graphwalks BFS >128k80.0%Aut.
SWE-Bench ProPrinceton NLP (2024)77.8%Aut.
Humanity's Last Exam64.7%Aut.
SWE-Bench Multimodal59.0%Aut.

Safety

CyBench100.0%Aut.
FigQA89.0%Aut.

Índices de evaluación AA

(Artificial Analysis)

No hay datos de evaluación AA disponibles

Puntuaciones por categoría LLM Stats

(LLM Stats (zeroeval))
Language
90
Physics
90
Safety
90
Search
90
Frontend Development
90
Healthcare
90
Biology
90
Chemistry
90
Long Context
80
Math
80
Multimodal
80
Reasoning
80
Spatial Reasoning
80
General
80
Agents
80
Code
80
Tool Calling
80
Vision
80

Precios

No hay datos de precios disponibles

Velocidad

No hay datos de velocidad disponibles

Ranking de Precios por Proveedor

No hay datos de proveedores disponibles

Fuentes externas