Saltar al contenido principal

Llama 3.1 Instruct 405B

MetaLlamaOpen WeightLlama 3.1 Community License

Descripción

Llama 3.1 405B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks. The model supports 8 languages and has a 128K token context length.

Fecha de lanzamiento
2024-07-23
Parámetros
405.0B
Longitud del contexto
—
Modalidades
text

Radar de capacidades

27
general
31
coding
23
reasoning
37
science
70
agents
0
multimodal

Rankings

Dominio#PosiciónPuntuaciónFuente
Ranking de codificación472
24.0
AA
Ranking general453
31.0
AA
Ciencia494
27.0
AA

Puntuaciones de benchmarks (LLM Stats)

(LLM Stats (zeroeval))

Chat

IFEvalGoogle Research (2023)88.6%Aut.

General

BFCL88.5%Aut.
MMLU87.3%Aut.
Multipl-E MBPP65.7%Aut.
Nexus58.7%Aut.

Language

MMLU (CoT)88.6%Aut.
Multipl-E HumanEval75.2%Aut.
MMLU-Pro73.3%Aut.

Math

GSM8k96.8%Aut.
Multilingual MGSM (CoT)91.6%Aut.
MATH73.8%Aut.

Reasoning

ARC-C96.9%Aut.
API-Bank92.0%Aut.
HumanEvalOpenAI (2021)89.0%Aut.
MBPP EvalPlus88.6%Aut.
DROP84.8%Aut.
GPQANYU + Cohere + Anthropic (2023)50.7%Aut.
Gorilla Benchmark API Bench35.3%Aut.

Índices de evaluación AA

(Artificial Analysis)
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
73.2
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
70.3
Gpqa(NYU + Cohere + Anthropic (2023))
51.5
Ifbench(Google Research (2023))
39.0
Livecodebench(UC Berkeley + MIT + Cornell (2024))
30.5
Lcr(Artificial Analysis)
25.3
Aime(MAA (Mathematical Association of America))
21.3
Tau2(Sierra + U Toronto + Vector Institute (2025))
19.0
Intelligence Index(Artificial Analysis)
7.3
Terminalbench Hard(Stanford × Laude Institute (2026))
6.8
Hle(Center for AI Safety + Scale AI (2025))
4.0
Math Index(Artificial Analysis)
3.0
Aime 25(MAA (Mathematical Association of America))
3.0

Puntuaciones por categoría LLM Stats

(LLM Stats (zeroeval))
Chat
90
Instruction Following
90
Math
90
Structured Output
90
Language
80
Legal
80
Reasoning
80
Finance
80
General
80
Healthcare
80
Tool Calling
70
Code
60
Physics
50
Biology
50
Chemistry
50

Precios

Precio de entradaGratis
Precio de salidaGratis
Precio mixto (3:1)Gratis

Velocidad

Tokens/seg0.0
Retraso del primer token0.00s
Tiempo hasta la respuesta0.00s

Ranking de Precios por Proveedor

No hay datos de proveedores disponibles

Fuentes externas