跳转到主要内容

GLM-5.1 (Reasoning)

Z AIGLM开源权重MIT · 商用许可

描述

GLM-5.1 is Z.AI's next-generation flagship foundation model designed for long-horizon agentic engineering tasks. Built on a 754B MoE architecture (40B active parameters), it can work continuously and autonomously on a single task for up to 8 hours, completing the full loop from planning and execution to iterative optimization and delivery. GLM-5.1 achieves state-of-the-art on SWE-Bench Pro (58.4) and demonstrates strong performance across coding, reasoning, and agentic benchmarks. It supports 200K context length, 128K max output tokens, thinking mode, function calling, structured output, context caching, and MCP integration. Overall performance is aligned with Claude Opus 4.6 with particular strengths in sustained execution and complex engineering optimization.

发布日期
2026-04-07
参数规模
754.0B
上下文长度
200K
支持模态
text

能力雷达图

27
general
54
coding
87
reasoning
61
science
60
agents
0
multimodal

排行榜排名

领域#排名分数来源
智能体能力模型榜44
49.0
LS
代码能力榜171
71.0
AA
通用能力榜92
68.0
AA
科学能力145
67.0
AA

基准测试分数 (LLM Stats)

(LLM Stats (zeroeval))

Agents

TAU3-Bench70.6%自报
Finance Agent v244.8%
Toolathlon40.7%自报

Code

CyberGym68.7%自报
NL2Repo42.7%自报
FrontierSWE31.0%

Math

AIME 202695.3%自报
HMMT 202594.0%自报
IMO-AnswerBench83.8%自报
HMMT Feb 2682.6%自报
LiveBench70.2%

Reasoning

Vending-Bench 2563441.0%自报
GPQANYU + Cohere + Anthropic (2023)86.2%自报
BrowseCompOpenAI (2025)79.3%自报
MCP Atlas71.8%自报
Terminal-Bench 2.0Stanford × Laude Institute (2026)69.0%自报
SWE-Bench ProPrinceton NLP (2024)58.4%自报
Humanity's Last Exam52.3%自报

AA 评测指数

(Artificial Analysis)
Tau2(Sierra + U Toronto + Vector Institute (2025))
97.7
Gpqa(NYU + Cohere + Anthropic (2023))
86.8
Ifbench(Google Research (2023))
76.3
Lcr(Artificial Analysis)
73.7
Terminalbench V2 1
61.8
Coding Index(Artificial Analysis)
55.8
Scicode(UIUC + Argonne National Lab (2024))
44.8
Terminalbench Hard(Stanford × Laude Institute (2026))
43.2
Hle(Center for AI Safety + Scale AI (2025))
30.1
Intelligence Index(Artificial Analysis)
26.1
Tau Banking
13.6
Terminalbench V4 0
2.0

LLM Stats 分类评分

(LLM Stats (zeroeval))
Agents
100
Reasoning
100
General
100
Physics
90
Biology
90
Chemistry
90
Math
80
Search
80
Safety
70
Code
60
Tool Calling
60
Vision
50
Finance
40

定价

输入价格$1.285 / 1M tokens
输出价格$4.07 / 1M tokens
混合价格(3:1)$1.981 / 1M tokens
缓存读取价格$0.26 / 1M tokens
缓存写入价格免费

速度

Tokens/秒0.0
首Token延迟0.00s
首回答延迟0.00s

供应商价格排行

供应商价格排行

5 个供应商

最便宜: DeepInfra最贵: Z AI
供应商输入输出
1DeepInfra最便宜
$0
$0
2ZAI
$0
$0
3FriendliAI
$0
$0
4EmpirioLabs AI
$0.825
$3.301
5Z AI主要
$1.285
$4.07

比较该模型在不同 API 供应商之间的定价。

外部链接