Skip to main content

Hy3

TencentOpen WeightApache 2.0 · Commercial OK

Description

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and a 3.8B MTP layer, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, the team scaled up post-training with higher-quality data and RL, gathering feedback from 50+ products. Hy3 outperforms similar-size models and rivals flagship open-source models with 2-5x the parameters, with strong gains in reasoning, agentic, and long-context tasks. It uses 80 layers (plus 1 MTP layer), 64 GQA attention heads (8 KV heads, head dim 128), a 4096 hidden size, 192 experts with top-8 activated, a 256K context window, and BF16 precision. Hy3 is a hybrid-thinking model supporting configurable reasoning effort (no_think, low, high), and emphasizes production-grade tool-call and output-format stability, reduced hallucination, and reliable multi-turn intent tracking.

Release Date
2026-07-06
Parameters
295.0B
Context Length
262K
Modalities
text

Capability Radar

27
general
57
coding
90
reasoning
64
science
70
agents
0
multimodal

Rankings

Domain#RankScoreSource
Agentic Capability53
48.0
LS
Code Ranking128
78.0
AA
General Ranking321
41.0
AA
Science111
72.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Agents

WildClawBench53.6%SR
Toolathlon48.5%SR

Chemistry

SuperChem54.9%SR

Code

Claw-Eval68.5%SR
SkillsBench55.3%SR
NL2Repo45.6%SR
DeepSWE28.0%SR
CL-bench23.8%SR
CL-bench (Life)17.0%SR

Math

USAMO 202630.24 / 42SR
IMO-AnswerBench90.0%SR
ArXivMath52.2%SR
MathArena Apex38.7%SR
HorizonMath7.1%SR

Physics

PHYBench77.4%SR
CMT-Benchmark37.9%SR

Reasoning

GPQANYU + Cohere + Anthropic (2023)90.4%SR
BrowseCompOpenAI (2025)84.2%SR
MCP Atlas79.1%SR
SWE-Bench Verified78.0%SR
SWE-bench Multilingual75.8%SR
FrontierScience Olympiad74.8%SR
AA-LCR73.4%SR
Terminal-Bench 2.171.7%SR
SWE-Bench ProPrinceton NLP (2024)57.9%SR
Humanity's Last Exam (with tools, text-only)53.2%SR
Humanity's Last Exam (no tools, text-only)47.0%SR
APEX-Agents25.6%SR
FrontierScience Research21.3%SR

Search

DeepSearchQA91.0%SR
WideSearch76.4%SR

AA Evaluation Indices

(Artificial Analysis)
Gpqa(NYU + Cohere + Anthropic (2023))
89.7
Lcr(Artificial Analysis)
79.0
Terminalbench V2 1
64.4
Coding Index(Artificial Analysis)
58.8
Scicode(UIUC + Argonne National Lab (2024))
48.6
Hle(Center for AI Safety + Scale AI (2025))
33.5
Intelligence Index(Artificial Analysis)
25.3
Tau Banking
22.9
Terminalbench V4 0
0.5

LLM Stats Category Scores

(LLM Stats (zeroeval))
Math
4
Reasoning
2
General
2
Biology
90
Physics
80
Search
80
Frontend Development
80
Long Context
70
Chemistry
70
Tool Calling
70
Science
60
Agents
60
Code
60
Knowledge
50
Coding
50

Pricing

Input Price$0.136 / 1M tokens
Output Price$0.555 / 1M tokens
Blended Price (3:1)$0.241 / 1M tokens
Cache Read Price$0.033 / 1M tokens

Speed

Tokens/sec87.3
Time to First Token2.07s
Time to Answer24.97s

Provider Price Ranking

Provider Price Ranking

14 providers

Cheapest: DeepInfraMost Expensive: OrcaRouter
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2NanoGPT
$0.066
$0.26
3Kilo Gateway
$0.13
$0.53
4OpenRouter
$0.132
$0.528
5DevPass (LLM Gateway)
$0.132
$0.528
6LLM Gateway
$0.132
$0.528
7TencentPRIMARY
$0.136
$0.555
8OpenCode Go
$0.14
$0.58
9Requesty
$0.14
$0.58
10Vercel AI Gateway
$0.14
$0.58
11Jalapeno Cloud
$0.14
$0.58
12AIHubMix
$0.1562
$0.6248
13CrossModel
$0.16
$0.64
14OrcaRouter
$0.18
$0.59

Compare pricing across different API providers for this model.

External Sources