Gemma 3 12B Instruct
GoogleGemmaOpen WeightGemma · Commercial OK
Description
Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.
Release Date
2025-03-12
Parameters
12.0B
Context Length
131K
Modalities
image, text
Capability Radar
21
general
10
coding
31
reasoning
22
science
25
agents
80
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 598 | 8.0 | AA |
| General Ranking | 564 | 22.0 | AA |
| Multimodal Ranking | 119 | 38.0 | LS |
| Science | 598 | 17.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Chat
IFEvalGoogle Research (2023)
88.9%SR
Factuality
SimpleQA
6.3%SR
General
Global-MMLU-Lite
69.5%SR
Language
MMLU-Pro
60.6%SR
WMT24++
51.6%SR
ECLeKTic
10.3%SR
Math
GSM8k
94.4%SR
MATH
83.8%SR
MathVista-Mini
62.9%SR
HiddenMath
54.5%SR
Reasoning
BIG-Bench Hard
85.7%SR
HumanEvalOpenAI (2021)
85.4%SR
Natural2Code
80.7%SR
FACTS Grounding
75.8%SR
ChartQAMasry et al. (2022)
75.7%SR
MBPP
0.73 / 100SR
Bird-SQL (dev)
47.9%SR
GPQANYU + Cohere + Anthropic (2023)
40.9%SR
LiveCodeBench
24.6%SR
BIG-Bench Extra Hard
16.3%SR
Vision
DocVQADocVQA (2020)
87.1%SR
AI2D
84.2%SR
VQAv2 (val)
71.6%SR
TextVQA
67.7%SR
InfoVQA
64.9%SR
MMMU (val)
59.6%SR
AA Evaluation Indices
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))85.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))59.5
Ifbench(Google Research (2023))36.7
Gpqa(NYU + Cohere + Anthropic (2023))34.9
Aime(MAA (Mathematical Association of America))22.0
Aime 25(MAA (Mathematical Association of America))18.3
Math Index(Artificial Analysis)18.3
Scicode(UIUC + Argonne National Lab (2024))16.4
Livecodebench(UC Berkeley + MIT + Cornell (2024))13.7
Tau2(Sierra + U Toronto + Vector Institute (2025))10.8
Lcr(Artificial Analysis)8.3
Coding Index(Artificial Analysis)5.8
Hle(Center for AI Safety + Scale AI (2025))4.2
Intelligence Index(Artificial Analysis)3.8
Tau Banking0.8
Terminalbench Hard(Stanford × Laude Institute (2026))0.8
Terminalbench V2 10.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Chat90
Instruction Following90
Structured Output90
Image To Text80
Grounding80
Math70
Multimodal70
Vision70
Legal60
Reasoning60
Finance60
General60
Healthcare60
Code60
Language50
Physics40
Factuality40
Biology40
Chemistry40
Pricing
Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Speed
Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s
Provider Price Ranking
Provider Price Ranking
2 providers
Cheapest: DeepInfraMost Expensive: Neon
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2Neon
$0.15
$0.5
Compare pricing across different API providers for this model.