Skip to main content

Gemma 3 12B Instruct

GoogleGemmaOpen WeightGemma · Commercial OK

Description

Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

Release Date
2025-03-12
Parameters
12.0B
Context Length
131K
Modalities
image, text

Capability Radar

21
general
10
coding
31
reasoning
22
science
25
agents
80
multimodal

Rankings

Domain#RankScoreSource
Code Ranking598
8.0
AA
General Ranking564
22.0
AA
Multimodal Ranking119
38.0
LS
Science598
17.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Chat

IFEvalGoogle Research (2023)88.9%SR

Factuality

SimpleQA6.3%SR

General

Global-MMLU-Lite69.5%SR

Language

MMLU-Pro60.6%SR
WMT24++51.6%SR
ECLeKTic10.3%SR

Math

GSM8k94.4%SR
MATH83.8%SR
MathVista-Mini62.9%SR
HiddenMath54.5%SR

Reasoning

BIG-Bench Hard85.7%SR
HumanEvalOpenAI (2021)85.4%SR
Natural2Code80.7%SR
FACTS Grounding75.8%SR
ChartQAMasry et al. (2022)75.7%SR
MBPP0.73 / 100SR
Bird-SQL (dev)47.9%SR
GPQANYU + Cohere + Anthropic (2023)40.9%SR
LiveCodeBench24.6%SR
BIG-Bench Extra Hard16.3%SR

Vision

DocVQADocVQA (2020)87.1%SR
AI2D84.2%SR
VQAv2 (val)71.6%SR
TextVQA67.7%SR
InfoVQA64.9%SR
MMMU (val)59.6%SR

AA Evaluation Indices

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
85.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
59.5
Ifbench(Google Research (2023))
36.7
Gpqa(NYU + Cohere + Anthropic (2023))
34.9
Aime(MAA (Mathematical Association of America))
22.0
Aime 25(MAA (Mathematical Association of America))
18.3
Math Index(Artificial Analysis)
18.3
Scicode(UIUC + Argonne National Lab (2024))
16.4
Livecodebench(UC Berkeley + MIT + Cornell (2024))
13.7
Tau2(Sierra + U Toronto + Vector Institute (2025))
10.8
Lcr(Artificial Analysis)
8.3
Coding Index(Artificial Analysis)
5.8
Hle(Center for AI Safety + Scale AI (2025))
4.2
Intelligence Index(Artificial Analysis)
3.8
Tau Banking
0.8
Terminalbench Hard(Stanford × Laude Institute (2026))
0.8
Terminalbench V2 1
0.0

LLM Stats Category Scores

(LLM Stats (zeroeval))
Chat
90
Instruction Following
90
Structured Output
90
Image To Text
80
Grounding
80
Math
70
Multimodal
70
Vision
70
Legal
60
Reasoning
60
Finance
60
General
60
Healthcare
60
Code
60
Language
50
Physics
40
Factuality
40
Biology
40
Chemistry
40

Pricing

Input PriceFree
Output PriceFree
Blended Price (3:1)Free

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

2 providers

Cheapest: DeepInfraMost Expensive: Neon
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2Neon
$0.15
$0.5

Compare pricing across different API providers for this model.

External Sources