Skip to main content

Gemma 3 4B Instruct

GoogleGemmaOpen WeightGemma · Commercial OK

Description

Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

Release Date
2025-03-12
Parameters
4.0B
Context Length
131K
Modalities
image, text

Capability Radar

16
general
6
coding
22
reasoning
22
science
17
agents
70
multimodal

Rankings

Domain#RankScoreSource
Code Ranking616
5.0
AA
General Ranking631
15.0
AA
Multimodal Ranking133
32.0
LS
Science623
15.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Chat

IFEvalGoogle Research (2023)90.2%SR

Factuality

SimpleQA4.0%SR

General

Global-MMLU-Lite54.5%SR

Language

WMT24++46.8%SR
MMLU-Pro43.6%SR
ECLeKTic4.6%SR

Math

GSM8k89.2%SR
MATH75.6%SR
MathVista-Mini50.0%SR
HiddenMath43.0%SR

Reasoning

BIG-Bench Hard72.2%SR
HumanEvalOpenAI (2021)71.3%SR
Natural2Code70.3%SR
FACTS Grounding70.1%SR
ChartQAMasry et al. (2022)68.8%SR
MBPP0.63 / 100SR
Bird-SQL (dev)36.3%SR
GPQANYU + Cohere + Anthropic (2023)30.8%SR
LiveCodeBench12.6%SR
BIG-Bench Extra Hard11.0%SR

Vision

DocVQADocVQA (2020)75.8%SR
AI2D74.8%SR
VQAv2 (val)62.4%SR
TextVQA57.8%SR
InfoVQA50.0%SR
MMMU (val)48.8%SR

AA Evaluation Indices

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
76.6
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
41.7
Gpqa(NYU + Cohere + Anthropic (2023))
29.1
Ifbench(Google Research (2023))
28.3
Math Index(Artificial Analysis)
12.7
Aime 25(MAA (Mathematical Association of America))
12.7
Livecodebench(UC Berkeley + MIT + Cornell (2024))
11.2
Lcr(Artificial Analysis)
6.7
Aime(MAA (Mathematical Association of America))
6.3
Hle(Center for AI Safety + Scale AI (2025))
5.3
Tau2(Sierra + U Toronto + Vector Institute (2025))
5.0
Intelligence Index(Artificial Analysis)
4.8
Coding Index(Artificial Analysis)
2.7
Terminalbench Hard(Stanford × Laude Institute (2026))
0.8
Tau Banking
0.4
Terminalbench V2 1
0.4

LLM Stats Category Scores

(LLM Stats (zeroeval))
Chat
90
Instruction Following
90
Structured Output
90
Image To Text
70
Grounding
70
Math
60
Multimodal
60
Vision
60
Reasoning
50
General
50
Healthcare
50
Language
40
Legal
40
Factuality
40
Finance
40
Code
40
Physics
30
Biology
30
Chemistry
30

Pricing

Input PriceFree
Output PriceFree
Blended Price (3:1)Free

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

6 providers

Cheapest: DeepInfraMost Expensive: Kilo Gateway
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2Merge Gateway
$0.04
$0.08
3OpenRouter
$0.05
$0.1
4Hugging Face
$0.05
$0.1
5Deep Infra
$0.05
$0.1
6Kilo Gateway
$0.05
$0.1

Compare pricing across different API providers for this model.

External Sources