跳转到主要内容

North Micro Vision Instruct

Cohere开源权重Apache 2.0 · 商用许可

描述

Compact open-weight vision-language model (2.4B) with native-resolution image support for VQA, captioning, grounding, OCR, charts, and documents. Custom 400M SigLIP 2-based vision encoder + 2B North Micro LLM backbone (Command A+ style). Multilingual and multi-image. LM context 128K; multimodal training validated to 8K. Not a reasoning/tool-calling model. Intended for prototyping and fine-tuning.

发布日期
2026-08-12
参数规模
2.4B
上下文长度
支持模态

能力雷达图

60
general
0
coding
40
reasoning
34
science估算
28
agents
70
multimodal

缺少专门科学评测时,Science 由 LLM Stats 科学得分或推理能力估算。

排行榜排名

领域#排名分数来源
多模态榜86
18.0
LS

基准测试分数 (LLM Stats)

(LLM Stats (zeroeval))

3d

BLINK52.7%自报

Communication

Multi-IF37.3%自报

Finance

MMLU50.4%自报
MMLU-Pro30.7%自报

General

IFEvalGoogle Research (2023)74.9%自报
MMStar51.8%自报
MMMU (val)32.9%自报

Grounding

RefCOCO-avg0.73 / 100自报

Image To Text

DocVQADocVQA (2020)92.1%自报
OCRBench79.2%自报
OCRBench-V2 (en)36.7%自报

Multimodal

ChartQAMasry et al. (2022)80.8%自报
AI2D77.5%自报
MMBench-V1.168.7%自报
InfoVQA65.2%自报
CharXiv-D60.0%自报

Reasoning

CountBench0.72 / 100自报
Hallusion Bench61.5%自报

Spatial Reasoning

RealWorldQA62.2%自报

AA 评测指数

(Artificial Analysis)

暂无 AA 评测数据

LLM Stats 分类评分

(LLM Stats (zeroeval))
Spatial Reasoning
70
Image To Text
70
Grounding
70
Multimodal
60
Reasoning
60
Structured Output
60
Instruction Following
60
General
60
Vision
60
3d
50
Legal
40
Math
40
Language
40
Finance
40
Healthcare
40
Communication
40

定价

暂无定价数据

速度

暂无速度数据

供应商价格排行

暂无提供商数据

外部链接