メインコンテンツへスキップ

North Micro Vision Instruct

CohereオープンウエイトApache 2.0 · 商用利用可

説明

Compact open-weight vision-language model (2.4B) with native-resolution image support for VQA, captioning, grounding, OCR, charts, and documents. Custom 400M SigLIP 2-based vision encoder + 2B North Micro LLM backbone (Command A+ style). Multilingual and multi-image. LM context 128K; multimodal training validated to 8K. Not a reasoning/tool-calling model. Intended for prototyping and fine-tuning.

リリース日
2026-08-12
パラメータ
2.4B
コンテキスト長
—
モダリティ
—

能力レーダー

60
general
0
coding
40
reasoning
34
science推定
28
agents
70
multimodal

専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。

ランキング

ドメイン#順位スコアソース
マルチモーダルランキング148
12.0
LS

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Chat

IFEvalGoogle Research (2023)74.9%自己申告
Multi-IF37.3%自己申告

General

MMLU50.4%自己申告

Language

MMLU-Pro30.7%自己申告

Reasoning

ChartQAMasry et al. (2022)80.8%自己申告
CharXiv-D60.0%自己申告

Vision

DocVQADocVQA (2020)92.1%自己申告
OCRBench79.2%自己申告
AI2D77.5%自己申告
RefCOCO-avg0.73 / 100自己申告
CountBench0.72 / 100自己申告
MMBench-V1.168.7%自己申告
InfoVQA65.2%自己申告
RealWorldQA62.2%自己申告
Hallusion Bench61.5%自己申告
BLINK52.7%自己申告
MMStar51.8%自己申告
OCRBench-V2 (en)36.7%自己申告
MMMU (val)32.9%自己申告

AA評価指数

(Artificial Analysis)

AA評価データがありません

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Image To Text
70
Spatial Reasoning
70
Grounding
70
Chat
60
Instruction Following
60
Multimodal
60
Reasoning
60
Structured Output
60
General
60
Vision
60
3d
50
Language
40
Legal
40
Math
40
Finance
40
Healthcare
40
Communication
40

価格設定

価格データがありません

速度

速度データがありません

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク