Inkling Small
描述
Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning, variable thinking effort, and a context window up to 1M tokens (Tinker exposes 64K and 256K configurations). Hugging Face weights: thinkingmachines/Inkling-Small and thinkingmachines/Inkling-Small-NVFP4. Vendor self-host VRAM: BF16 ≥ ~600 GB aggregated; NVFP4 ≥ ~180 GB aggregated. SWE-Bench Verified 80.2% vs Inkling 77.6% (same bash-only harness). Fine-tuning and playground chat are available via Tinker. Official Tinker serverless inference (256K, list): $0.30 / $1.20 per 1M input/output tokens ($0.06 cached).
能力雷達圖
排行榜排名
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))Agents
Audio
Chat
Factuality
General
Instruction Following
Math
Reasoning
Vision
AA 評測指數
(Artificial Analysis)LLM Stats 分類評分
(LLM Stats (zeroeval))定價
速度
供應商價格排行
供應商價格排行
8 個供應商
比較該模型在不同 API 供應商之間的定價。