Inkling Small
설명
Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning, variable thinking effort, and a context window up to 1M tokens (Tinker exposes 64K and 256K configurations). Hugging Face weights: thinkingmachines/Inkling-Small and thinkingmachines/Inkling-Small-NVFP4. Vendor self-host VRAM: BF16 ≥ ~600 GB aggregated; NVFP4 ≥ ~180 GB aggregated. SWE-Bench Verified 80.2% vs Inkling 77.6% (same bash-only harness). Fine-tuning and playground chat are available via Tinker. Official Tinker serverless inference (256K, list): $0.30 / $1.20 per 1M input/output tokens ($0.06 cached).
능력 레이더
랭킹
벤치마크 점수 (LLM Stats)
(LLM Stats (zeroeval))Agents
Audio
Chat
Factuality
General
Instruction Following
Math
Reasoning
Vision
AA 평가 지수
(Artificial Analysis)LLM Stats 카테고리 점수
(LLM Stats (zeroeval))가격
속도
공급자 가격 순위
공급자 가격 순위
8개 공급자
이 모델의 다양한 API 공급자 간 가격 비교.