Inkling Small
Descripción
Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning, variable thinking effort, and a context window up to 1M tokens (Tinker exposes 64K and 256K configurations). Hugging Face weights: thinkingmachines/Inkling-Small and thinkingmachines/Inkling-Small-NVFP4. Vendor self-host VRAM: BF16 ≥ ~600 GB aggregated; NVFP4 ≥ ~180 GB aggregated. SWE-Bench Verified 80.2% vs Inkling 77.6% (same bash-only harness). Fine-tuning and playground chat are available via Tinker. Official Tinker serverless inference (256K, list): $0.30 / $1.20 per 1M input/output tokens ($0.06 cached).
Radar de capacidades
Rankings
| Dominio | #Posición | Puntuación | Fuente |
|---|---|---|---|
| Capacidad agéntica | 43 | 49.0 | LS |
| Ranking de codificación | 127 | 72.0 | AA |
| Ranking general | 260 | 45.0 | AA |
| Ranking multimodal | 82 | 49.0 | LS |
| Ciencia | 90 | 73.0 | AA |
Puntuaciones de benchmarks (LLM Stats)
(LLM Stats (zeroeval))Agents
Audio
Chat
Factuality
General
Instruction Following
Math
Reasoning
Vision
Índices de evaluación AA
(Artificial Analysis)Puntuaciones por categoría LLM Stats
(LLM Stats (zeroeval))Precios
Velocidad
Ranking de Precios por Proveedor
Ranking de Precios por Proveedor
10 proveedores
Comparar precios entre diferentes proveedores de API para este modelo.