MiniCPM-SALA
Descripción
MiniCPM-SALA (Sparse Attention and Linear Attention) is a 9B hybrid model built from a MiniCPM-4.0 checkpoint via continual training (~2T tokens, 25% of training-from-scratch cost). It interleaves 25% InfLLM-V2 sparse attention and 75% Lightning Attention layers, achieving up to 3.5x inference speed over dense baselines at 256K tokens. With HyPE (Hybrid Positional Encoding) and NoPE in sparse layers, the model extrapolates to 2048K tokens despite a 520K training length, enabling 1M-token inference on consumer GPUs like the RTX 5090.
Radar de capacidades
Science se estima a partir de las puntuaciones científicas de LLM Stats o del razonamiento cuando no hay benchmarks científicos dedicados.
Rankings
No hay datos de ranking disponibles
Puntuaciones de benchmarks (LLM Stats)
(LLM Stats (zeroeval))Code
Finance
General
Language
Long Context
Math
Índices de evaluación AA
(Artificial Analysis)No hay datos de evaluación AA disponibles
Puntuaciones por categoría LLM Stats
(LLM Stats (zeroeval))Precios
No hay datos de precios disponibles
Velocidad
No hay datos de velocidad disponibles
Ranking de Precios por Proveedor
No hay datos de proveedores disponibles