Upstream UpdateNullstack contributes Spectrum diffusion & CDNA3 attention upstream to SGLangRead PR log
Nullstackby NullStack
Back to all posts
ResearchSGLang

Optimizing DeepSeek Sparse MLA for 1M Token Contexts

April 14, 2026โ€ขBy NullStack Research

Hardware-aware indexing and cache budgeting for sparse multi-head latent attention kernels.

MM

NullStack Research

@nullstack_in

๐•

Low-level algorithms and hardware-aware systems power private, high-throughput AI inference for coding agents. Stay tuned as we release further benchmarks.