Nullstack is built from the ground up by NullStack.in—a team of mathematicians and systems engineers designing low-level algorithms, custom attention kernels, and hardware-aware compression.
Click each layer to explore our end-to-end inference pipeline from developer agents down to silicon.
Standard OpenAI and Anthropic client SDK adapters, CLI subshell launchers, and IDE plugins (Claude Code, Cursor, Codex CLI, Cline, OpenCode, Hermes, OpenClaw, Pi).
Deep dives on attention kernels, GPU memory hierarchies, and quantization algorithms.
Live record of upstream PRs to SGLang including Spectrum diffusion acceleration, JAX profiling, and Kimi-K3 on CDNA3.
How we achieved near-peak memory bandwidth on AMD CDNA3 with clean C++ HIP code, beating hand-tuned assembly.
Achieving 3.2-bit effective weights with zero perplexity degradation on 70B+ open models.
Benchmarking NullStack MLA against upstream backends on AMD gfx942 CDNA3 hardware with multi-head latent verification.