Upstream UpdateNullstack contributes Spectrum diffusion & CDNA3 attention upstream to SGLangRead PR log
Nullstackby NullStack
Open-Model Inference

Private inference for coding agents.

Fast multi-region endpoint for open-weight models. Zero data retention, zero training, and optimized for long-context workloads.

Code: Private
Models: Open Weights
Setup: Minutes
Inference: Multi-region
Nullstack Core0 Retention
Live Playground

Experience the speed in real time.

Test low-latency completions across served coding models with live token counters and guaranteed volatile VRAM processing.

Live Inference Simulator

Zero Retention Active

Simulate instantaneous token generation across NullStack hardware clusters.

Sample Prompt:
model: glm-5.2 • region: EU Sovereign
Time to First Token

42 ms

Generation Throughput

~182 tok/s

Data Retention

0 Bytes Logged

VRAM Contraction

5.3 TB/s CDNA3

Sovereign Architecture

Multi-Region Infrastructure

Engineered with strict privacy boundaries from the network ingress to the hardware accelerator.

Location

Multi-region (EU & US)

Every inference request is authenticated and routed by our LiteLLM gateway in the European Union. You select whether model inference runs in Europe, the US, or either region on a per-key basis.

Retention

Zero by default

We do not log or retain prompt text, completions, or request bodies. Content exists only in volatile GPU memory during active token generation. Logging is physically disabled at the gateway.

Training

Never

Your prompts, completions, and code are never used for model training, fine-tuning, evaluation datasets, or product analytics. Upstream providers are bound by strict non-retention agreements.

NullStack Systems

Tuned for long-context, multi-turn coding sessions.

Nullstack runs on custom compression and attention kernels developed by NullStack.in, optimized specifically for agentic coding patterns with large system prompts, multiple file context windows, and rapid iterations.

HyperQuant Matrix Compression

Rate-distortion optimal quantization achieving 3.2-bit effective weights with zero loss in coding precision.

Custom Attention Kernels (HIP / CDNA3)

Handcrafted multi-query verification kernels unlocking near-theoretical peak memory bandwidth on AMD MI300X and NVIDIA Hopper.

Hardware-Aware Acceleration

Dedicated high-throughput deployments across AMD MI300X, NVIDIA H100 Hopper, and Google Cloud TPU v5p pods.

NULLSTACK SYSTEM STACKv2.4 Kernel
01Agent Layer
Claude Code • Codex • Cursor • Cline
02API & Router Layer
LiteLLM Gateway • EU Sovereign Ingress
03Privacy Layer
Zero Retention Engine • Ephemeral VRAM
04Compression Layer
HyperQuant 3.2b Lattice Quantization
05Kernel Layer
NullStack MLA • Fused Sparse Attention
06Hardware Silicon
AMD MI300X • NVIDIA Hopper • TPU v5p
Developer Quickstart

One command to run your coding agents privately.

The Nullstack CLI launcher automatically injects ephemeral authentication, routes requests through private endpoints, and disables upstream telemetry.

nullstack-terminal ~ zsh
# 1. Install Nullstack CLI globally
$npm install -g @nullstack-ai/nullstack
# 2. Login to your account
$nullstack login
# 3. Launch your favorite coding agent with private zero-retention endpoint
$nullstack launch claude
Zero Data Retention Guaranteed EU Gateway Routed Telemetry Blocked
Install globally: npm install -g @nullstack-ai/nullstackView full CLI reference
Served Models

Open models tuned for coding agents.

Transparent per-token rates with generous prefix caching discounts.

Artificial Analysis coding benchmarks
Z.aiActive Endpoint

GLM-5.2

glm-5.2

Z.ai's leading open-weight coding powerhouse. Highly performant across complex multi-file refactoring, agent tool calls, and deep algorithmic synthesis.

Context Window524K
~87 code filesZero VRAM retention
Input / 1M

$1.10

Output / 1M

$4.00

DeepSeekActive Endpoint

DeepSeek V4 Flash 0731

deepseek-v4-flash-0731

Extremely cost-effective and lightning-fast inference for high-frequency agent loops, automated code reviews, and large repository scanning.

Context Window1M
~167 code filesZero VRAM retention
Input / 1M

$0.14

Output / 1M

$0.28

Moonshot AIActive Endpoint

Kimi K3

kimi-k3

State-of-the-art long-context reasoning with deep architectural multi-turn memory. Excels at whole-codebase comprehension and autonomous debugging.

Context Window1M
~167 code filesZero VRAM retention
Input / 1M

$2.50

Output / 1M

$12.00

NullStackComing Soon

Laguna 2.1

laguna-2.1

Next-generation 2M context inference engine featuring non-linear lattice quantization and sub-millisecond prefix caching. Currently in private benchmarking.

Context Window2M
~333 code filesZero VRAM retention
Input / 1M

Output / 1M

Developer Proof

Trusted by engineers building with coding agents.

AM

Alexandre Mercier

Staff Infrastructure Engineer • DevOps Core

GLM-5.2

Nullstack's GLM-5.2 endpoint is ridiculously fast. Our Claude Code multi-turn loops complete in a third of the time compared to generic cloud gateways.

ER

Elena Rostova

Head of Security • FinScale

Kimi K3

Zero data retention wasn't just a nice-to-have for our enterprise security compliance—it was mandatory. Nullstack made open models viable for our core repositories.

MC

Marcus Chen

Founder & CTO • Synthetix AI

DeepSeek V4 Flash

DeepSeek V4 Flash at $0.14 / 1M input tokens paired with sub-50ms TTFT is an absolute gamechanger for high-frequency agent scanning.

SN

Siddharth Nair

Principal Systems Architect • VectorScale

Claude Code via Nullstack

The CLI launcher is brilliant. Literally `nullstack launch claude` and I'm pair-programming against high-spec MI300X clusters without touching API configs.

JV

Dr. Julian Vance

AI Research Lead • KernelLabs

NullStack MLA

NullStack's custom HIP kernels squeeze every ounce of memory bandwidth out of AMD silicon. The speed difference in 100k+ token sessions is night and day.

SL

Sofia Lindqvist

VP of Engineering • NordicTech

Cursor + Nullstack

We migrated our Cursor team seats to Nullstack's EU endpoint. Lower latency, deterministic pricing, and guaranteed privacy for all developer workstations.

AM

Alexandre Mercier

Staff Infrastructure Engineer • DevOps Core

GLM-5.2

Nullstack's GLM-5.2 endpoint is ridiculously fast. Our Claude Code multi-turn loops complete in a third of the time compared to generic cloud gateways.

ER

Elena Rostova

Head of Security • FinScale

Kimi K3

Zero data retention wasn't just a nice-to-have for our enterprise security compliance—it was mandatory. Nullstack made open models viable for our core repositories.

MC

Marcus Chen

Founder & CTO • Synthetix AI

DeepSeek V4 Flash

DeepSeek V4 Flash at $0.14 / 1M input tokens paired with sub-50ms TTFT is an absolute gamechanger for high-frequency agent scanning.

SN

Siddharth Nair

Principal Systems Architect • VectorScale

Claude Code via Nullstack

The CLI launcher is brilliant. Literally `nullstack launch claude` and I'm pair-programming against high-spec MI300X clusters without touching API configs.

JV

Dr. Julian Vance

AI Research Lead • KernelLabs

NullStack MLA

NullStack's custom HIP kernels squeeze every ounce of memory bandwidth out of AMD silicon. The speed difference in 100k+ token sessions is night and day.

SL

Sofia Lindqvist

VP of Engineering • NordicTech

Cursor + Nullstack

We migrated our Cursor team seats to Nullstack's EU endpoint. Lower latency, deterministic pricing, and guaranteed privacy for all developer workstations.

FAQ

Frequently Asked Questions

Everything you need to know about Nullstack's architecture, privacy, and billing.

Nullstack is a high-speed, private inference platform specifically tuned for coding agents and developer tooling. Built on custom NullStack kernels and quantization algorithms, it delivers low-latency open-weight model completions across EU and US regions with guaranteed zero data retention and zero model training.

Start running private coding agents today.

Get an instant API key, configure your inference region, and pair program with zero request retention.