Fast multi-region endpoint for open-weight models. Zero data retention, zero training, and optimized for long-context workloads.
Test low-latency completions across served coding models with live token counters and guaranteed volatile VRAM processing.
Simulate instantaneous token generation across NullStack hardware clusters.
42 ms
~182 tok/s
0 Bytes Logged
5.3 TB/s CDNA3
Engineered with strict privacy boundaries from the network ingress to the hardware accelerator.
Multi-region (EU & US)
Every inference request is authenticated and routed by our LiteLLM gateway in the European Union. You select whether model inference runs in Europe, the US, or either region on a per-key basis.
Zero by default
We do not log or retain prompt text, completions, or request bodies. Content exists only in volatile GPU memory during active token generation. Logging is physically disabled at the gateway.
Never
Your prompts, completions, and code are never used for model training, fine-tuning, evaluation datasets, or product analytics. Upstream providers are bound by strict non-retention agreements.
Nullstack runs on custom compression and attention kernels developed by NullStack.in, optimized specifically for agentic coding patterns with large system prompts, multiple file context windows, and rapid iterations.
Rate-distortion optimal quantization achieving 3.2-bit effective weights with zero loss in coding precision.
Handcrafted multi-query verification kernels unlocking near-theoretical peak memory bandwidth on AMD MI300X and NVIDIA Hopper.
Dedicated high-throughput deployments across AMD MI300X, NVIDIA H100 Hopper, and Google Cloud TPU v5p pods.
The Nullstack CLI launcher automatically injects ephemeral authentication, routes requests through private endpoints, and disables upstream telemetry.
# 1. Install Nullstack CLI globally$npm install -g @nullstack-ai/nullstack# 2. Login to your account$nullstack login# 3. Launch your favorite coding agent with private zero-retention endpoint$nullstack launch claude
npm install -g @nullstack-ai/nullstackView full CLI referenceTransparent per-token rates with generous prefix caching discounts.
glm-5.2Z.ai's leading open-weight coding powerhouse. Highly performant across complex multi-file refactoring, agent tool calls, and deep algorithmic synthesis.
$1.10
$4.00
deepseek-v4-flash-0731Extremely cost-effective and lightning-fast inference for high-frequency agent loops, automated code reviews, and large repository scanning.
$0.14
$0.28
kimi-k3State-of-the-art long-context reasoning with deep architectural multi-turn memory. Excels at whole-codebase comprehension and autonomous debugging.
$2.50
$12.00
laguna-2.1Next-generation 2M context inference engine featuring non-linear lattice quantization and sub-millisecond prefix caching. Currently in private benchmarking.
—
—
Staff Infrastructure Engineer • DevOps Core
“Nullstack's GLM-5.2 endpoint is ridiculously fast. Our Claude Code multi-turn loops complete in a third of the time compared to generic cloud gateways.”
Head of Security • FinScale
“Zero data retention wasn't just a nice-to-have for our enterprise security compliance—it was mandatory. Nullstack made open models viable for our core repositories.”
Founder & CTO • Synthetix AI
“DeepSeek V4 Flash at $0.14 / 1M input tokens paired with sub-50ms TTFT is an absolute gamechanger for high-frequency agent scanning.”
Principal Systems Architect • VectorScale
“The CLI launcher is brilliant. Literally `nullstack launch claude` and I'm pair-programming against high-spec MI300X clusters without touching API configs.”
AI Research Lead • KernelLabs
“NullStack's custom HIP kernels squeeze every ounce of memory bandwidth out of AMD silicon. The speed difference in 100k+ token sessions is night and day.”
VP of Engineering • NordicTech
“We migrated our Cursor team seats to Nullstack's EU endpoint. Lower latency, deterministic pricing, and guaranteed privacy for all developer workstations.”
Staff Infrastructure Engineer • DevOps Core
“Nullstack's GLM-5.2 endpoint is ridiculously fast. Our Claude Code multi-turn loops complete in a third of the time compared to generic cloud gateways.”
Head of Security • FinScale
“Zero data retention wasn't just a nice-to-have for our enterprise security compliance—it was mandatory. Nullstack made open models viable for our core repositories.”
Founder & CTO • Synthetix AI
“DeepSeek V4 Flash at $0.14 / 1M input tokens paired with sub-50ms TTFT is an absolute gamechanger for high-frequency agent scanning.”
Principal Systems Architect • VectorScale
“The CLI launcher is brilliant. Literally `nullstack launch claude` and I'm pair-programming against high-spec MI300X clusters without touching API configs.”
AI Research Lead • KernelLabs
“NullStack's custom HIP kernels squeeze every ounce of memory bandwidth out of AMD silicon. The speed difference in 100k+ token sessions is night and day.”
VP of Engineering • NordicTech
“We migrated our Cursor team seats to Nullstack's EU endpoint. Lower latency, deterministic pricing, and guaranteed privacy for all developer workstations.”
Everything you need to know about Nullstack's architecture, privacy, and billing.
Get an instant API key, configure your inference region, and pair program with zero request retention.