MM
NullStack Research
@nullstack_in
Low-level algorithms and hardware-aware systems power private, high-throughput AI inference for coding agents. Stay tuned as we release further benchmarks.
Achieving 3.2-bit effective weights with zero perplexity degradation on 70B+ open models using non-uniform lattice vector quantization.
@nullstack_in
Low-level algorithms and hardware-aware systems power private, high-throughput AI inference for coding agents. Stay tuned as we release further benchmarks.