MM
NullStack Research
@nullstack_in
Low-level algorithms and hardware-aware systems power private, high-throughput AI inference for coding agents. Stay tuned as we release further benchmarks.
W4A16 and W8A16 matrix multiplication kernels designed to bypass dequantization overhead on NVIDIA Hopper and AMD CDNA3.
@nullstack_in
Low-level algorithms and hardware-aware systems power private, high-throughput AI inference for coding agents. Stay tuned as we release further benchmarks.