Design a high-throughput fused W4A16 GEMM kernel targeting NVIDIA H100 Hopper Tensor Cores.
Maximize asynchronous Tensor Memory Accelerator (TMA) and warp-group matrix multiply and accumulate (WGMMA) instructions to eliminate memory stall cycles during 4-bit weight dequantization.
Weight matrix W is stored in packed 4-bit format with FP8/FP16 group scales.
Activation matrix A is stored in FP16 row-major layout.
Accumulate results in FP32 and store output in FP16 format.