蚂蚁 inclusionAI:GitHub 新仓库· 蚂蚁 inclusionAI:GitHub 新仓库·· 2026-02-11精选AI 评分61
inclusionAI 发布高性能量化推理 GEMM 内核库 Humming
AI 导读
inclusionAI 开源了 Humming,这是一个专为量化推理设计的高性能、轻量级即时编译 GEMM 内核库。它支持在 FP16、BF16、FP8 等多种激活数据类型下进行 8 比特以下任意权重类型的推理,兼容多种量化策略与缩放类型,并同时支持稠密 GEMM 和混合专家 GEMM 运算。该库兼容 SM75+ 及以上的所有 NVIDIA GPU,在多种计算场景下能提供业界领先的吞吐量和效率。其依赖极简,仅需 PyTorch 和 NVCC,软件包大小仅约 100 KB,便于超轻量化部署。
推荐理由
蚂蚁 inclusionAI 开源了一个 100KB 级的量化 GEMM 库,支持从 INT1 到 FP8 全家桶,SM75+ 全覆盖,做推理部署的工程师值得花半小时跑一下 benchmark,看看能不能替换掉现有的 Marlin 方案。
正文
Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.
Key Features
- High Flexibility
- Supports inference for any weight type under 8-bit across FP16 / BF16 / FP8 / FP4 / INT8 / INT4 activations (provided the activation's dynamic range covers the weight type).
- Supports various quantization strategies.
- Supports various scale types (BF16, FP16, E4M3, E5M2, and UE8M0).
- Supports both Dense GEMM and MoE GEMM.
- High Compatibility: supports all NVIDIA GPUs from SM75+ (Turing architecture) and beyond.
- High Performance
- Delivers State-of-the-Art (SOTA) throughput and efficiency across a wide range of computational scenarios.
- Ultra-Lightweight
- Minimal dependencies: Requires only PyTorch and NVCC.
- Compact footprint: The package size is only 100+KB.
Support Matrix
| Activation Type | Supported Devices | Supported Weight Types |
|---|---|---|
| FP16 (e5m10) | SM75+ | • Symmetric INT1-8 • INT1-8 with dynamic zero point • Arbitrary signed FP (kBits ≤ 8, kExp ≤ 5) |
| BF16 (e8m7) | SM80+ | • Symmetric INT1-8 • INT1-8 with dynamic zero point • Arbitrary signed FP (kBits ≤ 8) |
| FP8 (e4m3) | SM89+ | • Symmetric INT1-5 • INT1-4 with dynamic zero point • Arbitrary signed FP (kExp ≤ 4, kMan ≤ 3) |
| FP8 (e5m2) | SM89+ | • Symmetric INT1-4 • INT1-3 with dynamic zero point • Arbitrary signed FP (kExp ≤ 5, kMan ≤ 2) |
| FP4 (e2m1) | SM120+ | • Symmetric INT1-3 • INT1-2 with dynamic zero point • Arbitrary signed FP (kExp ≤ 2, kMan ≤ 1) |
| INT8 | SM75+ | • Symmetric INT1-8 • INT1-7 with dynamic zero point |
| INT4 | SM80+ | • Symmetric INT1-4 • INT1-3 with dynamic zero point |
Getting Started
Installation
pip install humming-kernels
To also install CUDA dependencies, choose the extra matching your PyTorch CUDA version:
pip install "humming-kernels[cu12]" # CUDA 12 pip install "humming-kernels[cu13]" # CUDA 13
Or install from source:
pip install git+https://github.com/vllm-project/humming.git
Usage Example
import torch from humming.layer import HummingLayer layer = HummingLayer( shape_n=8192, shape_k=8192, weight_config={"dtype": "int6"}, torch_dtype=torch.float16, ).cuda() weight = torch.randn((8192, 8192), dtype=torch.float16, device="cuda:0") inputs = torch.randn((128, 8192), dtype=torch.float16, device="cuda:0") # Load unquantized weight and quantize to layer quantization format layer.load_from_unquantized(weight) # Transform weight to humming format and prepare default kernels layer.transform() # Run quantized GEMM (tuning_config is optional, auto-selected by default) output = layer(inputs) print("Quantized GEMM Output:") print(output) print("\nReference Output:") print(inputs.matmul(weight.T))
Acknowledgement
This project is highly inspired by
- DeepGEMM
- Marlin Kernel and vLLM Marlin Kernel
- lmdeploy GEMM kernel
- CUTLASS
来源:蚂蚁 inclusionAI:GitHub 新仓库 · github.com