🐛 Describe the bug
On an NVIDIA B300 (SM103), the Blackwell Liger SwiGLU path is substantially
slower than the equivalent torch.compile expression for a Qwen3-shaped BF16
tensor.
With shape (3072, 12288), including forward and backward:
| implementation |
mean step time |
torch.compile(lambda a, b: F.silu(a) * b) |
2.169 ms |
torch.compile(lambda a, b: LigerSiLUMulFunction.apply(a, b)) |
6.831 ms |
The Liger path takes approximately 3.15x as long. The result reproduces on the
current Liger main (91ae44ae659ebaf40d4a851777f4f02515e9ba65), so it does
not appear to be specific to the 0.8.1 release.
This was initially observed in Qwen3-8B DFlash training. At sequence length
8192 and max_anchors=3072, enabling Liger increased the end-to-end training
step from 176.20 ms to 223.37 ms (+26.77%) and reduced throughput from 46.54k
to 36.70k tokens/s (-21.14%). Three alternating runs were used for each mode,
with 10 warmup and 50 measured steps per run.
The outputs and gradients remain numerically consistent at BF16 tolerances, so
this report is about performance rather than correctness.
Reproduce
import json
import statistics
import torch
import torch.nn.functional as F
from liger_kernel.ops.swiglu import LigerSiLUMulFunction
ROWS = 3072
COLS = 12288
WARMUP = 10
STEPS = 50
def benchmark(function):
compiled = torch.compile(function)
torch.manual_seed(42)
a = torch.randn(
(ROWS, COLS), device="cuda", dtype=torch.bfloat16, requires_grad=True
)
b = torch.randn(
(ROWS, COLS), device="cuda", dtype=torch.bfloat16, requires_grad=True
)
torch.manual_seed(43)
gradient = torch.randn_like(a)
def step():
a.grad = None
b.grad = None
compiled(a, b).backward(gradient)
for _ in range(WARMUP):
step()
torch.cuda.synchronize()
samples = []
for _ in range(STEPS):
start = torch.cuda.Event(enable_timing=True)
end = torch.cuda.Event(enable_timing=True)
start.record()
step()
end.record()
end.synchronize()
samples.append(start.elapsed_time(end))
return {
"mean_ms": statistics.mean(samples),
"median_ms": statistics.median(samples),
"stdev_ms": statistics.stdev(samples),
}
native = benchmark(lambda a, b: F.silu(a) * b)
liger = benchmark(lambda a, b: LigerSiLUMulFunction.apply(a, b))
print(
json.dumps(
{
"native": native,
"liger": liger,
"liger_over_native": liger["mean_ms"] / native["mean_ms"],
},
indent=2,
)
)
Profiling attributes most of the difference to the tiled Liger activation
kernels rather than the surrounding GEMMs. Disabling the Blackwell tiled path
made the isolated Liger result slower still (about 7.88 ms), so tiling helps
relative to Liger's one-row path but does not close the gap with the compiled
native expression.
Versions
- Liger Kernel: main at
91ae44ae659ebaf40d4a851777f4f02515e9ba65
- GPU: NVIDIA B300 SXM6 AC
- Compute capability: 10.3 (SM103)
- PyTorch:
2.11.0+cu130
- CUDA runtime:
13.0
- OS: Linux 6.8.0, x86_64
- dtype: BF16
🐛 Describe the bug
On an NVIDIA B300 (SM103), the Blackwell Liger SwiGLU path is substantially
slower than the equivalent
torch.compileexpression for a Qwen3-shaped BF16tensor.
With shape
(3072, 12288), including forward and backward:torch.compile(lambda a, b: F.silu(a) * b)torch.compile(lambda a, b: LigerSiLUMulFunction.apply(a, b))The Liger path takes approximately 3.15x as long. The result reproduces on the
current Liger main (
91ae44ae659ebaf40d4a851777f4f02515e9ba65), so it doesnot appear to be specific to the 0.8.1 release.
This was initially observed in Qwen3-8B DFlash training. At sequence length
8192 and
max_anchors=3072, enabling Liger increased the end-to-end trainingstep from 176.20 ms to 223.37 ms (+26.77%) and reduced throughput from 46.54k
to 36.70k tokens/s (-21.14%). Three alternating runs were used for each mode,
with 10 warmup and 50 measured steps per run.
The outputs and gradients remain numerically consistent at BF16 tolerances, so
this report is about performance rather than correctness.
Reproduce
Profiling attributes most of the difference to the tiled Liger activation
kernels rather than the surrounding GEMMs. Disabling the Blackwell tiled path
made the isolated Liger result slower still (about 7.88 ms), so tiling helps
relative to Liger's one-row path but does not close the gap with the compiled
native expression.
Versions
91ae44ae659ebaf40d4a851777f4f02515e9ba652.11.0+cu13013.0