Linear Attention & SSM¶
2 ops, 10 workloads.
One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.
Scan¶
CumprodFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-scan - W2x: [2048, 4096]dtype=f16
hidden-state-scan - W3x: [64, 32768]dtype=bf16
long-seq-scan
dtype=f16
- W4input: [2, 5, 7]dim=1
non-last-axis-host-permute - W5input: [4097]
nondivisible-4097
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 16.39× | 8.1635 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 133.9033 134.0408 | · | · | · |
| W2 | 16.44× | 8.1628 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 134.0186 134.1221 | · | · | · |
| W3 | 13.70× | 2.0857 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 28.5701 28.5780 | · | · | · |
| W4 | 1.37× | 0.0232 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0333 0.0309 | · | · | · |
| W5 | 3.36× | 0.0267 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0899 0.0882 | · | · | · |
CumsumFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-scan - W2x: [2048, 4096]dtype=f16
hidden-state-scan - W3x: [64, 32768]dtype=bf16
long-seq-scan
dtype=f16
- W4input: [2, 5, 7]dim=1
non-last-axis-host-permute - W5input: [4097]
nondivisible-4097
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.05× | 8.1691 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.4340 0.3981 0.3949 | · | · | · |
| W2 | 0.05× | 8.1645 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.4339 0.3956 0.3940 | · | · | · |
| W3 | 0.07× | 2.0892 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1686 0.1424 0.1446 | · | · | · |
| W4 | 0.86× | 0.0238 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0120 0.0107 0.0107 | · | · | · |
| W5 | 0.96× | 0.0283 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0272 0.0125 0.0185 | · | · | · |