Skip to content

Linear Attention & SSM

2 ops, 10 workloads.

One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

Scan

CumprodFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-scan
  • W2x: [2048, 4096]dtype=f16hidden-state-scan
  • W3x: [64, 32768]dtype=bf16long-seq-scan

dtype=f16

  • W4input: [2, 5, 7]dim=1non-last-axis-host-permute
  • W5input: [4097]nondivisible-4097
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W116.39×8.1635torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
133.9033
134.0408
···
W216.44×8.1628torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
134.0186
134.1221
···
W313.70×2.0857torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
28.5701
28.5780
···
W41.37×0.0232torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0333
0.0309
···
W53.36×0.0267torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0899
0.0882
···

CumsumFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-scan
  • W2x: [2048, 4096]dtype=f16hidden-state-scan
  • W3x: [64, 32768]dtype=bf16long-seq-scan

dtype=f16

  • W4input: [2, 5, 7]dim=1non-last-axis-host-permute
  • W5input: [4097]nondivisible-4097
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.05×8.1691torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.4340
0.3981
0.3949
···
W20.05×8.1645torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.4339
0.3956
0.3940
···
W30.07×2.0892torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1686
0.1424
0.1446
···
W40.86×0.0238torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0120
0.0107
0.0107
···
W50.96×0.0283torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0272
0.0125
0.0185
···