跳转至

Linear Attention 与 SSM

2 个算子,10 个工作负载。

每个算子一张表,每个工作负载一行。比值 是同一测量口径下基准的耗时除以我们的耗时,所以 绿色 表示我们更快,无色 表示持平,红色 表示我们更慢。时间单位是 ms。这些数字是怎么来的。

扫描(Scan)

CumprodFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-scan
  • W2x: [2048, 4096]dtype=f16hidden-state-scan
  • W3x: [64, 32768]dtype=bf16long-seq-scan

dtype=f16

  • W4input: [2, 5, 7]dim=1non-last-axis-host-permute
  • W5input: [4097]nondivisible-4097
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W116.39×8.1635torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库
133.9033
134.0408
···
W216.44×8.1628torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库
134.0186
134.1221
···
W313.70×2.0857torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库
28.5701
28.5780
···
W41.37×0.0232torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0333
0.0309
···
W53.36×0.0267torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0899
0.0882
···

CumsumFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-scan
  • W2x: [2048, 4096]dtype=f16hidden-state-scan
  • W3x: [64, 32768]dtype=bf16long-seq-scan

dtype=f16

  • W4input: [2, 5, 7]dim=1non-last-axis-host-permute
  • W5input: [4097]nondivisible-4097
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.05×8.1691torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.4340
0.3981
0.3949
···
W20.05×8.1645torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.4339
0.3956
0.3940
···
W30.07×2.0892torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.1686
0.1424
0.1446
···
W40.86×0.0238torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0120
0.0107
0.0107
···
W50.96×0.0283torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0272
0.0125
0.0185
···