跳转至

第二阶段算子

50 个算子,236 个工作负载 —— Attention 9 · Linear Attention 与 SSM 2 · 混合专家(MoE) 1 · 量化 3 · 逐元素运算 22 · 归约 1 · 归一化 2 · 池化 2 · 位置编码 6 · MHC 2。

每个算子一张表,每个工作负载一行。比值 是同一测量口径下基准的耗时除以我们的耗时,所以 绿色 表示我们更快,无色 表示持平,红色 表示我们更慢。时间单位是 ms。这些数字是怎么来的。

独立的一组,独立的分母

这些算子不在 oplist150 这份第一阶段验收清单(docs/reports/R267-data/oplist150.json)里 —— 总览页和各家族数据页都只覆盖这份清单,对上汇报的进度也以它为准。它们出现在这里,是因为 harness 已经为它们实测到了一个真实对手,不是因为第一阶段变大了。

本页的任何数字都不会被加进其它页面的任何计数里。总览页的覆盖率、各家族小计、达标数,全部只统计第一阶段,并且继续只统计第一阶段。下面那行本页自己的小计,只数这一组算子,不数别的。

只有快照真的为一个算子发布了比值,它才会出现在这一页;没有实测到对手的算子会被留在页面之外,而不是渲染成一行空白。

Attention

GroupedQueryAttentionDecodePagedWithKVCacheFwd

dtype=f16

  • W1q: [4, 128, 128]kv: [4096, 8, 128]page_size=256serving-405b-p256
  • W2q: [8, 64, 128]kv: [4096, 8, 128]page_size=256serving-70b-p256
  • W3q: [8, 64, 128]kv: [4096, 8, 128]page_size=64serving-70b-p64
  • W4q: [8, 32, 128]kv: [32768, 8, 128]page_size=64serving-8b-long-p64
  • W5q: [32, 32, 128]kv: [4096, 8, 128]page_size=256serving-8b-p256
  • W6q: [32, 32, 128]kv: [4096, 8, 128]page_size=64serving-8b-p64
  • W7q: [64, 32, 128]kv: [2048, 8, 128]page_size=64throughput-8b-p64
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.45×33.4937torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager python reference 厂商库
48.4476
48.5021
48.5410
···
W21.19×33.4759torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager python reference 厂商库
39.9049
39.9079
44.9059
···
W3—·torch_compile_ge torch_compile 基准
torch_npu eager python reference 厂商库
39.6264
44.8341
···
W4—·torch_compile_ge torch_compile 基准
torch_npu eager python reference 厂商库
136.3076
165.6499
···
W51.18×65.3830torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager python reference 厂商库
77.1264
77.1175
87.2837
···
W6—·torch_compile_ge torch_compile 基准
torch_npu eager python reference 厂商库
76.0240
88.2759
···
W7—·torch_compile_ge torch_compile 基准
torch_npu eager python reference 厂商库
144.2580
161.4332
···

GroupedQueryAttentionSlidingWindowVarlenFwd

dtype=f16batch=4dim=128

  • W1total_q=8192total_k=8192heads=64/8window_size_left=1024max_seqlen_q=2048llama-70b-long-w1024
  • W2total_q=2048total_k=2048heads=64/8window_size_left=256max_seqlen_q=512llama-70b-short-w256
  • W3total_q=8192total_k=8192heads=32/8window_size_left=1024max_seqlen_q=2048llama-8b-long-w1024
  • W4total_q=2048total_k=2048heads=32/8window_size_left=256max_seqlen_q=512llama-8b-short-w256
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.81×6.2363torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库 基准
5.0464
5.0319
···
W21.16×0.5373torch_compile_aclgraph torch_compile 基准
torch_npu eager python reference 厂商库
0.6249
0.6279
···
W30.80×3.3505torch_compile_aclgraph torch_compile 基准
torch_npu eager python reference 厂商库
2.6565
2.7251
···
W41.26×0.3385torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库 基准
0.4236
0.4211
···

MultiHeadAttentionFwd

  • W1q, kv: [1, 2048, 64, 128]dtype=bf16llama-70b-long
  • W2q, kv: [1, 2048, 64, 128]dtype=f16llama-70b-long
  • W3q, kv: [2, 512, 64, 128]dtype=bf16llama-70b-short
  • W4q, kv: [2, 512, 64, 128]dtype=f16llama-70b-short
  • W5q, kv: [2, 2048, 32, 128]dtype=bf16llama-8b-long
  • W6q, kv: [2, 2048, 32, 128]dtype=f16llama-8b-long
  • W7q, kv: [4, 512, 32, 128]dtype=bf16llama-8b-short
  • W8q, kv: [4, 512, 32, 128]dtype=f16llama-8b-short
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.42×1.5755ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager SDPA 厂商库
0.6620
0.8296
0.8330
0.8325
···
W20.42×1.5684ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.6521
0.8214
0.8235
···
W30.63×0.2579ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1611
0.2915
0.2884
···
W40.65×0.2534ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1608
0.2904
0.2855
···
W50.42×1.5901ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.6626
0.8350
0.8306
···
W60.41×1.5860ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.6498
0.8173
0.8165
···
W70.64×0.2560ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1616
0.2949
0.2945
···
W80.66×0.2555ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1638
0.2984
0.2961
···

GroupedQueryAttentionFwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]llama-70b-long
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]llama-70b-short
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]llama-8b-long
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]llama-8b-short
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.41×1.5300ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.6172
0.7180
0.7185
···
W20.60×0.2390ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1411
0.2612
0.2616
···
W30.41×1.5433ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.6211
0.7411
0.7349
···
W40.60×0.2439ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库
0.1449
0.2695
0.2679
···

GroupedQueryAttentionSlidingWindowFwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]window_size_left=1024llama-70b-long-w1024
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]window_size_left=256llama-70b-short-w256
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]window_size_left=1024llama-8b-long-w1024
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]window_size_left=256llama-8b-short-w256
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.30×1.8444ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库
0.5539
1.5036
1.5085
···
W20.43×0.3362ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库
0.1391
0.3902
0.3925
···
W30.30×1.8560ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库
0.5573
1.5281
1.5203
···
W40.43×0.3429ops-transformer-b1:aclnnFlashAttentionScore 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager python reference 厂商库
0.1434
0.3946
0.3921
···

MultiHeadAttentionDecodeWithKVCacheFwd

  • W1q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=bf16llama-70b-32k
  • W2q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=f16llama-70b-32k
  • W3q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=bf16llama-70b-4k
  • W4q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=f16llama-70b-4k
  • W5q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=bf16llama-8b-32k
  • W6q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=f16llama-8b-32k
  • W7q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=bf16llama-8b-4k
  • W8q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=f16llama-8b-4k
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.13×139.8074torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager SDPA 厂商库 基准
18.6446
19.0927
18.4155
···
W20.13×139.8650torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库 基准
18.7369
18.6160
···
W30.14×65.0734torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库 基准
9.2641
9.2630
···
W40.14×65.1240torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库 基准
9.2516
9.2351
···
W50.15×139.7050torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA 厂商库
20.7512
20.7583
···
W60.15×139.7111torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA 厂商库
20.6946
20.7183
···
W70.14×65.0439torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库 基准
9.0760
9.0411
···
W80.14×65.1153torch_compile_aclgraph torch_compile
torch_npu eager SDPA 厂商库 基准
9.0676
9.0475
···

GroupedQueryAttentionBwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]llama-70b-long
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]llama-70b-short
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]llama-8b-long
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]llama-8b-short
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×270.8478torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
2.3700
2.3945
···
W20.02×44.8023torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
0.7984
0.8086
···
W30.01×272.6906torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
2.4531
2.4635
···
W40.02×45.1514torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
0.7960
0.8096
···

MultiHeadAttentionBwd

dtype=f16

  • W1q, kv: [1, 2048, 64, 128]llama-70b-long
  • W2q, kv: [2, 512, 64, 128]llama-70b-short
  • W3q, kv: [2, 2048, 32, 128]llama-8b-long
  • W4q, kv: [4, 512, 32, 128]llama-8b-short
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×365.0621torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
2.9143
2.9274
···
W20.01×91.6334torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
0.7804
0.7946
···
W30.01×365.5979torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
2.8984
2.9120
···
W40.01×91.8776torch_compile_aclgraph torch_compile 基准
torch_npu eager SDPA autograd 厂商库
0.7937
0.7959
···

MultiHeadLatentAttentionDecodeWithKVCacheFwd

dtype=f16pe_dim=64

  • W1q: [8, 128, 128]q_pe: [8, 128, 64]kv: [8, 32768, 1, 128]k_pe: [8, 32768, 1, 64]deepseek-v2-32k
  • W2q: [32, 128, 128]q_pe: [32, 128, 64]kv: [32, 4096, 1, 128]k_pe: [32, 4096, 1, 64]deepseek-v2-4k
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.00×521.9301torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager python reference 厂商库 基准
0.6825
0.6180
0.6683
···
W20.00×260.9430torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager python reference 厂商库 基准
0.3005
0.2662
0.2871
···

Linear Attention 与 SSM

EngramGateConvBwd

dY, H, k, v, vhat: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]alpha, rrms_h, rrms_k, rrms_v: [M, seq_len], f32

  • W1dtype=bf16M=1seq_len=128d=256bwd-b1-s128-d256
  • W2dtype=f16M=1seq_len=32d=256bwd-b1-s32-d256
  • W3dtype=f16M=2seq_len=64d=512bwd-b2-s64-d512
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W111.01×0.0831Torch-NPU eager exact reference0.9155···
W217.36×0.0408Torch-NPU eager exact reference0.7074···
W312.66×0.0825Torch-NPU eager exact reference1.0442···

EngramGateConvFwd

H, k, v: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]

  • W1dtype=bf16M=1seq_len=128d=256fwd-b1-s128-d256
  • W2dtype=f16M=1seq_len=32d=256fwd-b1-s32-d256
  • W3dtype=f16M=2seq_len=64d=512fwd-b2-s64-d512
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W18.71×0.0410Torch-NPU eager exact reference0.3571···
W29.95×0.0280Torch-NPU eager exact reference0.2785···
W38.76×0.0452Torch-NPU eager exact reference0.3962···

混合专家(MoE)

MoeUnpermuteFwd

dtype=bf16top_k=8

  • W1total_tokens=1hidden_size=7168large-hidden-decode
  • W2total_tokens=512hidden_size=7168large-hidden-medium
  • W3total_tokens=4096hidden_size=7168large-hidden-prefill
  • W4total_tokens=32hidden_size=7168large-hidden-small
  • W5total_tokens=1hidden_size=3072small-hidden-decode
  • W6total_tokens=512hidden_size=3072small-hidden-medium
  • W7total_tokens=4096hidden_size=3072small-hidden-prefill
  • W8total_tokens=32hidden_size=3072small-hidden-small
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.84×0.00625ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.00425
0.0445
0.0245
···
W20.51×0.1747ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.0870
0.3869
0.3757
···
W30.41×1.3121ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.5310
2.8803
2.8481
···
W40.64×0.0180ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.00975
0.0600
0.0454
···
W50.90×0.005ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.00313
0.0380
0.0180
···
W60.35×0.1300ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.0420
0.1510
0.1355
···
W70.29×0.9976ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.2871
1.2434
1.2379
···
W80.64×0.0152ops-transformer-b2:aclnnMoeTokenUnpermute 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.00675
0.0561
0.0400
···

量化

FP8LightningIndexerFwd

index_q: [batch, seq_len, heads, index_dim]index_k: [batch, seq_len_kv, kv_group, index_dim]weights: [seq_len, heads], f32cu_seqlen_ks, cu_seqlen_ke: [seq_len], i32index_k_scale: [batch, seq_len_kv, kv_group], f32

batch=1seq_len=8192heads=32index_dim=64seq_len_kv=32768kv_group=1

  • W1dtype=bf16lightning-indexer-s8k-h32-d64
  • W2dtype=fp8e4m3lightning-indexer-s8k-h32-d64

index_k_scale: [batch, seq_len_kv, kv_group], f32

dtype=fp8e4m3batch=1seq_len_kv=32768kv_group=1seq_len=8192heads=32index_dim=64

  • W3lightning-indexer-s8k-h32-d64-kscale
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W15.88×24.0535torch_npu eager python reference 厂商库 基准 ⓘ141.4828···
W25.82×24.2429torch_npu eager python reference 厂商库 基准 ⓘ141.1920···
W35.70×25.1346torch_npu eager python reference 厂商库 基准 ⓘ143.3777···

BmmFp8NKFwd

dtype=fp8e4m3

  • W1b=32m=128n=128k=2048mha-decode-b32-pv-per-tensor
  • W2b=64m=128n=2048k=128mha-decode-b64-qk-per-tensor
  • W3b=128m=512n=512k=2048moe-prefill-b128-per-tensor
  • W4b=4m=1024n=1024k=1024square-b4-1k-per-tensor
  • W5b=8m=2048n=2048k=2048square-b8-2k-per-tensor
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W14.83×0.1133torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
0.5479
0.5954
···
W23.65×0.2238torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
0.8139
0.9902
···
W39.28×2.3949torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
22.3016
22.2772
···
W42.42×0.1383torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.3374
0.3259
···
W53.48×1.4383torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
5.0110
5.5925
···

BmmFp8KNFwd

dtype=fp8e4m3

  • W1b=32m=128n=128k=2048mha-decode-b32-pv-per-tensor
  • W2b=64m=128n=2048k=128mha-decode-b64-qk-per-tensor
  • W3b=128m=512n=512k=2048moe-prefill-b128-per-tensor
  • W4b=4m=1024n=1024k=1024square-b4-1k-per-tensor
  • W5b=8m=2048n=2048k=2048square-b8-2k-per-tensor
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W14.12×0.1305torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
0.5369
0.5880
···
W23.87×0.2070torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
0.8036
0.9409
···
W39.28×2.3847torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
22.2211
22.1105
···
W42.36×0.1400torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.3287
0.3186
···
W53.50×1.4418torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
5.0496
5.6059
···

逐元素运算

ClampScalarFwd

min=-0.5max=0.5

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W14.18×0.0691ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_2200100000.2890···
W22.23×0.0683ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_2200100000.1522···
W32.50×0.1150ClipByValueV2_54f1d1caf8fe36c496f699fefe98fbca_high_performance_2200100000.2878···
W44.73×0.9459ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_2200100004.4703···
W52.35×0.9565ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_2200100002.2517···

BitwiseNotFwd

  • W1input: [4096, 4096]dtype=i32elementwise-16M
  • W2input: [4096, 4096]dtype=i64elementwise-16M
  • W3input: [16384, 16384]dtype=i32elementwise-256M
  • W4input: [4097]dtype=i32legacy-probe-1
  • W5input: [4097]dtype=i64legacy-probe-2
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W14.53×5.0855torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
22.9669
22.9959
···
W26.08×5.1049torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
31.4594
31.4969
···
W34.53×80.5434torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库 ⓘ
364.8630
364.9903
···
W40.77×0.0410torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0325
0.0284
···
W50.83×0.0431torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0333
0.0297
···

AlibiFwd

num_heads=32

  • W1dtype=bf16seq_len=2048llama-prefill-2k
  • W2dtype=f16seq_len=2048llama-prefill-2k
  • W3dtype=bf16seq_len=4096llama-prefill-4k
  • W4dtype=f16seq_len=4096llama-prefill-4k
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W12.25×0.5493Torch-NPU eager reference1.2385···
W22.25×0.5505Torch-NPU eager reference1.2361···
W32.16×2.1581Torch-NPU eager reference4.6643···
W42.16×2.1609Torch-NPU eager reference4.6753···

PreluFwd

  • W1input: [1, 4, 1025]weight: [4]dtype=f16channel-tail-nondivisible
  • W2input: [16, 256, 56, 56]weight: [256]dtype=bf16cnn-feat-per-channel
  • W3input: [16, 512, 28, 28]weight: [512]dtype=bf16cnn-feat-per-channel-deep
  • W4input: [16, 512, 28, 28]weight: [512]dtype=f16cnn-feat-per-channel-deep
  • W5input: [16, 256, 56, 56]weight: [256]dtype=f16cnn-feat-per-channel
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W10.61×0.0248Torch-NPU eager exact reference0.0150···
W22.43×0.0620Torch-NPU eager exact reference0.1509···
W32.41×0.0379Torch-NPU eager exact reference0.0912···
W42.41×0.0380Torch-NPU eager exact reference0.0917···
W52.30×0.0670Torch-NPU eager exact reference0.1541···

MaskedFillScalarFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W11.67×0.0843Torch-NPU eager exact reference0.1410···
W21.25×0.0848Torch-NPU eager exact reference0.1060···
W31.38×0.1395Torch-NPU eager exact reference0.1925···
W42.18×1.1220Torch-NPU eager exact reference2.4454···
W51.85×1.1454Torch-NPU eager exact reference2.1162···

LogicalOrFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W4input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W5input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W7input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W8input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.73×0.0612torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0998
0.0998
0.0998
···
W20.94×0.0452torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager 厂商库
0.0359
0.0356
0.0360
···
W31.83×0.0512torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0835
0.0833
0.0835
···
W41.60×0.0895torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.1360
0.1363
0.1360
···
W51.71×0.0576torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0990
0.0931
0.0830
···
W61.00×0.0356torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0355
0.0325
0.0297
···
W71.53×0.0559torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0828
0.0770
0.0674
···
W81.37×0.1034torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.1415
0.1325
0.1224
···

NeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.81×0.0658torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1077
0.1077
0.1077
···
W21.01×0.0563torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0560
0.0548
0.0530
···
W31.29×0.0945torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1163
0.1160
0.1163
···
W41.37×0.0644torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0885
0.0794
0.0793
···
W50.96×0.0616torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0592
0.0607
0.0553
···
W61.16×0.1062torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1217
0.1123
0.1153
···

GtFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.22×0.0595torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0653
0.0650
0.0651
···
W21.06×0.0512torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0540
0.0678
0.0464
···
W31.10×0.0835torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0833
0.0833
0.0831
···
W41.15×0.0545torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0626
0.0550
0.0548
···
W50.93×0.0568torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0530
0.0510
0.0485
···
W61.04×0.0943torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0980
0.0998
0.0935
···

LtFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.21×0.0578torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0602
0.0602
0.0600
···
W20.98×0.0455torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0480
0.0633
0.0415
···
W31.18×0.0789torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0789
0.0788
0.0786
···
W41.16×0.0504torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0580
0.0532
0.0493
···
W50.95×0.0478torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0451
0.0430
0.0393
···
W61.01×0.0917torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0960
0.0944
0.0902
···

LeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.16×0.0553torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager 厂商库 ⓘ
0.0573
0.0575
0.0574
···
W21.01×0.0457torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0460
0.0558
0.0405
···
W31.17×0.0756torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0747
0.0751
0.0746
···
W41.15×0.0483torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0553
0.0503
0.0480
···
W50.96×0.0483torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0465
0.0455
0.0428
···
W61.03×0.0880torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0877
0.0876
0.0846
···

GeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.16×0.0633torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager 厂商库 ⓘ
0.0673
0.0673
0.0674
···
W21.10×0.0452torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0498
0.0658
0.0430
···
W31.08×0.0830torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0833
0.0835
0.0833
···
W41.10×0.0537torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0575
0.0537
0.0514
···
W50.96×0.0505torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0498
0.0491
0.0447
···
W61.04×0.0958torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1000
0.1010
0.0970
···

EqFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.13×0.0665torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0653
0.0653
0.0651
···
W21.03×0.0560torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0573
0.0575
0.0527
···
W31.09×0.0932torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0927
0.0927
0.0925
···
W41.10×0.0633torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0700
0.0633
0.0636
···
W50.94×0.0620torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0587
0.0580
0.0550
···
W61.04×0.1057torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1082
0.1065
0.1045
···

NanToNumFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W11.11×0.0798NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_2100100000.0885···
W20.86×0.0786NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_2100100000.0673···
W30.95×0.1295NanToNum_5703d6e339578b11bc6d39a70205aa62_high_performance_2100100000.1229···
W41.08×1.0601NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_2100100001.1493···
W50.88×1.0551NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_2100100000.9250···

SinusoidalFwd

d_model=4096

  • W1dtype=bf16seq_len=2048transformer-2k-4k
  • W2dtype=f16seq_len=2048transformer-2k-4k
  • W3dtype=bf16seq_len=4096transformer-4k-4k
  • W4dtype=f16seq_len=4096transformer-4k-4k
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W11.02×0.4160Torch-NPU eager reference0.4228···
W21.03×0.4193Torch-NPU eager reference0.4328···
W30.91×0.8217Torch-NPU eager reference0.7468···
W40.91×0.8220Torch-NPU eager reference0.7458···

SeluFwd

  • W1input: [2048, 4096]dtype=bf16snn-fc
  • W2input: [2048, 4096]dtype=f16snn-fc
  • W3input: [2048, 8192]dtype=bf16snn-fc-wide
  • W4input: [2048, 8192]dtype=f16snn-fc-wide
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.84×0.0440torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0367
0.0367
0.0318
···
W20.78×0.0460torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0360
0.0376
0.0323
0.0296
···
W30.84×0.0779torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0663
0.0661
0.0612
···
W40.84×0.0770torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0643
0.0648
0.0607
0.0586
···

FloorDivideFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.52×0.1578torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0814
0.0650
···
W20.44×0.1555torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0688
0.0586
···
W30.61×0.1580torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0993
0.0917
···
W40.61×0.1052torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0620
0.0563
···
W50.60×0.1054torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0639
0.0573
···
W60.87×0.1227torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1062
0.1015
···
W70.69×0.00325torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.00225
0.0015
···
W80.75×0.003torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.00225
0.0015
···
W90.75×0.003torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0015
0.0015
···

BitwiseXorFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×4.3217torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0602
0.0455
···
W20.02×7.0530torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1075
0.0986
···
W30.02×10.1686torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1926
0.1872
···
W41.15×0.0333torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0379
0.0306
···
W51.04×0.1020torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1062
0.1027
···
W61.07×0.1956torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.2102
0.2025
···
W70.89×0.00225torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.002
0.0015
···

BitwiseAndFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×4.3252torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0483
0.0350
···
W20.01×7.0652torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0993
0.0912
···
W30.02×10.1616torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1905
0.1815
···
W41.08×0.0328torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0362
0.0312
···
W51.00×0.1051torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1057
0.1025
···
W61.04×0.1993torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.2055
0.2004
···
W71.06×0.002torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.00225
0.0015
···

BitwiseOrFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×4.3249torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0504
0.0357
···
W20.01×7.0569torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0998
0.0920
···
W30.02×10.1641torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1883
0.1815
···
W41.07×0.0318torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0339
0.0296
···
W51.01×0.1040torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1070
0.1010
···
W61.04×0.1968torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.2065
0.1980
···
W71.00×0.002torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.00225
0.0015
···

IsinfFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.02×5.9417torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0930
0.0571
0.0809
···
W20.01×6.0650torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0770
0.0555
0.0678
···
W30.02×5.9460torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.1360
0.1035
0.1303
···
W40.02×94.2352torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
1.7075
1.7066
1.7060
···
W50.02×94.1644torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
1.6007
0.7289
1.5947
···

IsnanFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.02×6.5216torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1148
0.1067
···
W20.01×8.8604torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0587
0.0542
···
W30.02×6.5269torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1309
0.1212
···
W40.02×103.4591torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
1.6560
1.6482
···
W50.01×103.3778torch_compile_aclgraph torch_compile 基准
torch_npu eager 厂商库
0.7666
0.7720
···

IsfiniteFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.01×5.9420torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.2279
7.7950
0.0555
···
W20.01×6.0667torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1160
7.7988
0.0545
···
W30.02×5.9465torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.2697
7.8072
0.0985
···
W40.01×94.2362torch_compile_aclgraph torch_compile 基准
torch_compile_ge torch_compile
torch_npu eager 厂商库 ⓘ
0.7360
0.7370
0.7366
···
W50.01×94.1655torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
3.1088
122.3452
0.7299
···

归约

VarMeanFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-var-mean
  • W3x: [2048, 4096]dtype=f16hidden-state-var-mean
  • W4x: [64, 32768]dtype=bf16long-seq-var-mean
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W14.45×0.0264torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1175
0.1090
···
W22.69×0.0535torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1442
0.1361
···
W32.80×0.0512torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.1427
0.1368
···
W42.92×0.0275torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
0.0810
0.0746
···

归一化

AdaLayerNormZeroFwd

  • W1x: [1024, 1152]dtype=bf16dit-xl-2
  • W2x: [1024, 1152]dtype=f16dit-xl-2
  • W3x: [1, 4096]dtype=bf16llama-8b-decode
  • W4x: [2048, 4096]dtype=bf16llama-8b-prefill
  • W5x: [2048, 4096]dtype=f16llama-8b-prefill
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W12.33×0.0269torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0622
0.0549
···
W22.38×0.0272torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.0617
0.0578
···
W33.12×0.004torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准
tilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0123
0.0110
0.00424
···
W41.29×0.1231torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1593
0.1464
···
W51.51×0.1205torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.1816
0.1737
···

InfNormFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-inf
  • W3x: [2048, 4096]dtype=f16hidden-state-inf
  • W4x: [64, 32768]dtype=bf16long-seq-inf
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W10.87×0.0357Torch-NPU eager reference0.0310···
W21.64×0.0360Torch-NPU eager reference0.0591···
W31.57×0.0375Torch-NPU eager reference0.0589···
W41.55×0.0187Torch-NPU eager reference0.0290···

池化

AdaptiveMaxPool2dIndicesFwd

input: [N, C, H_in, W_in]

  • W1dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]global-1x1
  • W2dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]nondiv-7x7
  • W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]spp-6x6
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W10.06×1.7141Torch-NPU eager reference0.0973···
W20.10×1.3225Torch-NPU eager reference0.1296···
W30.05×3.0819Torch-NPU eager reference0.1675···

MaxPool3dIndicesFwd

input: [N, C, D_in, H_in, W_in]

  • W1dtype=f16N=8C=64D_in=16H_in=112W_in=112kernel_size=[1,2,2]stride=[1,2,2]padding=[0,0,0]c3d-pool1
  • W2dtype=f16N=4C=128D_in=16H_in=56W_in=56kernel_size=[2,2,2]stride=[2,2,2]padding=[0,0,0]c3d-pool2
  • W3dtype=bf16N=2C=64D_in=32H_in=112W_in=112kernel_size=[3,3,3]stride=[2,2,2]padding=[1,1,1]ceil_mode=truemedicalnet-stem
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.00×240.7824torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.9711
0.7434
···
W20.00×64.8989torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
0.3008
0.1850
···
W30.00×558.8976torch_compile_ge torch_compile
torch_npu eager 厂商库 基准 ⓘ
2.4289
1.2630
···

位置编码

RopeNeoxPositionIdsFwd

x: [num_tokens, num_heads, head_dim]position_ids: [num_tokens]

num_heads=32head_dim=128

  • W1dtype=f16num_tokens=2048max_position=4096position-ids-s2k-h32-d128
  • W2dtype=bf16num_tokens=4096max_position=8192position-ids-s4k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.85×0.1086ops-transformer-b2:aclnnRopeWithSinCosCache 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.2000
0.2979
0.2835
···
W21.91×0.2024ops-transformer-b2:aclnnRopeWithSinCosCache 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.3875
0.4447
0.4719
···

RopeNeoxFwd

  • W1dtype=f16seq_len=2048head_dim=64neox-1d-2k-d64
  • W2dtype=bf16seq_len=4096head_dim=128neox-1d-4k-d128
  • W3dtype=f16batch=2seq_len=2048num_heads=32head_dim=128layout=2dneox-2d-b2-s2k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.14×0.0143ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 ⓘ
0.0118
0.0698
0.0638
···
W20.96×0.0173ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 ⓘ
0.0105
0.0757
0.0719
···
W30.91×0.1145ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 ⓘ
0.0949
0.3815
0.3954
···

RopeLlama31Fwd

seq_len=8192head_dim=128

  • W1dtype=bf16llama31-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dllama31-2d-b1-s8k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W11.04×0.0206ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 ⓘ
0.0155
0.0921
0.0780
···
W20.86×0.2246ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库 ⓘ
0.1874
1.0653
1.1926
···

RopeYarnFwd

seq_len=8192head_dim=128

  • W1dtype=bf16yarn-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dyarn-2d-b1-s8k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.99×0.0209ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.0152
0.0925
0.0772
···
W20.86×0.2286ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.1898
1.0504
1.1845
···

RopeLongRopeFwd

seq_len=8192head_dim=128

  • W1dtype=bf16longrope-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dlongrope-2d-b1-s8k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.94×0.0203ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.0150
0.0920
0.0790
···
W20.86×0.2300ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.1908
1.0408
1.2095
···

RopeNonNeoxFwd

seq_len=2048

  • W1dtype=f16head_dim=64non-neox-1d-2k-d64
  • W2dtype=bf16batch=2num_heads=32head_dim=128layout=2dnon-neox-2d-b2-s2k-h32-d128
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W10.91×0.0170ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.00813
0.0938
0.1526
···
W20.72×0.1965ops-transformer-b1:aclnnRotaryPositionEmbedding 开源库 ⓘ 基准
torch_compile_aclgraph torch_compile
torch_npu eager 厂商库
0.1358
0.4402
0.5316
···

MHC

MHCPostFwd

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 10240]

dtype=bf16batch=4c_x=2560n_expand=4

  • W1post-large

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 7680]

dtype=bf16batch=2c_x=1920n_expand=4

  • W2post-medium

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 5120]

dtype=bf16batch=1c_x=1280n_expand=4

  • W3post-small
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 · 档位 ms TFLOP/s 占天花板 受限于
W13.22×0.0112torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0362
0.0278
0.0296
···
W22.82×0.0085torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0240
0.0195
0.0170
···
W33.83×0.0045torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager 厂商库 基准
0.0195
0.0123
0.0123
···

MHCPreFwd

phi: [10240, 24], f32x: [batch, 10240]b: [24], f32

dtype=bf16batch=4alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W1pre-large

phi: [7680, 24], f32x: [batch, 7680]b: [24], f32

dtype=bf16batch=2alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W2pre-medium

phi: [5120, 24], f32x: [batch, 5120]b: [24], f32

dtype=bf16batch=1alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W3pre-small
工作负载 比值 耗时 对照实现 吞吐 SOL 瓶颈
对照 / 我们 ms 名称 ms TFLOP/s 占天花板 受限于
W11.06×0.2026Torch-NPU eager exact reference0.2150···
W21.03×0.1253Torch-NPU eager exact reference0.1295···
W31.14×0.0757Torch-NPU eager exact reference0.0860···