Skip to content

Norm, Conv, Pool & Other

45 ops, 414 workloads — Normalization 23 · Convolution 13 · Pooling 8 · Top-k 1.

One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

Normalization

AdaLayerNormFwd

  • W1x: [1024, 1152]dtype=bf16dit-xl-2
  • W2x: [1024, 1152]dtype=f16dit-xl-2
  • W3x: [1, 4096]dtype=bf16llama-8b-decode
  • W4x: [2048, 4096]dtype=bf16llama-8b-prefill
  • W5x: [2048, 4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.73×0.0222torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0615
0.0521
···
W22.09×0.0255torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0540
0.0503
···
W33.14×0.0035torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0110
0.00962
0.00389
···
W41.34×0.1030torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1393
0.1192
···
W51.52×0.1005torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1625
0.1522
···

BatchNormBwd

dtype=f16

  • W1x: [4, 128, 1024, 1024]large-spatial
  • W2x: [32, 64]resnet50-fc
  • W3x: [8, 64, 32, 32]resnet50-stage1
  • W4x: [4, 128, 32, 32]resnet50-stage2
  • W5x: [4, 256, 28, 28]resnet50-stage3
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.17×11.4718torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
13.4796
13.4795
13.4889
···
W21.64×0.0377torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0511
0.0508
0.0506
···
W32.06×0.0318torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0732
0.0644
0.0650
···
W42.15×0.0297torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0712
0.0630
0.0616
···
W51.64×0.0525torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0935
0.0855
0.0825
···

L1NormFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-l1
  • W3x: [2048, 4096]dtype=f16hidden-state-l1
  • W4x: [64, 32768]dtype=bf16long-seq-l1
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.93×0.0357torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0334
0.0318
0.0198
···
W21.96×0.0403torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0798
0.0695
0.0508
···
W31.98×0.0399torch_npu eager vendor library basis0.0790···
W42.05×0.0227torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0481
0.0362
0.0200
···

L2NormFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-l2
  • W3x: [2048, 4096]dtype=f16hidden-state-l2
  • W4x: [64, 32768]dtype=bf16long-seq-l2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.14×0.0355torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0401
0.0208
···
W21.76×0.0420torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0742
0.0678
···
W31.79×0.0411torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0747
0.0670
···
W41.88×0.0220torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0400
0.0350
···

FusedAddLayerNormFwd

  • W1x: [1, 8192]dtype=bf16llama-70b-decode
  • W2x: [2048, 8192]dtype=bf16llama-70b-prefill
  • W3x: [2048, 8192]dtype=f16llama-70b-prefill
  • W4x: [1, 4096]dtype=bf16llama-8b-decode
  • W5x: [2048, 4096]dtype=bf16llama-8b-prefill
  • W6x: [2048, 4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.38×0.00525torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0125
0.0103
···
W21.26×0.2077torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2644
0.2492
···
W31.30×0.1886torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2459
0.2288
···
W42.09×0.004torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0085
0.00725
···
W51.16×0.1185torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1309
0.1222
···
W61.39×0.1135torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1584
0.1405
···

RMSNormQuantFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.84×0.0622torch_npu eager vendor library basis0.1143···
W21.35×4.5127torch_npu eager vendor library basis6.0799···
W31.36×4.6085torch_npu eager vendor library basis6.2891···
W41.74×0.0511torch_npu eager vendor library basis0.0887···
W51.12×1.9303torch_npu eager vendor library basis2.1546···
W61.14×1.9325torch_npu eager vendor library basis2.2107···
W71.71×0.0418torch_npu eager vendor library basis0.0712···
W81.14×0.9872torch_npu eager vendor library basis1.1266···
W91.12×0.9789torch_npu eager vendor library basis1.0958···

LayerNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.66×0.00725torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0116
0.0110
···
W21.22×0.2909torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.3538
0.3422
···
W31.20×0.2868torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.3441
0.3369
···
W41.76×0.00425torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.00725
0.00725
···
W50.78×0.1628torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1280
0.1192
···
W60.79×0.1653torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1300
0.1180
···
W71.29×0.0035ops-nn:aclnnLayerNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0035
0.00475
0.0035
0.0335
···
W80.71×0.0775ops-nn:aclnnLayerNorm open-source ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0537
0.0553
0.0529
0.0935
···
W91.24×0.0757torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0931
0.0793
···

BatchNormFwd

dtype=f16

  • W1x: [4, 128, 1024, 1024]large-spatial
  • W2x: [32, 64]resnet50-fc
  • W3x: [8, 64, 32, 32]resnet50-stage1
  • W4x: [4, 128, 32, 32]resnet50-stage2
  • W5x: [4, 256, 28, 28]resnet50-stage3
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.27×2.6470torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
3.3580
3.3549
3.3563
···
W21.13×0.00575torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0075
0.0065
0.0035
···
W31.02×0.0150torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0152
0.0100
0.0107
···
W41.21×0.0127torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0152
0.0120
0.0120
···
W50.95×0.0185torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0175
0.0110
0.0136
···

InstanceNormFwd

  • W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16image-affine
  • W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16image-affine
  • W5x: [4, 64, 30, 30]weight, bias: [64]dtype=f16tail-spatial-affine
  • W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16wider-channel-affine
  • W3x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=bf16image
  • W4x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=f16image
  • W6x: [4, 64, 30, 30]running_mean, running_var: [64], f32dtype=f16tail-spatial
  • W8x: [4, 256, 28, 28]running_mean, running_var: [256], f32dtype=f16wider-channel
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.21×0.0217ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_npu eager vendor library basis
0.0222
0.0201
···
W21.32×0.0198ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_npu eager vendor library basis
0.0231
0.0216
···
W30.91×0.0532torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0480
0.0405
···
W40.92×0.0525torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0486
0.0408
···
W51.13×0.0163ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_npu eager vendor library
0.0123
0.0168
···
W60.70×0.0447torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0320
0.0253
···
W71.51×0.0177ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_npu eager vendor library basis
0.0220
0.0200
···
W81.00×0.0516torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0542
0.0480
···

RMSNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.34×0.0055ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.00537
0.0240
0.0232
0.0226
···
W20.86×0.2157ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.1745
0.8171
0.8174
0.8177
···
W30.85×0.2120ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.1694
0.8265
0.8249
0.8260
17.4019
···
W41.38×0.00325ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.00275
0.0217
0.0210
0.0208
···
W50.82×0.0924ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.0675
0.2757
0.2769
0.2761
···
W60.82×0.0907ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0663
0.2767
0.3035
0.2741
0.0642
···
W71.45×0.00275ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0025
0.0302
0.00737
0.0195
0.00192
···
W80.98×0.0545ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0441
0.1700
0.1775
0.1685
0.0347
···
W90.93×0.0584ops-nn:aclnnRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0452
0.1731
0.2300
0.1726
0.0343
···

LayerNormQuantFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.29×0.0777torch_npu eager vendor library basis0.1004···
W21.00×5.9374torch_npu eager vendor library basis5.9467···
W31.05×5.9260torch_npu eager vendor library basis6.2252···
W41.13×0.0605torch_npu eager vendor library basis0.0683···
W50.91×2.3194torch_npu eager vendor library basis2.1175···
W60.89×2.3469torch_npu eager vendor library basis2.0978···
W71.08×0.0506torch_npu eager vendor library basis0.0545···
W80.87×1.2195torch_npu eager vendor library basis1.0617···
W90.88×1.2189torch_npu eager vendor library basis1.0720···

GroupNormFwd

  • W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16num_groups=32image-g32-affine
  • W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16num_groups=32image-g32-affine
  • W5x: [4, 128, 30, 30]weight, bias: [128]dtype=f16num_groups=16tail-spatial-g16-affine
  • W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16num_groups=32wider-channel-g32-affine
  • W3x: [8, 128, 32, 32]dtype=bf16num_groups=32image-g32
  • W4x: [8, 128, 32, 32]dtype=f16num_groups=32image-g32
  • W6x: [4, 128, 30, 30]dtype=f16num_groups=16tail-spatial-g16
  • W8x: [4, 256, 28, 28]dtype=f16num_groups=32wider-channel-g32
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.82×0.0262ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0140
0.0140
0.0140
0.0140
···
W20.69×0.0267ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0132
0.0130
0.0132
0.0132
···
W30.98×0.0220ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0132
0.0235
0.0232
0.0160
···
W40.96×0.0210ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0132
0.0225
0.0195
0.0158
···
W50.90×0.0190ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0107
0.0106
0.0105
0.0105
···
W61.07×0.0160ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0115
0.0198
0.0173
0.0138
···
W71.00×0.0180ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0114
0.0112
0.0112
0.0110
···
W81.16×0.0160ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0118
0.0205
0.0185
0.0135
···

FusedAddRMSNormFwd

  • W1x: [1, 16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.86×0.00725ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.00525
0.0295
0.0288
···
W20.90×0.3139ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.2755
1.1335
1.1365
·
···
W30.82×0.2913ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.2377
1.1364
1.0870
···
W40.78×0.0045ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.00325
0.0413
0.0272
···
W50.74×0.1713ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1235
0.4515
0.4954
···
W60.80×0.1590ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1153
0.3962
0.4716
···
W70.76×0.00363ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0025
0.0374
0.0250
·
···
W80.87×0.0907ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0750
0.2382
0.2285
·
···
W90.86×0.0872ops-nn:aclnnAddRmsNorm open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0602
0.2209
0.2171
·
···

InstanceNormBwd

  • W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16image-affine
  • W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16image-affine
  • W5x: [4, 64, 30, 30]weight, bias: [64]dtype=f16tail-spatial-affine
  • W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16wider-channel-affine
  • W3x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=bf16image
  • W4x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=f16image
  • W6x: [4, 64, 30, 30]running_mean, running_var: [64], f32dtype=f16tail-spatial
  • W8x: [4, 256, 28, 28]running_mean, running_var: [256], f32dtype=f16wider-channel
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.56×0.4622torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2585
0.2526
···
W20.55×0.4680torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2582
0.2540
···
W30.43×0.3782torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1749
0.1487
···
W40.43×0.3805torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1771
0.1496
···
W50.41×0.5150torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2130
0.2070
···
W60.36×0.3488torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1363
0.1060
···
W70.28×0.6202torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1790
0.1693
···
W80.21×0.4335torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0963
0.0710
···

GroupNormBwd

  • W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16num_groups=32image-g32-affine
  • W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16num_groups=32image-g32-affine
  • W5x: [4, 128, 30, 30]weight, bias: [128]dtype=f16num_groups=16tail-spatial-g16-affine
  • W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16num_groups=32wider-channel-g32-affine
  • W3x: [8, 128, 32, 32]dtype=bf16num_groups=32image-g32
  • W4x: [8, 128, 32, 32]dtype=f16num_groups=32image-g32
  • W6x: [4, 128, 30, 30]dtype=f16num_groups=16tail-spatial-g16
  • W8x: [4, 256, 28, 28]dtype=f16num_groups=32wider-channel-g32
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.47×0.5555torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2587
0.2497
···
W20.48×0.5400torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2619
0.2511
···
W30.47×0.5058torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2396
0.2290
···
W40.48×0.5104torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2411
0.2299
···
W50.38×0.5814torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2177
0.2127
···
W60.35×0.5711torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1968
0.1940
···
W70.24×0.7140torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1779
0.1695
···
W80.23×0.6837torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1540
0.1497
···

WeightNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.85×0.0624torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0250
0.0245
···
W20.26×2.6923torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.7278
0.6945
···
W30.26×2.7535torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.7109
0.7335
···
W40.53×0.0633torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0331
0.0210
···
W50.27×1.0950torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3050
0.2989
···
W60.27×1.1033torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3020
0.2988
···
W70.61×0.0498torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0302
0.0182
···
W80.31×0.5753torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1797
0.1787
···
W90.32×0.5763torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1831
0.1814
···

GroupRMSNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.34×0.1215torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0214
0.0211
···
W20.43×2.6197torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.1327
1.1786
···
W30.44×2.5910torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.1342
1.1487
···
W40.33×0.1089torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0355
0.0226
···
W50.28×1.1604torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3225
0.3919
···
W60.28×1.1596torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3234
0.3954
···
W70.29×0.1096torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0314
0.0214
···
W80.35×0.5785torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.2028
0.2076
···
W90.35×0.5807torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.2030
0.2094
···

GemmaRMSNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.35×0.1289torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0272
0.0272
···
W20.24×2.8287torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.7970
0.7252
···
W30.26×2.8723torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8133
0.7340
···
W40.32×0.1190torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0367
0.0255
···
W50.24×1.1386torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2721
0.2652
···
W60.24×1.1420torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2715
0.2646
···
W70.27×0.1215torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0331
0.0235
···
W80.30×0.5690torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1720
0.1747
···
W90.31×0.5677torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1745
0.1754
···

QKNormFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.30×0.2245torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0455
0.0437
···
W20.27×5.3171torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
1.6066
1.5517
···
W30.27×5.3952torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
1.6015
1.4194
···
W40.28×0.2050torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0579
0.0430
···
W50.22×2.3981torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.5231
0.4988
···
W60.22×2.3940torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.5114
0.4986
···
W70.26×0.1995torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0505
0.0405
···
W80.29×1.1075torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3285
0.3304
···
W90.29×1.1240torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3311
0.3319
···

RMSNormBwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.12×0.4083torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0348
0.0339
···
W20.32×6.3724torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
2.0200
2.2234
···
W30.31×6.4097torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
2.0053
2.2535
···
W40.15×0.2492torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0387
0.0305
···
W50.29×2.7129torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.7750
0.9525
···
W60.29×2.7153torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.7927
0.9451
···
W70.19×0.1713torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0318
0.0267
···
W80.29×1.2224torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3574
0.3711
···
W90.29×1.2143torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3554
0.3805
···

SpectralNormPowerIterFwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.11×0.6388torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0725
0.0729
···
W20.19×4.6319torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8709
0.8558
···
W30.19×4.6324torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8629
0.8546
···
W40.14×0.4525torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0607
0.0548
···
W50.21×1.7875torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3654
0.3498
···
W60.20×1.7924torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3663
0.3508
···
W70.08×0.3643torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0285
0.0260
···
W80.22×0.9865torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2136
0.2000
···
W90.22×0.9859torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2130
0.2003
···

BatchNormInferenceFwd

dtype=f16

  • W1x: [4, 128, 1024, 1024]large-spatial
  • W2x: [32, 64]resnet50-fc
  • W3x: [8, 64, 32, 32]resnet50-stage1
  • W4x: [4, 128, 32, 32]resnet50-stage2
  • W5x: [4, 256, 28, 28]resnet50-stage3
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.07×42.3053torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
3.1336
3.1236
···
W20.08×0.0796torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00775
0.0035
···
W30.11×0.1669torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0168
0.0110
···
W40.10×0.1704torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0195
0.0120
···
W50.10×0.1817torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0190
0.0135
···

LayerNormBwd

  • W1x: [1, 16384]normalized: [16384]dtype=bf16llama-405b-decode
  • W2x: [2048, 16384]normalized: [16384]dtype=bf16llama-405b-prefill
  • W3x: [2048, 16384]normalized: [16384]dtype=f16llama-405b-prefill
  • W4x: [1, 8192]normalized: [8192]dtype=bf16llama-70b-decode
  • W5x: [2048, 8192]normalized: [8192]dtype=bf16llama-70b-prefill
  • W6x: [2048, 8192]normalized: [8192]dtype=f16llama-70b-prefill
  • W7x: [1, 4096]normalized: [4096]dtype=bf16llama-8b-decode
  • W8x: [2048, 4096]normalized: [4096]dtype=bf16llama-8b-prefill
  • W9x: [2048, 4096]normalized: [4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.06×0.7544torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0314
0.0311
···
W20.05×9.2674torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.4285
0.4290
···
W30.05×9.2705torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4310
0.4308
···
W40.08×0.4495torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0222
0.0220
···
W50.04×3.7468torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1665
0.1669
···
W60.04×3.7363torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1586
0.1593
···
W70.09×0.2985torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0198
0.0201
···
W80.06×1.7646torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1080
0.1084
···
W90.06×1.7297torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1021
0.1020
···

Convolution

DepthwiseConv2d

input: [N, C_in, H, W]

dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]

  • W1mobilenetv2-depthwise
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.10×0.0602torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0668
0.0501
···

PointwiseConv2d

input: [N, C_in, H, W]bias: [C_out]

stride=[1,1]padding=[0,0]

  • W1dtype=f16N=2C_in=128H=28W=28C_out=512bottleneck-expand-1x1-bias
  • W3dtype=f16N=2C_in=512H=28W=28C_out=128bottleneck-reduce-1x1-bias
  • W5dtype=f16N=1C_in=512H=7W=7C_out=2048classifier-1x1-bias
  • W7dtype=f16N=1C_in=256H=14W=14C_out=1024late-stage-1x1-bias
  • W10dtype=bf16N=2C_in=64H=56W=56C_out=256resnet-1x1-bias
  • W11dtype=f16N=2C_in=64H=56W=56C_out=256resnet-1x1-bias

input: [N, C_in, H, W]

stride=[1,1]padding=[0,0]

  • W2dtype=f16N=2C_in=128H=28W=28C_out=512bottleneck-expand-1x1
  • W4dtype=f16N=2C_in=512H=28W=28C_out=128bottleneck-reduce-1x1
  • W6dtype=f16N=1C_in=512H=7W=7C_out=2048classifier-1x1
  • W8dtype=f16N=1C_in=256H=14W=14C_out=1024late-stage-1x1
  • W9dtype=bf16N=2C_in=64H=56W=56C_out=256resnet-1x1
  • W12dtype=f16N=2C_in=64H=56W=56C_out=256resnet-1x1
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.35×0.1226torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0289
0.0291
0.0286
···
W21.33×0.0345torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0460
0.0270
0.0304
···
W30.67×0.0645torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0300
0.0299
0.0297
···
W41.24×0.0396torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0441
0.0340
0.0316
···
W50.34×0.2185torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0648
0.0648
0.0648
···
W62.04×0.0340torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0649
0.0653
0.0649
···
W70.50×0.1060torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0319
0.0320
0.0316
···
W82.34×0.0213torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0331
0.0330
0.0328
···
W91.75×0.0300torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0527
0.0310
0.0362
···
W100.66×0.0795torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0377
0.0375
0.0377
···
W110.66×0.0788torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0379
0.0376
0.0379
···
W121.96×0.0310torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0535
0.0288
0.0355
···

GroupedConv2d

input: [N, C_in, H, W]

dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32

  • W1resnext-grouped-3x3
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.57×0.1125torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0643
0.0330
0.0478
···

Conv2dWgrad

  • W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias

input: [N, C_in, H, W]

  • W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1
  • W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1
  • W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1
  • W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]deeplabv3-aspp-3x3-rate12
  • W9dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1
  • W11dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1
  • W13dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1
  • W14dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32mobilenetv2-depthwise
  • W15dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W18dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W19dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W22dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W23dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32resnext-grouped-3x3
  • W25dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2
  • W27dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2
  • W29dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2
  • W30dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.25×0.0668torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0767
0.0776
···
W21.37×0.0620torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0780
0.0770
···
W30.86×0.0890torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0622
0.0624
···
W40.98×0.0849torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0626
0.0622
···
W51.08×0.1025torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1098
0.1101
···
W61.07×0.1030torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1096
0.1101
···
W70.03×7.2436torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1727
0.1722
···
W80.41×0.5530torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2454
0.2139
···
W90.40×0.5524torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2249
0.2129
···
W102.21×0.0390torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0749
0.0750
···
W112.26×0.0410torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0747
0.0736
···
W120.39×0.2480torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1085
0.0809
···
W130.38×0.2477torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0960
0.0808
···
W140.28×0.3775torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0881
0.0880
···
W150.99×0.0737torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0640
0.0648
···
W160.93×0.0785torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0636
0.0643
···
W171.08×0.0772torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0640
0.0640
···
W181.01×0.0727torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0635
0.0638
···
W190.42×0.1955torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0841
0.0683
···
W200.45×0.1970torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0989
0.0671
···
W210.45×0.1975torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1035
0.0675
···
W220.45×0.1958torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0910
0.0674
···
W230.43×0.2410torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0869
0.0858
···
W240.03×2.9286torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1056
0.0752
···
W250.03×2.9285torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0935
0.0762
···
W260.01×8.2952torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1131
0.0810
···
W270.01×8.2994torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1015
0.0801
···
W280.24×0.3060torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0831
0.0546
···
W290.24×0.3055torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0710
0.0541
···
W300.11×0.7768torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0855
0.0689
···
W310.11×0.7791torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0989
0.0688
···

Conv2dFp8Fwd

input: [N, C_in, H, W]

dtype=fp8e4m3

  • W1N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W2N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W3N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.32×0.0415torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0973
0.1960
0.0858
···
W20.08×1.2181torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0965
0.2034
0.0853
···
W30.05×1.8235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0902
0.1270
0.0810
···

Conv2dFwd

input: [N, C_in, H, W]bias: [C_out]

  • W1dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W3dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W5dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W14dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W17dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W19dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W23dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W24dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W27dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W28dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W31dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W33dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W35dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W39dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias

input: [N, C_in, H, W]

  • W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1
  • W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1
  • W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1
  • W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]deeplabv3-aspp-3x3-rate12
  • W15dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1
  • W18dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1
  • W20dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1
  • W21dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32mobilenetv2-depthwise
  • W22dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W25dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W26dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W29dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W30dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32resnext-grouped-3x3
  • W32dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2
  • W34dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2
  • W36dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2
  • W38dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2

dtype=f16

  • W8input_shape: [1, 16, 8, 8]weight_shape: [32, 16, 3, 3]default-dense-cube-fp16
  • W16input_shape: [16, 256, 56, 56]weight_shape: [256, 1, 1, 1]groups=256large-grid-depthwise-fp16
  • W9input_shape: [1, 16, 8, 8]padding: [1, 1]weight_shape: [16, 16, 3, 3]dtype=bf16dense-cube-bf16
  • W10input_shape: [1, 4, 8, 9]padding: [1, 1]weight_shape: [4, 1, 3, 3]dtype=f16groups=4depthwise-direct-fp16
  • W12input_shape: [1, 2, 5, 6]padding: [1, 1]weight_shape: [3, 2, 3, 3]dtype=f32fp32-direct
  • W13input_shape: [1, 8, 9, 10]padding: [1, 1]weight_shape: [8, 2, 3, 3]dtype=f16groups=4grouped-direct-fp16

dtype=f16

  • W11dilation, padding: [2, 2]input_shape: [1, 16, 9, 10]weight_shape: [16, 16, 3, 3]dilation-cube-fp16

dtype=f16

  • W37input_shape: [1, 3, 10, 11]padding: [1, 1]stride: [2, 2]weight_shape: [16, 3, 3, 3]stride-padding-nondiv-cube-fp16
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.38×0.1185torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0306
0.0305
···
W21.33×0.0354torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0318
0.0314
···
W30.71×0.0635torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0314
0.0309
···
W41.18×0.0395torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0310
0.0309
···
W50.33×0.2105torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0645
0.0640
···
W61.97×0.0330torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0640
0.0643
···
W70.00×79.0235torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1179
0.1181
···
W80.38×0.0881torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0267
0.0264
···
W90.57×0.0540torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0219
0.0206
···
W100.66×0.0840torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0444
0.0432
···
W110.28×0.1180torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0214
0.0200
···
W120.21×0.1457torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0262
0.0259
···
W130.31×0.1595torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0372
0.0360
···
W140.01×29.0843torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2147
0.2147
···
W150.01×29.0799torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2286
0.2157
···
W160.36×0.4460torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1510
0.1514
···
W170.41×0.1062torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0324
0.0326
···
W182.78×0.0200torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0345
0.0345
···
W190.02×4.2020torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0460
0.0459
···
W200.01×4.2009torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0589
0.0461
···
W210.94×0.0597torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0486
0.0494
···
W221.66×0.0294torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0377
0.0374
···
W230.71×0.0707torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0380
0.0380
···
W240.72×0.0688torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0380
0.0375
···
W251.63×0.0276torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0372
0.0375
···
W260.04×1.2043torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0485
0.0384
···
W270.05×1.2067torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0382
0.0381
···
W280.05×1.2054torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0396
0.0390
···
W290.05×1.2069torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0498
0.0385
···
W300.45×0.1130torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0473
0.0475
···
W310.01×8.3829torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0387
0.0380
···
W320.01×8.3692torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0515
0.0380
···
W330.00×24.0599torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0558
0.0551
···
W340.00×24.0326torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0590
0.0462
···
W350.13×0.4036torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0320
0.0318
···
W360.11×0.3975torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0454
0.0326
···
W370.95×0.0367torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0265
0.0257
···
W380.03×1.8193torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0460
0.0335
···
W390.02×1.8265torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0335
0.0328
···

Conv1dFwd

dtype=f16

  • W1input_shape: [1, 4, 17]weight_shape: [8, 2, 3]dilation=2groups=2padding=2dilation-groups
  • W10input_shape: [1, 2, 17]weight_shape: [4, 2, 3]padding=1stride=2stride-padding

input: [N, C_in, L_in]

N=1

  • W2dtype=bf16C_in=128L_in=600C_out=256kW=10encodec-deep
  • W5dtype=f16C_in=128L_in=600C_out=256kW=10encodec-deep
  • W6dtype=bf16C_in=1L_in=24000C_out=32kW=7encodec-init
  • W9dtype=f16C_in=1L_in=24000C_out=32kW=7encodec-init
  • W11dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1
  • W14dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1
  • W15dtype=bf16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1
  • W18dtype=f16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1

input: [N, C_in, L_in]bias: [C_out]

N=1

  • W3dtype=bf16C_in=128L_in=600C_out=256kW=10encodec-deep-bias
  • W4dtype=f16C_in=128L_in=600C_out=256kW=10encodec-deep-bias
  • W7dtype=bf16C_in=1L_in=24000C_out=32kW=7encodec-init-bias
  • W8dtype=f16C_in=1L_in=24000C_out=32kW=7encodec-init-bias
  • W12dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1-bias
  • W13dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1-bias
  • W16dtype=bf16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1-bias
  • W17dtype=f16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1-bias
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.99×0.0527torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0424
0.0415
···
W20.04×1.3505torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0470
0.0365
···
W30.03×1.3445torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0384
0.0386
···
W40.04×1.3594torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0377
0.0371
···
W50.03×1.3445torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0473
0.0375
···
W61.03×0.0320torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0328
0.0203
···
W71.02×0.0323torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0213
0.0210
···
W81.60×0.0304torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0328
0.0325
···
W91.79×0.0269torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0478
0.0326
···
W101.21×0.0320torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0302
0.0290
···
W110.00×41.5343torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0350
0.0245
···
W120.00×42.7620torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0257
0.0249
···
W130.00×42.7612torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0425
0.0424
···
W140.00×41.5373torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0532
0.0425
···
W150.04×1.6562torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0721
0.0609
···
W160.04×1.6618torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0622
0.0626
···
W170.05×1.6615torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0607
0.0605
···
W180.04×1.6499torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0715
0.0606
···

Conv2dDgrad

  • W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias
  • W2input: [2, 128, 28, 28]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1
  • W4input: [2, 512, 28, 28]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1
  • W6input: [1, 512, 7, 7]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1
  • W7input: [1, 2048, 32, 32]dtype=f16C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]deeplabv3-aspp-3x3-rate12
  • W9input: [1, 256, 112, 112]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1
  • W11input: [1, 256, 14, 14]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1
  • W13input: [1, 64, 56, 56]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1
  • W14input: [1, 32, 56, 56]dtype=f16C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32mobilenetv2-depthwise
  • W15input: [2, 64, 56, 56]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W18input: [2, 64, 56, 56]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W19input: [2, 64, 56, 56]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W22input: [2, 64, 56, 56]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W23input: [1, 128, 28, 28]dtype=f16C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32resnext-grouped-3x3
  • W25input: [1, 128, 56, 56]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2
  • W27input: [1, 128, 56, 56]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2
  • W29input: [1, 3, 112, 112]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2
  • W30input: [1, 128, 28, 28]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.20×0.3254torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0398
0.0393
0.0390
···
W20.22×0.3274torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0399
0.0398
0.0393
···
W30.21×0.3247torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0447
0.0444
0.0447
···
W40.23×0.3244torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0449
0.0445
0.0445
···
W50.02×4.0235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0861
0.0865
0.0861
···
W60.02×4.0235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0865
0.0866
0.0865
···
W70.00×81.0637torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1273
0.1280
0.1276
···
W80.01×37.9130torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.2029
0.2030
0.2030
···
W90.01×37.9222torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.2020
0.2019
0.2023
···
W100.07×1.0453torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0423
0.0423
0.0420
···
W110.07×1.0428torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0421
0.0420
0.0421
···
W120.01×7.8991torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0594
0.0592
0.0596
···
W130.01×7.9009torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0597
0.0594
0.0594
···
W140.74×0.0963torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0646
0.0646
0.0650
···
W150.65×0.1133torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0548
0.0542
0.0546
···
W160.59×0.1138torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0545
0.0548
0.0542
···
W170.61×0.1123torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0548
0.0550
0.0548
···
W180.67×0.1123torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0548
0.0548
0.0545
···
W190.05×1.3177torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0689
0.0305
0.0537
···
W200.05×1.3157torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0833
0.0305
0.0539
···
W210.05×1.3166torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0785
0.0307
0.0540
···
W220.05×1.3150torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0680
0.0310
0.0548
···
W230.39×0.2032torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0626
0.0628
0.0617
···
W240.01×6.9001torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0900
0.0750
0.0638
···
W250.01×6.9005torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0785
0.0747
0.0639
···
W260.00×20.2512torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0829
0.0830
0.0828
···
W270.00×20.2491torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0828
0.0829
0.0831
···
W280.02×2.7498torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0747
0.0532
0.0455
···
W290.02×2.7490torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0635
0.0535
0.0457
···
W300.03×2.2116torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0437
0.0430
0.0437
···
W310.03×2.2049torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0437
0.0435
0.0435
···

Conv2dTranspose

input: [N, C_out, out_H, out_W]bias: [C_in]

  • W1dtype=f16N=2C_out=128out_H=28out_W=28C_in=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W3dtype=f16N=2C_out=512out_H=28out_W=28C_in=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W5dtype=f16N=1C_out=512out_H=7out_W=7C_in=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W8dtype=f16N=1C_out=256out_H=112out_W=112C_in=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W10dtype=f16N=1C_out=256out_H=14out_W=14C_in=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W12dtype=f16N=1C_out=64out_H=56out_W=56C_in=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W16dtype=bf16N=2C_out=64out_H=56out_W=56C_in=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W17dtype=f16N=2C_out=64out_H=56out_W=56C_in=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W20dtype=bf16N=2C_out=64out_H=56out_W=56C_in=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W21dtype=f16N=2C_out=64out_H=56out_W=56C_in=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W24dtype=f16N=1C_out=128out_H=56out_W=56C_in=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W26dtype=f16N=1C_out=128out_H=56out_W=56C_in=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W28dtype=f16N=1C_out=3out_H=112out_W=112C_in=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W31dtype=bf16N=1C_out=128out_H=28out_W=28C_in=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias

input: [N, C_out, out_H, out_W]

  • W2dtype=f16N=2C_out=128out_H=28out_W=28kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1
  • W4dtype=f16N=2C_out=512out_H=28out_W=28kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1
  • W6dtype=f16N=1C_out=512out_H=7out_W=7kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1
  • W7dtype=f16N=1C_out=2048out_H=32out_W=32kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]deeplabv3-aspp-3x3-rate12
  • W9dtype=f16N=1C_out=256out_H=112out_W=112kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1
  • W11dtype=f16N=1C_out=256out_H=14out_W=14kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1
  • W13dtype=f16N=1C_out=64out_H=56out_W=56kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1
  • W14dtype=f16N=1C_out=32out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]groups=32mobilenetv2-depthwise
  • W15dtype=bf16N=2C_out=64out_H=56out_W=56kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W18dtype=f16N=2C_out=64out_H=56out_W=56kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W19dtype=bf16N=2C_out=64out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W22dtype=f16N=2C_out=64out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W23dtype=f16N=1C_out=128out_H=28out_W=28kH=3kW=3stride=[1,1]padding=[1,1]groups=32resnext-grouped-3x3
  • W25dtype=f16N=1C_out=128out_H=56out_W=56kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2
  • W27dtype=f16N=1C_out=128out_H=56out_W=56kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2
  • W29dtype=f16N=1C_out=3out_H=112out_W=112kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2
  • W30dtype=bf16N=1C_out=128out_H=28out_W=28kH=3kW=3stride=[2,2]padding=[1,1]stride2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.14×0.3599torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0323
0.0324
0.0320
···
W20.15×0.3319torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0307
0.0305
0.0306
···
W30.12×0.4286torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0336
0.0329
0.0324
···
W40.15×0.3354torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0314
0.0304
0.0301
···
W50.02×4.0781torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0712
0.0712
0.0707
···
W60.02×4.0293torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0740
0.0742
0.0741
···
W70.00×81.0684torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1030
0.1031
0.1035
···
W80.01×37.9281torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.2189
0.2188
0.2185
···
W90.01×37.9181torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1830
0.1830
0.1832
···
W100.05×1.0715torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0359
0.0360
0.0352
···
W110.06×1.0448torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0360
0.0357
0.0359
···
W120.01×7.9056torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0574
0.0570
0.0573
···
W130.01×7.8979torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0455
0.0455
0.0457
···
W140.71×0.0935torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0503
0.0498
0.0500
···
W150.47×0.1206torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0398
0.0395
0.0396
···
W160.37×0.1385torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0419
0.0420
0.0424
···
W170.41×0.1368torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0404
0.0403
0.0400
···
W180.40×0.1215torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0403
0.0403
0.0406
···
W190.04×1.3161torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0564
0.0312
0.0425
···
W200.04×1.3156torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0441
0.0440
0.0442
···
W210.04×1.3186torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0450
0.0442
0.0447
···
W220.05×1.3177torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0590
0.0307
0.0405
···
W230.31×0.2030torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0539
0.0539
0.0541
···
W240.01×6.9045torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0542
0.0542
0.0542
···
W250.01×6.8994torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0620
0.0382
0.0461
···
W260.01×20.2496torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1381
0.1375
0.1376
···
W270.00×20.2524torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0659
0.0656
0.0653
···
W280.02×2.7501torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0370
0.0371
0.0371
···
W290.02×2.7475torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0541
0.0297
0.0374
···
W300.02×2.2070torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0375
0.0370
0.0371
···
W310.02×2.2314torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0365
0.0356
0.0359
···

Conv2dBiasRelu

input: [N, C_in, H, W]bias: [C_out]

  • W1dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W2dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W3dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W4dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W5dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W6dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W7dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W8dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W9dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W10dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W11dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W12dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W13dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W14dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.03×1.6334torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0360
0.0356
0.0357
···
W20.11×0.4725torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0357
0.0354
0.0354
···
W30.19×0.4158torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0775
0.0775
0.0772
···
W40.01×41.1983torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.2266
0.2264
0.2268
···
W50.12×0.5000torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0441
0.0437
0.0434
···
W60.01×4.9813torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0585
0.0586
0.0590
···
W70.02×3.0994torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0432
0.0435
0.0434
···
W80.02×3.0996torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0431
0.0430
0.0431
···
W90.03×1.9804torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0690
0.0300
0.0503
···
W100.03×1.9825torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0629
0.0305
0.0503
···
W110.01×8.7918torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0607
0.0475
0.0495
···
W120.00×24.4659torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0668
0.0665
0.0665
···
W130.07×0.8106torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0563
0.0260
0.0450
···
W140.03×1.9054torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0377
0.0375
0.0375
···

Conv1dCausal

input: [N, C_in, L_in]

N=1

  • W1dtype=bf16C_in=128L_in=600C_out=256kW=10encodec-deep
  • W4dtype=f16C_in=128L_in=600C_out=256kW=10encodec-deep
  • W5dtype=bf16C_in=1L_in=24000C_out=32kW=7encodec-init
  • W8dtype=f16C_in=1L_in=24000C_out=32kW=7encodec-init
  • W9dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1
  • W12dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1
  • W13dtype=bf16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1
  • W16dtype=f16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1

input: [N, C_in, L_in]bias: [C_out]

N=1

  • W2dtype=bf16C_in=128L_in=600C_out=256kW=10encodec-deep-bias
  • W3dtype=f16C_in=128L_in=600C_out=256kW=10encodec-deep-bias
  • W6dtype=bf16C_in=1L_in=24000C_out=32kW=7encodec-init-bias
  • W7dtype=f16C_in=1L_in=24000C_out=32kW=7encodec-init-bias
  • W10dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1-bias
  • W11dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5wav2vec2-layer1-bias
  • W14dtype=bf16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1-bias
  • W15dtype=f16C_in=80L_in=3000C_out=1280kW=3whisper-large-conv1-bias
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×5.2867torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0537
0.0545
0.0542
···
W20.01×5.3150torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0553
0.0553
0.0559
···
W30.01×5.3111torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0545
0.0542
0.0545
···
W40.01×5.2912torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0544
0.0548
0.0540
···
W50.09×0.6435torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0437
0.0437
0.0440
···
W60.10×0.6607torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0447
0.0447
0.0447
···
W70.12×0.6617torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0591
0.0592
0.0586
···
W80.11×0.6438torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0587
0.0589
0.0592
···
W90.06×1.0870torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0703
0.0425
0.0510
···
W100.06×1.0490torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0529
0.0530
0.0529
···
W110.08×1.0466torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0693
0.0688
0.0689
···
W120.08×1.1024torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0969
0.0452
0.0690
···
W130.00×24.3873torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1008
0.0573
0.0853
···
W140.00×24.4970torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0886
0.0885
0.0885
···
W150.00×24.4935torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0875
0.0872
0.0876
···
W160.00×24.3831torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1013
0.0560
0.0854
···

Conv3dFwd

input: [N, C_in, D, H, W]

N=1

  • W1dtype=f16C_in=64D=8H=28W=28C_out=128stride=[1,1,1]padding=[1,1,1]groups=323d-resnext-grouped-k3
  • W2dtype=f16C_in=256D=8H=16W=16C_out=256stride=[1,1,1]padding=[6,6,6]dilation=[6,6,6]3d-unet-aspp-3x3x3-rate6
  • W6dtype=f16C_in=3D=16H=112W=112C_out=64stride=[1,1,1]padding=[1,1,1]r3d-stem-k3-s1
  • W7dtype=bf16C_in=32D=32H=64W=64C_out=64stride=[1,1,1]padding=[1,1,1]unet-encoder-k3-s1
  • W10dtype=f16C_in=64D=8H=56W=56C_out=128stride=[2,2,2]padding=[1,1,1]video-stage-downsample-k3-s2

dtype=f16

  • W3input_shape: [1, 2, 4, 5, 6]weight_shape: [4, 2, 3, 3, 3]padding=1default-3d
  • W4input_shape: [1, 1, 258, 256, 128]weight_shape: [1, 1, 1, 1, 1]grid-stride-sentinel

input: [N, C_in, D, H, W]bias: [C_out]

N=1padding=[1,1,1]

  • W5dtype=f16C_in=3D=16H=112W=112C_out=64stride=[1,1,1]r3d-stem-k3-s1-bias
  • W8dtype=bf16C_in=32D=32H=64W=64C_out=64stride=[1,1,1]unet-encoder-k3-s1-bias
  • W9dtype=f16C_in=64D=8H=56W=56C_out=128stride=[2,2,2]video-stage-downsample-k3-s2-bias
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.07×1.8140torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1353
0.1104
···
W20.00×169.8936torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1353
0.1249
···
W30.10×0.4829torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0447
0.0446
···
W44.03×0.2365torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.9497
0.9509
···
W50.04×5.1566torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2110
0.2042
···
W60.04×5.1413torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2028
0.2014
···
W70.00×39.7941torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1767
0.1659
···
W80.00×39.7875torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1770
0.1681
···
W90.00×2022.7671torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0886
0.0824
···
W100.00×2021.8212torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0914
0.0833
···

DilatedConv2d

input: [N, C_in, H, W]

dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]

  • W1deeplabv3-aspp-3x3-rate12
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.00×78.9945torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1340
0.1072
0.1190
···

Pooling

AvgPool3dFwd

input: [N, C, D_in, H_in, W_in]

stride=[2,2,2]

  • W1dtype=f16N=2C=64D_in=8H_in=28W_in=28kernel_size=[2,3,3]padding=[1,1,1]ceil_mode=truecount_include_pad=falseceil-video
  • W2dtype=bf16N=2C=24D_in=10H_in=20W_in=22kernel_size=[2,2,3]padding=[0,1,1]divisor
  • W3dtype=f16N=1C=32D_in=16H_in=56W_in=56kernel_size=[2,2,2]padding=[0,0,0]video-2x2x2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W11.75×0.0457Torch-NPU eager reference0.0800···
W23.15×0.0220Torch-NPU eager reference0.0693···
W32.37×0.0269Torch-NPU eager reference0.0638···

AvgPool1dFwd

input: [N, C, L_in]

  • W1dtype=f16N=4C=128L_in=4096kernel_size=3stride=2padding=1audio-downsample
  • W2dtype=bf16N=2C=128L_in=2048kernel_size=4stride=2padding=1ceil_mode=truecount_include_pad=falseceil
  • W3dtype=f16N=2C=256L_in=32000kernel_size=5stride=4padding=2long-temporal
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.55×0.0333torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0515
0.0558
···
W21.87×0.0213torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0395
0.0424
···
W31.05×0.1635torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1714
0.1742
···

AdaptiveMaxPool2dFwd

input: [N, C, H_in, W_in]

  • W1dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]global-1x1
  • W2dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]nondiv-7x7
  • W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]spp-6x6
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.57×0.0628torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0961
0.0921
···
W21.24×0.1020ops-nn-b2:aclnnAdaptiveMaxPool3d open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1220
0.1308
0.1227
···
W31.20×0.1400ops-nn-b2:aclnnAdaptiveMaxPool3d open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1601
0.1727
0.1631
···

AvgPool2dFwd

input: [N, C, H_in, W_in]

stride=[2,2]

  • W1dtype=bf16N=3C=96H_in=55W_in=57kernel_size=[3,5]padding=[1,2]ceil_mode=truecount_include_pad=falseceil-divisor
  • W2dtype=f16N=2C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]vision-3x3-s2
  • W3dtype=f16N=2C=128H_in=56W_in=56kernel_size=[5,5]padding=[2,2]vision-5x5-s2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.70×0.0370torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0625
0.0617
···
W20.91×0.0360torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0325
0.0209
···
W30.77×0.0495torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0376
0.0290
···

MaxPool3dFwd

input: [N, C, D_in, H_in, W_in]

  • W1dtype=f16N=8C=64D_in=16H_in=112W_in=112kernel_size=[1,2,2]stride=[1,2,2]padding=[0,0,0]c3d-pool1
  • W2dtype=f16N=4C=128D_in=16H_in=56W_in=56kernel_size=[2,2,2]stride=[2,2,2]padding=[0,0,0]c3d-pool2
  • W3dtype=bf16N=2C=64D_in=32H_in=112W_in=112kernel_size=[3,3,3]stride=[2,2,2]padding=[1,1,1]ceil_mode=truemedicalnet-stem
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.96×0.5791ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.5477
0.5506
0.5475
···
W20.75×0.2328ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1705
0.1732
0.1715
···
W31.45×0.8512ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘ
torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.2424
1.2369
1.2424
···

MaxPool2dFwd

input: [N, C, H_in, W_in]

stride=[2,2]

  • W1dtype=bf16N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=truealexnet-ceil
  • W2dtype=f16N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=truealexnet-ceil
  • W3dtype=f32N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=truealexnet-ceil
  • W4dtype=bf16N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]resnet-stem
  • W5dtype=f16N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]resnet-stem
  • W6dtype=f32N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]resnet-stem
  • W7dtype=bf16N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]vgg-block
  • W8dtype=f16N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]vgg-block
  • W9dtype=f32N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]vgg-block
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.1235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0975
0.0970
0.0973
···
W20.59×0.1245torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0670
0.0668
0.0666
···
W30.81×0.1265torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0936
0.0936
0.0935
···
W41.10×0.2938torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.3235
0.7558
0.3215
···
W51.01×0.2940torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.3075
0.2472
0.3065
···
W61.33×0.2774torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.3701
0.8113
0.3698
···
W71.16×0.0737torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0803
0.0801
0.0805
···
W80.94×0.0712torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0660
0.0707
0.0609
···
W90.91×0.0767torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0669
0.0665
0.0665
···

AdaptiveAvgPool2dFwd

input: [N, C, H_in, W_in]

  • W1dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]nondiv-7x7
  • W2dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]resnet-global
  • W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]spp-6x6
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.56×0.1017ops-nn-b2:aclnnAdaptiveAvgPool3d open-source ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0474
0.0553
0.0473
···
W20.54×0.0630torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0340
0.0213
···
W30.36×0.1410ops-nn-b2:aclnnAdaptiveAvgPool3d open-source ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0419
0.0555
0.0416
···

MaxPool1dFwd

input: [N, C, L_in]

  • W1dtype=bf16N=16C=128L_in=2048kernel_size=5stride=2padding=2dilation=2ceil_mode=trueecg-cnn-dilated
  • W2dtype=f16N=32C=80L_in=4096kernel_size=3stride=3sincnet-speaker-local
  • W3dtype=f16N=64C=256L_in=128kernel_size=128stride=1textcnn-global
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W1—·torch_npu eager vendor library basis····
W20.00×23.4675torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0910
0.0850
···
W30.01×5.0949torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0779
0.0689
···

Top-k

TopkSelectorFwd

index_score: [batch, seq_len, seq_len_kv, kv_group]starts, ends: [batch, seq_len], i32

dtype=f32batch=1seq_len=32768seq_len_kv=65536kv_group=1

  • W1topk=1024topk1024-s32k-kv64k
  • W2topk=2048topk2048-s32k-kv64k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.76×117.7379torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
89.9394
89.9271
···
W20.60×164.3116torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
98.9299
98.8865
···