Norm, Conv, Pool & Other¶
45 ops, 414 workloads — Normalization 23 · Convolution 13 · Pooling 8 · Top-k 1.
One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.
Normalization¶
AdaLayerNormFwd¶
- W1x: [1024, 1152]dtype=bf16
dit-xl-2 - W2x: [1024, 1152]dtype=f16
dit-xl-2 - W3x: [1, 4096]dtype=bf16
llama-8b-decode - W4x: [2048, 4096]dtype=bf16
llama-8b-prefill - W5x: [2048, 4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.73× | 0.0222 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0615 0.0521 | · | · | · |
| W2 | 2.09× | 0.0255 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0540 0.0503 | · | · | · |
| W3 | 3.14× | 0.0035 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0110 0.00962 0.00389 | · | · | · |
| W4 | 1.34× | 0.1030 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1393 0.1192 | · | · | · |
| W5 | 1.52× | 0.1005 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1625 0.1522 | · | · | · |
BatchNormBwd¶
dtype=f16
- W1x: [4, 128, 1024, 1024]
large-spatial - W2x: [32, 64]
resnet50-fc - W3x: [8, 64, 32, 32]
resnet50-stage1 - W4x: [4, 128, 32, 32]
resnet50-stage2 - W5x: [4, 256, 28, 28]
resnet50-stage3
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.17× | 11.4718 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 13.4796 13.4795 13.4889 | · | · | · |
| W2 | 1.64× | 0.0377 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0511 0.0508 0.0506 | · | · | · |
| W3 | 2.06× | 0.0318 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0732 0.0644 0.0650 | · | · | · |
| W4 | 2.15× | 0.0297 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0712 0.0630 0.0616 | · | · | · |
| W5 | 1.64× | 0.0525 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0935 0.0855 0.0825 | · | · | · |
L1NormFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-l1 - W3x: [2048, 4096]dtype=f16
hidden-state-l1 - W4x: [64, 32768]dtype=bf16
long-seq-l1
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.93× | 0.0357 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0334 0.0318 0.0198 | · | · | · |
| W2 | 1.96× | 0.0403 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0798 0.0695 0.0508 | · | · | · |
| W3 | 1.98× | 0.0399 | torch_npu eager vendor library basis | 0.0790 | · | · | · |
| W4 | 2.05× | 0.0227 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0481 0.0362 0.0200 | · | · | · |
L2NormFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-l2 - W3x: [2048, 4096]dtype=f16
hidden-state-l2 - W4x: [64, 32768]dtype=bf16
long-seq-l2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.14× | 0.0355 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0401 0.0208 | · | · | · |
| W2 | 1.76× | 0.0420 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0742 0.0678 | · | · | · |
| W3 | 1.79× | 0.0411 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0747 0.0670 | · | · | · |
| W4 | 1.88× | 0.0220 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0400 0.0350 | · | · | · |
FusedAddLayerNormFwd¶
- W1x: [1, 8192]dtype=bf16
llama-70b-decode - W2x: [2048, 8192]dtype=bf16
llama-70b-prefill - W3x: [2048, 8192]dtype=f16
llama-70b-prefill - W4x: [1, 4096]dtype=bf16
llama-8b-decode - W5x: [2048, 4096]dtype=bf16
llama-8b-prefill - W6x: [2048, 4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.38× | 0.00525 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0125 0.0103 | · | · | · |
| W2 | 1.26× | 0.2077 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2644 0.2492 | · | · | · |
| W3 | 1.30× | 0.1886 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2459 0.2288 | · | · | · |
| W4 | 2.09× | 0.004 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0085 0.00725 | · | · | · |
| W5 | 1.16× | 0.1185 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1309 0.1222 | · | · | · |
| W6 | 1.39× | 0.1135 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1584 0.1405 | · | · | · |
RMSNormQuantFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.84× | 0.0622 | torch_npu eager vendor library basis | 0.1143 | · | · | · |
| W2 | 1.35× | 4.5127 | torch_npu eager vendor library basis | 6.0799 | · | · | · |
| W3 | 1.36× | 4.6085 | torch_npu eager vendor library basis | 6.2891 | · | · | · |
| W4 | 1.74× | 0.0511 | torch_npu eager vendor library basis | 0.0887 | · | · | · |
| W5 | 1.12× | 1.9303 | torch_npu eager vendor library basis | 2.1546 | · | · | · |
| W6 | 1.14× | 1.9325 | torch_npu eager vendor library basis | 2.2107 | · | · | · |
| W7 | 1.71× | 0.0418 | torch_npu eager vendor library basis | 0.0712 | · | · | · |
| W8 | 1.14× | 0.9872 | torch_npu eager vendor library basis | 1.1266 | · | · | · |
| W9 | 1.12× | 0.9789 | torch_npu eager vendor library basis | 1.0958 | · | · | · |
LayerNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.66× | 0.00725 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0116 0.0110 | · | · | · |
| W2 | 1.22× | 0.2909 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.3538 0.3422 | · | · | · |
| W3 | 1.20× | 0.2868 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.3441 0.3369 | · | · | · |
| W4 | 1.76× | 0.00425 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.00725 0.00725 | · | · | · |
| W5 | 0.78× | 0.1628 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1280 0.1192 | · | · | · |
| W6 | 0.79× | 0.1653 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1300 0.1180 | · | · | · |
| W7 | 1.29× | 0.0035 | ops-nn:aclnnLayerNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0035 0.00475 0.0035 0.0335 | · | · | · |
| W8 | 0.71× | 0.0775 | ops-nn:aclnnLayerNorm open-source ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0537 0.0553 0.0529 0.0935 | · | · | · |
| W9 | 1.24× | 0.0757 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0931 0.0793 | · | · | · |
BatchNormFwd¶
dtype=f16
- W1x: [4, 128, 1024, 1024]
large-spatial - W2x: [32, 64]
resnet50-fc - W3x: [8, 64, 32, 32]
resnet50-stage1 - W4x: [4, 128, 32, 32]
resnet50-stage2 - W5x: [4, 256, 28, 28]
resnet50-stage3
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.27× | 2.6470 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 3.3580 3.3549 3.3563 | · | · | · |
| W2 | 1.13× | 0.00575 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0075 0.0065 0.0035 | · | · | · |
| W3 | 1.02× | 0.0150 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0152 0.0100 0.0107 | · | · | · |
| W4 | 1.21× | 0.0127 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0152 0.0120 0.0120 | · | · | · |
| W5 | 0.95× | 0.0185 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0175 0.0110 0.0136 | · | · | · |
InstanceNormFwd¶
- W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16
image-affine - W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16
image-affine - W5x: [4, 64, 30, 30]weight, bias: [64]dtype=f16
tail-spatial-affine - W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16
wider-channel-affine
- W3x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=bf16
image - W4x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=f16
image - W6x: [4, 64, 30, 30]running_mean, running_var: [64], f32dtype=f16
tail-spatial - W8x: [4, 256, 28, 28]running_mean, running_var: [256], f32dtype=f16
wider-channel
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.21× | 0.0217 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_npu eager vendor library basis | 0.0222 0.0201 | · | · | · |
| W2 | 1.32× | 0.0198 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_npu eager vendor library basis | 0.0231 0.0216 | · | · | · |
| W3 | 0.91× | 0.0532 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0480 0.0405 | · | · | · |
| W4 | 0.92× | 0.0525 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0486 0.0408 | · | · | · |
| W5 | 1.13× | 0.0163 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_npu eager vendor library | 0.0123 0.0168 | · | · | · |
| W6 | 0.70× | 0.0447 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0320 0.0253 | · | · | · |
| W7 | 1.51× | 0.0177 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_npu eager vendor library basis | 0.0220 0.0200 | · | · | · |
| W8 | 1.00× | 0.0516 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0542 0.0480 | · | · | · |
RMSNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.34× | 0.0055 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.00537 0.0240 0.0232 0.0226 | · | · | · |
| W2 | 0.86× | 0.2157 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.1745 0.8171 0.8174 0.8177 | · | · | · |
| W3 | 0.85× | 0.2120 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.1694 0.8265 0.8249 0.8260 17.4019 | · | · | · |
| W4 | 1.38× | 0.00325 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.00275 0.0217 0.0210 0.0208 | · | · | · |
| W5 | 0.82× | 0.0924 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.0675 0.2757 0.2769 0.2761 | · | · | · |
| W6 | 0.82× | 0.0907 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0663 0.2767 0.3035 0.2741 0.0642 | · | · | · |
| W7 | 1.45× | 0.00275 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0025 0.0302 0.00737 0.0195 0.00192 | · | · | · |
| W8 | 0.98× | 0.0545 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0441 0.1700 0.1775 0.1685 0.0347 | · | · | · |
| W9 | 0.93× | 0.0584 | ops-nn:aclnnRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:plain tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0452 0.1731 0.2300 0.1726 0.0343 | · | · | · |
LayerNormQuantFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.29× | 0.0777 | torch_npu eager vendor library basis | 0.1004 | · | · | · |
| W2 | 1.00× | 5.9374 | torch_npu eager vendor library basis | 5.9467 | · | · | · |
| W3 | 1.05× | 5.9260 | torch_npu eager vendor library basis | 6.2252 | · | · | · |
| W4 | 1.13× | 0.0605 | torch_npu eager vendor library basis | 0.0683 | · | · | · |
| W5 | 0.91× | 2.3194 | torch_npu eager vendor library basis | 2.1175 | · | · | · |
| W6 | 0.89× | 2.3469 | torch_npu eager vendor library basis | 2.0978 | · | · | · |
| W7 | 1.08× | 0.0506 | torch_npu eager vendor library basis | 0.0545 | · | · | · |
| W8 | 0.87× | 1.2195 | torch_npu eager vendor library basis | 1.0617 | · | · | · |
| W9 | 0.88× | 1.2189 | torch_npu eager vendor library basis | 1.0720 | · | · | · |
GroupNormFwd¶
- W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16num_groups=32
image-g32-affine - W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16num_groups=32
image-g32-affine - W5x: [4, 128, 30, 30]weight, bias: [128]dtype=f16num_groups=16
tail-spatial-g16-affine - W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16num_groups=32
wider-channel-g32-affine
- W3x: [8, 128, 32, 32]dtype=bf16num_groups=32
image-g32 - W4x: [8, 128, 32, 32]dtype=f16num_groups=32
image-g32 - W6x: [4, 128, 30, 30]dtype=f16num_groups=16
tail-spatial-g16 - W8x: [4, 256, 28, 28]dtype=f16num_groups=32
wider-channel-g32
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.82× | 0.0262 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0140 0.0140 0.0140 0.0140 | · | · | · |
| W2 | 0.69× | 0.0267 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0132 0.0130 0.0132 0.0132 | · | · | · |
| W3 | 0.98× | 0.0220 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0132 0.0235 0.0232 0.0160 | · | · | · |
| W4 | 0.96× | 0.0210 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0132 0.0225 0.0195 0.0158 | · | · | · |
| W5 | 0.90× | 0.0190 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0107 0.0106 0.0105 0.0105 | · | · | · |
| W6 | 1.07× | 0.0160 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0115 0.0198 0.0173 0.0138 | · | · | · |
| W7 | 1.00× | 0.0180 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0114 0.0112 0.0112 0.0110 | · | · | · |
| W8 | 1.16× | 0.0160 | ops-nn-b2:aclnnGroupNormSiluV2 open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0118 0.0205 0.0185 0.0135 | · | · | · |
FusedAddRMSNormFwd¶
- W1x: [1, 16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.86× | 0.00725 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.00525 0.0295 0.0288 | · | · | · |
| W2 | 0.90× | 0.3139 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.2755 1.1335 1.1365 · | · | · | · |
| W3 | 0.82× | 0.2913 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.2377 1.1364 1.0870 | · | · | · |
| W4 | 0.78× | 0.0045 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.00325 0.0413 0.0272 | · | · | · |
| W5 | 0.74× | 0.1713 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1235 0.4515 0.4954 | · | · | · |
| W6 | 0.80× | 0.1590 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1153 0.3962 0.4716 | · | · | · |
| W7 | 0.76× | 0.00363 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0025 0.0374 0.0250 · | · | · | · |
| W8 | 0.87× | 0.0907 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0750 0.2382 0.2285 · | · | · | · |
| W9 | 0.86× | 0.0872 | ops-nn:aclnnAddRmsNorm open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:selector tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0602 0.2209 0.2171 · | · | · | · |
InstanceNormBwd¶
- W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16
image-affine - W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16
image-affine - W5x: [4, 64, 30, 30]weight, bias: [64]dtype=f16
tail-spatial-affine - W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16
wider-channel-affine
- W3x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=bf16
image - W4x: [8, 128, 32, 32]running_mean, running_var: [128], f32dtype=f16
image - W6x: [4, 64, 30, 30]running_mean, running_var: [64], f32dtype=f16
tail-spatial - W8x: [4, 256, 28, 28]running_mean, running_var: [256], f32dtype=f16
wider-channel
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.56× | 0.4622 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2585 0.2526 | · | · | · |
| W2 | 0.55× | 0.4680 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2582 0.2540 | · | · | · |
| W3 | 0.43× | 0.3782 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1749 0.1487 | · | · | · |
| W4 | 0.43× | 0.3805 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1771 0.1496 | · | · | · |
| W5 | 0.41× | 0.5150 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2130 0.2070 | · | · | · |
| W6 | 0.36× | 0.3488 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1363 0.1060 | · | · | · |
| W7 | 0.28× | 0.6202 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1790 0.1693 | · | · | · |
| W8 | 0.21× | 0.4335 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0963 0.0710 | · | · | · |
GroupNormBwd¶
- W1x: [8, 128, 32, 32]weight, bias: [128]dtype=bf16num_groups=32
image-g32-affine - W2x: [8, 128, 32, 32]weight, bias: [128]dtype=f16num_groups=32
image-g32-affine - W5x: [4, 128, 30, 30]weight, bias: [128]dtype=f16num_groups=16
tail-spatial-g16-affine - W7x: [4, 256, 28, 28]weight, bias: [256]dtype=f16num_groups=32
wider-channel-g32-affine
- W3x: [8, 128, 32, 32]dtype=bf16num_groups=32
image-g32 - W4x: [8, 128, 32, 32]dtype=f16num_groups=32
image-g32 - W6x: [4, 128, 30, 30]dtype=f16num_groups=16
tail-spatial-g16 - W8x: [4, 256, 28, 28]dtype=f16num_groups=32
wider-channel-g32
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.47× | 0.5555 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2587 0.2497 | · | · | · |
| W2 | 0.48× | 0.5400 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2619 0.2511 | · | · | · |
| W3 | 0.47× | 0.5058 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2396 0.2290 | · | · | · |
| W4 | 0.48× | 0.5104 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2411 0.2299 | · | · | · |
| W5 | 0.38× | 0.5814 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2177 0.2127 | · | · | · |
| W6 | 0.35× | 0.5711 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1968 0.1940 | · | · | · |
| W7 | 0.24× | 0.7140 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1779 0.1695 | · | · | · |
| W8 | 0.23× | 0.6837 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1540 0.1497 | · | · | · |
WeightNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.85× | 0.0624 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0250 0.0245 | · | · | · |
| W2 | 0.26× | 2.6923 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.7278 0.6945 | · | · | · |
| W3 | 0.26× | 2.7535 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.7109 0.7335 | · | · | · |
| W4 | 0.53× | 0.0633 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0331 0.0210 | · | · | · |
| W5 | 0.27× | 1.0950 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3050 0.2989 | · | · | · |
| W6 | 0.27× | 1.1033 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3020 0.2988 | · | · | · |
| W7 | 0.61× | 0.0498 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0302 0.0182 | · | · | · |
| W8 | 0.31× | 0.5753 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1797 0.1787 | · | · | · |
| W9 | 0.32× | 0.5763 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1831 0.1814 | · | · | · |
GroupRMSNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.34× | 0.1215 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0214 0.0211 | · | · | · |
| W2 | 0.43× | 2.6197 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.1327 1.1786 | · | · | · |
| W3 | 0.44× | 2.5910 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.1342 1.1487 | · | · | · |
| W4 | 0.33× | 0.1089 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0355 0.0226 | · | · | · |
| W5 | 0.28× | 1.1604 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3225 0.3919 | · | · | · |
| W6 | 0.28× | 1.1596 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3234 0.3954 | · | · | · |
| W7 | 0.29× | 0.1096 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0314 0.0214 | · | · | · |
| W8 | 0.35× | 0.5785 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.2028 0.2076 | · | · | · |
| W9 | 0.35× | 0.5807 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.2030 0.2094 | · | · | · |
GemmaRMSNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.35× | 0.1289 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0272 0.0272 | · | · | · |
| W2 | 0.24× | 2.8287 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.7970 0.7252 | · | · | · |
| W3 | 0.26× | 2.8723 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8133 0.7340 | · | · | · |
| W4 | 0.32× | 0.1190 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0367 0.0255 | · | · | · |
| W5 | 0.24× | 1.1386 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2721 0.2652 | · | · | · |
| W6 | 0.24× | 1.1420 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2715 0.2646 | · | · | · |
| W7 | 0.27× | 0.1215 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0331 0.0235 | · | · | · |
| W8 | 0.30× | 0.5690 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1720 0.1747 | · | · | · |
| W9 | 0.31× | 0.5677 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1745 0.1754 | · | · | · |
QKNormFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.30× | 0.2245 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0455 0.0437 | · | · | · |
| W2 | 0.27× | 5.3171 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 1.6066 1.5517 | · | · | · |
| W3 | 0.27× | 5.3952 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 1.6015 1.4194 | · | · | · |
| W4 | 0.28× | 0.2050 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0579 0.0430 | · | · | · |
| W5 | 0.22× | 2.3981 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.5231 0.4988 | · | · | · |
| W6 | 0.22× | 2.3940 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.5114 0.4986 | · | · | · |
| W7 | 0.26× | 0.1995 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0505 0.0405 | · | · | · |
| W8 | 0.29× | 1.1075 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3285 0.3304 | · | · | · |
| W9 | 0.29× | 1.1240 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3311 0.3319 | · | · | · |
RMSNormBwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.12× | 0.4083 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0348 0.0339 | · | · | · |
| W2 | 0.32× | 6.3724 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 2.0200 2.2234 | · | · | · |
| W3 | 0.31× | 6.4097 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 2.0053 2.2535 | · | · | · |
| W4 | 0.15× | 0.2492 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0387 0.0305 | · | · | · |
| W5 | 0.29× | 2.7129 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.7750 0.9525 | · | · | · |
| W6 | 0.29× | 2.7153 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.7927 0.9451 | · | · | · |
| W7 | 0.19× | 0.1713 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0318 0.0267 | · | · | · |
| W8 | 0.29× | 1.2224 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3574 0.3711 | · | · | · |
| W9 | 0.29× | 1.2143 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3554 0.3805 | · | · | · |
SpectralNormPowerIterFwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.11× | 0.6388 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0725 0.0729 | · | · | · |
| W2 | 0.19× | 4.6319 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8709 0.8558 | · | · | · |
| W3 | 0.19× | 4.6324 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8629 0.8546 | · | · | · |
| W4 | 0.14× | 0.4525 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0607 0.0548 | · | · | · |
| W5 | 0.21× | 1.7875 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3654 0.3498 | · | · | · |
| W6 | 0.20× | 1.7924 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3663 0.3508 | · | · | · |
| W7 | 0.08× | 0.3643 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0285 0.0260 | · | · | · |
| W8 | 0.22× | 0.9865 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2136 0.2000 | · | · | · |
| W9 | 0.22× | 0.9859 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2130 0.2003 | · | · | · |
BatchNormInferenceFwd¶
dtype=f16
- W1x: [4, 128, 1024, 1024]
large-spatial - W2x: [32, 64]
resnet50-fc - W3x: [8, 64, 32, 32]
resnet50-stage1 - W4x: [4, 128, 32, 32]
resnet50-stage2 - W5x: [4, 256, 28, 28]
resnet50-stage3
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.07× | 42.3053 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 3.1336 3.1236 | · | · | · |
| W2 | 0.08× | 0.0796 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00775 0.0035 | · | · | · |
| W3 | 0.11× | 0.1669 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0168 0.0110 | · | · | · |
| W4 | 0.10× | 0.1704 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0195 0.0120 | · | · | · |
| W5 | 0.10× | 0.1817 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0190 0.0135 | · | · | · |
LayerNormBwd¶
- W1x: [1, 16384]normalized: [16384]dtype=bf16
llama-405b-decode - W2x: [2048, 16384]normalized: [16384]dtype=bf16
llama-405b-prefill - W3x: [2048, 16384]normalized: [16384]dtype=f16
llama-405b-prefill - W4x: [1, 8192]normalized: [8192]dtype=bf16
llama-70b-decode - W5x: [2048, 8192]normalized: [8192]dtype=bf16
llama-70b-prefill - W6x: [2048, 8192]normalized: [8192]dtype=f16
llama-70b-prefill - W7x: [1, 4096]normalized: [4096]dtype=bf16
llama-8b-decode - W8x: [2048, 4096]normalized: [4096]dtype=bf16
llama-8b-prefill - W9x: [2048, 4096]normalized: [4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.06× | 0.7544 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0314 0.0311 | · | · | · |
| W2 | 0.05× | 9.2674 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.4285 0.4290 | · | · | · |
| W3 | 0.05× | 9.2705 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4310 0.4308 | · | · | · |
| W4 | 0.08× | 0.4495 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0222 0.0220 | · | · | · |
| W5 | 0.04× | 3.7468 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1665 0.1669 | · | · | · |
| W6 | 0.04× | 3.7363 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1586 0.1593 | · | · | · |
| W7 | 0.09× | 0.2985 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0198 0.0201 | · | · | · |
| W8 | 0.06× | 1.7646 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1080 0.1084 | · | · | · |
| W9 | 0.06× | 1.7297 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1021 0.1020 | · | · | · |
Convolution¶
DepthwiseConv2d¶
input: [N, C_in, H, W]
dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]
- W1
mobilenetv2-depthwise
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.10× | 0.0602 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0668 0.0501 | · | · | · |
PointwiseConv2d¶
input: [N, C_in, H, W]bias: [C_out]
stride=[1,1]padding=[0,0]
- W1dtype=f16N=2C_in=128H=28W=28C_out=512
bottleneck-expand-1x1-bias - W3dtype=f16N=2C_in=512H=28W=28C_out=128
bottleneck-reduce-1x1-bias - W5dtype=f16N=1C_in=512H=7W=7C_out=2048
classifier-1x1-bias - W7dtype=f16N=1C_in=256H=14W=14C_out=1024
late-stage-1x1-bias - W10dtype=bf16N=2C_in=64H=56W=56C_out=256
resnet-1x1-bias - W11dtype=f16N=2C_in=64H=56W=56C_out=256
resnet-1x1-bias
input: [N, C_in, H, W]
stride=[1,1]padding=[0,0]
- W2dtype=f16N=2C_in=128H=28W=28C_out=512
bottleneck-expand-1x1 - W4dtype=f16N=2C_in=512H=28W=28C_out=128
bottleneck-reduce-1x1 - W6dtype=f16N=1C_in=512H=7W=7C_out=2048
classifier-1x1 - W8dtype=f16N=1C_in=256H=14W=14C_out=1024
late-stage-1x1 - W9dtype=bf16N=2C_in=64H=56W=56C_out=256
resnet-1x1 - W12dtype=f16N=2C_in=64H=56W=56C_out=256
resnet-1x1
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.35× | 0.1226 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0289 0.0291 0.0286 | · | · | · |
| W2 | 1.33× | 0.0345 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0460 0.0270 0.0304 | · | · | · |
| W3 | 0.67× | 0.0645 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0300 0.0299 0.0297 | · | · | · |
| W4 | 1.24× | 0.0396 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0441 0.0340 0.0316 | · | · | · |
| W5 | 0.34× | 0.2185 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0648 0.0648 0.0648 | · | · | · |
| W6 | 2.04× | 0.0340 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0649 0.0653 0.0649 | · | · | · |
| W7 | 0.50× | 0.1060 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0319 0.0320 0.0316 | · | · | · |
| W8 | 2.34× | 0.0213 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0331 0.0330 0.0328 | · | · | · |
| W9 | 1.75× | 0.0300 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0527 0.0310 0.0362 | · | · | · |
| W10 | 0.66× | 0.0795 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0377 0.0375 0.0377 | · | · | · |
| W11 | 0.66× | 0.0788 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0379 0.0376 0.0379 | · | · | · |
| W12 | 1.96× | 0.0310 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0535 0.0288 0.0355 | · | · | · |
GroupedConv2d¶
input: [N, C_in, H, W]
dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32
- W1
resnext-grouped-3x3
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.57× | 0.1125 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0643 0.0330 0.0478 | · | · | · |
Conv2dWgrad¶
- W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
input: [N, C_in, H, W]
- W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1 - W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1 - W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1 - W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
deeplabv3-aspp-3x3-rate12 - W9dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1 - W11dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1 - W13dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1 - W14dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32
mobilenetv2-depthwise - W15dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W18dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W19dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W22dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W23dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32
resnext-grouped-3x3 - W25dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2 - W27dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2 - W29dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2 - W30dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.25× | 0.0668 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0767 0.0776 | · | · | · |
| W2 | 1.37× | 0.0620 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0780 0.0770 | · | · | · |
| W3 | 0.86× | 0.0890 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0622 0.0624 | · | · | · |
| W4 | 0.98× | 0.0849 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0626 0.0622 | · | · | · |
| W5 | 1.08× | 0.1025 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1098 0.1101 | · | · | · |
| W6 | 1.07× | 0.1030 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1096 0.1101 | · | · | · |
| W7 | 0.03× | 7.2436 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1727 0.1722 | · | · | · |
| W8 | 0.41× | 0.5530 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2454 0.2139 | · | · | · |
| W9 | 0.40× | 0.5524 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2249 0.2129 | · | · | · |
| W10 | 2.21× | 0.0390 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0749 0.0750 | · | · | · |
| W11 | 2.26× | 0.0410 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0747 0.0736 | · | · | · |
| W12 | 0.39× | 0.2480 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1085 0.0809 | · | · | · |
| W13 | 0.38× | 0.2477 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0960 0.0808 | · | · | · |
| W14 | 0.28× | 0.3775 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0881 0.0880 | · | · | · |
| W15 | 0.99× | 0.0737 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0640 0.0648 | · | · | · |
| W16 | 0.93× | 0.0785 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0636 0.0643 | · | · | · |
| W17 | 1.08× | 0.0772 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0640 0.0640 | · | · | · |
| W18 | 1.01× | 0.0727 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0635 0.0638 | · | · | · |
| W19 | 0.42× | 0.1955 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0841 0.0683 | · | · | · |
| W20 | 0.45× | 0.1970 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0989 0.0671 | · | · | · |
| W21 | 0.45× | 0.1975 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1035 0.0675 | · | · | · |
| W22 | 0.45× | 0.1958 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0910 0.0674 | · | · | · |
| W23 | 0.43× | 0.2410 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0869 0.0858 | · | · | · |
| W24 | 0.03× | 2.9286 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1056 0.0752 | · | · | · |
| W25 | 0.03× | 2.9285 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0935 0.0762 | · | · | · |
| W26 | 0.01× | 8.2952 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1131 0.0810 | · | · | · |
| W27 | 0.01× | 8.2994 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1015 0.0801 | · | · | · |
| W28 | 0.24× | 0.3060 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0831 0.0546 | · | · | · |
| W29 | 0.24× | 0.3055 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0710 0.0541 | · | · | · |
| W30 | 0.11× | 0.7768 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0855 0.0689 | · | · | · |
| W31 | 0.11× | 0.7791 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0989 0.0688 | · | · | · |
Conv2dFp8Fwd¶
input: [N, C_in, H, W]
dtype=fp8e4m3
- W1N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W2N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W3N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.32× | 0.0415 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0973 0.1960 0.0858 | · | · | · |
| W2 | 0.08× | 1.2181 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0965 0.2034 0.0853 | · | · | · |
| W3 | 0.05× | 1.8235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0902 0.1270 0.0810 | · | · | · |
Conv2dFwd¶
input: [N, C_in, H, W]bias: [C_out]
- W1dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W3dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W5dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W14dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W17dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W19dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W23dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W24dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W27dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W28dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W31dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W33dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W35dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W39dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
input: [N, C_in, H, W]
- W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1 - W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1 - W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1 - W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
deeplabv3-aspp-3x3-rate12 - W15dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1 - W18dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1 - W20dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1 - W21dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32
mobilenetv2-depthwise - W22dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W25dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W26dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W29dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W30dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32
resnext-grouped-3x3 - W32dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2 - W34dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2 - W36dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2 - W38dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2
dtype=f16
- W8input_shape: [1, 16, 8, 8]weight_shape: [32, 16, 3, 3]
default-dense-cube-fp16 - W16input_shape: [16, 256, 56, 56]weight_shape: [256, 1, 1, 1]groups=256
large-grid-depthwise-fp16
- W9input_shape: [1, 16, 8, 8]padding: [1, 1]weight_shape: [16, 16, 3, 3]dtype=bf16
dense-cube-bf16 - W10input_shape: [1, 4, 8, 9]padding: [1, 1]weight_shape: [4, 1, 3, 3]dtype=f16groups=4
depthwise-direct-fp16 - W12input_shape: [1, 2, 5, 6]padding: [1, 1]weight_shape: [3, 2, 3, 3]dtype=f32
fp32-direct - W13input_shape: [1, 8, 9, 10]padding: [1, 1]weight_shape: [8, 2, 3, 3]dtype=f16groups=4
grouped-direct-fp16
dtype=f16
- W11dilation, padding: [2, 2]input_shape: [1, 16, 9, 10]weight_shape: [16, 16, 3, 3]
dilation-cube-fp16
dtype=f16
- W37input_shape: [1, 3, 10, 11]padding: [1, 1]stride: [2, 2]weight_shape: [16, 3, 3, 3]
stride-padding-nondiv-cube-fp16
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.38× | 0.1185 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0306 0.0305 | · | · | · |
| W2 | 1.33× | 0.0354 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0318 0.0314 | · | · | · |
| W3 | 0.71× | 0.0635 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0314 0.0309 | · | · | · |
| W4 | 1.18× | 0.0395 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0310 0.0309 | · | · | · |
| W5 | 0.33× | 0.2105 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0645 0.0640 | · | · | · |
| W6 | 1.97× | 0.0330 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0640 0.0643 | · | · | · |
| W7 | 0.00× | 79.0235 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1179 0.1181 | · | · | · |
| W8 | 0.38× | 0.0881 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0267 0.0264 | · | · | · |
| W9 | 0.57× | 0.0540 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0219 0.0206 | · | · | · |
| W10 | 0.66× | 0.0840 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0444 0.0432 | · | · | · |
| W11 | 0.28× | 0.1180 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0214 0.0200 | · | · | · |
| W12 | 0.21× | 0.1457 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0262 0.0259 | · | · | · |
| W13 | 0.31× | 0.1595 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0372 0.0360 | · | · | · |
| W14 | 0.01× | 29.0843 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2147 0.2147 | · | · | · |
| W15 | 0.01× | 29.0799 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2286 0.2157 | · | · | · |
| W16 | 0.36× | 0.4460 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1510 0.1514 | · | · | · |
| W17 | 0.41× | 0.1062 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0324 0.0326 | · | · | · |
| W18 | 2.78× | 0.0200 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0345 0.0345 | · | · | · |
| W19 | 0.02× | 4.2020 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0460 0.0459 | · | · | · |
| W20 | 0.01× | 4.2009 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0589 0.0461 | · | · | · |
| W21 | 0.94× | 0.0597 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0486 0.0494 | · | · | · |
| W22 | 1.66× | 0.0294 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0377 0.0374 | · | · | · |
| W23 | 0.71× | 0.0707 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0380 0.0380 | · | · | · |
| W24 | 0.72× | 0.0688 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0380 0.0375 | · | · | · |
| W25 | 1.63× | 0.0276 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0372 0.0375 | · | · | · |
| W26 | 0.04× | 1.2043 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0485 0.0384 | · | · | · |
| W27 | 0.05× | 1.2067 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0382 0.0381 | · | · | · |
| W28 | 0.05× | 1.2054 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0396 0.0390 | · | · | · |
| W29 | 0.05× | 1.2069 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0498 0.0385 | · | · | · |
| W30 | 0.45× | 0.1130 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0473 0.0475 | · | · | · |
| W31 | 0.01× | 8.3829 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0387 0.0380 | · | · | · |
| W32 | 0.01× | 8.3692 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0515 0.0380 | · | · | · |
| W33 | 0.00× | 24.0599 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0558 0.0551 | · | · | · |
| W34 | 0.00× | 24.0326 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0590 0.0462 | · | · | · |
| W35 | 0.13× | 0.4036 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0320 0.0318 | · | · | · |
| W36 | 0.11× | 0.3975 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0454 0.0326 | · | · | · |
| W37 | 0.95× | 0.0367 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0265 0.0257 | · | · | · |
| W38 | 0.03× | 1.8193 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0460 0.0335 | · | · | · |
| W39 | 0.02× | 1.8265 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0335 0.0328 | · | · | · |
Conv1dFwd¶
dtype=f16
- W1input_shape: [1, 4, 17]weight_shape: [8, 2, 3]dilation=2groups=2padding=2
dilation-groups - W10input_shape: [1, 2, 17]weight_shape: [4, 2, 3]padding=1stride=2
stride-padding
input: [N, C_in, L_in]
N=1
- W2dtype=bf16C_in=128L_in=600C_out=256kW=10
encodec-deep - W5dtype=f16C_in=128L_in=600C_out=256kW=10
encodec-deep - W6dtype=bf16C_in=1L_in=24000C_out=32kW=7
encodec-init - W9dtype=f16C_in=1L_in=24000C_out=32kW=7
encodec-init - W11dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1 - W14dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1 - W15dtype=bf16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1 - W18dtype=f16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1
input: [N, C_in, L_in]bias: [C_out]
N=1
- W3dtype=bf16C_in=128L_in=600C_out=256kW=10
encodec-deep-bias - W4dtype=f16C_in=128L_in=600C_out=256kW=10
encodec-deep-bias - W7dtype=bf16C_in=1L_in=24000C_out=32kW=7
encodec-init-bias - W8dtype=f16C_in=1L_in=24000C_out=32kW=7
encodec-init-bias - W12dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1-bias - W13dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1-bias - W16dtype=bf16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1-bias - W17dtype=f16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1-bias
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.99× | 0.0527 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0424 0.0415 | · | · | · |
| W2 | 0.04× | 1.3505 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0470 0.0365 | · | · | · |
| W3 | 0.03× | 1.3445 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0384 0.0386 | · | · | · |
| W4 | 0.04× | 1.3594 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0377 0.0371 | · | · | · |
| W5 | 0.03× | 1.3445 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0473 0.0375 | · | · | · |
| W6 | 1.03× | 0.0320 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0328 0.0203 | · | · | · |
| W7 | 1.02× | 0.0323 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0213 0.0210 | · | · | · |
| W8 | 1.60× | 0.0304 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0328 0.0325 | · | · | · |
| W9 | 1.79× | 0.0269 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0478 0.0326 | · | · | · |
| W10 | 1.21× | 0.0320 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0302 0.0290 | · | · | · |
| W11 | 0.00× | 41.5343 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0350 0.0245 | · | · | · |
| W12 | 0.00× | 42.7620 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0257 0.0249 | · | · | · |
| W13 | 0.00× | 42.7612 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0425 0.0424 | · | · | · |
| W14 | 0.00× | 41.5373 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0532 0.0425 | · | · | · |
| W15 | 0.04× | 1.6562 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0721 0.0609 | · | · | · |
| W16 | 0.04× | 1.6618 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0622 0.0626 | · | · | · |
| W17 | 0.05× | 1.6615 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0607 0.0605 | · | · | · |
| W18 | 0.04× | 1.6499 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0715 0.0606 | · | · | · |
Conv2dDgrad¶
- W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
- W2input: [2, 128, 28, 28]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1 - W4input: [2, 512, 28, 28]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1 - W6input: [1, 512, 7, 7]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1 - W7input: [1, 2048, 32, 32]dtype=f16C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
deeplabv3-aspp-3x3-rate12 - W9input: [1, 256, 112, 112]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1 - W11input: [1, 256, 14, 14]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1 - W13input: [1, 64, 56, 56]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1 - W14input: [1, 32, 56, 56]dtype=f16C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32
mobilenetv2-depthwise - W15input: [2, 64, 56, 56]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W18input: [2, 64, 56, 56]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W19input: [2, 64, 56, 56]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W22input: [2, 64, 56, 56]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W23input: [1, 128, 28, 28]dtype=f16C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32
resnext-grouped-3x3 - W25input: [1, 128, 56, 56]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2 - W27input: [1, 128, 56, 56]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2 - W29input: [1, 3, 112, 112]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2 - W30input: [1, 128, 28, 28]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.20× | 0.3254 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0398 0.0393 0.0390 | · | · | · |
| W2 | 0.22× | 0.3274 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0399 0.0398 0.0393 | · | · | · |
| W3 | 0.21× | 0.3247 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0447 0.0444 0.0447 | · | · | · |
| W4 | 0.23× | 0.3244 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0449 0.0445 0.0445 | · | · | · |
| W5 | 0.02× | 4.0235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0861 0.0865 0.0861 | · | · | · |
| W6 | 0.02× | 4.0235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0865 0.0866 0.0865 | · | · | · |
| W7 | 0.00× | 81.0637 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1273 0.1280 0.1276 | · | · | · |
| W8 | 0.01× | 37.9130 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.2029 0.2030 0.2030 | · | · | · |
| W9 | 0.01× | 37.9222 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.2020 0.2019 0.2023 | · | · | · |
| W10 | 0.07× | 1.0453 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0423 0.0423 0.0420 | · | · | · |
| W11 | 0.07× | 1.0428 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0421 0.0420 0.0421 | · | · | · |
| W12 | 0.01× | 7.8991 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0594 0.0592 0.0596 | · | · | · |
| W13 | 0.01× | 7.9009 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0597 0.0594 0.0594 | · | · | · |
| W14 | 0.74× | 0.0963 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0646 0.0646 0.0650 | · | · | · |
| W15 | 0.65× | 0.1133 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0548 0.0542 0.0546 | · | · | · |
| W16 | 0.59× | 0.1138 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0545 0.0548 0.0542 | · | · | · |
| W17 | 0.61× | 0.1123 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0548 0.0550 0.0548 | · | · | · |
| W18 | 0.67× | 0.1123 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0548 0.0548 0.0545 | · | · | · |
| W19 | 0.05× | 1.3177 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0689 0.0305 0.0537 | · | · | · |
| W20 | 0.05× | 1.3157 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0833 0.0305 0.0539 | · | · | · |
| W21 | 0.05× | 1.3166 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0785 0.0307 0.0540 | · | · | · |
| W22 | 0.05× | 1.3150 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0680 0.0310 0.0548 | · | · | · |
| W23 | 0.39× | 0.2032 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0626 0.0628 0.0617 | · | · | · |
| W24 | 0.01× | 6.9001 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0900 0.0750 0.0638 | · | · | · |
| W25 | 0.01× | 6.9005 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0785 0.0747 0.0639 | · | · | · |
| W26 | 0.00× | 20.2512 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0829 0.0830 0.0828 | · | · | · |
| W27 | 0.00× | 20.2491 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0828 0.0829 0.0831 | · | · | · |
| W28 | 0.02× | 2.7498 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0747 0.0532 0.0455 | · | · | · |
| W29 | 0.02× | 2.7490 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0635 0.0535 0.0457 | · | · | · |
| W30 | 0.03× | 2.2116 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0437 0.0430 0.0437 | · | · | · |
| W31 | 0.03× | 2.2049 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0437 0.0435 0.0435 | · | · | · |
Conv2dTranspose¶
input: [N, C_out, out_H, out_W]bias: [C_in]
- W1dtype=f16N=2C_out=128out_H=28out_W=28C_in=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W3dtype=f16N=2C_out=512out_H=28out_W=28C_in=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W5dtype=f16N=1C_out=512out_H=7out_W=7C_in=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W8dtype=f16N=1C_out=256out_H=112out_W=112C_in=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W10dtype=f16N=1C_out=256out_H=14out_W=14C_in=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W12dtype=f16N=1C_out=64out_H=56out_W=56C_in=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W16dtype=bf16N=2C_out=64out_H=56out_W=56C_in=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W17dtype=f16N=2C_out=64out_H=56out_W=56C_in=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W20dtype=bf16N=2C_out=64out_H=56out_W=56C_in=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W21dtype=f16N=2C_out=64out_H=56out_W=56C_in=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W24dtype=f16N=1C_out=128out_H=56out_W=56C_in=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W26dtype=f16N=1C_out=128out_H=56out_W=56C_in=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W28dtype=f16N=1C_out=3out_H=112out_W=112C_in=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W31dtype=bf16N=1C_out=128out_H=28out_W=28C_in=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
input: [N, C_out, out_H, out_W]
- W2dtype=f16N=2C_out=128out_H=28out_W=28kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1 - W4dtype=f16N=2C_out=512out_H=28out_W=28kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1 - W6dtype=f16N=1C_out=512out_H=7out_W=7kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1 - W7dtype=f16N=1C_out=2048out_H=32out_W=32kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
deeplabv3-aspp-3x3-rate12 - W9dtype=f16N=1C_out=256out_H=112out_W=112kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1 - W11dtype=f16N=1C_out=256out_H=14out_W=14kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1 - W13dtype=f16N=1C_out=64out_H=56out_W=56kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1 - W14dtype=f16N=1C_out=32out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]groups=32
mobilenetv2-depthwise - W15dtype=bf16N=2C_out=64out_H=56out_W=56kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W18dtype=f16N=2C_out=64out_H=56out_W=56kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W19dtype=bf16N=2C_out=64out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W22dtype=f16N=2C_out=64out_H=56out_W=56kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W23dtype=f16N=1C_out=128out_H=28out_W=28kH=3kW=3stride=[1,1]padding=[1,1]groups=32
resnext-grouped-3x3 - W25dtype=f16N=1C_out=128out_H=56out_W=56kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2 - W27dtype=f16N=1C_out=128out_H=56out_W=56kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2 - W29dtype=f16N=1C_out=3out_H=112out_W=112kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2 - W30dtype=bf16N=1C_out=128out_H=28out_W=28kH=3kW=3stride=[2,2]padding=[1,1]
stride2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.14× | 0.3599 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0323 0.0324 0.0320 | · | · | · |
| W2 | 0.15× | 0.3319 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0307 0.0305 0.0306 | · | · | · |
| W3 | 0.12× | 0.4286 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0336 0.0329 0.0324 | · | · | · |
| W4 | 0.15× | 0.3354 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0314 0.0304 0.0301 | · | · | · |
| W5 | 0.02× | 4.0781 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0712 0.0712 0.0707 | · | · | · |
| W6 | 0.02× | 4.0293 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0740 0.0742 0.0741 | · | · | · |
| W7 | 0.00× | 81.0684 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1030 0.1031 0.1035 | · | · | · |
| W8 | 0.01× | 37.9281 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.2189 0.2188 0.2185 | · | · | · |
| W9 | 0.01× | 37.9181 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1830 0.1830 0.1832 | · | · | · |
| W10 | 0.05× | 1.0715 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0359 0.0360 0.0352 | · | · | · |
| W11 | 0.06× | 1.0448 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0360 0.0357 0.0359 | · | · | · |
| W12 | 0.01× | 7.9056 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0574 0.0570 0.0573 | · | · | · |
| W13 | 0.01× | 7.8979 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0455 0.0455 0.0457 | · | · | · |
| W14 | 0.71× | 0.0935 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0503 0.0498 0.0500 | · | · | · |
| W15 | 0.47× | 0.1206 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0398 0.0395 0.0396 | · | · | · |
| W16 | 0.37× | 0.1385 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0419 0.0420 0.0424 | · | · | · |
| W17 | 0.41× | 0.1368 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0404 0.0403 0.0400 | · | · | · |
| W18 | 0.40× | 0.1215 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0403 0.0403 0.0406 | · | · | · |
| W19 | 0.04× | 1.3161 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0564 0.0312 0.0425 | · | · | · |
| W20 | 0.04× | 1.3156 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0441 0.0440 0.0442 | · | · | · |
| W21 | 0.04× | 1.3186 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0450 0.0442 0.0447 | · | · | · |
| W22 | 0.05× | 1.3177 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0590 0.0307 0.0405 | · | · | · |
| W23 | 0.31× | 0.2030 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0539 0.0539 0.0541 | · | · | · |
| W24 | 0.01× | 6.9045 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0542 0.0542 0.0542 | · | · | · |
| W25 | 0.01× | 6.8994 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0620 0.0382 0.0461 | · | · | · |
| W26 | 0.01× | 20.2496 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1381 0.1375 0.1376 | · | · | · |
| W27 | 0.00× | 20.2524 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0659 0.0656 0.0653 | · | · | · |
| W28 | 0.02× | 2.7501 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0370 0.0371 0.0371 | · | · | · |
| W29 | 0.02× | 2.7475 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0541 0.0297 0.0374 | · | · | · |
| W30 | 0.02× | 2.2070 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0375 0.0370 0.0371 | · | · | · |
| W31 | 0.02× | 2.2314 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0365 0.0356 0.0359 | · | · | · |
Conv2dBiasRelu¶
input: [N, C_in, H, W]bias: [C_out]
- W1dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W2dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W3dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W4dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W5dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W6dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W7dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W8dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W9dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W10dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W11dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W12dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W13dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W14dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.03× | 1.6334 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0360 0.0356 0.0357 | · | · | · |
| W2 | 0.11× | 0.4725 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0357 0.0354 0.0354 | · | · | · |
| W3 | 0.19× | 0.4158 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0775 0.0775 0.0772 | · | · | · |
| W4 | 0.01× | 41.1983 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.2266 0.2264 0.2268 | · | · | · |
| W5 | 0.12× | 0.5000 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0441 0.0437 0.0434 | · | · | · |
| W6 | 0.01× | 4.9813 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0585 0.0586 0.0590 | · | · | · |
| W7 | 0.02× | 3.0994 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0432 0.0435 0.0434 | · | · | · |
| W8 | 0.02× | 3.0996 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0431 0.0430 0.0431 | · | · | · |
| W9 | 0.03× | 1.9804 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0690 0.0300 0.0503 | · | · | · |
| W10 | 0.03× | 1.9825 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0629 0.0305 0.0503 | · | · | · |
| W11 | 0.01× | 8.7918 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0607 0.0475 0.0495 | · | · | · |
| W12 | 0.00× | 24.4659 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0668 0.0665 0.0665 | · | · | · |
| W13 | 0.07× | 0.8106 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0563 0.0260 0.0450 | · | · | · |
| W14 | 0.03× | 1.9054 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0377 0.0375 0.0375 | · | · | · |
Conv1dCausal¶
input: [N, C_in, L_in]
N=1
- W1dtype=bf16C_in=128L_in=600C_out=256kW=10
encodec-deep - W4dtype=f16C_in=128L_in=600C_out=256kW=10
encodec-deep - W5dtype=bf16C_in=1L_in=24000C_out=32kW=7
encodec-init - W8dtype=f16C_in=1L_in=24000C_out=32kW=7
encodec-init - W9dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1 - W12dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1 - W13dtype=bf16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1 - W16dtype=f16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1
input: [N, C_in, L_in]bias: [C_out]
N=1
- W2dtype=bf16C_in=128L_in=600C_out=256kW=10
encodec-deep-bias - W3dtype=f16C_in=128L_in=600C_out=256kW=10
encodec-deep-bias - W6dtype=bf16C_in=1L_in=24000C_out=32kW=7
encodec-init-bias - W7dtype=f16C_in=1L_in=24000C_out=32kW=7
encodec-init-bias - W10dtype=bf16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1-bias - W11dtype=f16C_in=1L_in=16000C_out=512kW=10stride=5
wav2vec2-layer1-bias - W14dtype=bf16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1-bias - W15dtype=f16C_in=80L_in=3000C_out=1280kW=3
whisper-large-conv1-bias
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 5.2867 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0537 0.0545 0.0542 | · | · | · |
| W2 | 0.01× | 5.3150 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0553 0.0553 0.0559 | · | · | · |
| W3 | 0.01× | 5.3111 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0545 0.0542 0.0545 | · | · | · |
| W4 | 0.01× | 5.2912 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0544 0.0548 0.0540 | · | · | · |
| W5 | 0.09× | 0.6435 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0437 0.0437 0.0440 | · | · | · |
| W6 | 0.10× | 0.6607 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0447 0.0447 0.0447 | · | · | · |
| W7 | 0.12× | 0.6617 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0591 0.0592 0.0586 | · | · | · |
| W8 | 0.11× | 0.6438 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0587 0.0589 0.0592 | · | · | · |
| W9 | 0.06× | 1.0870 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0703 0.0425 0.0510 | · | · | · |
| W10 | 0.06× | 1.0490 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0529 0.0530 0.0529 | · | · | · |
| W11 | 0.08× | 1.0466 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0693 0.0688 0.0689 | · | · | · |
| W12 | 0.08× | 1.1024 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0969 0.0452 0.0690 | · | · | · |
| W13 | 0.00× | 24.3873 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1008 0.0573 0.0853 | · | · | · |
| W14 | 0.00× | 24.4970 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0886 0.0885 0.0885 | · | · | · |
| W15 | 0.00× | 24.4935 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0875 0.0872 0.0876 | · | · | · |
| W16 | 0.00× | 24.3831 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1013 0.0560 0.0854 | · | · | · |
Conv3dFwd¶
input: [N, C_in, D, H, W]
N=1
- W1dtype=f16C_in=64D=8H=28W=28C_out=128stride=[1,1,1]padding=[1,1,1]groups=32
3d-resnext-grouped-k3 - W2dtype=f16C_in=256D=8H=16W=16C_out=256stride=[1,1,1]padding=[6,6,6]dilation=[6,6,6]
3d-unet-aspp-3x3x3-rate6 - W6dtype=f16C_in=3D=16H=112W=112C_out=64stride=[1,1,1]padding=[1,1,1]
r3d-stem-k3-s1 - W7dtype=bf16C_in=32D=32H=64W=64C_out=64stride=[1,1,1]padding=[1,1,1]
unet-encoder-k3-s1 - W10dtype=f16C_in=64D=8H=56W=56C_out=128stride=[2,2,2]padding=[1,1,1]
video-stage-downsample-k3-s2
dtype=f16
- W3input_shape: [1, 2, 4, 5, 6]weight_shape: [4, 2, 3, 3, 3]padding=1
default-3d - W4input_shape: [1, 1, 258, 256, 128]weight_shape: [1, 1, 1, 1, 1]
grid-stride-sentinel
input: [N, C_in, D, H, W]bias: [C_out]
N=1padding=[1,1,1]
- W5dtype=f16C_in=3D=16H=112W=112C_out=64stride=[1,1,1]
r3d-stem-k3-s1-bias - W8dtype=bf16C_in=32D=32H=64W=64C_out=64stride=[1,1,1]
unet-encoder-k3-s1-bias - W9dtype=f16C_in=64D=8H=56W=56C_out=128stride=[2,2,2]
video-stage-downsample-k3-s2-bias
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.07× | 1.8140 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1353 0.1104 | · | · | · |
| W2 | 0.00× | 169.8936 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1353 0.1249 | · | · | · |
| W3 | 0.10× | 0.4829 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0447 0.0446 | · | · | · |
| W4 | 4.03× | 0.2365 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.9497 0.9509 | · | · | · |
| W5 | 0.04× | 5.1566 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2110 0.2042 | · | · | · |
| W6 | 0.04× | 5.1413 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2028 0.2014 | · | · | · |
| W7 | 0.00× | 39.7941 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1767 0.1659 | · | · | · |
| W8 | 0.00× | 39.7875 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1770 0.1681 | · | · | · |
| W9 | 0.00× | 2022.7671 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0886 0.0824 | · | · | · |
| W10 | 0.00× | 2021.8212 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0914 0.0833 | · | · | · |
DilatedConv2d¶
input: [N, C_in, H, W]
dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
- W1
deeplabv3-aspp-3x3-rate12
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.00× | 78.9945 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1340 0.1072 0.1190 | · | · | · |
Pooling¶
AvgPool3dFwd¶
input: [N, C, D_in, H_in, W_in]
stride=[2,2,2]
- W1dtype=f16N=2C=64D_in=8H_in=28W_in=28kernel_size=[2,3,3]padding=[1,1,1]ceil_mode=truecount_include_pad=false
ceil-video - W2dtype=bf16N=2C=24D_in=10H_in=20W_in=22kernel_size=[2,2,3]padding=[0,1,1]
divisor - W3dtype=f16N=1C=32D_in=16H_in=56W_in=56kernel_size=[2,2,2]padding=[0,0,0]
video-2x2x2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.75× | 0.0457 | Torch-NPU eager reference | 0.0800 | · | · | · |
| W2 | 3.15× | 0.0220 | Torch-NPU eager reference | 0.0693 | · | · | · |
| W3 | 2.37× | 0.0269 | Torch-NPU eager reference | 0.0638 | · | · | · |
AvgPool1dFwd¶
input: [N, C, L_in]
- W1dtype=f16N=4C=128L_in=4096kernel_size=3stride=2padding=1
audio-downsample - W2dtype=bf16N=2C=128L_in=2048kernel_size=4stride=2padding=1ceil_mode=truecount_include_pad=false
ceil - W3dtype=f16N=2C=256L_in=32000kernel_size=5stride=4padding=2
long-temporal
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.55× | 0.0333 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0515 0.0558 | · | · | · |
| W2 | 1.87× | 0.0213 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0395 0.0424 | · | · | · |
| W3 | 1.05× | 0.1635 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1714 0.1742 | · | · | · |
AdaptiveMaxPool2dFwd¶
input: [N, C, H_in, W_in]
- W1dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]
global-1x1 - W2dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]
nondiv-7x7 - W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]
spp-6x6
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.57× | 0.0628 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0961 0.0921 | · | · | · |
| W2 | 1.24× | 0.1020 | ops-nn-b2:aclnnAdaptiveMaxPool3d open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1220 0.1308 0.1227 | · | · | · |
| W3 | 1.20× | 0.1400 | ops-nn-b2:aclnnAdaptiveMaxPool3d open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1601 0.1727 0.1631 | · | · | · |
AvgPool2dFwd¶
input: [N, C, H_in, W_in]
stride=[2,2]
- W1dtype=bf16N=3C=96H_in=55W_in=57kernel_size=[3,5]padding=[1,2]ceil_mode=truecount_include_pad=false
ceil-divisor - W2dtype=f16N=2C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]
vision-3x3-s2 - W3dtype=f16N=2C=128H_in=56W_in=56kernel_size=[5,5]padding=[2,2]
vision-5x5-s2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.70× | 0.0370 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0625 0.0617 | · | · | · |
| W2 | 0.91× | 0.0360 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0325 0.0209 | · | · | · |
| W3 | 0.77× | 0.0495 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0376 0.0290 | · | · | · |
MaxPool3dFwd¶
input: [N, C, D_in, H_in, W_in]
- W1dtype=f16N=8C=64D_in=16H_in=112W_in=112kernel_size=[1,2,2]stride=[1,2,2]padding=[0,0,0]
c3d-pool1 - W2dtype=f16N=4C=128D_in=16H_in=56W_in=56kernel_size=[2,2,2]stride=[2,2,2]padding=[0,0,0]
c3d-pool2 - W3dtype=bf16N=2C=64D_in=32H_in=112W_in=112kernel_size=[3,3,3]stride=[2,2,2]padding=[1,1,1]ceil_mode=true
medicalnet-stem
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.96× | 0.5791 | ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.5477 0.5506 0.5475 | · | · | · |
| W2 | 0.75× | 0.2328 | ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1705 0.1732 0.1715 | · | · | · |
| W3 | 1.45× | 0.8512 | ops-nn:aclnnMaxPool3dWithArgmax open-source ⓘtorch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.2424 1.2369 1.2424 | · | · | · |
MaxPool2dFwd¶
input: [N, C, H_in, W_in]
stride=[2,2]
- W1dtype=bf16N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=true
alexnet-ceil - W2dtype=f16N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=true
alexnet-ceil - W3dtype=f32N=32C=64H_in=55W_in=55kernel_size=[3,3]padding=[0,0]ceil_mode=true
alexnet-ceil - W4dtype=bf16N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]
resnet-stem - W5dtype=f16N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]
resnet-stem - W6dtype=f32N=32C=64H_in=112W_in=112kernel_size=[3,3]padding=[1,1]
resnet-stem - W7dtype=bf16N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]
vgg-block - W8dtype=f16N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]
vgg-block - W9dtype=f32N=16C=128H_in=56W_in=56kernel_size=[2,2]padding=[0,0]
vgg-block
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.1235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0975 0.0970 0.0973 | · | · | · |
| W2 | 0.59× | 0.1245 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0670 0.0668 0.0666 | · | · | · |
| W3 | 0.81× | 0.1265 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0936 0.0936 0.0935 | · | · | · |
| W4 | 1.10× | 0.2938 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.3235 0.7558 0.3215 | · | · | · |
| W5 | 1.01× | 0.2940 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.3075 0.2472 0.3065 | · | · | · |
| W6 | 1.33× | 0.2774 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.3701 0.8113 0.3698 | · | · | · |
| W7 | 1.16× | 0.0737 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0803 0.0801 0.0805 | · | · | · |
| W8 | 0.94× | 0.0712 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0660 0.0707 0.0609 | · | · | · |
| W9 | 0.91× | 0.0767 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0669 0.0665 0.0665 | · | · | · |
AdaptiveAvgPool2dFwd¶
input: [N, C, H_in, W_in]
- W1dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]
nondiv-7x7 - W2dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]
resnet-global - W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]
spp-6x6
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.56× | 0.1017 | ops-nn-b2:aclnnAdaptiveAvgPool3d open-source ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0474 0.0553 0.0473 | · | · | · |
| W2 | 0.54× | 0.0630 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0340 0.0213 | · | · | · |
| W3 | 0.36× | 0.1410 | ops-nn-b2:aclnnAdaptiveAvgPool3d open-source ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0419 0.0555 0.0416 | · | · | · |
MaxPool1dFwd¶
input: [N, C, L_in]
- W1dtype=bf16N=16C=128L_in=2048kernel_size=5stride=2padding=2dilation=2ceil_mode=true
ecg-cnn-dilated - W2dtype=f16N=32C=80L_in=4096kernel_size=3stride=3
sincnet-speaker-local - W3dtype=f16N=64C=256L_in=128kernel_size=128stride=1
textcnn-global
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | — | · | torch_npu eager vendor library basis | · | · | · | · |
| W2 | 0.00× | 23.4675 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0910 0.0850 | · | · | · |
| W3 | 0.01× | 5.0949 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0779 0.0689 | · | · | · |
Top-k¶
TopkSelectorFwd¶
index_score: [batch, seq_len, seq_len_kv, kv_group]starts, ends: [batch, seq_len], i32
dtype=f32batch=1seq_len=32768seq_len_kv=65536kv_group=1
- W1topk=1024
topk1024-s32k-kv64k - W2topk=2048
topk2048-s32k-kv64k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.76× | 117.7379 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 89.9394 89.9271 | · | · | · |
| W2 | 0.60× | 164.3116 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 98.9299 98.8865 | · | · | · |