Tuning Memory-Bound Kernels¶
Upstream NVIDIA / H200 tuning reference
This NVIDIA-platform tutorial is retained from tile-ai/TileOPs.github.io. All measurements, figures, hardware parameters and rules involving SMs, warps and shared memory belong to the original H200 / CUDA context. They are not Ascend 910B1 results or hardware specifications. These cases illustrate tuning methods; hardware rules and performance must be validated again when porting to Ascend. See Timing for this fork's Ascend measurements.
What memory-bound means¶
On a GPU, kernel speed is determined by whether compute or bandwidth becomes the bottleneck first. TileOPs measures a calibration factor with a macro benchmark: the hardware peak from the specification, multiplied by that factor, gives the effective peak a kernel can actually reach. That effective peak is the reference point for tuning. On an H200, the measured values are 57.27 TFLOP/s for fp32 FMA and 4.07 TB/s of memory bandwidth. The ridge point of the roofline is where the bandwidth slope meets the compute ceiling. Dividing compute by bandwidth places it at an arithmetic intensity of 14.07 flop/byte: the number of floating-point operations per byte moved when compute and bandwidth are saturated at the same time:
silu performs 5 operations per 4 bytes moved, for an arithmetic intensity of 1.25 flop/byte, one eleventh of the ridge. Even with bandwidth fully saturated, it can reach only 9% of the compute ceiling.