Skip to content

Tuning Memory-Bound Kernels

Upstream NVIDIA / H200 tuning reference

This NVIDIA-platform tutorial is retained from tile-ai/TileOPs.github.io. All measurements, figures, hardware parameters and rules involving SMs, warps and shared memory belong to the original H200 / CUDA context. They are not Ascend 910B1 results or hardware specifications. These cases illustrate tuning methods; hardware rules and performance must be validated again when porting to Ascend. See Timing for this fork's Ascend measurements.

What memory-bound means

On a GPU, kernel speed is determined by whether compute or bandwidth becomes the bottleneck first. TileOPs measures a calibration factor with a macro benchmark: the hardware peak from the specification, multiplied by that factor, gives the effective peak a kernel can actually reach. That effective peak is the reference point for tuning. On an H200, the measured values are 57.27 TFLOP/s for fp32 FMA and 4.07 TB/s of memory bandwidth. The ridge point of the roofline is where the bandwidth slope meets the compute ceiling. Dividing compute by bandwidth places it at an arithmetic intensity of 14.07 flop/byte: the number of floating-point operations per byte moved when compute and bandwidth are saturated at the same time:

1 2 4 8 16 32 64 1 2 4 8 16 32 64 Attainable TFLOP/s (log) Arithmetic intensity: flop per byte (log) Ridge point 14 flop/byte silu (fp16) 5 flop / 4 bytes — ceiling 5.1 bandwidth roof 4.07 TB/s compute roof 57.3 TFLOP/s

The bent line is the roofline. Every kernel's performance point lies below it. To the left of the ridge, the ceiling is arithmetic intensity times bandwidth, so attainable compute rises linearly with intensity. To the right, the ceiling is the compute peak, and higher intensity no longer raises it. silu performs 5 operations per 4 bytes moved, for an arithmetic intensity of 1.25 flop/byte, one eleventh of the ridge. Even with bandwidth fully saturated, it can reach only 9% of the compute ceiling.
  1. Optimizing Global Memory Access
  2. Optimizing Shared Memory Access