跳转至

Quantization Operators

本页暂无中文版。以下为英文原文。

Every op on this page is used the same way: construct it once, then call it. The constructor takes what the kernel is compiled with; the call takes the tensors. Both are documented under each op — __init__ and forward, where forward is what runs when you call op(...).

FP8 quantization

tileops.ops.fp8_quant.FP8QuantFwdOp

__init__

__init__(
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • tune (bool, default: False ) –

    Whether to autotune, applied when a kernel is first built.

forward

forward(
    input_tensor,
)

Run the op on the inputs the manifest declares.

Parameters:

  • input_tensor –

    Input tensor, dtype float16 | bfloat16 | float32.

Returns:

  • Tuple[Tensor, Tensor] –

    scale_tensor, output_tensor, as the manifest declares.