跳转至

RoPE Operators

本页暂无中文版。以下为英文原文。

Every op on this page is used the same way: construct it once, then call it. The constructor takes what the kernel is compiled with; the call takes the tensors. Both are documented under each op — __init__ and forward, where forward is what runs when you call op(...).

NeoX layout

tileops.ops.rope.RopeNeoxFwdOp

GPT-NeoX style RoPE op with standard theta frequencies.

Computes cos/sin tables at construction using standard theta = base^(-2k/d).

Reference: GPT-NeoX / HuggingFace transformers RotaryEmbedding.

__init__

__init__(
    layout="1d",
    base=10000.0,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • layout (str, default: '1d' ) –

    "1d" or "2d".

  • base (float, default: 10000.0 ) –

    Frequency base (default 10000).

  • tune (bool, default: False ) –

    Whether to autotune.

forward

forward(
    x,
)

Apply RoPE rotation using internally computed cos/sin tables.

Parameters:

  • x (Tensor) –

    Input tensor. Shape depends on layout: - 1D: (seq_len, head_dim) - 2D: (batch, seq_len, num_heads, head_dim)

Returns:

  • Tensor –

    Rotated output tensor with same shape as x.

tileops.ops.rope.RopeNeoxPositionIdsFwdOp

GPT-NeoX style RoPE for packed THD tensors with explicit positions.

__init__

__init__(
    max_position,
    base=10000.0,
    rotary_dim=None,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • max_position (int) –

    Manifest params.max_position, int.

  • base (float, default: 10000.0 ) –

    Manifest params.base, float, default 10000.0.

  • rotary_dim (Optional[int], default: None ) –

    Manifest params.rotary_dim, int | None, default None.

  • tune (bool, default: False ) –

    Whether to autotune, applied when a kernel is first built.

forward

forward(
    x,
    position_ids,
)

Run the op on the inputs the manifest declares.

Parameters:

  • x (Tensor) –

    Input tensor, dtype float16 | bfloat16 | float32.

  • position_ids (Tensor) –

    Input tensor, dtype int32 | int64.

Returns:

  • Tensor –

    output, as the manifest declares. Shape rules: output.shape == x.shape.

Interleaved layout

tileops.ops.rope.RopeNonNeoxFwdOp

Original RoFormer RoPE op with adjacent-pair rotation.

Computes cos/sin tables at construction using standard theta = base^(-2k/d).

Reference: Su et al., "RoFormer: Enhanced Transformer with Rotary Position Embedding".

__init__

__init__(
    layout="1d",
    base=10000.0,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • layout (str, default: '1d' ) –

    "1d" or "2d".

  • base (float, default: 10000.0 ) –

    Frequency base (default 10000).

  • tune (bool, default: False ) –

    Whether to autotune.

forward

forward(
    x,
)

Apply RoPE rotation using internally computed cos/sin tables.

Parameters:

  • x (Tensor) –

    Input tensor. Shape depends on layout: - 1D: (seq_len, head_dim) - 2D: (batch, seq_len, num_heads, head_dim)

Returns:

  • Tensor –

    Rotated output tensor with same shape as x.

Scaled frequencies

tileops.ops.rope.RopeLlama31FwdOp

Llama 3.1 RoPE op with piecewise frequency scaling.

Computes cos/sin tables at construction using Llama 3.1 piecewise-scaled frequencies based on wavelength thresholds.

Reference: Meta Llama 3.1 model implementation.

__init__

__init__(
    layout="1d",
    base=10000.0,
    scale_factor=8.0,
    low_freq_factor=1.0,
    high_freq_factor=4.0,
    original_max_position=8192,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • layout (str, default: '1d' ) –

    "1d" or "2d".

  • base (float, default: 10000.0 ) –

    Frequency base (default 10000).

  • scale_factor (float, default: 8.0 ) –

    Scaling factor for low frequencies (default 8.0).

  • low_freq_factor (float, default: 1.0 ) –

    Low-frequency wavelen threshold (default 1.0).

  • high_freq_factor (float, default: 4.0 ) –

    High-frequency wavelen threshold (default 4.0).

  • original_max_position (int, default: 8192 ) –

    Original max position (default 8192).

  • tune (bool, default: False ) –

    Whether to autotune.

forward

forward(
    x,
)

Apply RoPE rotation using internally computed cos/sin tables.

Parameters:

  • x (Tensor) –

    Input tensor. Shape depends on layout: - 1D: (seq_len, head_dim) - 2D: (batch, seq_len, num_heads, head_dim)

Returns:

  • Tensor –

    Rotated output tensor with same shape as x.

tileops.ops.rope.RopeYarnFwdOp

YaRN RoPE op with linear-ramp frequency interpolation.

Computes cos/sin tables at construction using YaRN linear-ramp interpolation between scaled and original frequencies.

Reference: Peng et al., "YaRN: Efficient Context Window Extension of LLMs".

__init__

__init__(
    layout="1d",
    base=10000.0,
    scale=16.0,
    original_max_position=4096,
    beta_fast=32.0,
    beta_slow=1.0,
    attn_factor=1.0,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • layout (str, default: '1d' ) –

    "1d" or "2d".

  • base (float, default: 10000.0 ) –

    Frequency base (default 10000).

  • scale (float, default: 16.0 ) –

    Context extension scale (default 16.0).

  • original_max_position (int, default: 4096 ) –

    Original max position (default 4096).

  • beta_fast (float, default: 32.0 ) –

    Fast decay boundary (default 32.0).

  • beta_slow (float, default: 1.0 ) –

    Slow decay boundary (default 1.0).

  • attn_factor (float, default: 1.0 ) –

    Attention scaling factor (default 1.0).

  • tune (bool, default: False ) –

    Whether to autotune.

forward

forward(
    x,
)

Apply RoPE rotation using internally computed cos/sin tables.

Parameters:

  • x (Tensor) –

    Input tensor. Shape depends on layout: - 1D: (seq_len, head_dim) - 2D: (batch, seq_len, num_heads, head_dim)

Returns:

  • Tensor –

    Rotated output tensor with same shape as x.

tileops.ops.rope.RopeLongRopeFwdOp

LongRoPE op with per-dimension frequency rescaling.

Computes cos/sin tables at construction using per-dimension rescale factors (ext_factors) that multiply the divisor, plus a scale-dependent amplitude factor applied to cos/sin output.

Reference: TVM rope_freq_longrope in position_embedding.py; Ding et al., "LongRoPE: Extending LLM Context Window Beyond 2M Tokens".

__init__

__init__(
    layout="1d",
    base=10000.0,
    rescale_factors=None,
    max_position_embeddings=4096,
    original_max_position_embeddings=4096,
    tune=False,
)

Build the op. Shapes and dtype are taken from the first call.

Parameters:

  • layout (str, default: '1d' ) –

    "1d" or "2d".

  • base (float, default: 10000.0 ) –

    Frequency base (default 10000).

  • rescale_factors (Optional[Tensor], default: None ) –

    Per-dimension rescale factors (ext_factors) of shape (head_dim // 2,). These multiply the divisor.

  • max_position_embeddings (int, default: 4096 ) –

    Extended max position length (default 4096).

  • original_max_position_embeddings (int, default: 4096 ) –

    Original max position length (default 4096).

  • tune (bool, default: False ) –

    Whether to autotune.

forward

forward(
    x,
)

Apply RoPE rotation using internally computed cos/sin tables.

Parameters:

  • x (Tensor) –

    Input tensor. Shape depends on layout: - 1D: (seq_len, head_dim) - 2D: (batch, seq_len, num_heads, head_dim)

Returns:

  • Tensor –

    Rotated output tensor with same shape as x.