RoPE Operators¶
Every op on this page is used the same way: construct it once, then call it. The
constructor takes what the kernel is compiled with; the call takes the tensors.
Both are documented under each op — __init__ and forward, where forward is
what runs when you call op(...).
NeoX layout¶
tileops.ops.rope.RopeNeoxFwdOp
¶
GPT-NeoX style RoPE op with standard theta frequencies.
Computes cos/sin tables at construction using standard theta = base^(-2k/d).
Reference: GPT-NeoX / HuggingFace transformers RotaryEmbedding.
__init__
¶
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
layout(str, default:'1d') –"1d" or "2d".
-
base(float, default:10000.0) –Frequency base (default 10000).
-
tune(bool, default:False) –Whether to autotune.
forward
¶
Apply RoPE rotation using internally computed cos/sin tables.
Parameters:
-
x(Tensor) –Input tensor. Shape depends on layout: - 1D:
(seq_len, head_dim)- 2D:(batch, seq_len, num_heads, head_dim)
Returns:
-
Tensor–Rotated output tensor with same shape as x.
tileops.ops.rope.RopeNeoxPositionIdsFwdOp
¶
GPT-NeoX style RoPE for packed THD tensors with explicit positions.
__init__
¶
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
max_position(int) –Manifest
params.max_position,int. -
base(float, default:10000.0) –Manifest
params.base,float, default10000.0. -
rotary_dim(Optional[int], default:None) –Manifest
params.rotary_dim,int | None, defaultNone. -
tune(bool, default:False) –Whether to autotune, applied when a kernel is first built.
forward
¶
Run the op on the inputs the manifest declares.
Parameters:
-
x(Tensor) –Input tensor, dtype
float16 | bfloat16 | float32. -
position_ids(Tensor) –Input tensor, dtype
int32 | int64.
Returns:
-
Tensor–output, as the manifest declares. Shape rules:output.shape == x.shape.
Interleaved layout¶
tileops.ops.rope.RopeNonNeoxFwdOp
¶
Original RoFormer RoPE op with adjacent-pair rotation.
Computes cos/sin tables at construction using standard theta = base^(-2k/d).
Reference: Su et al., "RoFormer: Enhanced Transformer with Rotary Position Embedding".
__init__
¶
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
layout(str, default:'1d') –"1d" or "2d".
-
base(float, default:10000.0) –Frequency base (default 10000).
-
tune(bool, default:False) –Whether to autotune.
forward
¶
Apply RoPE rotation using internally computed cos/sin tables.
Parameters:
-
x(Tensor) –Input tensor. Shape depends on layout: - 1D:
(seq_len, head_dim)- 2D:(batch, seq_len, num_heads, head_dim)
Returns:
-
Tensor–Rotated output tensor with same shape as x.
Scaled frequencies¶
tileops.ops.rope.RopeLlama31FwdOp
¶
Llama 3.1 RoPE op with piecewise frequency scaling.
Computes cos/sin tables at construction using Llama 3.1 piecewise-scaled frequencies based on wavelength thresholds.
Reference: Meta Llama 3.1 model implementation.
__init__
¶
__init__(
layout="1d",
base=10000.0,
scale_factor=8.0,
low_freq_factor=1.0,
high_freq_factor=4.0,
original_max_position=8192,
tune=False,
)
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
layout(str, default:'1d') –"1d" or "2d".
-
base(float, default:10000.0) –Frequency base (default 10000).
-
scale_factor(float, default:8.0) –Scaling factor for low frequencies (default 8.0).
-
low_freq_factor(float, default:1.0) –Low-frequency wavelen threshold (default 1.0).
-
high_freq_factor(float, default:4.0) –High-frequency wavelen threshold (default 4.0).
-
original_max_position(int, default:8192) –Original max position (default 8192).
-
tune(bool, default:False) –Whether to autotune.
forward
¶
Apply RoPE rotation using internally computed cos/sin tables.
Parameters:
-
x(Tensor) –Input tensor. Shape depends on layout: - 1D:
(seq_len, head_dim)- 2D:(batch, seq_len, num_heads, head_dim)
Returns:
-
Tensor–Rotated output tensor with same shape as x.
tileops.ops.rope.RopeYarnFwdOp
¶
YaRN RoPE op with linear-ramp frequency interpolation.
Computes cos/sin tables at construction using YaRN linear-ramp interpolation between scaled and original frequencies.
Reference: Peng et al., "YaRN: Efficient Context Window Extension of LLMs".
__init__
¶
__init__(
layout="1d",
base=10000.0,
scale=16.0,
original_max_position=4096,
beta_fast=32.0,
beta_slow=1.0,
attn_factor=1.0,
tune=False,
)
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
layout(str, default:'1d') –"1d" or "2d".
-
base(float, default:10000.0) –Frequency base (default 10000).
-
scale(float, default:16.0) –Context extension scale (default 16.0).
-
original_max_position(int, default:4096) –Original max position (default 4096).
-
beta_fast(float, default:32.0) –Fast decay boundary (default 32.0).
-
beta_slow(float, default:1.0) –Slow decay boundary (default 1.0).
-
attn_factor(float, default:1.0) –Attention scaling factor (default 1.0).
-
tune(bool, default:False) –Whether to autotune.
forward
¶
Apply RoPE rotation using internally computed cos/sin tables.
Parameters:
-
x(Tensor) –Input tensor. Shape depends on layout: - 1D:
(seq_len, head_dim)- 2D:(batch, seq_len, num_heads, head_dim)
Returns:
-
Tensor–Rotated output tensor with same shape as x.
tileops.ops.rope.RopeLongRopeFwdOp
¶
LongRoPE op with per-dimension frequency rescaling.
Computes cos/sin tables at construction using per-dimension rescale factors (ext_factors) that multiply the divisor, plus a scale-dependent amplitude factor applied to cos/sin output.
Reference: TVM rope_freq_longrope in position_embedding.py;
Ding et al., "LongRoPE: Extending LLM Context Window Beyond 2M Tokens".
__init__
¶
__init__(
layout="1d",
base=10000.0,
rescale_factors=None,
max_position_embeddings=4096,
original_max_position_embeddings=4096,
tune=False,
)
Build the op. Shapes and dtype are taken from the first call.
Parameters:
-
layout(str, default:'1d') –"1d" or "2d".
-
base(float, default:10000.0) –Frequency base (default 10000).
-
rescale_factors(Optional[Tensor], default:None) –Per-dimension rescale factors (ext_factors) of shape (head_dim // 2,). These multiply the divisor.
-
max_position_embeddings(int, default:4096) –Extended max position length (default 4096).
-
original_max_position_embeddings(int, default:4096) –Original max position length (default 4096).
-
tune(bool, default:False) –Whether to autotune.
forward
¶
Apply RoPE rotation using internally computed cos/sin tables.
Parameters:
-
x(Tensor) –Input tensor. Shape depends on layout: - 1D:
(seq_len, head_dim)- 2D:(batch, seq_len, num_heads, head_dim)
Returns:
-
Tensor–Rotated output tensor with same shape as x.