One image patch branching into three Gaussian supports in world coordinates

WorldRoPEProbabilistic World-Centric Rotary Embedding for Multi-View Attention

Mijeong Kim1Kyujin Lim2Bohyung Han1,2

1ECE  ·  2IPAI, Seoul National University

Under Review

Paper ↗

TL;DR: Multi-view attention, grounded in 3D.

Abstract

Rotary Positional Embedding (RoPE) encodes relative positions in Transformers, but image coordinates are defined separately in each camera view. We introduce WorldRoPE, a probabilistic world-centric RoPE for multi-view attention. At each layer, patch tokens predict mixtures of 3D Gaussians in a shared world frame. These primitives capture alternative depths, uncertainty, and the spatial extent of a patch.

Projected into a query view, the Gaussians define where a key token may appear. WorldRoPE integrates the rotary operator over this distribution in closed form, using mixture weights and covariance-dependent attenuation to guide attention without rasterizing token features. Updated features refine the geometry at the next layer. Experiments in novel view synthesis, stereo depth estimation, and feed-forward 3D reconstruction show its effectiveness across multi-view Transformers.

A 2D patch token becomes world-centric 3D Gaussian support; RoPE is integrated over its projected 2D distribution.

Why world-centric 3D Gaussians?

A patch covers a viewing frustum, often crossing more than one depth. A mixture of volumetric Gaussians can retain those alternative surfaces, their uncertainty, and the finite footprint of the patch.

Figure 1 from the paper: ray-based geometry, world-centric Gaussian token geometry, and probabilistic rotary embedding
Figure 1. WorldRoPE overview. Multi-hypothesis token support is projected into the query view, where RoPE is integrated over possible key positions.

How does this enter attention?

Token features predict the Gaussians. Projecting them into the query view places possible key positions alongside the query; a closed-form rotary expectation then guides attention. Updated features predict geometry again at the next layer.

Figure 2 from the paper: WorldRoPE inside multi-view attention, from Gaussian prediction to the next attention layer
Figure 2. Multi-view attention with WorldRoPE. Geometry guides attention and updated features inform geometry prediction in the next layer.

The expected operator combines rotations at the projected Gaussian centers with covariance-dependent attenuation. It acts on positional evidence without rasterizing or warping token features.

Interactive Projected Distribution

Move the key camera in 3D and change the possible depths of its center patch. The right panel shows the combined projected distribution in the query view.

3D camera pose and depth

Drag the blue K camera up or down to change Y, including below zero. Drag the XZ handle to change depth and horizontal position. X 0.70 · Y +0.35 · Z 0.10

Camera geometry interactive 3D view

Orange: fixed query camera. Blue: movable key camera. Ellipses: possible 3D support.

Projected Probability Distribution (Query View)

01
Equal-weight mixture of two projected Gaussians; the dashed orange line is the epipolar line.

Color encodes density relative to the visible peak (0–1). The equal-weight mixture integrates to one over the full query plane.

Demo assumptions

The center key patch has two possible depths with equal mixture weight and independently adjustable depth standard deviations, as in the paper's per-hypothesis mixture. A fixed 3D depth σ can look narrower after projection at a greater depth because perspective reduces disparity. Its depth uncertainty and optional patch footprint are projected through the fixed query camera. The dashed epipolar line is the projection of the center key ray over varying depth. To make spatial extent visible at this display size, the illustrative patch-footprint width is enlarged and grows with depth; it is not a learned covariance. A half-pixel regularization keeps the displayed density finite when spatial extent is off. Probability outside the query image frame is not shown.

Does the world-centric 3D support improve synthesis?

Within LVSM, WorldRoPE improves all three reported image-quality metrics on RE10K and Objaverse. In the examples below, thin structures remain more distinct.

Figure 3 from the paper: reference views, ground truth, six positional encoding baselines, and WorldRoPE on RE10K and Objaverse
Figure 3. WorldRoPE better preserves thin structures and fine details, including stair railings and bicycle wheels.
Novel view synthesis with LVSM. Higher PSNR and SSIM, lower LPIPS.
MethodRE10KObjaverse
PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓
Plücker ray24.620.7900.10815.150.8380.355
GTA25.050.7990.10222.220.8920.123
PRoPE25.440.8110.09522.410.8940.117
RayRoPE26.170.8330.08122.250.8920.121
URoPE26.260.8340.07821.990.8900.128
WorldRoPE26.750.8460.07322.520.8960.113

PSNR is in dB. These are the values reported in Table 2 of the paper.

What do the world-centric Gaussians capture?

At a boundary between vegetation and sky, one token can have two depth peaks. Across attention layers, the predicted support becomes more structured in the paper's visualizations.

Figure 4 from the paper: image tokens with multiple predicted depth peaks and two example depth distributions
Figure 4. A patch spanning vegetation and sky has two distinct depth peaks, while a patch within vegetation has one dominant peak.
Figure 7 from the paper: world-space Gaussian primitives across layers and their projection into the query view
Figure 7. World-space primitives across attention layers and their projection into the query view. RGB rasterization is used here only for visualization.

Does it transfer beyond synthesis?

WorldRoPE is also evaluated in UniMatch for stereo depth estimation and in VGGT for feed-forward 3D reconstruction.

Stereo depth estimation

On the unseen ScanNet++ dataset, WorldRoPE records the lowest AbsRel and RMSE in the paper's UniMatch comparison.

Stereo depth estimation with UniMatch. Lower is better.
MethodScenes11
RMSE ↓
SUN3D
RMSE ↓
RGBD
RMSE ↓
ScanNet++
AbsRel ↓
ScanNet++
RMSE ↓
RayRoPE0.7430.3700.6770.2570.426
URoPE0.6860.3740.6370.2340.398
WorldRoPE0.6610.3560.6300.2240.386

Feed-forward 3D reconstruction

In the VGGT comparison on ScanNet++, WorldRoPE has the lowest reported depth, camera-space point, and camera rotation errors.

Feed-forward 3D reconstruction with VGGT on ScanNet++. Lower is better.
MethodDepth
AbsRel ↓
Norm.
RMSE ↓
Camera points
AbsRel ↓
Rotation
error ↓
RayRoPE0.05790.06240.05751.04
URoPE0.05810.06260.05761.02
WorldRoPE0.05670.06190.05650.98

For derivations, ablations, and implementation details, read the full paper ↗.