RayRoPE: Projective Ray Positional Encoding for Multi-View Attention

We study the spatial encoding of multi-view converters that process tokens from a set of input images, and seek a method that applies patches differently, allows consistent attention to SE(3) for multi-frequency matching, and can adapt to the geometry of the sub-scene. We find that previous coding schemes (absolute or relative) for multi-view attention do not meet the above requirements, and we introduce RayRoPE to address this gap. RayRoPE represents a patch orientation based on correlated rays but uses a projected point on the rays instead of the direction of the geometry encoding. To achieve SE(3) invariance, RayRoPE combines query-frame projective coordinates to compute multi-frequency matching. Finally, since the 'predicted' 3D point in the ray may be imprecise, RayRoPE introduces an analytical calculation method for the encoding state under uncertainty. We validate RayRoPE on novel visual integration tasks and stereo depth estimation and show that it consistently improves over other spatial encoding schemes (e.g. 15% improvement relative to LPIPS in CO3D). We also show that RayRoPE can seamlessly encode RGB-D input, resulting in an even greater advantage over other methods that cannot spatially encode this information.
- † Carnegie Mellon University


