The ORBIT structure from motion dataset. ORBIT includes a variety of clips that are challenging for SfM methods. Scenes with many dynamic objects (a) or large textureless areas (b) are challenging for traditional approaches like COLMAP.

Abstract

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes. Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress or to pinpoint where improvements are most needed. To address this gap, we introduce a new benchmark for evaluating camera pose estimation. Our key insight is to leverage online panoramic 360° video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery. The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects. After tracking camera motion across full 360° videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark, called ORBIT. Experiments show that COLMAP, as well as recent optimization-based and feed-forward SfM methods struggle to accurately estimate camera poses on our benchmark. Hence, ORBIT provides a valuable testbed where researchers can meaningfully measure progress on truly challenging, real-world SfM problems.

Sample 360 videos and benchmark clips

Each frame at the top shows the 360 frame from the original online video, in the middle the four cube faces for cross validation are shown, at the buttom the test clip is shown which is included in our benchmark.

Frame ATE Comparisons

In order to show how different SFM methods, mainly MegaSaM, Colmap, and VGGT-Long perform on our benchmark we showcase sample clips of ORBIT. For each clip, we select the median ATE of all methods as a success threshold to differentiate between method's performances. The small box at the top left corner corresponding each method is green if the ATE of that frame for that method is less than the median ATE. The median ATE is written as the fourth box.

Estimated Camera Trajectories and Sample Clip Frames of ORBIT

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Description

camera comparison. GT: , Estimate:

Input Video

Challenge Categorization

Each clip manifests a range of challenges in ORBIT. We provide sub-category evaluations on 6 of the challenges, namely High Speed, Low Texture, Low light, Presence of Crowd --Independent of camera moving Objects--, Presence of Egocentric Parallel to camera moving Objects --Ego--, and presence of Fluids. Please note that sub categories have overlap and checkout the paper for more details.

Sample Clips Exhibiting Challenges
Challenges and Sample Clips manifesting them

We report the ATE and RPE-R for each method on each subcategory. Based on the results, the most challenging categories are the lack of texture, affecting MegaSaM and COLMAP significantly and high speed for other methods. In the table below the red indicates the hardest challenge for each method and blue indicates the easiest challenge. We observe that MegaSaM improves significantly on the crowd challenge in compare to Colmap for example. Overall, the following table shows that ORBIT exposes a diverse set of challenges and is a valuable tool for analyzing the current state-of-the-art by highlighting their failure modes.

Challenge Comparisons
Table of method performances --ATE and RPE-R-- on challenging sub-categories.

Previous Publication

A smaller version of ORBIT was published at CVPR 2026. You can find the CVPR paper here.

BibTeX

If you use our dataset or benchmark, please cite our paper:

@article{sabour2026orbit,
  title={ORBIT++: Benchmarking SfM in the Wild with 360$^{\circ}$ Video},
  author={Sabour, Sara and Jin, Linyi and Tucker, Richard and Hertz, Amir and Brubaker, Marcus and Saxena, Saurabh and Hur, Junhwa and Tagliasacchi, Andrea and Sun, Deqing and Fleet, David J. and Szeliski, Richard and Snavely, Noah},
  journal={arXiv preprint arXiv:2608.22039},
  year={2026}
}