Files
spirula-studio/docs/datasets.md
T

10 KiB

Datasets

Supported layouts

Three formats, auto-detected in this order: Nerfstudio, then COLMAP, then Metashape.

The native parsers (src/data/parsers/*Parser.cpp) are the one implementation, shared by the CLI trainer, the GUI and the WASM viewer.

format inputs parser
COLMAP cameras/images/points3D in .bin or .txt ColmapParser.cpp
Nerfstudio transforms.json + a PLY point cloud NerfstudioParser.cpp (PLY reader lives here)
Metashape camera-export .xml + .ply, optionally a .psx project for filename disambiguation MetashapeParser.cpp (XML via app/Xml.h, zips via external/miniz)

Default subdirectory names: images/, masks/, depths/, normals/. The COLMAP reconstruction directory is auto-detected over {sparse/0, colmap/sparse/0, sparse, colmap, .} unless recon_dir is set.

Camera models

Perspective (with full radial / tangential / thin-prism distortion), equidistant and equisolid fisheye (including >180° FOV as produced by 360 cameras), and equirectangular/spherical. See src/core/CameraModel.h — it is plain C++17 with no CUDA dependency, which is why the WASM viewer can reuse it.

COLMAP writes 18 camera models and ColmapParser.cpp accepts every one. EQUIRECTANGULAR (id 17) is the spherical one: its params are (w, h) rather than a calibration, because the image is the calibration. It reaches the same CameraModelType::EQUIRECTANGULAR as a Metashape spherical sensor, with the same convention — +Z forward at the image centre, azimuth wrapping at the left/right edge — so the two formats describe an identical camera and bake_post_split treats them identically. Note the engine's canonical panorama intrinsics assume a 2:1 (360°×180°) image; the parser warns when one is not.

Lens distortion tiers

Distortion is a separate axis from the camera model, and a COMPILE-TIME one: a template argument in CUDA, a kDistortion specialization constant on Vulkan. The four tiers, cheapest first, are None, OpenCV (k1 k2 p1 p2), ThinPrism (k1 k2 k3 k4 p1 p2 sx1 sy1, COLMAP's THIN_PRISM_FISHEYE) and Rational (k1..k6 p1 p2, COLMAP's FULL_OPENCV, where k4..k6 divide). A slot index does NOT mean the same thing across tiers. core/CameraModel.h is the one definition; shaders/projection_utils.slang is the one implementation, generic over an ICameraDistortion.

The parser picks the CHEAPEST tier that represents the source camera exactly, so a PINHOLE dataset costs no distortion registers at all and a FULL_OPENCV camera whose k4..k6 are zero demotes to ThinPrism (camera_distortion_demote). Only eleven (model, tier) pairs are compiled — no COLMAP fisheye model is rational, and EQUIRECTANGULAR carries no distortion. camera_distortion_is_compiled() is that list, and it must stay in step with kCameraVariants in tools/codegen/generate_kernel_instantiation.py and the export list in shaders/primitive_3dgs.slang.

A transforms.json is ambiguous about what k4 means -- OpenCV's first rational DENOMINATOR term, or Kannala-Brandt's / Metashape's fourth RADIAL one. An explicit camera_distortion key settles it (MetashapeParser always writes one); failing that, a fisheye camera is Kannala-Brandt, k5/k6 mean rational on their own, and the mere PRESENCE of b1/b2/sx1/sy1 -- keys a rational camera never carries -- makes k4 radial. That last rule is what a Metashape-converted transforms.json needs: reading its k4 as a denominator is a different lens, not a small error.

FOV, SIMPLE_DIVISION, DIVISION, EUCM and RAD_TAN_THIN_PRISM_FISHEYE have no exact tier, and neither does a Metashape sensor skew (b2), which is an off-diagonal pixel term where every tier's pixel map is diagonal. They are fitted onto a (model, ThinPrism) pair by near-minimax regression (data/DistortionFit.h).

fit_camera_auto chooses the camera model from the source's MEASURED field of view, not from its name: a COLMAP FOV lens is perspective at omega 0.3 and a 180-degree fisheye at omega 0.87, and forcing the second onto a pinhole target puts tan(85 deg) = 11.4 into a degree-8 polynomial and fails. It then walks a coefficient ladder (all eight, no thin prism, no k3/k4, ..., none) until the fitted distortion is invertible everywhere sampled. It never fails: a dataset that took hours to reconstruct must not refuse to load over a lens model, so the worst case is a plain fisheye and a warning, not an exception.

A fit closer than dsfit::kExactFitPx (0.1 px) is left alone -- the fitted camera already reproduces the source to better than bilinear resampling can resolve, so re-distorting would only cost VRAM and blur. In practice that covers FOV, both division models and EUCM; RAD_TAN_THIN_PRISM_FISHEYE and a real b2 do not reach it. When it is not reached, the source model goes into ParsedDataset::redistort, which makes bake_post_split set any_warp even at K = 1 so the images route through the warp path's staging and get resampled.

The resampling reads the TRUE source projection (shaders/camera_source.slang, the one place those models live on device; data/SourceCamera.h is the host mirror, and the two must agree exactly) rather than the fit -- going through the fit would be a no-op. Source model ids 0..17 are COLMAP's own CameraModelId values with COLMAP's parameter array verbatim; ours start at 1000 so COLMAP can keep appending to its enum.

A source model is sampled only where its image still grows outward as the ray tilts off axis. Past that it has folded -- at the lens border for a polynomial fisheye, well inside the frame for the division and unified models -- and two directions share a pixel, so a warped face would be ringed by a mirrored copy of the image instead of ending at the lens. Those rays are dropped, and the synthesized FOV mask drops them by the same test. The fit domain stops earlier still, where the radial rate falls below a quarter of its on-axis value: at the fold itself no fitted distortion is invertible, so the coefficient ladder would degrade a camera the tier otherwise reproduces to a fraction of a pixel.

Two paths, both gathering per destination pixel:

  • K = 1 (kernels/pixelwise/ImageRedistort.cu): destination pixel -> ray through the fitted camera -> source pixel. The fit leaves the pose alone, so a destination pixel and its source pixel are the SAME ray: depth (linear or ray) and camera-frame normals transfer unchanged and only the sampling coordinate moves. None of GtDepthNormalWarp.cu's point-space handling applies.
  • K > 1 (warp_to_pinhole): the cubemap face ray projects STRAIGHT through the source camera, so the fitted camera is never materialized and the two passes cost one kernel and no intermediate image. RayToPixel<D, kFromSource> is the seam; kFromSource is a template argument (a specialization constant on Vulkan) so an ordinary dataset pays neither the branch nor the 16 registers the source parameters occupy.

Two-stage parse

  1. parse_dataset → ParsedDataset: per-input cameras in the raw train_frame="points" frame — poses and points exactly as stored. The normalized-frame similarity is computed only to obtain the train_frame_scale scalar.
  2. bake_post_split → the post-split arrays engine_setup_data_manager consumes. This is either an identity (K=1) pass-through, or the cubemap split when warp_to_pinhole is enabled: 5 faces for fisheye/equisolid, 6 faces for equirectangular.

That normalized-frame similarity is where orientation_method and center_method act — and only up / poses is implemented natively; anything else is approximated with a warning. Since train_frame="points" leaves splats in the raw frame, the choice moves train_frame_scale and the viewer's default camera, not the training coordinates. The unported methods (pca, vertical, gsplat, focus) still have a working Python reference: notes/pose-normalization.md.

Train/eval split

eval_mode selects the strategy:

  • all (default) — every image trains.
  • fraction — linspace-spread ceil(N * train_split_fraction) train images.
  • interval — index % eval_interval == 0 is eval, rest train.
  • filename — basename contains train / eval.

validation_fraction additionally holds out a linspace-spread slice for validation. Frames whose camera position exceeds outlier_threshold MADs from the geometric median of all camera positions are rejected (default: off).

require_image_files=false keeps frames whose image file is missing — the standalone viewer uses this to load camera poses from a dataset shipped without pixels.

Downscaling

Stored intrinsics can be divided by a fixed factor (the Mip-NeRF 360 images_2 / images_4 convention). Note: the auto-detect mode that probes the first image's actual resolution is implemented on the Python side only; the native parser currently requires an explicit factor.

Preprocessing tools

reference/scripts/ holds standalone Python utilities that produce these layouts: frame extraction with blur skipping, COLMAP/GLOMAP driving, Metashape conversion, downscaling, undistortion, masking (including a SAM2 GUI), monocular depth/normal prediction, and raw conversion. See reference/scripts/README.md. These are preprocessing, separate from the training data path, and stay on the Python side.

reference/scripts/batch_process_data.bash needs a COLMAP vocabulary tree; set SS_VOCAB_TREE to its path.

Benchmarking

reference/python/benchmark.py drives multi-scene runs over standard academic sets (Mip-NeRF 360 360_v2, ZipNeRF) by calling spirula train. Dataset roots are passed as arguments — no paths are hardcoded. Each scene runs in its own subprocess: the engine is a process-global singleton, and running several scenes in one process leaks state between them and silently degrades metrics.

Do not hardcode local paths

Dataset locations belong in arguments or environment variables. See the "Do not commit" section of ../AGENTS.md and tools/check_private_paths.sh.