10 KiB
Datasets
Supported layouts
Three formats, auto-detected in this order: Nerfstudio, then COLMAP, then Metashape.
The native parsers (src/data/parsers/*Parser.cpp) are the one
implementation, shared by the CLI trainer, the GUI and the WASM viewer.
| format | inputs | parser |
|---|---|---|
| COLMAP | cameras/images/points3D in .bin or .txt |
ColmapParser.cpp |
| Nerfstudio | transforms.json + a PLY point cloud |
NerfstudioParser.cpp (PLY reader lives here) |
| Metashape | camera-export .xml + .ply, optionally a .psx project for filename disambiguation |
MetashapeParser.cpp (XML via app/Xml.h, zips via external/miniz) |
Default subdirectory names: images/, masks/, depths/, normals/.
The COLMAP reconstruction directory is auto-detected over
{sparse/0, colmap/sparse/0, sparse, colmap, .} unless recon_dir is set.
Camera models
Perspective (with full radial / tangential / thin-prism distortion),
equidistant and equisolid fisheye (including >180° FOV as produced by 360
cameras), and equirectangular/spherical. See src/core/CameraModel.h — it is
plain C++17 with no CUDA dependency, which is why the WASM viewer can reuse
it.
COLMAP writes 18 camera models and ColmapParser.cpp accepts every one.
EQUIRECTANGULAR (id 17) is the spherical one: its params are (w, h)
rather than a calibration, because the image is the calibration. It reaches
the same CameraModelType::EQUIRECTANGULAR as a Metashape spherical sensor,
with the same convention — +Z forward at the image centre, azimuth wrapping at
the left/right edge — so the two formats describe an identical camera and
bake_post_split treats them identically. Note the engine's canonical panorama
intrinsics assume a 2:1 (360°×180°) image; the parser warns when one is not.
Lens distortion tiers
Distortion is a separate axis from the camera model, and a COMPILE-TIME one:
a template argument in CUDA, a kDistortion specialization constant on Vulkan.
The four tiers, cheapest first, are None, OpenCV (k1 k2 p1 p2),
ThinPrism (k1 k2 k3 k4 p1 p2 sx1 sy1, COLMAP's THIN_PRISM_FISHEYE) and
Rational (k1..k6 p1 p2, COLMAP's FULL_OPENCV, where k4..k6 divide).
A slot index does NOT mean the same thing across tiers. core/CameraModel.h
is the one definition; shaders/projection_utils.slang is the one
implementation, generic over an ICameraDistortion.
The parser picks the CHEAPEST tier that represents the source camera exactly,
so a PINHOLE dataset costs no distortion registers at all and a FULL_OPENCV
camera whose k4..k6 are zero demotes to ThinPrism
(camera_distortion_demote). Only eleven (model, tier) pairs are compiled —
no COLMAP fisheye model is rational, and EQUIRECTANGULAR carries no distortion.
camera_distortion_is_compiled() is that list, and it must stay in step with
kCameraVariants in tools/codegen/generate_kernel_instantiation.py and the
export list in shaders/primitive_3dgs.slang.
A transforms.json is ambiguous about what k4 means -- OpenCV's first
rational DENOMINATOR term, or Kannala-Brandt's / Metashape's fourth RADIAL one.
An explicit camera_distortion key settles it (MetashapeParser always writes
one); failing that, a fisheye camera is Kannala-Brandt, k5/k6 mean rational
on their own, and the mere PRESENCE of b1/b2/sx1/sy1 -- keys a rational
camera never carries -- makes k4 radial. That last rule is what a
Metashape-converted transforms.json needs: reading its k4 as a denominator
is a different lens, not a small error.
FOV, SIMPLE_DIVISION, DIVISION, EUCM and RAD_TAN_THIN_PRISM_FISHEYE
have no exact tier, and neither does a Metashape sensor skew (b2), which is
an off-diagonal pixel term where every tier's pixel map is diagonal. They are
fitted onto a (model, ThinPrism) pair by near-minimax regression
(data/DistortionFit.h).
fit_camera_auto chooses the camera model from the source's MEASURED field of
view, not from its name: a COLMAP FOV lens is perspective at omega 0.3 and a
180-degree fisheye at omega 0.87, and forcing the second onto a pinhole target
puts tan(85 deg) = 11.4 into a degree-8 polynomial and fails. It then walks a
coefficient ladder (all eight, no thin prism, no k3/k4, ..., none) until the
fitted distortion is invertible everywhere sampled. It never fails: a
dataset that took hours to reconstruct must not refuse to load over a lens
model, so the worst case is a plain fisheye and a warning, not an exception.
A fit closer than dsfit::kExactFitPx (0.1 px) is left alone -- the fitted
camera already reproduces the source to better than bilinear resampling can
resolve, so re-distorting would only cost VRAM and blur. In practice that
covers FOV, both division models and EUCM; RAD_TAN_THIN_PRISM_FISHEYE and
a real b2 do not reach it. When it is not reached, the source model goes into
ParsedDataset::redistort, which makes bake_post_split set any_warp even at
K = 1 so the images route through the warp path's staging and get resampled.
The resampling reads the TRUE source projection (shaders/camera_source.slang,
the one place those models live on device; data/SourceCamera.h is the host
mirror, and the two must agree exactly) rather than the fit -- going through the
fit would be a no-op. Source model ids 0..17 are COLMAP's own CameraModelId
values with COLMAP's parameter array verbatim; ours start at 1000 so COLMAP can
keep appending to its enum.
A source model is sampled only where its image still grows outward as the ray tilts off axis. Past that it has folded -- at the lens border for a polynomial fisheye, well inside the frame for the division and unified models -- and two directions share a pixel, so a warped face would be ringed by a mirrored copy of the image instead of ending at the lens. Those rays are dropped, and the synthesized FOV mask drops them by the same test. The fit domain stops earlier still, where the radial rate falls below a quarter of its on-axis value: at the fold itself no fitted distortion is invertible, so the coefficient ladder would degrade a camera the tier otherwise reproduces to a fraction of a pixel.
Two paths, both gathering per destination pixel:
- K = 1 (
kernels/pixelwise/ImageRedistort.cu): destination pixel -> ray through the fitted camera -> source pixel. The fit leaves the pose alone, so a destination pixel and its source pixel are the SAME ray: depth (linear or ray) and camera-frame normals transfer unchanged and only the sampling coordinate moves. None of GtDepthNormalWarp.cu's point-space handling applies. - K > 1 (warp_to_pinhole): the cubemap face ray projects STRAIGHT through
the source camera, so the fitted camera is never materialized and the two
passes cost one kernel and no intermediate image.
RayToPixel<D, kFromSource>is the seam;kFromSourceis a template argument (a specialization constant on Vulkan) so an ordinary dataset pays neither the branch nor the 16 registers the source parameters occupy.
Two-stage parse
parse_dataset→ParsedDataset: per-input cameras in the rawtrain_frame="points"frame — poses and points exactly as stored. The normalized-frame similarity is computed only to obtain thetrain_frame_scalescalar.bake_post_split→ the post-split arraysengine_setup_data_managerconsumes. This is either an identity (K=1) pass-through, or the cubemap split whenwarp_to_pinholeis enabled: 5 faces for fisheye/equisolid, 6 faces for equirectangular.
That normalized-frame similarity is where orientation_method and
center_method act — and only up / poses is implemented natively;
anything else is approximated with a warning. Since train_frame="points"
leaves splats in the raw frame, the choice moves train_frame_scale and the
viewer's default camera, not the training coordinates. The unported methods
(pca, vertical, gsplat, focus) still have a working Python reference:
notes/pose-normalization.md.
Train/eval split
eval_mode selects the strategy:
all(default) — every image trains.fraction— linspace-spreadceil(N * train_split_fraction)train images.interval— index% eval_interval == 0is eval, rest train.filename— basename containstrain/eval.
validation_fraction additionally holds out a linspace-spread slice for
validation. Frames whose camera position exceeds outlier_threshold MADs from
the geometric median of all camera positions are rejected (default: off).
require_image_files=false keeps frames whose image file is missing — the
standalone viewer uses this to load camera poses from a dataset shipped
without pixels.
Downscaling
Stored intrinsics can be divided by a fixed factor (the Mip-NeRF 360
images_2 / images_4 convention). Note: the auto-detect mode that probes
the first image's actual resolution is implemented on the Python side only;
the native parser currently requires an explicit factor.
Preprocessing tools
reference/scripts/ holds standalone Python utilities that produce these
layouts: frame extraction with blur skipping, COLMAP/GLOMAP driving,
Metashape conversion, downscaling, undistortion, masking (including a SAM2
GUI), monocular depth/normal prediction, and raw conversion. See
reference/scripts/README.md. These are preprocessing, separate from the
training data path, and stay on the Python side.
reference/scripts/batch_process_data.bash needs a COLMAP vocabulary tree; set
SS_VOCAB_TREE to its path.
Benchmarking
reference/python/benchmark.py drives multi-scene runs over standard
academic sets (Mip-NeRF 360 360_v2, ZipNeRF) by calling spirula train.
Dataset roots are passed as arguments — no paths are hardcoded. Each scene runs in its own subprocess: the engine is
a process-global singleton, and running several scenes in one process leaks
state between them and silently degrades metrics.
Do not hardcode local paths
Dataset locations belong in arguments or environment variables. See the
"Do not commit" section of ../AGENTS.md and
tools/check_private_paths.sh.