Files
Debashish SahuandHarry Chen 2a19c6887f sfm: DJI Avata 360 telemetry, attitude and factory lens calibration (#128)
* telemetry: read the Avata 360 djmd with its own field numbers

read_dji applied the Osmo 360 field numbers to dvtm_AVATA360. That printed
an IMU rate of 2^64 (1.10.1 is an int64 -555 on the Avata), put every
quaternion of a frame on one time (sensor fps was read from 1.11, which is
not it), and found no GPS (the Avata writes it at 3.4.4, in degrees, with no
unit field).

A DjiLayout table now carries the numbers per proto:

- clip rate 1.8, fps 1.9, no focal, no per-frame accelerometer
- GPS at 3.4.4.1.{2,3} in degrees, altitude at 3.4.4.2 (mm)
- relative altitude at 3.4.5.1 (f32 mm, new TelemetryGps::rel_alt)
- per-frame exposure: ISO 3.2.3.1 (f32), shutter 3.2.4.1 (two varints n, d,
  n/d seconds), colour temperature 3.2.6.1 (varint K)

The Osmo keeps its numbers. An unknown proto keeps them too and is still
flagged as unverified. f-number and EV are not read: both SRTs carry one
constant value on every frame, so no field can be told apart. Nothing
downstream consumes the exposure yet.

The clip note no longer prints "focal 0.0 px" when the clip has no focal
field. The report gains an exposure line and a relative altitude range on
the GPS line, each only when present, so Osmo reports are unchanged.
docs/notes/imu-gps-for-sfm.md has the field map.

Tests in sfm_telemetry_test: synthetic Avata frames, an Osmo frame carrying
Avata numbers, and an unknown proto. Two more run on real clips when
SS_TEST_AVATA_OSV, SS_TEST_AVATA_OSV_HOVER and SS_TEST_AVATA_LRF_HOVER name
them, and print SKIP otherwise.

The test helper near() becomes close_to() in sfm_telemetry_test. near is a
Windows macro and check_winmacro flags the file otherwise.

Checked against the SRT of a 139 s flight (8354 frames, packet p is FrameCnt
p+1): lat/lon within 5e-7 deg, altitudes within 1 mm, ISO and colour
temperature exact, shutter within the SRT's 1/3-stop rounding, 0 misses.
Shifted by one frame the same comparison misses 2695 colour temperatures
and 662 altitudes, so the alignment is real. A hover clip's .LRF proxy,
read as an oracle, pairs 317 frames by exact time and matches bit for bit.
The oracle shares the reader, so the SRT comparison is the independent one.

An Osmo 360 clip's report is byte-identical before and after.

Written by Claude (Sonnet 5.5), with review by a second Claude session.

* sfm: attitude sense and vertical without an accelerometer (Avata 360)

The Avata 360 writes a 4 kHz fused attitude and no accelerometer. With no
accelerometer, telemetry_check never measures which way the quaternion maps,
and the timeline conjugated it by default. On a real flight that is the wrong
sense: the hand-eye fit gives sig_rot 5.37 deg against 0.315 deg for the
quaternion itself. The eigen gap is 2.0 against 620, and the old refit also
comes out mirrored (det -1). So every run logged that the gyro and the poses
disagree, and no rotation prior was used.

- SensorTimeline: sign -1 conjugates an attitude whose sense no
  accelerometer settled (attitudeSenseOpen), in orientationAt and in upAt's
  attitude branch.
- ImuExtrinsic: the existing sign-hypothesis loop tries that conjugate
  instead of breaking, and recomputes the up votes per hypothesis. A source
  with an accelerometer keeps break-at-minus-one and its own votes.
- Telemetry::attitude_world_up: a carrier may declare the vertical of the
  attitude world. The DJI reader declares z-down for dvtm_AVATA360 only. With
  no accelerometer the timeline uses it as the up source, so hasUp() is true
  and the gravity refit, the up factors and the sensor gauge's up all run.
  An accelerometer-less attitude with no declaration now answers no up vote
  instead of guessing world +Z.

Tests:

- sfm_sensor_prior_test: no accelerometer, declared and undeclared; an
  accelerometer wins over a declared vertical; attitudeSenseOpen() is open
  exactly when there is no accelerometer.
- sfm_sensor_gauge_test T8 and T8b.
- sfm_telemetry_test: the declaration is set for the Avata and not for the
  Osmo or an unknown proto.
- sfm_attitude_replay_test, gated on SS_TEST_AVATA_MODEL and SS_TEST_AVATA_OSV
  (SKIP when unset): runs a finished sparse model and its clip's telemetry
  through the pair calibration, the factors, the registration check and the
  gauge, so the fix is checked without mapping again. It requires det(X) +1
  on every refit, which is what catches a flipped declared vertical: the fit
  settles X's sign by agreement with the up votes, so a flip leaves the up
  consensus at 0.227 deg and mirrors X instead.

On a 2090-image GPS-levelled model of one flight: sig_rot 0.315 deg, up
0.227 deg from the model's +Z, 2090 of 2090 images get a prior, 10 held.
With the old sense restored the gate rejects the pair set. One flight and
one clip only.

Osmo 360 is unchanged: attitudeSenseOpen() is false, hasUp() is true through
the accelerometer, and the declared vertical is never read.

Written by Claude (Sonnet 5.5), with review by a second Claude session.

* sfm: factory lens calibration from the Avata 360 OSV, through #119's params

The Avata 360 clip header carries each lens's calibration at StreamMeta.5
(PanoDewarpParams). Entry 4 (native_refine_far_master) is video track 0 and
entry 3 (slave) is track 1. The model is Kannala-Brandt with five radial
terms (k5 at field 15), tangential p1 p2 at field 20 on the equidistant
coordinates, in pixels of the 3840x3840 frame, lens_model 8.

- Telemetry: LensCalibration, djmd_lenses, video_lenses.
- LensCalibration.h: refits the five-term curve to the fisheye models' four
  (1.97 px worst on the master lens, 1.22 px on the slave, +0.1 px median on
  real matches), and appends one CameraOverride with params per lens folder
  (camN, under the capture prefix), unless a setting a person gave already
  covers it: a per-folder or bare --focal or --distortion, manifest params,
  or a dataset-wide focal, distortion or params. A model-only entry keeps its
  model.
- Pipeline: applied in `auto` before the match signature. The [run] lines
  say which lens got what, or why not. `--sensor-gauge none` does not turn
  it off.
- 7 messages in 13 languages.

Measured on real frames (docs/notes/imu-gps-for-sfm.md section 2.3). Against
DJI Studio's equirect export of a hover clip, the master lens on track 0 fits
at 0.55 px median and the slave's set at 2.2 px. Lens to lens on a flight,
0.88 and 0.90 px, swapped 6.2 and 8.4 px, without p 2.0 to 2.2 px, without k5
unusable (the four-term polynomial turns over before 90 deg).

Tests: sfm_telemetry_test (decoding, and real clips when SS_TEST_AVATA_OSV and
SS_TEST_AVATA_OSV_HOVER are set), and the new sfm_lens_calib_test (the refit
against an independent five-term projection, the bearing() round trip, which
is shown able to fail, precedence, folder sizing).

No accuracy gain over a hand-set focal was measured on one flight (median
distance to GPS 0.77 m with the factory calibration, 0.66 m with a hand
focal, one run each, so not shown to differ from run-to-run noise). What it gives is no
hand-tuned focal and a faster match stage (4:06 against 6:07).

Written by Claude (Sonnet 5.5), with review by a second Claude session.

* sfm: a person's lens setting at any prefix beats the factory calibration

applyLensPlans read only the longest matching override. A focal, extra or
params given for the capture was ignored as soon as a longer override named
one of its lenses with only a model. With `--focal clip=900` and
`--camera-model clip/cam0=opencv-fisheye`, clip/cam0 got the factory
calibration and the person's focal was overridden. The model had the same
flaw: it came from the longest match only, even when that entry set none.

It now reads every matching override. Any that sets a focal, extra or params
makes the lens a person's own. The model is the longest one that names a
model, and is also written into an existing entry for the lens folder, so the
params are read in the model they were cut for.

Tests, each run against the old code first:

- a capture focal with a model-only entry for one lens: the lens kept by the
  person (fails on the old code), and the other lens, which the old code
  already handled
- a capture calibration with a model-only entry for one lens (fails on the
  old code)
- a capture model with a nearer entry that names none: the params are cut in
  the capture's model, and buildCameras builds that lens in it (both fail on
  the old code)
- a dataset-wide manifest calibration, and a frame size that differs in only
  one of width and height. Both pin behaviour the old code already had, so
  they pass on it. They fail on a build with the dataset-wide params check
  removed, and with either size comparison removed.

Written by Claude (Sonnet 5.5), with review by a second Claude session.

* rebuild telemetry wasm module

---------

Co-authored-by: Harry Chen <harry7557558@gmail.com>
2026-09-30 20:35:03 -04:00
..
2026-09-24 19:56:52 -04:00
2026-09-24 19:56:52 -04:00
2026-07-11 23:46:26 -04:00
2026-09-24 19:56:52 -04:00
2026-09-24 19:56:52 -04:00
2026-09-24 19:56:52 -04:00

Spirula Studio Standalone Viewer

A dependency-free, client-side WebGL2 viewer for 3D Gaussian Splats, meshes, and MVS datasets (COLMAP / Nerfstudio / Metashape). It runs entirely in the browser (no server) and can be hosted as a static site (e.g. GitHub Pages). Performance-critical parsing, depth sorting, and statistics run in C++ compiled to WebAssembly (built with CMake + Emscripten); rendering is WebGL2.

The directory holds two independent pages: index.html, the model viewer described below, and telemetry.html, an IMU/GPS viewer for captures. Each has its own WASM module.

The viewer's own runtime code is independent from the rest of the trainer — nothing here imports from the training code at runtime. The one deliberate exception is at build time: the WASM module compiles the trainer's dataset parsers (src/data/parsers/*Parser.cpp) in place — referenced by relative path, not copied — so COLMAP/Nerfstudio/ Metashape parsing has exactly one implementation in the repo. Those files are plain C++17 with no CUDA dependency (see csrc/CameraModel.h).

Features

  • 3D Gaussian Splatting (.ply, INRIA and Spirula Studio layouts, including very large files — binary PLY is parsed streaming through a small chunk buffer, non-position attributes are stored as half floats, rest SH coefficients are 8-bit quantized (Gaussian-wise scale, RGB8_SNORM textures — half the VRAM and heap of f16, visually lossless), and the SH texture / mesh index buffers are split into chunks below per-resource GPU limits; tested with a 6.4 GB, 39M-splat SH2 scan (~3 GB VRAM). Splats are Morton-reordered at load so the depth-sorted draw order touches attribute textures cache-coherently (several-fold frame-rate gain on multi-10M-splat models); the async sort worker keeps only an xyz copy. Deep-linked ?model= URLs stream straight into the parser (no Blob buffering, so multi-GB hosted models load). If the GPU runs out of memory on the SH textures, the viewer drops one SH degree at a time and retries instead of failing.
    • Primitive select: 3DGS, Mip (antialiased), 3DGUT (unscented transform projection; fragments evaluate the 3D Gaussian along per-pixel rays, "eval3d"). 3DGS/Mip use analytic projection Jacobians with the conventional out-of-bounds Jacobian clipping.
    • Color gamut (Rec.709 / DCI-P3 / Rec.2020 / AdobeRGB / ACEScg / ACES2065-1) and a linear-color toggle, matching the training pipeline's rgb_to_srgb conventions.
    • Spherical harmonics up to degree 4, exposure.
    • Depth sorting matches the training code's get_sorting_depth: planar for perspective/orthographic, distance for equirectangular, and a smooth |z|/radial blend for fisheye — content behind the camera composites correctly in >180° views. Sorting runs asynchronously in a Web Worker that hosts its own instance of the WASM module (native counting sort over a positions copy), so looking around never blocks the render loop; results apply latest-wins when ready.
    • GPU-friendly: attributes live in TEXTURE_2D_ARRAYs laid out so arbitrarily large models render even under a small MAX_TEXTURE_SIZE (tiles into array layers).
  • Meshes (.ply, .obj, .gltf, .glb, .stl) as produced by the meshing code: vertex colors or a base-color texture atlas, shading toggles (shaded / unshaded, flat / interpolated normals, color on/off) with a view-following headlight so the surface reads from every angle. GLTF/GLB are parsed in JS; drop the model together with its external .bin / .mtl / image files for textured .gltf / .obj.
  • Orbit / trackball / first-person / free-fly navigation (mouse, touch, keyboard, gamepad — matched to the training viewer), Y-up ↔ Z-up toggle (switching keeps the current view), and camera models perspective / orthographic / fisheye (equidistant) / fisheye (equisolid) / equirectangular 360° — all with analytic Jacobians / sigma-point projection; mesh triangles crossing a projection discontinuity (the equirect seam, the fisheye backward point) are discarded, and fragments outside a fisheye image circle are clipped.
  • A rasterized, depth-tested axes + grid overlay in the model's native frame (power-of-10 cells that adapt to the zoom level; the line patch follows the orbit target while staying on the global lattice) and a configurable background.
  • Statistics (splat count / vertices / edges / faces) and on-demand parameter histograms (opacity, scale, effective rank, RGB, anisotropy for splats; edge length, triangle area, coordinates for meshes), computed in WASM and cached.
  • MVS datasets — drag & drop a COLMAP reconstruction (cameras/images/points3D, binary .bin or text .txt), a Nerfstudio transforms.json (+ point cloud .ply), or a Metashape camera export (.xml + .ply, optional .psx) — as a folder or as individual files:
    • Renders the seed point cloud and camera frustums with the true per-camera projection: the image border is discretized and unprojected through the camera model (pinhole / fisheye / equisolid / equirectangular) including lens distortion (Newton undistort over the parser's distortion tier: OpenCV k1 k2 p1 p2 or thin-prism k1–k4 p1 p2 sx1 sy1), so a fisheye camera's frustum visibly bulges. Wide cameras get an image-aligned wire dome (fisheye) or a lat/long wire globe (equirectangular) instead of a lone border ring. Point size and frustum size are adjustable; either layer can be hidden.
    • A compact info panel: source format, image / point counts, camera groups, per-model image counts.
    • Component picker when the dataset has several reconstructions (sparse/0, sparse/1, multiple Metashape <component> chunks, …).
    • Hover anywhere on a camera frustum (ray-picked, not just its apex) to see its intrinsics (model, fx/fy/cx/cy, distortion); double-click a camera to view the scene from it (average focal, cx/cy/distortion omitted); double-click the point cloud to recenter. Picking is depth-ordered: whatever is visually in front at the cursor — frustum or point cloud — wins.
    • Degrades gracefully: dropping only sparse/, only a transforms.json, only a Metashape .xml, or only a point-cloud .ply shows whatever is available (images are never required, or read at all — only the metadata files are parsed). Load failures and partial loads pop a visible error toast (no need to open the console).
  • Double-click the viewport to center the view on the point under the cursor (MeshLab-style: the point becomes the orbit pivot and slides onto the optical axis). Splats use a one-pixel GPU depth pass; meshes use a WASM raycast; datasets pick the nearest point along the view ray. Works with every camera model.
  • The Center menu (Scene section, next to the up-axis toggle) picks what the view orbits about and fits to: the model's origin, the point cloud's geometric median or mean, or over a dataset the camera positions' median or mean or the point the cameras look at. Camera position median is the default; over a splat or mesh file, which has no cameras, the camera entries are disabled and the point statistics stand in. The same six modes are the trainer's --scene-center and the native GUI viewport's menu (src/data/SceneCenter.h is the one implementation).
  • One model at a time — dropping another replaces it and frees the previous GPU buffers. Replacing keeps the current viewpoint (the camera is only fitted for the first model; refresh the page to start over).

Telemetry viewer — telemetry.html

A second, independent page in the same directory: drop videos, photos or whole folders and see the IMU and GPS they carry. It reuses the trainer's reader (src/sfm/core/Telemetry.cpp, compiled in place into its own WASM module) so the browser sees exactly what the reconstruction pipeline sees — docs/notes/imu-gps-for-sfm.md is what that is for.

  • Detection is by content, not extension. A GoPro gpmd track (GPMF), the Insta360 trailer that follows the MP4, a DJI djmd protobuf track, a Google CAMM track, or EXIF GPS in a JPEG. Anything else is skipped silently, so a folder of a few thousand mixed files is a reasonable thing to drop.
  • Files stay on the machine and are never buffered whole. The scan worker answers the parser's reads with FileReaderSync over a File slice behind a 1 MB block cache, so a 4 GB capture is read in place — tested end to end on 4 GB .360 and .insv files, which no ArrayBuffer would hold. Up to three workers scan in parallel, each with its own module instance.
  • 3D viewport in a local metric frame (east / north / up, metres at the scene origin) with a raster basemap on the ground plane — OpenStreetMap, Esri satellite, CARTO dark, OpenTopoMap, or a custom {z}/{x}/{y} template for a keyed provider. Tiles load lazily by view and zoom, and a tile that has not arrived is drawn from its nearest loaded ancestor, so panning never shows holes. Altitude is real (with an exaggeration slider and optional drop lines to the ground), and the frame is Web Mercator scaled by cos(lat₀), which is the one frame where imagery and track agree without warping either.
  • The IMU is drawn where it was measured: an attitude triad at the playhead, optional acceleration and gravity vectors, and an attitude trail every n seconds along the track. The track itself can be coloured by time, speed, altitude, DOP, or gyro/accel magnitude.
  • Stacked plots under the viewport share a time axis with the viewport's playhead (drag to scrub, shift-drag to pan, wheel to zoom, space to play). Each channel keeps a min/max pyramid, each level 8× coarser than the one below, so a 400 000-sample gyro redraws in the cost of the canvas width rather than the capture length.
  • Screenshot-safe mode (one toggle, also in the header) hides the basemap, latitude/longitude, absolute altitude, wall-clock times and file names, and redacts the fix line from the report — shapes, distances and every IMU stream stay, so a plot can be shared without publishing where it was taken.
  • Per-file verdicts come from the same checks the trainer runs: sample-rate regularity, the gravity norm, cross-stream gravity agreement, and whether the GPS moves rather than repeating one stale fix. The full text report is in the panel.

English only, by design — this is a diagnostic, and the model viewer's i18n would be overkill.

Build

Requires the Emscripten SDK on PATH (emcc, emcmake) and CMake ≥ 3.16.

source ~/emsdk/emsdk_env.sh      # activate emsdk
./build.sh                       # both modules
./build.sh ssv_telemetry         # just the telemetry one

The two modules are separate CMake targets, so editing one page's sources rebuilds only that page. This produces js/ssv_wasm.{js,wasm} and js/ssv_telemetry.{js,wasm} (committed so the site is directly hostable without a build step).

Run

Serve the directory over HTTP (ES modules + WASM require it — file:// will not work):

python3 -m http.server -d .      # then open http://localhost:8000/

Drag a model or a dataset folder onto the canvas, or click to browse (the picker takes multiple files; folder drops walk the directory tree). You can also deep-link a hosted model: index.html?model=<url> (external .bin/.mtl/image siblings are fetched automatically).

telemetry.html is the IMU/GPS page. Its basemap fetches tiles from a public provider at view time, so that page needs network access and the provider's usage policy applies (the built-ins are fine for interactive use; for anything heavier, pick Custom URL… and paste your own {z}/{x}/{y} template with a key). Choosing None — which screenshot-safe mode does anyway — leaves the page fully offline.

Test

test/run.sh generates synthetic COLMAP / Nerfstudio / Metashape datasets and drives the viewer end-to-end in headless Chrome (needs google-chrome and node ≥ 20):

test/run.sh            # all cases
test/run.sh metashape  # one case; see test/test_ds.html for the list

The telemetry page's parser is covered by the trainer's sfm_telemetry_test (synthetic GPMF / Insta360 / DJI / CAMM fixtures). test/telemetry_probe.mjs checks the page itself — WASM boot, the worker's File reads, the clip model — against captures of your own, which is the only way to exercise the multi-GB path:

python3 -m http.server 8098 &
node test/telemetry_probe.mjs http://localhost:8098/telemetry.html VIDEO.360 PHOTO.jpg

Layout

index.html          model viewer: UI + panel
telemetry.html      IMU/GPS viewer: UI + panel
css/style.css        theme (adapted from the training viewer)
css/telemetry.css    the telemetry page's own layout and overlays
js/
  main.js            app wiring, input, render loop, sorting
  renderer.js        WebGL2 renderer (splat HDR pass, mesh, dataset, grid lines, tonemap)
  shaders.js         GLSL (splat / tonemap / mesh / points / frustum+grid lines)
  camera.js          camera + navigation modes
  dataset.js         dataset load orchestration + frustum geometry (undistort)
  sortworker.js      async depth-sort worker (own WASM instance)
  wasm.js            WASM bridge + streaming loader + GLTF/GLB (JS) + MEMFS mount
  colors.js          gamut matrices
  linalg.js          vec/quat/mat helpers
  histogram.js       canvas bar chart
  ssv_wasm.{js,wasm} built WASM module
  telemetry/
    main.js          telemetry app wiring, options, viewport render loop
    scanworker.js    per-file scan worker (own WASM instance, FileReaderSync)
    scene.js         WebGL2 viewport (instanced polylines, tiles, points)
    tiles.js         basemap tile cache with ancestor fallback
    charts.js        stacked plots with min/max decimation pyramids
    geo.js           Mercator frame, per-clip model, derived series
  ssv_telemetry.{js,wasm}  built telemetry WASM module
src/viewer.cpp       parsers (PLY/OBJ), depth sort, histograms
src/dataset_bridge.cpp  dataset C ABI over MEMFS (drives the trainer's parsers)
src/telemetry_bridge.cpp  telemetry C ABI (drives the trainer's Telemetry.cpp)
test/                end-to-end harness (headless Chrome) + data generators
CMakeLists.txt       Emscripten build (also compiles ../src/data/ parsers)

Frustum lines use the same seam handling as meshes: fragments whose interpolated camera-space position re-projects far from the rasterized position (a segment wrapped across the equirect ±180° seam or the fisheye backward point) are discarded, and fisheye display models clip to the image circle. Double-clicking a dataset camera switches the display projection to that camera's model (pinhole → perspective, fisheye → equidistant, equisolid, equirectangular) with the matching field of view.