* telemetry: read the Avata 360 djmd with its own field numbers
read_dji applied the Osmo 360 field numbers to dvtm_AVATA360. That printed
an IMU rate of 2^64 (1.10.1 is an int64 -555 on the Avata), put every
quaternion of a frame on one time (sensor fps was read from 1.11, which is
not it), and found no GPS (the Avata writes it at 3.4.4, in degrees, with no
unit field).
A DjiLayout table now carries the numbers per proto:
- clip rate 1.8, fps 1.9, no focal, no per-frame accelerometer
- GPS at 3.4.4.1.{2,3} in degrees, altitude at 3.4.4.2 (mm)
- relative altitude at 3.4.5.1 (f32 mm, new TelemetryGps::rel_alt)
- per-frame exposure: ISO 3.2.3.1 (f32), shutter 3.2.4.1 (two varints n, d,
n/d seconds), colour temperature 3.2.6.1 (varint K)
The Osmo keeps its numbers. An unknown proto keeps them too and is still
flagged as unverified. f-number and EV are not read: both SRTs carry one
constant value on every frame, so no field can be told apart. Nothing
downstream consumes the exposure yet.
The clip note no longer prints "focal 0.0 px" when the clip has no focal
field. The report gains an exposure line and a relative altitude range on
the GPS line, each only when present, so Osmo reports are unchanged.
docs/notes/imu-gps-for-sfm.md has the field map.
Tests in sfm_telemetry_test: synthetic Avata frames, an Osmo frame carrying
Avata numbers, and an unknown proto. Two more run on real clips when
SS_TEST_AVATA_OSV, SS_TEST_AVATA_OSV_HOVER and SS_TEST_AVATA_LRF_HOVER name
them, and print SKIP otherwise.
The test helper near() becomes close_to() in sfm_telemetry_test. near is a
Windows macro and check_winmacro flags the file otherwise.
Checked against the SRT of a 139 s flight (8354 frames, packet p is FrameCnt
p+1): lat/lon within 5e-7 deg, altitudes within 1 mm, ISO and colour
temperature exact, shutter within the SRT's 1/3-stop rounding, 0 misses.
Shifted by one frame the same comparison misses 2695 colour temperatures
and 662 altitudes, so the alignment is real. A hover clip's .LRF proxy,
read as an oracle, pairs 317 frames by exact time and matches bit for bit.
The oracle shares the reader, so the SRT comparison is the independent one.
An Osmo 360 clip's report is byte-identical before and after.
Written by Claude (Sonnet 5.5), with review by a second Claude session.
* sfm: attitude sense and vertical without an accelerometer (Avata 360)
The Avata 360 writes a 4 kHz fused attitude and no accelerometer. With no
accelerometer, telemetry_check never measures which way the quaternion maps,
and the timeline conjugated it by default. On a real flight that is the wrong
sense: the hand-eye fit gives sig_rot 5.37 deg against 0.315 deg for the
quaternion itself. The eigen gap is 2.0 against 620, and the old refit also
comes out mirrored (det -1). So every run logged that the gyro and the poses
disagree, and no rotation prior was used.
- SensorTimeline: sign -1 conjugates an attitude whose sense no
accelerometer settled (attitudeSenseOpen), in orientationAt and in upAt's
attitude branch.
- ImuExtrinsic: the existing sign-hypothesis loop tries that conjugate
instead of breaking, and recomputes the up votes per hypothesis. A source
with an accelerometer keeps break-at-minus-one and its own votes.
- Telemetry::attitude_world_up: a carrier may declare the vertical of the
attitude world. The DJI reader declares z-down for dvtm_AVATA360 only. With
no accelerometer the timeline uses it as the up source, so hasUp() is true
and the gravity refit, the up factors and the sensor gauge's up all run.
An accelerometer-less attitude with no declaration now answers no up vote
instead of guessing world +Z.
Tests:
- sfm_sensor_prior_test: no accelerometer, declared and undeclared; an
accelerometer wins over a declared vertical; attitudeSenseOpen() is open
exactly when there is no accelerometer.
- sfm_sensor_gauge_test T8 and T8b.
- sfm_telemetry_test: the declaration is set for the Avata and not for the
Osmo or an unknown proto.
- sfm_attitude_replay_test, gated on SS_TEST_AVATA_MODEL and SS_TEST_AVATA_OSV
(SKIP when unset): runs a finished sparse model and its clip's telemetry
through the pair calibration, the factors, the registration check and the
gauge, so the fix is checked without mapping again. It requires det(X) +1
on every refit, which is what catches a flipped declared vertical: the fit
settles X's sign by agreement with the up votes, so a flip leaves the up
consensus at 0.227 deg and mirrors X instead.
On a 2090-image GPS-levelled model of one flight: sig_rot 0.315 deg, up
0.227 deg from the model's +Z, 2090 of 2090 images get a prior, 10 held.
With the old sense restored the gate rejects the pair set. One flight and
one clip only.
Osmo 360 is unchanged: attitudeSenseOpen() is false, hasUp() is true through
the accelerometer, and the declared vertical is never read.
Written by Claude (Sonnet 5.5), with review by a second Claude session.
* sfm: factory lens calibration from the Avata 360 OSV, through #119's params
The Avata 360 clip header carries each lens's calibration at StreamMeta.5
(PanoDewarpParams). Entry 4 (native_refine_far_master) is video track 0 and
entry 3 (slave) is track 1. The model is Kannala-Brandt with five radial
terms (k5 at field 15), tangential p1 p2 at field 20 on the equidistant
coordinates, in pixels of the 3840x3840 frame, lens_model 8.
- Telemetry: LensCalibration, djmd_lenses, video_lenses.
- LensCalibration.h: refits the five-term curve to the fisheye models' four
(1.97 px worst on the master lens, 1.22 px on the slave, +0.1 px median on
real matches), and appends one CameraOverride with params per lens folder
(camN, under the capture prefix), unless a setting a person gave already
covers it: a per-folder or bare --focal or --distortion, manifest params,
or a dataset-wide focal, distortion or params. A model-only entry keeps its
model.
- Pipeline: applied in `auto` before the match signature. The [run] lines
say which lens got what, or why not. `--sensor-gauge none` does not turn
it off.
- 7 messages in 13 languages.
Measured on real frames (docs/notes/imu-gps-for-sfm.md section 2.3). Against
DJI Studio's equirect export of a hover clip, the master lens on track 0 fits
at 0.55 px median and the slave's set at 2.2 px. Lens to lens on a flight,
0.88 and 0.90 px, swapped 6.2 and 8.4 px, without p 2.0 to 2.2 px, without k5
unusable (the four-term polynomial turns over before 90 deg).
Tests: sfm_telemetry_test (decoding, and real clips when SS_TEST_AVATA_OSV and
SS_TEST_AVATA_OSV_HOVER are set), and the new sfm_lens_calib_test (the refit
against an independent five-term projection, the bearing() round trip, which
is shown able to fail, precedence, folder sizing).
No accuracy gain over a hand-set focal was measured on one flight (median
distance to GPS 0.77 m with the factory calibration, 0.66 m with a hand
focal, one run each, so not shown to differ from run-to-run noise). What it gives is no
hand-tuned focal and a faster match stage (4:06 against 6:07).
Written by Claude (Sonnet 5.5), with review by a second Claude session.
* sfm: a person's lens setting at any prefix beats the factory calibration
applyLensPlans read only the longest matching override. A focal, extra or
params given for the capture was ignored as soon as a longer override named
one of its lenses with only a model. With `--focal clip=900` and
`--camera-model clip/cam0=opencv-fisheye`, clip/cam0 got the factory
calibration and the person's focal was overridden. The model had the same
flaw: it came from the longest match only, even when that entry set none.
It now reads every matching override. Any that sets a focal, extra or params
makes the lens a person's own. The model is the longest one that names a
model, and is also written into an existing entry for the lens folder, so the
params are read in the model they were cut for.
Tests, each run against the old code first:
- a capture focal with a model-only entry for one lens: the lens kept by the
person (fails on the old code), and the other lens, which the old code
already handled
- a capture calibration with a model-only entry for one lens (fails on the
old code)
- a capture model with a nearer entry that names none: the params are cut in
the capture's model, and buildCameras builds that lens in it (both fail on
the old code)
- a dataset-wide manifest calibration, and a frame size that differs in only
one of width and height. Both pin behaviour the old code already had, so
they pass on it. They fail on a build with the dataset-wide params check
removed, and with either size comparison removed.
Written by Claude (Sonnet 5.5), with review by a second Claude session.
* rebuild telemetry wasm module
---------
Co-authored-by: Harry Chen <harry7557558@gmail.com>
Spirula Studio Standalone Viewer
A dependency-free, client-side WebGL2 viewer for 3D Gaussian Splats, meshes, and MVS datasets (COLMAP / Nerfstudio / Metashape). It runs entirely in the browser (no server) and can be hosted as a static site (e.g. GitHub Pages). Performance-critical parsing, depth sorting, and statistics run in C++ compiled to WebAssembly (built with CMake + Emscripten); rendering is WebGL2.
The directory holds two independent pages: index.html, the model viewer
described below, and telemetry.html, an
IMU/GPS viewer for captures. Each has its own WASM module.
The viewer's own runtime code is independent from the rest of
the trainer — nothing here imports from the training code at runtime. The
one deliberate exception is at build time: the WASM module compiles the
trainer's dataset parsers (src/data/parsers/*Parser.cpp)
in place — referenced by relative path, not copied — so COLMAP/Nerfstudio/
Metashape parsing has exactly one implementation in the repo. Those files are
plain C++17 with no CUDA dependency (see csrc/CameraModel.h).
Features
- 3D Gaussian Splatting (
.ply, INRIA and Spirula Studio layouts, including very large files — binary PLY is parsed streaming through a small chunk buffer, non-position attributes are stored as half floats, rest SH coefficients are 8-bit quantized (Gaussian-wise scale,RGB8_SNORMtextures — half the VRAM and heap of f16, visually lossless), and the SH texture / mesh index buffers are split into chunks below per-resource GPU limits; tested with a 6.4 GB, 39M-splat SH2 scan (~3 GB VRAM). Splats are Morton-reordered at load so the depth-sorted draw order touches attribute textures cache-coherently (several-fold frame-rate gain on multi-10M-splat models); the async sort worker keeps only an xyz copy. Deep-linked?model=URLs stream straight into the parser (no Blob buffering, so multi-GB hosted models load). If the GPU runs out of memory on the SH textures, the viewer drops one SH degree at a time and retries instead of failing.- Primitive select: 3DGS, Mip (antialiased), 3DGUT (unscented transform projection; fragments evaluate the 3D Gaussian along per-pixel rays, "eval3d"). 3DGS/Mip use analytic projection Jacobians with the conventional out-of-bounds Jacobian clipping.
- Color gamut (Rec.709 / DCI-P3 / Rec.2020 / AdobeRGB / ACEScg /
ACES2065-1) and a linear-color toggle, matching the training pipeline's
rgb_to_srgbconventions. - Spherical harmonics up to degree 4, exposure.
- Depth sorting matches the training code's
get_sorting_depth: planar for perspective/orthographic, distance for equirectangular, and a smooth |z|/radial blend for fisheye — content behind the camera composites correctly in >180° views. Sorting runs asynchronously in a Web Worker that hosts its own instance of the WASM module (native counting sort over a positions copy), so looking around never blocks the render loop; results apply latest-wins when ready. - GPU-friendly: attributes live in
TEXTURE_2D_ARRAYs laid out so arbitrarily large models render even under a smallMAX_TEXTURE_SIZE(tiles into array layers).
- Meshes (
.ply,.obj,.gltf,.glb,.stl) as produced by the meshing code: vertex colors or a base-color texture atlas, shading toggles (shaded / unshaded, flat / interpolated normals, color on/off) with a view-following headlight so the surface reads from every angle. GLTF/GLB are parsed in JS; drop the model together with its external.bin/.mtl/ image files for textured.gltf/.obj. - Orbit / trackball / first-person / free-fly navigation (mouse, touch, keyboard, gamepad — matched to the training viewer), Y-up ↔ Z-up toggle (switching keeps the current view), and camera models perspective / orthographic / fisheye (equidistant) / fisheye (equisolid) / equirectangular 360° — all with analytic Jacobians / sigma-point projection; mesh triangles crossing a projection discontinuity (the equirect seam, the fisheye backward point) are discarded, and fragments outside a fisheye image circle are clipped.
- A rasterized, depth-tested axes + grid overlay in the model's native frame (power-of-10 cells that adapt to the zoom level; the line patch follows the orbit target while staying on the global lattice) and a configurable background.
- Statistics (splat count / vertices / edges / faces) and on-demand parameter histograms (opacity, scale, effective rank, RGB, anisotropy for splats; edge length, triangle area, coordinates for meshes), computed in WASM and cached.
- MVS datasets — drag & drop a COLMAP reconstruction
(
cameras/images/points3D, binary.binor text.txt), a Nerfstudiotransforms.json(+ point cloud.ply), or a Metashape camera export (.xml+.ply, optional.psx) — as a folder or as individual files:- Renders the seed point cloud and camera frustums with the true
per-camera projection: the image border is discretized and unprojected
through the camera model (pinhole / fisheye / equisolid / equirectangular)
including lens distortion (Newton undistort over the parser's
distortion tier: OpenCV
k1 k2 p1 p2or thin-prismk1–k4 p1 p2 sx1 sy1), so a fisheye camera's frustum visibly bulges. Wide cameras get an image-aligned wire dome (fisheye) or a lat/long wire globe (equirectangular) instead of a lone border ring. Point size and frustum size are adjustable; either layer can be hidden. - A compact info panel: source format, image / point counts, camera groups, per-model image counts.
- Component picker when the dataset has several reconstructions
(
sparse/0,sparse/1, multiple Metashape<component>chunks, …). - Hover anywhere on a camera frustum (ray-picked, not just its apex) to see its intrinsics (model, fx/fy/cx/cy, distortion); double-click a camera to view the scene from it (average focal, cx/cy/distortion omitted); double-click the point cloud to recenter. Picking is depth-ordered: whatever is visually in front at the cursor — frustum or point cloud — wins.
- Degrades gracefully: dropping only
sparse/, only atransforms.json, only a Metashape.xml, or only a point-cloud.plyshows whatever is available (images are never required, or read at all — only the metadata files are parsed). Load failures and partial loads pop a visible error toast (no need to open the console).
- Renders the seed point cloud and camera frustums with the true
per-camera projection: the image border is discretized and unprojected
through the camera model (pinhole / fisheye / equisolid / equirectangular)
including lens distortion (Newton undistort over the parser's
distortion tier: OpenCV
- Double-click the viewport to center the view on the point under the cursor (MeshLab-style: the point becomes the orbit pivot and slides onto the optical axis). Splats use a one-pixel GPU depth pass; meshes use a WASM raycast; datasets pick the nearest point along the view ray. Works with every camera model.
- The Center menu (Scene section, next to the up-axis toggle) picks what
the view orbits about and fits to: the model's origin, the point cloud's
geometric median or mean, or over a dataset the camera positions' median or
mean or the point the cameras look at. Camera position median is the
default; over a splat or mesh file, which has no cameras, the camera entries
are disabled and the point statistics stand in. The same six modes are the
trainer's
--scene-centerand the native GUI viewport's menu (src/data/SceneCenter.his the one implementation). - One model at a time — dropping another replaces it and frees the previous GPU buffers. Replacing keeps the current viewpoint (the camera is only fitted for the first model; refresh the page to start over).
Telemetry viewer — telemetry.html
A second, independent page in the same directory: drop videos, photos or whole
folders and see the IMU and GPS they carry. It reuses the trainer's reader
(src/sfm/core/Telemetry.cpp, compiled in place into its own WASM module) so
the browser sees exactly what the reconstruction pipeline sees —
docs/notes/imu-gps-for-sfm.md is what that is for.
- Detection is by content, not extension. A GoPro
gpmdtrack (GPMF), the Insta360 trailer that follows the MP4, a DJIdjmdprotobuf track, a Google CAMM track, or EXIF GPS in a JPEG. Anything else is skipped silently, so a folder of a few thousand mixed files is a reasonable thing to drop. - Files stay on the machine and are never buffered whole. The scan worker
answers the parser's reads with
FileReaderSyncover aFileslice behind a 1 MB block cache, so a 4 GB capture is read in place — tested end to end on 4 GB.360and.insvfiles, which noArrayBufferwould hold. Up to three workers scan in parallel, each with its own module instance. - 3D viewport in a local metric frame (east / north / up, metres at the
scene origin) with a raster basemap on the ground plane —
OpenStreetMap, Esri satellite, CARTO dark, OpenTopoMap, or a custom
{z}/{x}/{y}template for a keyed provider. Tiles load lazily by view and zoom, and a tile that has not arrived is drawn from its nearest loaded ancestor, so panning never shows holes. Altitude is real (with an exaggeration slider and optional drop lines to the ground), and the frame is Web Mercator scaled by cos(lat₀), which is the one frame where imagery and track agree without warping either. - The IMU is drawn where it was measured: an attitude triad at the playhead, optional acceleration and gravity vectors, and an attitude trail every n seconds along the track. The track itself can be coloured by time, speed, altitude, DOP, or gyro/accel magnitude.
- Stacked plots under the viewport share a time axis with the viewport's playhead (drag to scrub, shift-drag to pan, wheel to zoom, space to play). Each channel keeps a min/max pyramid, each level 8× coarser than the one below, so a 400 000-sample gyro redraws in the cost of the canvas width rather than the capture length.
- Screenshot-safe mode (one toggle, also in the header) hides the basemap, latitude/longitude, absolute altitude, wall-clock times and file names, and redacts the fix line from the report — shapes, distances and every IMU stream stay, so a plot can be shared without publishing where it was taken.
- Per-file verdicts come from the same checks the trainer runs: sample-rate regularity, the gravity norm, cross-stream gravity agreement, and whether the GPS moves rather than repeating one stale fix. The full text report is in the panel.
English only, by design — this is a diagnostic, and the model viewer's i18n would be overkill.
Build
Requires the Emscripten SDK on PATH (emcc, emcmake) and CMake ≥ 3.16.
source ~/emsdk/emsdk_env.sh # activate emsdk
./build.sh # both modules
./build.sh ssv_telemetry # just the telemetry one
The two modules are separate CMake targets, so editing one page's sources
rebuilds only that page. This produces js/ssv_wasm.{js,wasm} and
js/ssv_telemetry.{js,wasm} (committed so the site is directly hostable
without a build step).
Run
Serve the directory over HTTP (ES modules + WASM require it — file:// will not
work):
python3 -m http.server -d . # then open http://localhost:8000/
Drag a model or a dataset folder onto the canvas, or click to browse (the
picker takes multiple files; folder drops walk the directory tree). You can
also deep-link a hosted model: index.html?model=<url> (external
.bin/.mtl/image siblings are fetched automatically).
telemetry.html is the IMU/GPS page. Its basemap fetches tiles from a public
provider at view time, so that page needs network access and the provider's
usage policy applies (the built-ins are fine for interactive use; for anything
heavier, pick Custom URL… and paste your own {z}/{x}/{y} template with a
key). Choosing None — which screenshot-safe mode does anyway — leaves the
page fully offline.
Test
test/run.sh generates synthetic COLMAP / Nerfstudio / Metashape datasets and
drives the viewer end-to-end in headless Chrome (needs google-chrome and
node ≥ 20):
test/run.sh # all cases
test/run.sh metashape # one case; see test/test_ds.html for the list
The telemetry page's parser is covered by the trainer's sfm_telemetry_test
(synthetic GPMF / Insta360 / DJI / CAMM fixtures). test/telemetry_probe.mjs
checks the page itself — WASM boot, the worker's File reads, the clip
model — against captures of your own, which is the only way to exercise the
multi-GB path:
python3 -m http.server 8098 &
node test/telemetry_probe.mjs http://localhost:8098/telemetry.html VIDEO.360 PHOTO.jpg
Layout
index.html model viewer: UI + panel
telemetry.html IMU/GPS viewer: UI + panel
css/style.css theme (adapted from the training viewer)
css/telemetry.css the telemetry page's own layout and overlays
js/
main.js app wiring, input, render loop, sorting
renderer.js WebGL2 renderer (splat HDR pass, mesh, dataset, grid lines, tonemap)
shaders.js GLSL (splat / tonemap / mesh / points / frustum+grid lines)
camera.js camera + navigation modes
dataset.js dataset load orchestration + frustum geometry (undistort)
sortworker.js async depth-sort worker (own WASM instance)
wasm.js WASM bridge + streaming loader + GLTF/GLB (JS) + MEMFS mount
colors.js gamut matrices
linalg.js vec/quat/mat helpers
histogram.js canvas bar chart
ssv_wasm.{js,wasm} built WASM module
telemetry/
main.js telemetry app wiring, options, viewport render loop
scanworker.js per-file scan worker (own WASM instance, FileReaderSync)
scene.js WebGL2 viewport (instanced polylines, tiles, points)
tiles.js basemap tile cache with ancestor fallback
charts.js stacked plots with min/max decimation pyramids
geo.js Mercator frame, per-clip model, derived series
ssv_telemetry.{js,wasm} built telemetry WASM module
src/viewer.cpp parsers (PLY/OBJ), depth sort, histograms
src/dataset_bridge.cpp dataset C ABI over MEMFS (drives the trainer's parsers)
src/telemetry_bridge.cpp telemetry C ABI (drives the trainer's Telemetry.cpp)
test/ end-to-end harness (headless Chrome) + data generators
CMakeLists.txt Emscripten build (also compiles ../src/data/ parsers)
Frustum lines use the same seam handling as meshes: fragments whose interpolated camera-space position re-projects far from the rasterized position (a segment wrapped across the equirect ±180° seam or the fisheye backward point) are discarded, and fisheye display models clip to the image circle. Double-clicking a dataset camera switches the display projection to that camera's model (pinhole → perspective, fisheye → equidistant, equisolid, equirectangular) with the matching field of view.