code restructure

This commit is contained in:
Harry Chen
2026-07-22 14:11:28 -04:00
parent 2992414892
commit 7908fb54a0
502 changed files with 2580 additions and 2238 deletions
+4 -4
View File
@@ -1,8 +1,8 @@
# Third-party dependencies (miniz, stb, npy, ...)
spirulae_splat/splat/cuda/csrc/external/** linguist-vendored
src/external/** linguist-vendored
# Auto-generated code (generate_headers.py, generate_kernel_instantiation.py,
# generate_cli_config.py)
spirulae_splat/splat/cuda/csrc/generated/** linguist-vendored linguist-generated
spirulae_splat/splat/cuda/csrc/app/generated/** linguist-vendored linguist-generated
spirulae_splat/splat/cuda/ins/** linguist-vendored linguist-generated
src/generated/** linguist-vendored linguist-generated
src/app/generated/** linguist-vendored linguist-generated
src/instantiations/** linguist-vendored linguist-generated
+65 -51
View File
@@ -24,53 +24,67 @@ Direction of travel, so you don't push the wrong way:
## Repo map
```
CMakeLists.txt all native builds (723 lines, 4 modes — see docs/build.md)
setup.py / pyproject.toml the pip/torch-extension build (independent of CMake)
CMakeLists.txt running order only; the logic is in cmake/
cmake/ build modules (options, backends, apps) + sources.txt,
the source list BOTH build systems read
setup.py / pyproject.toml the pip/torch-extension build (own flags; see docs/build.md)
build_develop.bash/.bat the dev build entry points — USE THESE, not pip install
AGENTS.md CLAUDE.md this file (CLAUDE.md is a pointer to it)
docs/ architecture, build, backends, codegen, testing, notes
src/ ALL native code (see below)
tools/codegen/ the five codegen tools (see "Codegen" below)
tests/python/ Python tests (may be stale; see docs/testing.md)
tests/native/ standalone native benchmarks
scripts/ dataset preprocessing CLI tools (Python, standalone)
tests/ Python tests (may be stale; see docs/testing.md)
viewer/ standalone WebGL2 + WASM viewer, independent of training
(build-time exception: compiles csrc/app/*Parser.cpp in place)
spirulae_splat/ Python package
(build-time exception: compiles src/data/parsers/*.cpp
in place)
spirulae_splat/ Python package — a *client* of the engine, on its way out
├── ss_{trainer,benchmark,viewer,meshing}.py console-script entry points
├── generate_*.py the five codegen tools (see "Codegen" below)
├── modules/ config dataclasses, training driver, eval metrics, resume
├── viewer/ Python HTTP viewer server + viewer.html (html is shared
│ with the C++ viewer, which embeds it at build time)
└── splat/cuda/
├── _backend.py _wrapper*.py the extension import + lazy function wrappers
├── slang/ Slang device math, shared by CUDA and Vulkan
│ └── vulkan/ Vulkan-only compute entry points
├── ins/ GENERATED kernel instantiations (111 files)
└── csrc/ ALL native code (see below)
└── splat/cuda/*.py the extension import + lazy function wrappers
```
`spirulae_splat/splat/cuda/csrc/` is where nearly everything lives. The path
is vestigial — `splat/` is a gsplat leftover, `cuda/` also holds the Slang
shaders and the Vulkan backend, and `csrc/` holds the engine, the dataset
parsers, the CLI, the GUI and the web viewer. Inside it:
`src/` is the include root: every local include is path-qualified relative to
it (`#include "core/Common.cuh"`, `#include "backend/api/BackendTypes.h"`), so
a file's include lines tell you which subsystems it depends on.
```
csrc/
├── Engine*.cpp/.h the torch-free training engine (process-global singleton)
├── <Kernel>.cu kernel launchers + __global__ kernels
├── <Kernel>.cuh GENERATED declaration section (see "Codegen")
├── <Kernel>_kernel.cuh the device-side kernel body, split out for reuse
├── Bilagrid* 43 files — bilateral grid; see docs/notes/
├── Primitive*.cuh 3DGS / Mip / 3DGUT primitive traits
├── DataManager.{cpp,h} image cache / prefetch / warp pipeline
├── Tensor.h Camera.h Common.cuh core types and device helpers
├── Mesh* Delaunay3D.* meshing pipeline
├── ext.cpp the pybind11 module (144 m.def's)
├── app/ ssplat-train CLI, dataset parsers, web viewer, gui/
src/
├── core/ Tensor.h, Camera.h, Common.cuh, GradQuant.cuh, …
│ the types and device helpers everything uses
├── primitives/ Primitive*.cuh — 3DGS / Mip / 3DGUT traits
│ (compile-time types, not runtime branches)
├── kernels/ one directory per family; each holds the launchers
│ │ (<Name>.cu), the GENERATED declaration header
│ │ (<Name>.cuh) and the device body (<Name>_kernel.cuh)
│ ├── projection/ raster/ tile/
│ ├── pixelwise/ ppisp/ bilagrid/ (bilagrid is 42 files, see docs/notes/)
│ ├── optim/ densify/ loss/ background/ visualize/
├── engine/ Engine*.cpp/.h — the torch-free training engine
│ (process-global singleton)
├── data/ DataManager (image cache / prefetch / warp) and
│ └── parsers/ COLMAP / Nerfstudio / Metashape readers
├── mesh/ meshing pipeline, Delaunay3D, UV, export
├── backend/ the backend seam — READ backend/README.md
│ ├── api/ backend-neutral launch declarations (GENERATED forwarders)
│ ├── cuda/ common/ CUDA runtime shim, SortScan, Profiler
│ └── vulkan/ Vulkan runtime + kernels/ — READ backend/vulkan/README.md
├── generated/ external/ do not hand-edit / vendored
└── tests/ backend/tests/ native parity tests
│ ├── vulkan/ the whole Vulkan backend: runtime + kernels/ +
│ │ shaders/ (its own entry points and SPIR-V build)
│ │ — READ backend/vulkan/README.md
│ └── tests/ native cross-backend parity tests
├── shaders/ Slang device math SHARED by both backends — compiled
│ twice, to src/generated/*.cuh and to SPIR-V
├── app/ the native applications
│ ├── cli/ main.cpp (ssplat-train), mesh_main.cpp (ssplat-mesh)
│ ├── gui/ Dear ImGui desktop app (ssplat-gui)
│ ├── webviewer/ HTTP server + render worker for the embedded viewer
│ └── TrainerCore.{h,cpp} the CLI/GUI training loop
├── bindings/ext.cpp the pybind11 module (144 m.def's)
├── generated/ app/generated/ instantiations/ GENERATED — do not hand-edit
└── external/ vendored (miniz, stb, npy)
```
## Building
@@ -97,17 +111,17 @@ Full matrix and per-platform notes: `docs/build.md`.
## Codegen — the invariants that bite
Generated trees are marked in `.gitattributes` (`csrc/generated/`,
`csrc/app/generated/`, `cuda/ins/`) and are **committed**, so a fresh checkout
Generated trees are marked in `.gitattributes` (`src/generated/`,
`src/app/generated/`, `src/instantiations/`) and are **committed**, so a fresh checkout
builds with no Python at all. Five generators, all run from the repo root:
| generator | reads | writes |
|---|---|---|
| `generate_headers.py` | `/*[AutoHeaderGeneratorExport]*/` markers in `csrc/*.cu` | the declaration section of `csrc/<Name>.cuh` |
| `generate_kernel_instantiation.py` | kernel decls in `csrc/*.cuh` | `cuda/ins/*.cu` |
| `generate_cli_config.py` | the Python config dataclasses (`ast`-parsed, no torch import) | `csrc/app/generated/cli_config.h` |
| `generate_backend_api.py` | per-kernel `.cuh` headers | `csrc/backend/api/*.h` forwarders |
| `generate_vulkan_stubs.py` | link-probes the Vulkan build | throwing stubs for unported kernels |
| `tools/codegen/generate_headers.py` | `/*[AutoHeaderGeneratorExport]*/` markers in `src/kernels/**/*.cu` | the declaration section of the matching `<Name>.cuh` |
| `tools/codegen/generate_kernel_instantiation.py` | kernel decls in `src/kernels/**/*_kernel.cuh` | `src/instantiations/*.cu` |
| `tools/codegen/generate_cli_config.py` | the Python config dataclasses (`ast`-parsed, no torch import) | `src/app/generated/cli_config.h` |
| `tools/codegen/generate_backend_api.py` | per-kernel `.cuh` headers | `src/backend/api/*.h` forwarders |
| `tools/codegen/generate_vulkan_stubs.py` | link-probes the Vulkan build | throwing stubs for unported kernels |
Rules:
@@ -116,13 +130,13 @@ Rules:
regenerated from the `.cu`.
2. To export a launch function to the engine, put
`/*[AutoHeaderGeneratorExport]*/` immediately above its definition and
rerun `generate_headers.py`.
rerun `tools/codegen/generate_headers.py`.
3. **A header can be fed by several `.cu` files**, listed explicitly in
`HEADER_SOURCES` in `generate_headers.py`. This is the sanctioned way to
`HEADER_SOURCES` in `tools/codegen/generate_headers.py`. This is the sanctioned way to
split a large `.cu`: name each part after **what it does**, not after the
header it feeds (e.g. `ImageWarp.cu`, `DepthGeometry.cu` → `PixelWise.cuh`),
and add it to the list. Both CMake and `setup.py` glob `*.cu`, so no build
file changes are needed. A listed file that doesn't exist is a hard error,
and add it to the list. Both build systems glob via `cmake/sources.txt`, so
no build file changes are needed. A listed file that doesn't exist is a hard error,
so a rename can't silently drop declarations.
4. The Python config dataclasses are the **single source of truth** for the
training config. Adding a field there makes it appear in the native CLI,
@@ -136,14 +150,14 @@ Rules:
Every kernel-level change needs **three** things, or the Vulkan build breaks
or silently diverges:
1. the CUDA implementation (`csrc/<Kernel>.cu` + `_kernel.cuh`),
1. the CUDA implementation (`src/kernels/<family>/<Kernel>.cu` + `_kernel.cuh`),
2. the Slang implementation (`cuda/slang/vulkan/*.slang`) and its launcher
(`csrc/backend/vulkan/kernels/*.cpp`),
3. a parity test in `csrc/backend/tests/` that runs both and compares.
(`src/backend/vulkan/kernels/*.cpp`),
3. a parity test in `src/backend/tests/` that runs both and compares.
If the Vulkan side isn't ready, `generate_vulkan_stubs.py` emits a throwing
If the Vulkan side isn't ready, `tools/codegen/generate_vulkan_stubs.py` emits a throwing
stub so the portable engine still links — that's a deliberate TODO marker, not
a finished state. Coverage is tracked in `csrc/backend/vulkan/README.md`,
a finished state. Coverage is tracked in `src/backend/vulkan/README.md`,
which is the authoritative and unusually detailed document on the Vulkan
design (device baseline, capability variants, memory model, atomics). Read it
before touching anything under `backend/vulkan/`.
@@ -154,10 +168,10 @@ before touching anything under `backend/vulkan/`.
# native parity tests (CUDA build)
bash build_develop.bash -DSSPLAT_BUILD_BACKEND_TESTS=ON && ./build/<test_name>
# the Vulkan build produces the same test binaries unconditionally
pytest tests/ # Python tests — currently unmaintained, may fail
pytest tests/python/ # Python tests — currently unmaintained, may fail
```
Each `csrc/backend/tests/*.cpp` becomes an executable of the same name.
Each `src/backend/tests/*.cpp` becomes an executable of the same name.
`backend/tests/engine/*` drive the real engine end to end. Details and the
CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
@@ -166,7 +180,7 @@ CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
- Files are named after **what they do**, not after the file they were split
out of, and not after the header they declare into.
- `<Name>_kernel.cuh` is reserved: it means "the device body that
`generate_kernel_instantiation.py` instantiates". Don't use the `_kernel`
`tools/codegen/generate_kernel_instantiation.py` instantiates". Don't use the `_kernel`
suffix for an ordinary shared-helper header — give it a descriptive name
(e.g. `BilinearSample.cuh`).
- A `<Family>Common.cuh` holds the shared preamble when a family spans several
@@ -179,7 +193,7 @@ CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
```
These are load-bearing for the split workflow — keep them meaningful.
- Slang device math is shared by both backends; don't fork it per backend.
- `csrc/*.cpp` (as opposed to `.cu`) means "portable, compiles for Vulkan
- A `.cpp` (as opposed to `.cu`) under `src/` means "portable, compiles for Vulkan
builds too". Keep the engine layer in `.cpp` and CUDA-free.
## Gotchas worth knowing before you hit them
+21 -716
View File
@@ -1,727 +1,32 @@
cmake_minimum_required(VERSION 3.18 FATAL_ERROR)
project(spirulae_splat LANGUAGES C CXX) # CUDA enabled below unless SSPLAT_BACKEND=vulkan
project(spirulae_splat LANGUAGES C CXX) # CUDA enabled by the cuda backend module
# This file is intended to be used with `build_develop.bash` for development purpose.
# ---------------------------------------------------------------------------
# Find Torch (optional)
# Development build (see build_develop.bash / build_develop.bat and
# docs/build.md). The pip build is separate -- setup.py -- and shares only the
# source list, cmake/sources.txt.
#
# Torch + Python are only needed for the Python extension module (ext.cpp).
# When either is missing -- or SSPLAT_NO_TORCH=ON -- the extension is skipped
# and only the standalone ssplat-train CLI is built (engine is torch-free).
# ---------------------------------------------------------------------------
option(SSPLAT_BUILD_CLI "Build the standalone ssplat-train CLI" OFF)
option(SSPLAT_BUILD_GUI "Build the native GUI app ssplat-gui (fetches GLFW + Dear ImGui)" OFF)
option(SSPLAT_NO_TORCH "Skip Torch/Python even if present; build only ssplat-train" OFF)
# Debug symbols / line info are OFF by default: they bloat the binaries
# massively (nvcc host -g, CUDA cubin lineinfo/source-in-ptx, and slangc -g2
# in the embedded SPIR-V). Turn on for profiling/debugging builds. The pip
# build (setup.py) is independent and unaffected (it has its own WITH_SYMBOLS
# and LINE_INFO env gates).
option(SSPLAT_DEBUG_SYMBOLS "Emit debug symbols / line info (host -g, CUDA cubin lineinfo, SPIR-V -g2)" OFF)
# ---------------------------------------------------------------------------
# Compute backend selection
# The logic lives in cmake/; this file is just the running order:
#
# cuda (default): the full build, exactly as before.
# vulkan: the portable engine layer (Engine*.cpp + host support) against the
# Vulkan compute runtime (csrc/backend/vulkan/, see its README.md), built
# WITHOUT the CUDA toolkit. Produces the backend tests, the ssplat-train CLI
# (always) and ssplat-gui (with SSPLAT_BUILD_GUI=ON); no Python extension.
# ---------------------------------------------------------------------------
set(SSPLAT_BACKEND "cuda" CACHE STRING "Compute backend: cuda | vulkan")
set_property(CACHE SSPLAT_BACKEND PROPERTY STRINGS cuda vulkan)
# SsplatOptions.cmake options, backend selection, tree paths
# SsplatSources.cmake the engine/kernel source list
# SsplatBackendCuda.cmake CUDA kernels + optional Torch extension } one
# SsplatBackendVulkan.cmake portable engine + Vulkan compute runtime } of
# SsplatApps.cmake ssplat-train / ssplat-mesh / ssplat-gui
#
# Each backend module leaves SSPLAT_WITH_TORCH and SSPLAT_APP_LIBS behind for
# the shared app targets.
list(APPEND CMAKE_MODULE_PATH ${CMAKE_CURRENT_SOURCE_DIR}/cmake)
include(SsplatOptions)
include(SsplatSources)
if(SSPLAT_BACKEND STREQUAL "vulkan")
message(STATUS "SSPLAT_BACKEND=vulkan: portable engine layer + "
"backend/vulkan runtime; kernel coverage is tracked in "
"csrc/backend/vulkan/README.md.")
# The Python extension is CUDA-only, so the CLI trainer is this build's
# primary artifact (same forcing as the CUDA no-torch build).
set(SSPLAT_BUILD_CLI ON)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(SSPLAT_CSRC ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc)
find_package(Vulkan REQUIRED)
# Portable engine layer (compile-checked; links against the Vulkan
# runtime once kernel launchers exist for the symbols it needs).
file(GLOB SSPLAT_PORTABLE_SOURCES CONFIGURE_DEPENDS ${SSPLAT_CSRC}/*.cpp)
list(REMOVE_ITEM SSPLAT_PORTABLE_SOURCES ${SSPLAT_CSRC}/ext.cpp)
add_library(csrc_portable OBJECT ${SSPLAT_PORTABLE_SOURCES})
target_compile_definitions(csrc_portable PUBLIC SSPLAT_BACKEND_VULKAN)
target_include_directories(csrc_portable PUBLIC ${SSPLAT_CSRC})
find_package(OpenMP)
if(OpenMP_CXX_FOUND)
target_link_libraries(csrc_portable PUBLIC OpenMP::OpenMP_CXX)
endif()
# ---- slangc: locate a matching compiler or fetch the pinned release ----
# SPIR-V blobs are never committed; they are compiled at build time (one
# slangc edge per blob, see slang/SpirvShaders.cmake) and embedded below.
# No Python is required for the Vulkan build.
set(SSPLAT_SLANG_VERSION "2026.12.0.1")
set(SSPLAT_SLANGC "" CACHE FILEPATH
"slangc to use (empty = find in PATH, fetch on miss/mismatch)")
function(_ssplat_slangc_version exe out_var)
set(${out_var} "" PARENT_SCOPE)
if(NOT exe OR NOT EXISTS ${exe})
return()
endif()
execute_process(COMMAND ${exe} -v
OUTPUT_VARIABLE _v ERROR_VARIABLE _v
OUTPUT_STRIP_TRAILING_WHITESPACE ERROR_STRIP_TRAILING_WHITESPACE
RESULT_VARIABLE _r)
if(_r EQUAL 0)
string(STRIP "${_v}" _v)
set(${out_var} "${_v}" PARENT_SCOPE)
endif()
endfunction()
set(_slangc "${SSPLAT_SLANGC}")
if(NOT _slangc)
find_program(_slangc_in_path slangc)
if(_slangc_in_path)
set(_slangc ${_slangc_in_path})
endif()
endif()
_ssplat_slangc_version("${_slangc}" _slangc_ver)
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
if(_slangc)
message(STATUS "slangc at ${_slangc} is version '${_slangc_ver}' "
"(need ${SSPLAT_SLANG_VERSION}); fetching the pinned release")
else()
message(STATUS "slangc not found; fetching the pinned release "
"${SSPLAT_SLANG_VERSION}")
endif()
if(WIN32)
set(_slang_plat "windows-x86_64")
set(_slang_ext "zip")
set(_slang_exe "bin/slangc.exe")
else()
if(CMAKE_HOST_SYSTEM_PROCESSOR MATCHES "aarch64|arm64")
set(_slang_arch "aarch64")
else()
set(_slang_arch "x86_64")
endif()
if(CMAKE_HOST_SYSTEM_NAME STREQUAL "Darwin")
set(_slang_plat "macos-${_slang_arch}")
else()
set(_slang_plat "linux-${_slang_arch}")
endif()
set(_slang_ext "tar.gz")
set(_slang_exe "bin/slangc")
endif()
set(_slang_name "slang-${SSPLAT_SLANG_VERSION}-${_slang_plat}")
set(_slang_dir ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name})
set(_slangc ${_slang_dir}/${_slang_exe})
if(NOT EXISTS ${_slangc})
set(_slang_url "https://github.com/shader-slang/slang/releases/download/v${SSPLAT_SLANG_VERSION}/${_slang_name}.${_slang_ext}")
set(_slang_archive ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name}.${_slang_ext})
message(STATUS "Downloading ${_slang_url}")
file(DOWNLOAD ${_slang_url} ${_slang_archive}
SHOW_PROGRESS STATUS _dl_status)
list(GET _dl_status 0 _dl_code)
if(NOT _dl_code EQUAL 0)
message(FATAL_ERROR "Failed to download slang release: "
"${_dl_status}. Install slang-${SSPLAT_SLANG_VERSION} "
"and pass -DSSPLAT_SLANGC=/path/to/slangc.")
endif()
file(ARCHIVE_EXTRACT INPUT ${_slang_archive}
DESTINATION ${_slang_dir})
file(REMOVE ${_slang_archive})
endif()
_ssplat_slangc_version("${_slangc}" _slangc_ver)
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
message(FATAL_ERROR "Fetched slangc reports version "
"'${_slangc_ver}', expected ${SSPLAT_SLANG_VERSION}")
endif()
endif()
message(STATUS "Using slangc ${_slangc} (${_slangc_ver})")
# ---- compile Slang -> SPIR-V (build tree) and embed ----
# Native (no Python): slang/spirv_tool.cpp is compiled by the module below,
# which enumerates one slangc custom command per blob (so -j bounds RAM) and
# embeds the result. SPIR-V debug info (slangc -g2) is opt-in via
# SSPLAT_DEBUG_SYMBOLS; it otherwise dominates the embedded blob size.
set(SSPLAT_CUDA_DIR ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda)
set(SSPLAT_SPIRV_DIR ${CMAKE_CURRENT_BINARY_DIR}/spirv)
set(SSPLAT_SPIRV_EMBED ${CMAKE_CURRENT_BINARY_DIR}/vk_shaders_embedded.cpp)
include(${SSPLAT_CUDA_DIR}/slang/SpirvShaders.cmake)
ssplat_setup_spirv(${SSPLAT_CUDA_DIR} ${_slangc} ${SSPLAT_SPIRV_DIR}
${SSPLAT_SPIRV_EMBED} ${SSPLAT_DEBUG_SYMBOLS})
# Vulkan backend runtime (device context, memory, streams, events) and
# kernel launchers.
file(GLOB SSPLAT_VULKAN_SOURCES CONFIGURE_DEPENDS ${SSPLAT_CSRC}/backend/vulkan/*.cpp
${SSPLAT_CSRC}/backend/vulkan/kernels/*.cpp)
list(APPEND SSPLAT_VULKAN_SOURCES ${SSPLAT_SPIRV_EMBED})
add_library(ssplat_backend_vulkan STATIC ${SSPLAT_VULKAN_SOURCES})
target_compile_definitions(ssplat_backend_vulkan PUBLIC SSPLAT_BACKEND_VULKAN)
target_include_directories(ssplat_backend_vulkan PUBLIC ${SSPLAT_CSRC})
target_link_libraries(ssplat_backend_vulkan PUBLIC Vulkan::Vulkan)
# Backend tests (run manually; see backend/vulkan/README.md).
file(GLOB SSPLAT_VULKAN_TESTS CONFIGURE_DEPENDS ${SSPLAT_CSRC}/backend/vulkan/tests/*.cpp)
foreach(test_src ${SSPLAT_VULKAN_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
endforeach()
# Cross-backend parity tools (same sources build in the CUDA branch).
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
${SSPLAT_CSRC}/backend/tests/*.cpp)
foreach(test_src ${SSPLAT_PARITY_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
endforeach()
# Engine-level parity tools: drive the real engine (forward_3dgs +
# background + blit) through the portable layer. The backend lib
# provides the ported kernels plus throwing stubs
# (kernels/TrainingStubs*.cpp) for everything still CUDA-only.
find_package(Threads REQUIRED)
file(GLOB SSPLAT_ENGINE_TESTS CONFIGURE_DEPENDS
${SSPLAT_CSRC}/backend/tests/engine/*.cpp)
foreach(test_src ${SSPLAT_ENGINE_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE csrc_portable
ssplat_backend_vulkan Threads::Threads)
endforeach()
# Torch/Python extension is CUDA-only; the app targets below build
# against the portable engine + Vulkan backend instead.
set(SSPLAT_WITH_TORCH OFF)
set(SSPLAT_APP_LIBS csrc_portable ssplat_backend_vulkan Threads::Threads)
include(SsplatBackendVulkan)
else()
enable_language(CUDA)
set(SSPLAT_WITH_TORCH OFF)
if(NOT SSPLAT_NO_TORCH)
find_package(Python3 QUIET COMPONENTS Interpreter)
if(Python3_Interpreter_FOUND)
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(torch.utils.cmake_prefix_path)"
OUTPUT_VARIABLE PYTORCH_CMAKE_PREFIX_PATH
RESULT_VARIABLE SSPLAT_TORCH_PROBE_RESULT
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "PyTorch CMake prefix path: ${PYTORCH_CMAKE_PREFIX_PATH}")
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import sysconfig; print(sysconfig.get_paths()['include'])"
OUTPUT_VARIABLE PYTHON_ADDITION_INCLUDE_PATH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "Python additional include path: ${PYTHON_ADDITION_INCLUDE_PATH}")
if(SSPLAT_TORCH_PROBE_RESULT EQUAL 0)
set(CMAKE_PREFIX_PATH ${PYTORCH_CMAKE_PREFIX_PATH})
find_package(Torch QUIET)
endif()
endif()
if(Torch_FOUND)
message(STATUS "Found Torch: ${TORCH_VERSION}")
find_library(TORCH_PYTHON_LIBRARY torch_python PATHS "${TORCH_INSTALL_PREFIX}/lib")
set(SSPLAT_WITH_TORCH ON)
endif()
include(SsplatBackendCuda)
endif()
if(NOT SSPLAT_WITH_TORCH)
message(STATUS "Torch/Python3 not available (or SSPLAT_NO_TORCH=ON) -- "
"skipping the Python extension; building the standalone ssplat-train CLI only.")
set(SSPLAT_BUILD_CLI ON)
endif()
find_package(CUDAToolkit REQUIRED)
# ---------------------------------------------------------------------------
# Find all sources
# ---------------------------------------------------------------------------
# CONFIGURE_DEPENDS: kernel families are routinely split into several TUs
# (Name.Part.cu -- see docs/codegen.md), and codegen adds/removes files under
# ins/. Without it a new source is silently not compiled and only surfaces as
# an undefined reference at link time.
file(GLOB SPLAT_SOURCES CONFIGURE_DEPENDS
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/*.cpp
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/*.cu
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/ins/*.cu
)
foreach(src ${SPLAT_SOURCES})
if(src MATCHES "hip")
list(REMOVE_ITEM SPLAT_SOURCES ${src})
endif()
endforeach()
if(NOT SSPLAT_WITH_TORCH)
# ext.cpp (the pybind11 module) is the only Torch-dependent source.
list(REMOVE_ITEM SPLAT_SOURCES
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/ext.cpp)
endif()
# message(STATUS "Extension sources:")
# foreach(src ${SPLAT_SOURCES})
# message(STATUS " ${src}")
# endforeach()
# ---------------------------------------------------------------------------
# Detect CUDA architectures
# ---------------------------------------------------------------------------
if(NOT DEFINED TORCH_CUDA_ARCH_LIST)
if(SSPLAT_WITH_TORCH)
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(' '.join(f'{a[0]}{a[1]}' for a in [torch.cuda.get_device_capability(i) for i in range(torch.cuda.device_count())]))"
OUTPUT_VARIABLE TORCH_CUDA_ARCH_LIST
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
endif()
if(NOT TORCH_CUDA_ARCH_LIST)
# No Torch (or Torch saw no device): ask the driver directly.
execute_process(
COMMAND nvidia-smi --query-gpu=compute_cap --format=csv,noheader
OUTPUT_VARIABLE SSPLAT_SMI_CAPS
RESULT_VARIABLE SSPLAT_SMI_RESULT
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
if(SSPLAT_SMI_RESULT EQUAL 0)
string(REPLACE "\r" "" SSPLAT_SMI_CAPS "${SSPLAT_SMI_CAPS}")
string(REPLACE "\n" " " TORCH_CUDA_ARCH_LIST "${SSPLAT_SMI_CAPS}")
endif()
endif()
if(NOT TORCH_CUDA_ARCH_LIST)
message(FATAL_ERROR "Could not detect a CUDA GPU. "
"Pass -DTORCH_CUDA_ARCH_LIST=\"8.6\" (or similar) to set the target architecture.")
endif()
endif()
# normalize: "8.6" / "8.6 7.5" / "86;75" all become a deduped list like "86;75"
string(REGEX REPLACE "\\." "" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
string(REPLACE " " ";" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
list(REMOVE_DUPLICATES TORCH_CUDA_ARCH_LIST)
message(STATUS "CUDA architecture(s): ${TORCH_CUDA_ARCH_LIST}")
set(CMAKE_CUDA_ARCHITECTURES "")
foreach(arch ${TORCH_CUDA_ARCH_LIST})
list(APPEND CMAKE_CUDA_ARCHITECTURES "${arch}")
endforeach()
# ---------------------------------------------------------------------------
# Compiler Flags
# ---------------------------------------------------------------------------
# CXX flags
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
if(MSVC)
set(SPLAT_CXX_FLAGS "/O2")
else()
set(SPLAT_CXX_FLAGS "-O3")
# set(SPLAT_CXX_FLAGS "-g")
list(APPEND SPLAT_CXX_FLAGS "-Wno-sign-compare")
endif()
if(WIN32)
add_compile_definitions(_USE_MATH_DEFINES NOMINMAX _CRT_SECURE_NO_WARNINGS)
endif()
if (MSVC)
add_compile_options(
$<$<COMPILE_LANGUAGE:CXX>:/Zc:preprocessor>
$<$<COMPILE_LANGUAGE:CUDA>:-Xcompiler=/Zc:preprocessor>
)
endif()
# Detect OpenMP
find_package(OpenMP)
if(OpenMP_CXX_FOUND AND NOT APPLE)
message(STATUS "Compiling with OpenMP")
list(APPEND SPLAT_CXX_FLAGS "-DAT_PARALLEL_OPENMP")
if(SSPLAT_WITH_TORCH)
# Compile flag only: the OpenMP runtime comes from torch's bundled
# libgomp. Linking the system one too would shadow it (CMake warns
# about the unsafe rpath). The no-torch build instead links the
# OpenMP::OpenMP_CXX target, whose flags are language-guarded so
# nvcc never sees them.
list(APPEND SPLAT_CXX_FLAGS ${OpenMP_CXX_FLAGS})
endif()
else()
message(STATUS "Compiling without OpenMP...")
endif()
# macOS ARM64 special case
if(APPLE AND CMAKE_SYSTEM_PROCESSOR MATCHES "arm64")
message(STATUS "Enabling macOS ARM64 flags")
add_compile_options("-arch" "arm64")
add_link_options("-arch" "arm64")
endif()
# nvcc flags (mirrors setup.py)
set(SPLAT_NVCC_FLAGS
"-O3" "--use_fast_math"
"--expt-relaxed-constexpr"
"-Xcudafe=--diag_suppress=20012"
"-Xcudafe=--diag_suppress=550"
"--threads" "0"
"-Xptxas" "--warn-on-double-precision-use"
)
# cubin/PTX line info (adds embedded source; needed for profiling but bloats
# the fatbin). Off unless SSPLAT_DEBUG_SYMBOLS.
if(SSPLAT_DEBUG_SYMBOLS)
list(APPEND SPLAT_NVCC_FLAGS "-lineinfo" "--generate-line-info" "--source-in-ptx")
endif()
# Add gencode flags
foreach(arch ${TORCH_CUDA_ARCH_LIST})
list(APPEND SPLAT_NVCC_FLAGS "-gencode" "arch=compute_${arch},code=sm_${arch}")
endforeach()
# Host side optimizations
if(NOT WIN32)
# list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-O3,-march=native")
# Host debug symbols (nvcc -Xcompiler=-g). Off unless SSPLAT_DEBUG_SYMBOLS.
if(SSPLAT_DEBUG_SYMBOLS)
list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-g")
endif()
endif()
# ---------------------------------------------------------------------------
# Define the extension module
# ---------------------------------------------------------------------------
if(SSPLAT_WITH_TORCH)
add_library(csrc SHARED ${SPLAT_SOURCES})
else()
# The engine API has no dllexport annotations (Linux .so exports
# everything by default); a static lib sidesteps that on Windows and
# makes ssplat-train a self-contained executable.
add_library(csrc STATIC ${SPLAT_SOURCES})
endif()
# Include directories
target_include_directories(csrc PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
# ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/glm
${CUDAToolkit_INCLUDE_DIRS}
)
# Definitions
if(WIN32)
target_compile_definitions(csrc PRIVATE spirulae_splat_EXPORTS)
endif()
if(SSPLAT_WITH_TORCH)
target_include_directories(csrc PRIVATE
${Python3_INCLUDE_DIRS}
${PYTHON_ADDITION_INCLUDE_PATH}
)
target_compile_definitions(csrc PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:TORCH_EXTENSION_NAME=csrc>
)
# Link Torch
target_link_libraries(csrc
${TORCH_LIBRARIES}
${TORCH_PYTHON_LIBRARY}
${Python3_LIBRARIES}
)
else()
# Torch normally drags these in for us
find_package(Threads REQUIRED)
target_link_libraries(csrc CUDA::cudart Threads::Threads)
if(OpenMP_CXX_FOUND AND NOT APPLE)
target_link_libraries(csrc OpenMP::OpenMP_CXX)
endif()
endif()
# Apply flags
target_compile_options(csrc PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
$<$<COMPILE_LANGUAGE:CUDA>:${SPLAT_NVCC_FLAGS}>
)
# Optional: strip symbols (setup.py uses '-s' if WITH_SYMBOLS==False)
if(DEFINED WITH_SYMBOLS AND NOT WITH_SYMBOLS)
if(NOT WIN32)
target_link_options(csrc PRIVATE "-s")
endif()
endif()
# Required by PyTorch C++ extensions
set_property(TARGET csrc PROPERTY CXX_STANDARD 17)
set_property(TARGET csrc PROPERTY CUDA_STANDARD 17)
set(SSPLAT_APP_LIBS csrc CUDA::cudart)
endif() # SSPLAT_BACKEND (vulkan | cuda core; app targets below are shared)
# ---------------------------------------------------------------------------
# Standalone CLI trainer -- no Python at runtime; links the csrc shared
# library for the engine. Off by default (unless Torch is unavailable, which
# forces it on) so the normal extension build is unaffected.
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_CLI OR SSPLAT_BUILD_GUI)
# Embed the (unchanged) web viewer client into the binary so the exe is
# self-contained. Regenerated at configure time when viewer.html changes.
set(SSPLAT_VIEWER_HTML ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/viewer/viewer.html)
set(SSPLAT_VIEWER_HTML_HDR ${CMAKE_BINARY_DIR}/app_generated/viewer_html.h)
file(READ ${SSPLAT_VIEWER_HTML} _ssplat_vh HEX)
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _ssplat_vh_bytes ${_ssplat_vh})
file(WRITE ${SSPLAT_VIEWER_HTML_HDR}
"#pragma once\n"
"// AUTO-GENERATED from spirulae_splat/viewer/viewer.html -- do not edit.\n"
"#include <cstddef>\n"
"inline const unsigned char kViewerHtml[] = {${_ssplat_vh_bytes}};\n"
"inline const size_t kViewerHtmlSize = sizeof(kViewerHtml);\n")
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS ${SSPLAT_VIEWER_HTML})
endif()
# Cross-backend parity tools (dump reference outputs; the Vulkan build's
# twin executable compares against them). See csrc/backend/vulkan/README.md.
# The Vulkan branch builds its comparing twins unconditionally above.
option(SSPLAT_BUILD_BACKEND_TESTS "Build backend parity test tools" OFF)
if(SSPLAT_BUILD_BACKEND_TESTS AND NOT SSPLAT_BACKEND STREQUAL "vulkan")
if(SSPLAT_WITH_TORCH)
# Same libpython + static-libstdc++ notes as ssplat-train below.
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
endif()
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
spirulae_splat/splat/cuda/csrc/backend/tests/*.cpp
spirulae_splat/splat/cuda/csrc/backend/tests/engine/*.cpp)
foreach(test_src ${SSPLAT_PARITY_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_include_directories(${test_name} PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
${CUDAToolkit_INCLUDE_DIRS})
target_link_libraries(${test_name} PRIVATE csrc CUDA::cudart)
if(SSPLAT_WITH_TORCH)
target_link_libraries(${test_name} PRIVATE Python3::Python)
if(NOT WIN32)
target_link_options(${test_name} PRIVATE
-static-libgcc -static-libstdc++)
endif()
endif()
endforeach()
endif()
if(SSPLAT_BUILD_CLI)
add_executable(ssplat-train
spirulae_splat/splat/cuda/csrc/app/main.cpp
spirulae_splat/splat/cuda/csrc/app/TrainerCore.cpp
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
spirulae_splat/splat/cuda/csrc/app/HttpServer.cpp
spirulae_splat/splat/cuda/csrc/app/RenderWorker.cpp
spirulae_splat/splat/cuda/csrc/app/Viewer.cpp
spirulae_splat/splat/cuda/csrc/external/miniz.c
)
target_include_directories(ssplat-train PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
${CMAKE_BINARY_DIR}
${CUDAToolkit_INCLUDE_DIRS}
)
target_link_libraries(ssplat-train PRIVATE ${SSPLAT_APP_LIBS})
if (WIN32)
target_link_libraries(ssplat-train PRIVATE ws2_32)
endif()
if(SSPLAT_WITH_TORCH)
# libpython is only needed because csrc's pybind/libtorch_python layer
# leaves CPython symbols undefined (a Python process normally provides
# them at import time). Drops out with the no-torch build.
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
target_link_libraries(ssplat-train PRIVATE Python3::Python)
endif()
target_compile_options(ssplat-train PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
)
# Statically link libstdc++ into the exe so its C++ runtime calls bind at
# link time. Without this, libtorch.so (pulled in via csrc) interposes
# its own bundled std::filesystem symbols at load time, and e.g.
# remove_all resolves to an ABI-incompatible copy that jumps to null.
if(SSPLAT_WITH_TORCH AND NOT WIN32)
target_link_options(ssplat-train PRIVATE -static-libgcc -static-libstdc++)
endif()
# libcsrc.so is renamed to spirulae_splat/csrc.so by build_develop.bash;
# keep the exe's lookup pointing at the build tree copy.
if(SSPLAT_WITH_TORCH)
set_property(TARGET ssplat-train PROPERTY BUILD_RPATH
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
else()
set_property(TARGET ssplat-train PROPERTY BUILD_RPATH "$ORIGIN")
endif()
set_property(TARGET ssplat-train PROPERTY CXX_STANDARD 17)
# ---- standalone mesh-extraction CLI (shares the dataset parsers) ----
# CUDA-only: the Vulkan backend stubs the meshing kernels
# (meshing::OccupancyEvaluator), so the tool would throw at startup.
if(NOT SSPLAT_BACKEND STREQUAL "vulkan")
add_executable(ssplat-mesh
spirulae_splat/splat/cuda/csrc/app/mesh_main.cpp
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
spirulae_splat/splat/cuda/csrc/external/miniz.c
)
target_include_directories(ssplat-mesh PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
${CUDAToolkit_INCLUDE_DIRS}
)
target_link_libraries(ssplat-mesh PRIVATE csrc CUDA::cudart)
if(SSPLAT_WITH_TORCH)
target_link_libraries(ssplat-mesh PRIVATE Python3::Python)
endif()
target_compile_options(ssplat-mesh PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
)
if(SSPLAT_WITH_TORCH AND NOT WIN32)
target_link_options(ssplat-mesh PRIVATE -static-libgcc -static-libstdc++)
endif()
if(SSPLAT_WITH_TORCH)
set_property(TARGET ssplat-mesh PROPERTY BUILD_RPATH
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
else()
set_property(TARGET ssplat-mesh PROPERTY BUILD_RPATH "$ORIGIN")
endif()
set_property(TARGET ssplat-mesh PROPERTY CXX_STANDARD 17)
endif() # NOT vulkan (ssplat-mesh)
endif()
# ---------------------------------------------------------------------------
# Native GUI (ssplat-gui) -- optional, off by default. Dear ImGui + GLFW +
# OpenGL 3 (imgui's embedded GL loader; no GLEW/GLAD). Dependencies are
# fetched at configure time via FetchContent and built statically, so the
# only system requirements are OpenGL dev headers (Linux: libgl1-mesa-dev +
# xorg-dev for GLFW; Windows/macOS: nothing extra).
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_GUI)
include(FetchContent)
set(GLFW_BUILD_DOCS OFF CACHE BOOL "" FORCE)
set(GLFW_BUILD_TESTS OFF CACHE BOOL "" FORCE)
set(GLFW_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
set(GLFW_INSTALL OFF CACHE BOOL "" FORCE)
FetchContent_Declare(glfw
GIT_REPOSITORY https://github.com/glfw/glfw.git
GIT_TAG 3.4
GIT_SHALLOW TRUE)
FetchContent_Declare(imgui
GIT_REPOSITORY https://github.com/ocornut/imgui.git
GIT_TAG v1.92.8
GIT_SHALLOW TRUE)
FetchContent_MakeAvailable(glfw imgui)
find_package(OpenGL REQUIRED)
# imgui ships no CMakeLists; build the core + glfw/gl3 backends +
# std::string InputText helper as a small static lib.
add_library(imgui_glfw STATIC
${imgui_SOURCE_DIR}/imgui.cpp
${imgui_SOURCE_DIR}/imgui_draw.cpp
${imgui_SOURCE_DIR}/imgui_tables.cpp
${imgui_SOURCE_DIR}/imgui_widgets.cpp
${imgui_SOURCE_DIR}/backends/imgui_impl_glfw.cpp
${imgui_SOURCE_DIR}/backends/imgui_impl_opengl3.cpp
${imgui_SOURCE_DIR}/misc/cpp/imgui_stdlib.cpp
)
target_include_directories(imgui_glfw PUBLIC
${imgui_SOURCE_DIR}
${imgui_SOURCE_DIR}/backends
${imgui_SOURCE_DIR}/misc/cpp
)
target_link_libraries(imgui_glfw PUBLIC glfw)
set_property(TARGET imgui_glfw PROPERTY CXX_STANDARD 17)
# Embed scripts/mask.py (AI masking helper, run via external Python) so
# the exe is self-contained. Same mechanism as the viewer.html embed.
set(SSPLAT_MASK_PY ${CMAKE_CURRENT_SOURCE_DIR}/scripts/mask.py)
set(SSPLAT_MASK_PY_HDR ${CMAKE_BINARY_DIR}/app_generated/mask_py.h)
file(READ ${SSPLAT_MASK_PY} _ssplat_mp HEX)
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _ssplat_mp_bytes ${_ssplat_mp})
file(WRITE ${SSPLAT_MASK_PY_HDR}
"#pragma once\n"
"// AUTO-GENERATED from scripts/mask.py -- do not edit.\n"
"#include <cstddef>\n"
"inline const unsigned char kMaskPy[] = {${_ssplat_mp_bytes}};\n"
"inline const size_t kMaskPySize = sizeof(kMaskPy);\n")
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS ${SSPLAT_MASK_PY})
add_executable(ssplat-gui
spirulae_splat/splat/cuda/csrc/app/gui/GuiMain.cpp
spirulae_splat/splat/cuda/csrc/app/gui/GuiApp.cpp
spirulae_splat/splat/cuda/csrc/app/gui/ConfigUI.cpp
spirulae_splat/splat/cuda/csrc/app/gui/FileDialog.cpp
spirulae_splat/splat/cuda/csrc/app/gui/Subprocess.cpp
spirulae_splat/splat/cuda/csrc/app/gui/ColmapRunner.cpp
spirulae_splat/splat/cuda/csrc/app/gui/FrameSelect.cpp
spirulae_splat/splat/cuda/csrc/app/gui/TrainRunner.cpp
spirulae_splat/splat/cuda/csrc/app/gui/GlLoader.cpp
spirulae_splat/splat/cuda/csrc/app/gui/NavCamera.cpp
spirulae_splat/splat/cuda/csrc/app/gui/PreviewRenderer.cpp
spirulae_splat/splat/cuda/csrc/app/gui/ViewportPanel.cpp
spirulae_splat/splat/cuda/csrc/app/TrainerCore.cpp
spirulae_splat/splat/cuda/csrc/app/RenderWorker.cpp
spirulae_splat/splat/cuda/csrc/app/HttpServer.cpp
spirulae_splat/splat/cuda/csrc/app/Viewer.cpp
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
spirulae_splat/splat/cuda/csrc/external/miniz.c
)
target_include_directories(ssplat-gui PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
${CMAKE_BINARY_DIR}
${CUDAToolkit_INCLUDE_DIRS}
)
target_link_libraries(ssplat-gui PRIVATE
${SSPLAT_APP_LIBS} imgui_glfw OpenGL::GL)
if(SSPLAT_WITH_TORCH)
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
target_link_libraries(ssplat-gui PRIVATE Python3::Python)
endif()
if (WIN32)
target_link_libraries(ssplat-gui PRIVATE ws2_32)
endif()
target_compile_options(ssplat-gui PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
)
# Same libtorch std::filesystem interposition workaround as ssplat-train.
if(SSPLAT_WITH_TORCH AND NOT WIN32)
target_link_options(ssplat-gui PRIVATE -static-libgcc -static-libstdc++)
endif()
if(SSPLAT_WITH_TORCH)
set_property(TARGET ssplat-gui PROPERTY BUILD_RPATH
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
else()
set_property(TARGET ssplat-gui PROPERTY BUILD_RPATH "$ORIGIN")
endif()
set_property(TARGET ssplat-gui PROPERTY CXX_STANDARD 17)
endif()
include(SsplatApps)
message(STATUS "spirulae_splat configured (backend: ${SSPLAT_BACKEND}).")
+11 -5
View File
@@ -7,14 +7,14 @@
# Regenerate headers/config. Skipped when python3 is unavailable -- the
# generated files are committed, so the build still works without it.
if command -v python3 >/dev/null 2>&1; then
python3 spirulae_splat/generate_headers.py
python3 spirulae_splat/generate_kernel_instantiation.py
python3 spirulae_splat/generate_cli_config.py
python3 tools/codegen/generate_headers.py
python3 tools/codegen/generate_kernel_instantiation.py
python3 tools/codegen/generate_cli_config.py
else
echo "python3 not found -- skipping codegen (using committed generated files)"
fi
cmake -G Ninja -B build "$@"
cmake -G Ninja -B build "$@" || exit $?
echo ""
@@ -29,7 +29,13 @@ echo "Available RAM : ${AVAILABLE_MB} MB"
echo "CPU cores : ${CPU_CORES}"
echo "Using jobs : ${JOBS}"
echo ""
cmake --build build --verbose -j"${JOBS}"
# Propagate the build's exit status: without this the script always exits 0
# (the trailing libcsrc.so move succeeds regardless) and a failed build reads
# as a green one.
if ! cmake --build build --verbose -j"${JOBS}"; then
echo "BUILD FAILED" >&2
exit 1
fi
# Torch build only: expose the extension as spirulae_splat/csrc.so, with a
# symlink back so ssplat-train's $ORIGIN lookup keeps working. A no-torch
+3 -3
View File
@@ -16,9 +16,9 @@ rem ---------------------------------------------------------------------------
rem Codegen (optional -- the generated files are committed)
rem ---------------------------------------------------------------------------
where python >nul 2>&1 || goto :skip_codegen
python spirulae_splat\generate_headers.py || goto :codegen_warn
python spirulae_splat\generate_kernel_instantiation.py || goto :codegen_warn
python spirulae_splat\generate_cli_config.py || goto :codegen_warn
python tools\codegen\generate_headers.py || goto :codegen_warn
python tools\codegen\generate_kernel_instantiation.py || goto :codegen_warn
python tools\codegen\generate_cli_config.py || goto :codegen_warn
goto :msvc_env
:codegen_warn
echo codegen failed -- continuing with committed generated files
+177
View File
@@ -0,0 +1,177 @@
# The native applications: ssplat-train (CLI), ssplat-mesh, ssplat-gui.
#
# Backend-agnostic: whichever backend module ran first left the engine to link
# in SSPLAT_APP_LIBS and set SSPLAT_WITH_TORCH. ssplat-mesh is the exception --
# the Vulkan backend stubs the meshing kernels, so it is CUDA-only.
include(SsplatEmbed)
# ---------------------------------------------------------------------------
# Source groups shared by several app targets
# ---------------------------------------------------------------------------
# Dataset parsing (COLMAP / Nerfstudio / Metashape) + miniz for the zipped
# Metashape .psx payloads.
set(SSPLAT_DATASET_SOURCES
${SSPLAT_SRC}/data/parsers/ColmapParser.cpp
${SSPLAT_SRC}/data/parsers/NerfstudioParser.cpp
${SSPLAT_SRC}/data/parsers/MetashapeParser.cpp
${SSPLAT_SRC}/data/parsers/DatasetCommon.cpp
${SSPLAT_SRC}/external/miniz.c
)
# The embedded web viewer: HTTP server + the render worker feeding it.
set(SSPLAT_VIEWER_SOURCES
${SSPLAT_SRC}/app/webviewer/HttpServer.cpp
${SSPLAT_SRC}/app/webviewer/RenderWorker.cpp
${SSPLAT_SRC}/app/webviewer/Viewer.cpp
)
# ---------------------------------------------------------------------------
# ssplat_configure_app(<target>)
#
# The settings every app target needs: include paths, the engine libraries,
# host flags, and the two link-time workarounds a Torch-enabled build requires.
# ---------------------------------------------------------------------------
function(ssplat_configure_app target)
target_include_directories(${target} PRIVATE
${SSPLAT_SRC}
${CMAKE_BINARY_DIR} # app_generated/{viewer_html,mask_py}.h
${CUDAToolkit_INCLUDE_DIRS}
)
target_link_libraries(${target} PRIVATE ${SSPLAT_APP_LIBS})
target_compile_options(${target} PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
)
set_property(TARGET ${target} PROPERTY CXX_STANDARD 17)
if(SSPLAT_WITH_TORCH)
# libpython is only needed because csrc's pybind/libtorch_python layer
# leaves CPython symbols undefined (a Python process normally provides
# them at import time). Drops out with the no-torch build.
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
target_link_libraries(${target} PRIVATE Python3::Python)
# Statically link libstdc++ into the exe so its C++ runtime calls bind
# at link time. Without this, libtorch.so (pulled in via csrc)
# interposes its own bundled std::filesystem symbols at load time, and
# e.g. remove_all resolves to an ABI-incompatible copy that jumps to
# null.
if(NOT WIN32)
target_link_options(${target} PRIVATE -static-libgcc -static-libstdc++)
endif()
# libcsrc.so is renamed to spirulae_splat/csrc.so by
# build_develop.bash; keep the exe's lookup pointing at the build tree
# copy (and at libtorch).
set_property(TARGET ${target} PROPERTY BUILD_RPATH
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
else()
set_property(TARGET ${target} PROPERTY BUILD_RPATH "$ORIGIN")
endif()
endfunction()
# ---------------------------------------------------------------------------
# Embedded web viewer client
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_CLI OR SSPLAT_BUILD_GUI)
# Embed the (unchanged) web viewer client into the binary so the exe is
# self-contained. Regenerated at configure time when viewer.html changes.
ssplat_embed_file(
${SSPLAT_ROOT}/spirulae_splat/viewer/viewer.html
${CMAKE_BINARY_DIR}/app_generated/viewer_html.h
ViewerHtml)
endif()
# ---------------------------------------------------------------------------
# ssplat-train -- standalone CLI trainer, no Python at runtime. Off by default
# (unless Torch is unavailable, or the backend is Vulkan, which force it on).
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_CLI)
add_executable(ssplat-train
${SSPLAT_SRC}/app/cli/main.cpp
${SSPLAT_SRC}/app/TrainerCore.cpp
${SSPLAT_DATASET_SOURCES}
${SSPLAT_VIEWER_SOURCES}
)
ssplat_configure_app(ssplat-train)
if(WIN32)
target_link_libraries(ssplat-train PRIVATE ws2_32)
endif()
# ---- standalone mesh-extraction CLI (shares the dataset parsers) ----
# CUDA-only: the Vulkan backend stubs the meshing kernels
# (meshing::OccupancyEvaluator), so the tool would throw at startup.
if(NOT SSPLAT_BACKEND STREQUAL "vulkan")
add_executable(ssplat-mesh
${SSPLAT_SRC}/app/cli/mesh_main.cpp
${SSPLAT_DATASET_SOURCES}
)
ssplat_configure_app(ssplat-mesh)
endif()
endif()
# ---------------------------------------------------------------------------
# ssplat-gui -- optional, off by default. Dear ImGui + GLFW + OpenGL 3
# (imgui's embedded GL loader; no GLEW/GLAD). Dependencies are fetched at
# configure time via FetchContent and built statically, so the only system
# requirements are OpenGL dev headers (Linux: libgl1-mesa-dev + xorg-dev for
# GLFW; Windows/macOS: nothing extra).
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_GUI)
include(FetchContent)
set(GLFW_BUILD_DOCS OFF CACHE BOOL "" FORCE)
set(GLFW_BUILD_TESTS OFF CACHE BOOL "" FORCE)
set(GLFW_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
set(GLFW_INSTALL OFF CACHE BOOL "" FORCE)
FetchContent_Declare(glfw
GIT_REPOSITORY https://github.com/glfw/glfw.git
GIT_TAG 3.4
GIT_SHALLOW TRUE)
FetchContent_Declare(imgui
GIT_REPOSITORY https://github.com/ocornut/imgui.git
GIT_TAG v1.92.8
GIT_SHALLOW TRUE)
FetchContent_MakeAvailable(glfw imgui)
find_package(OpenGL REQUIRED)
# imgui ships no CMakeLists; build the core + glfw/gl3 backends +
# std::string InputText helper as a small static lib.
add_library(imgui_glfw STATIC
${imgui_SOURCE_DIR}/imgui.cpp
${imgui_SOURCE_DIR}/imgui_draw.cpp
${imgui_SOURCE_DIR}/imgui_tables.cpp
${imgui_SOURCE_DIR}/imgui_widgets.cpp
${imgui_SOURCE_DIR}/backends/imgui_impl_glfw.cpp
${imgui_SOURCE_DIR}/backends/imgui_impl_opengl3.cpp
${imgui_SOURCE_DIR}/misc/cpp/imgui_stdlib.cpp
)
target_include_directories(imgui_glfw PUBLIC
${imgui_SOURCE_DIR}
${imgui_SOURCE_DIR}/backends
${imgui_SOURCE_DIR}/misc/cpp
)
target_link_libraries(imgui_glfw PUBLIC glfw)
set_property(TARGET imgui_glfw PROPERTY CXX_STANDARD 17)
# Embed scripts/mask.py (AI masking helper, run via external Python) so
# the exe is self-contained. Same mechanism as the viewer.html embed.
ssplat_embed_file(
${SSPLAT_ROOT}/scripts/mask.py
${CMAKE_BINARY_DIR}/app_generated/mask_py.h
MaskPy)
file(GLOB SSPLAT_GUI_SOURCES CONFIGURE_DEPENDS ${SSPLAT_SRC}/app/gui/*.cpp)
add_executable(ssplat-gui
${SSPLAT_GUI_SOURCES}
${SSPLAT_SRC}/app/TrainerCore.cpp
${SSPLAT_DATASET_SOURCES}
${SSPLAT_VIEWER_SOURCES}
)
ssplat_configure_app(ssplat-gui)
target_link_libraries(ssplat-gui PRIVATE imgui_glfw OpenGL::GL)
if(WIN32)
target_link_libraries(ssplat-gui PRIVATE ws2_32)
endif()
endif()
+275
View File
@@ -0,0 +1,275 @@
# SSPLAT_BACKEND=cuda (default): the CUDA kernels plus, when Torch and Python
# are available, the `csrc` Python extension module.
#
# Defines, for the shared app targets in SsplatApps.cmake:
# SSPLAT_WITH_TORCH ON when the Python extension is being built
# SSPLAT_APP_LIBS the engine library the apps link
# SPLAT_CXX_FLAGS host flags the app targets reuse
enable_language(CUDA)
# ---------------------------------------------------------------------------
# Find Torch (optional)
# ---------------------------------------------------------------------------
set(SSPLAT_WITH_TORCH OFF)
if(NOT SSPLAT_NO_TORCH)
find_package(Python3 QUIET COMPONENTS Interpreter)
if(Python3_Interpreter_FOUND)
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(torch.utils.cmake_prefix_path)"
OUTPUT_VARIABLE PYTORCH_CMAKE_PREFIX_PATH
RESULT_VARIABLE SSPLAT_TORCH_PROBE_RESULT
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "PyTorch CMake prefix path: ${PYTORCH_CMAKE_PREFIX_PATH}")
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import sysconfig; print(sysconfig.get_paths()['include'])"
OUTPUT_VARIABLE PYTHON_ADDITION_INCLUDE_PATH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "Python additional include path: ${PYTHON_ADDITION_INCLUDE_PATH}")
if(SSPLAT_TORCH_PROBE_RESULT EQUAL 0)
set(CMAKE_PREFIX_PATH ${PYTORCH_CMAKE_PREFIX_PATH})
find_package(Torch QUIET)
endif()
endif()
if(Torch_FOUND)
message(STATUS "Found Torch: ${TORCH_VERSION}")
find_library(TORCH_PYTHON_LIBRARY torch_python PATHS "${TORCH_INSTALL_PREFIX}/lib")
set(SSPLAT_WITH_TORCH ON)
endif()
endif()
if(NOT SSPLAT_WITH_TORCH)
message(STATUS "Torch/Python3 not available (or SSPLAT_NO_TORCH=ON) -- "
"skipping the Python extension; building the standalone ssplat-train CLI only.")
set(SSPLAT_BUILD_CLI ON)
endif()
find_package(CUDAToolkit REQUIRED)
# ---------------------------------------------------------------------------
# Sources
# ---------------------------------------------------------------------------
ssplat_collect_sources(SPLAT_SOURCES)
if(NOT SSPLAT_WITH_TORCH)
# bindings/ext.cpp (the pybind11 module) is the only Torch-dependent source.
list(REMOVE_ITEM SPLAT_SOURCES ${SSPLAT_SRC}/bindings/ext.cpp)
endif()
# ---------------------------------------------------------------------------
# Detect CUDA architectures
# ---------------------------------------------------------------------------
if(NOT DEFINED TORCH_CUDA_ARCH_LIST)
if(SSPLAT_WITH_TORCH)
execute_process(
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(' '.join(f'{a[0]}{a[1]}' for a in [torch.cuda.get_device_capability(i) for i in range(torch.cuda.device_count())]))"
OUTPUT_VARIABLE TORCH_CUDA_ARCH_LIST
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
endif()
if(NOT TORCH_CUDA_ARCH_LIST)
# No Torch (or Torch saw no device): ask the driver directly.
execute_process(
COMMAND nvidia-smi --query-gpu=compute_cap --format=csv,noheader
OUTPUT_VARIABLE SSPLAT_SMI_CAPS
RESULT_VARIABLE SSPLAT_SMI_RESULT
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
if(SSPLAT_SMI_RESULT EQUAL 0)
string(REPLACE "\r" "" SSPLAT_SMI_CAPS "${SSPLAT_SMI_CAPS}")
string(REPLACE "\n" " " TORCH_CUDA_ARCH_LIST "${SSPLAT_SMI_CAPS}")
endif()
endif()
if(NOT TORCH_CUDA_ARCH_LIST)
message(FATAL_ERROR "Could not detect a CUDA GPU. "
"Pass -DTORCH_CUDA_ARCH_LIST=\"8.6\" (or similar) to set the target architecture.")
endif()
endif()
# normalize: "8.6" / "8.6 7.5" / "86;75" all become a deduped list like "86;75"
string(REGEX REPLACE "\\." "" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
string(REPLACE " " ";" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
list(REMOVE_DUPLICATES TORCH_CUDA_ARCH_LIST)
message(STATUS "CUDA architecture(s): ${TORCH_CUDA_ARCH_LIST}")
set(CMAKE_CUDA_ARCHITECTURES "")
foreach(arch ${TORCH_CUDA_ARCH_LIST})
list(APPEND CMAKE_CUDA_ARCHITECTURES "${arch}")
endforeach()
# ---------------------------------------------------------------------------
# Compiler flags
#
# These intentionally mirror, but do not share, setup.py's flags: the pip build
# targets a released wheel (its own WITH_SYMBOLS / LINE_INFO env gates) while
# this one targets local development. Only the *source list* is shared, via
# cmake/sources.txt.
# ---------------------------------------------------------------------------
if(MSVC)
set(SPLAT_CXX_FLAGS "/O2")
else()
set(SPLAT_CXX_FLAGS "-O3")
list(APPEND SPLAT_CXX_FLAGS "-Wno-sign-compare")
endif()
if(MSVC)
add_compile_options(
$<$<COMPILE_LANGUAGE:CXX>:/Zc:preprocessor>
$<$<COMPILE_LANGUAGE:CUDA>:-Xcompiler=/Zc:preprocessor>
)
endif()
# Detect OpenMP
find_package(OpenMP)
if(OpenMP_CXX_FOUND AND NOT APPLE)
message(STATUS "Compiling with OpenMP")
list(APPEND SPLAT_CXX_FLAGS "-DAT_PARALLEL_OPENMP")
if(SSPLAT_WITH_TORCH)
# Compile flag only: the OpenMP runtime comes from torch's bundled
# libgomp. Linking the system one too would shadow it (CMake warns
# about the unsafe rpath). The no-torch build instead links the
# OpenMP::OpenMP_CXX target, whose flags are language-guarded so
# nvcc never sees them.
list(APPEND SPLAT_CXX_FLAGS ${OpenMP_CXX_FLAGS})
endif()
else()
message(STATUS "Compiling without OpenMP...")
endif()
# macOS ARM64 special case
if(APPLE AND CMAKE_SYSTEM_PROCESSOR MATCHES "arm64")
message(STATUS "Enabling macOS ARM64 flags")
add_compile_options("-arch" "arm64")
add_link_options("-arch" "arm64")
endif()
# nvcc flags
set(SPLAT_NVCC_FLAGS
"-O3" "--use_fast_math"
"--expt-relaxed-constexpr"
"-Xcudafe=--diag_suppress=20012"
"-Xcudafe=--diag_suppress=550"
"--threads" "0"
"-Xptxas" "--warn-on-double-precision-use"
)
# cubin/PTX line info (adds embedded source; needed for profiling but bloats
# the fatbin). Off unless SSPLAT_DEBUG_SYMBOLS.
if(SSPLAT_DEBUG_SYMBOLS)
list(APPEND SPLAT_NVCC_FLAGS "-lineinfo" "--generate-line-info" "--source-in-ptx")
endif()
# Add gencode flags
foreach(arch ${TORCH_CUDA_ARCH_LIST})
list(APPEND SPLAT_NVCC_FLAGS "-gencode" "arch=compute_${arch},code=sm_${arch}")
endforeach()
# Host side optimizations
if(NOT WIN32)
# list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-O3,-march=native")
# Host debug symbols (nvcc -Xcompiler=-g). Off unless SSPLAT_DEBUG_SYMBOLS.
if(SSPLAT_DEBUG_SYMBOLS)
list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-g")
endif()
endif()
# ---------------------------------------------------------------------------
# The engine library / Python extension module
# ---------------------------------------------------------------------------
if(SSPLAT_WITH_TORCH)
add_library(csrc SHARED ${SPLAT_SOURCES})
else()
# The engine API has no dllexport annotations (Linux .so exports
# everything by default); a static lib sidesteps that on Windows and
# makes ssplat-train a self-contained executable.
add_library(csrc STATIC ${SPLAT_SOURCES})
endif()
target_include_directories(csrc PRIVATE
${SSPLAT_SRC}
${CUDAToolkit_INCLUDE_DIRS}
)
if(WIN32)
target_compile_definitions(csrc PRIVATE spirulae_splat_EXPORTS)
endif()
if(SSPLAT_WITH_TORCH)
target_include_directories(csrc PRIVATE
${Python3_INCLUDE_DIRS}
${PYTHON_ADDITION_INCLUDE_PATH}
)
target_compile_definitions(csrc PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:TORCH_EXTENSION_NAME=csrc>
)
target_link_libraries(csrc
${TORCH_LIBRARIES}
${TORCH_PYTHON_LIBRARY}
${Python3_LIBRARIES}
)
else()
# Torch normally drags these in for us
find_package(Threads REQUIRED)
target_link_libraries(csrc CUDA::cudart Threads::Threads)
if(OpenMP_CXX_FOUND AND NOT APPLE)
target_link_libraries(csrc OpenMP::OpenMP_CXX)
endif()
endif()
target_compile_options(csrc PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
$<$<COMPILE_LANGUAGE:CUDA>:${SPLAT_NVCC_FLAGS}>
)
# Optional: strip symbols (setup.py uses '-s' if WITH_SYMBOLS==False)
if(DEFINED WITH_SYMBOLS AND NOT WITH_SYMBOLS)
if(NOT WIN32)
target_link_options(csrc PRIVATE "-s")
endif()
endif()
# Required by PyTorch C++ extensions
set_property(TARGET csrc PROPERTY CXX_STANDARD 17)
set_property(TARGET csrc PROPERTY CUDA_STANDARD 17)
set(SSPLAT_APP_LIBS csrc CUDA::cudart)
# ---------------------------------------------------------------------------
# Cross-backend parity tools -- CUDA side (dump reference outputs; the Vulkan
# build's twin executables compare against them). See backend/vulkan/README.md.
# The Vulkan branch builds its comparing twins unconditionally.
# ---------------------------------------------------------------------------
if(SSPLAT_BUILD_BACKEND_TESTS)
if(SSPLAT_WITH_TORCH)
# Same libpython + static-libstdc++ notes as ssplat-train.
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
endif()
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
${SSPLAT_SRC}/backend/tests/*.cpp
${SSPLAT_SRC}/backend/tests/engine/*.cpp)
foreach(test_src ${SSPLAT_PARITY_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_include_directories(${test_name} PRIVATE
${SSPLAT_SRC}
${CUDAToolkit_INCLUDE_DIRS})
target_link_libraries(${test_name} PRIVATE csrc CUDA::cudart)
if(SSPLAT_WITH_TORCH)
target_link_libraries(${test_name} PRIVATE Python3::Python)
if(NOT WIN32)
target_link_options(${test_name} PRIVATE
-static-libgcc -static-libstdc++)
endif()
endif()
endforeach()
endif()
+101
View File
@@ -0,0 +1,101 @@
# SSPLAT_BACKEND=vulkan: the portable engine layer built against the Vulkan
# compute runtime, without the CUDA toolkit. Kernel coverage is tracked in
# src/backend/vulkan/README.md.
#
# Defines, for the shared app targets in SsplatApps.cmake:
# SSPLAT_WITH_TORCH = OFF (the Python extension is CUDA-only)
# SSPLAT_APP_LIBS (portable engine + Vulkan runtime + threads)
message(STATUS "SSPLAT_BACKEND=vulkan: portable engine layer + "
"backend/vulkan runtime; kernel coverage is tracked in "
"src/backend/vulkan/README.md.")
# The Python extension is CUDA-only, so the CLI trainer is this build's
# primary artifact (same forcing as the CUDA no-torch build).
set(SSPLAT_BUILD_CLI ON)
find_package(Vulkan REQUIRED)
find_package(Threads REQUIRED)
# ---------------------------------------------------------------------------
# Portable engine layer
# ---------------------------------------------------------------------------
# Compile-checked here; links against the Vulkan runtime once kernel launchers
# exist for the symbols it needs.
ssplat_collect_portable_sources(SSPLAT_PORTABLE_SOURCES)
add_library(csrc_portable OBJECT ${SSPLAT_PORTABLE_SOURCES})
target_compile_definitions(csrc_portable PUBLIC SSPLAT_BACKEND_VULKAN)
target_include_directories(csrc_portable PUBLIC ${SSPLAT_SRC})
find_package(OpenMP)
if(OpenMP_CXX_FOUND)
target_link_libraries(csrc_portable PUBLIC OpenMP::OpenMP_CXX)
endif()
# ---------------------------------------------------------------------------
# Slang -> SPIR-V (build tree), then embed
# ---------------------------------------------------------------------------
# Native (no Python): backend/vulkan/shaders/spirv_tool.cpp is compiled by the
# module below,
# which enumerates one slangc custom command per blob (so -j bounds RAM) and
# embeds the result. SPIR-V debug info (slangc -g2) is opt-in via
# SSPLAT_DEBUG_SYMBOLS; it otherwise dominates the embedded blob size.
include(SsplatSlang)
ssplat_find_slangc(SSPLAT_SLANGC_EXE)
set(SSPLAT_SPIRV_DIR ${CMAKE_CURRENT_BINARY_DIR}/spirv)
set(SSPLAT_SPIRV_EMBED ${CMAKE_CURRENT_BINARY_DIR}/vk_shaders_embedded.cpp)
include(${SSPLAT_VK_SHADERS}/SpirvShaders.cmake)
ssplat_setup_spirv(${SSPLAT_SHADERS} ${SSPLAT_VK_SHADERS} ${SSPLAT_SLANGC_EXE}
${SSPLAT_SPIRV_DIR} ${SSPLAT_SPIRV_EMBED} ${SSPLAT_DEBUG_SYMBOLS})
# ---------------------------------------------------------------------------
# Vulkan backend runtime (device context, memory, streams, events) + kernels
# ---------------------------------------------------------------------------
file(GLOB SSPLAT_VULKAN_SOURCES CONFIGURE_DEPENDS
${SSPLAT_SRC}/backend/vulkan/*.cpp
${SSPLAT_SRC}/backend/vulkan/kernels/*.cpp)
list(APPEND SSPLAT_VULKAN_SOURCES ${SSPLAT_SPIRV_EMBED})
add_library(ssplat_backend_vulkan STATIC ${SSPLAT_VULKAN_SOURCES})
target_compile_definitions(ssplat_backend_vulkan PUBLIC SSPLAT_BACKEND_VULKAN)
target_include_directories(ssplat_backend_vulkan PUBLIC ${SSPLAT_SRC})
target_link_libraries(ssplat_backend_vulkan PUBLIC Vulkan::Vulkan)
# ---------------------------------------------------------------------------
# Tests (run manually; see backend/vulkan/README.md)
# ---------------------------------------------------------------------------
# Vulkan-only unit tests.
file(GLOB SSPLAT_VULKAN_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/vulkan/tests/*.cpp)
foreach(test_src ${SSPLAT_VULKAN_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
endforeach()
# Cross-backend parity tools (the same sources build in the CUDA branch, where
# they dump the reference outputs these compare against).
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/tests/*.cpp)
foreach(test_src ${SSPLAT_PARITY_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
endforeach()
# Engine-level parity tools: drive the real engine (forward_3dgs + background
# + blit) through the portable layer. The backend lib provides the ported
# kernels plus throwing stubs (kernels/TrainingStubs*.cpp) for everything still
# CUDA-only.
file(GLOB SSPLAT_ENGINE_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/tests/engine/*.cpp)
foreach(test_src ${SSPLAT_ENGINE_TESTS})
get_filename_component(test_name ${test_src} NAME_WE)
add_executable(${test_name} ${test_src})
target_link_libraries(${test_name} PRIVATE csrc_portable
ssplat_backend_vulkan Threads::Threads)
endforeach()
# Torch/Python extension is CUDA-only; the app targets build against the
# portable engine + Vulkan backend instead.
set(SSPLAT_WITH_TORCH OFF)
set(SSPLAT_APP_LIBS csrc_portable ssplat_backend_vulkan Threads::Threads)
+20
View File
@@ -0,0 +1,20 @@
# Embedding data files into the executables as byte arrays, so the apps are
# self-contained (no runtime lookup of viewer.html or scripts/mask.py).
# ssplat_embed_file(<input> <output_header> <symbol>)
#
# Writes a header defining `k<symbol>[]` / `k<symbol>Size` holding the bytes of
# <input>. Regenerated at configure time whenever <input> changes.
function(ssplat_embed_file input output_header symbol)
file(READ ${input} _hex HEX)
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _bytes ${_hex})
file(RELATIVE_PATH _rel ${SSPLAT_ROOT} ${input})
file(WRITE ${output_header}
"#pragma once\n"
"// AUTO-GENERATED from ${_rel} -- do not edit.\n"
"#include <cstddef>\n"
"inline const unsigned char k${symbol}[] = {${_bytes}};\n"
"inline const size_t k${symbol}Size = sizeof(k${symbol});\n")
set_property(DIRECTORY ${SSPLAT_ROOT} APPEND
PROPERTY CMAKE_CONFIGURE_DEPENDS ${input})
endfunction()
+61
View File
@@ -0,0 +1,61 @@
# Build options, backend selection, and the source-tree paths every other
# module builds on. Included first by the top-level CMakeLists.txt.
# ---------------------------------------------------------------------------
# Paths
# ---------------------------------------------------------------------------
# SSPLAT_ROOT is the repository root (the directory holding CMakeLists.txt).
# Everything else is expressed relative to it so the modules never care where
# they were included from.
# SSPLAT_SRC is also the include root: every local #include in the native
# tree is written relative to it (e.g. `#include "core/Common.cuh"`).
set(SSPLAT_ROOT ${CMAKE_CURRENT_SOURCE_DIR})
set(SSPLAT_SRC ${SSPLAT_ROOT}/src)
set(SSPLAT_SHADERS ${SSPLAT_SRC}/shaders) # shared Slang device math
set(SSPLAT_VK_SHADERS ${SSPLAT_SRC}/backend/vulkan/shaders) # Vulkan-only entry points
# ---------------------------------------------------------------------------
# Options
#
# Torch + Python are only needed for the Python extension module (ext.cpp).
# When either is missing -- or SSPLAT_NO_TORCH=ON -- the extension is skipped
# and only the standalone ssplat-train CLI is built (engine is torch-free).
# ---------------------------------------------------------------------------
option(SSPLAT_BUILD_CLI "Build the standalone ssplat-train CLI" OFF)
option(SSPLAT_BUILD_GUI "Build the native GUI app ssplat-gui (fetches GLFW + Dear ImGui)" OFF)
option(SSPLAT_NO_TORCH "Skip Torch/Python even if present; build only ssplat-train" OFF)
option(SSPLAT_BUILD_BACKEND_TESTS "Build backend parity test tools" OFF)
# Debug symbols / line info are OFF by default: they bloat the binaries
# massively (nvcc host -g, CUDA cubin lineinfo/source-in-ptx, and slangc -g2
# in the embedded SPIR-V). Turn on for profiling/debugging builds. The pip
# build (setup.py) is independent and unaffected (it has its own WITH_SYMBOLS
# and LINE_INFO env gates).
option(SSPLAT_DEBUG_SYMBOLS "Emit debug symbols / line info (host -g, CUDA cubin lineinfo, SPIR-V -g2)" OFF)
# ---------------------------------------------------------------------------
# Compute backend selection
#
# cuda (default): the full build -- CUDA kernels, optional Torch extension,
# and the app targets.
# vulkan: the portable engine layer (Engine*.cpp + host support) against the
# Vulkan compute runtime (src/backend/vulkan/, see its README.md), built
# WITHOUT the CUDA toolkit. Produces the backend tests, the ssplat-train CLI
# (always) and ssplat-gui (with SSPLAT_BUILD_GUI=ON); no Python extension.
# ---------------------------------------------------------------------------
set(SSPLAT_BACKEND "cuda" CACHE STRING "Compute backend: cuda | vulkan")
set_property(CACHE SSPLAT_BACKEND PROPERTY STRINGS cuda vulkan)
if(NOT SSPLAT_BACKEND STREQUAL "cuda" AND NOT SSPLAT_BACKEND STREQUAL "vulkan")
message(FATAL_ERROR "SSPLAT_BACKEND must be 'cuda' or 'vulkan', got '${SSPLAT_BACKEND}'")
endif()
# ---------------------------------------------------------------------------
# Shared C++ settings
# ---------------------------------------------------------------------------
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
if(WIN32)
add_compile_definitions(_USE_MATH_DEFINES NOMINMAX _CRT_SECURE_NO_WARNINGS)
endif()
+95
View File
@@ -0,0 +1,95 @@
# Locating slangc (the Slang -> SPIR-V compiler) for the Vulkan backend.
#
# SPIR-V blobs are never committed; they are compiled at build time (one slangc
# edge per blob, see src/backend/vulkan/shaders/SpirvShaders.cmake) and embedded into the
# binary. No Python is required for the Vulkan build.
set(SSPLAT_SLANG_VERSION "2026.12.0.1")
set(SSPLAT_SLANGC "" CACHE FILEPATH
"slangc to use (empty = find in PATH, fetch on miss/mismatch)")
function(_ssplat_slangc_version exe out_var)
set(${out_var} "" PARENT_SCOPE)
if(NOT exe OR NOT EXISTS ${exe})
return()
endif()
execute_process(COMMAND ${exe} -v
OUTPUT_VARIABLE _v ERROR_VARIABLE _v
OUTPUT_STRIP_TRAILING_WHITESPACE ERROR_STRIP_TRAILING_WHITESPACE
RESULT_VARIABLE _r)
if(_r EQUAL 0)
string(STRIP "${_v}" _v)
set(${out_var} "${_v}" PARENT_SCOPE)
endif()
endfunction()
# ssplat_find_slangc(<out_var>)
#
# Returns a slangc executable matching SSPLAT_SLANG_VERSION: the one named by
# -DSSPLAT_SLANGC, else one in PATH, else the pinned upstream release
# downloaded into the build tree.
function(ssplat_find_slangc out_var)
set(_slangc "${SSPLAT_SLANGC}")
if(NOT _slangc)
find_program(_slangc_in_path slangc)
if(_slangc_in_path)
set(_slangc ${_slangc_in_path})
endif()
endif()
_ssplat_slangc_version("${_slangc}" _slangc_ver)
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
if(_slangc)
message(STATUS "slangc at ${_slangc} is version '${_slangc_ver}' "
"(need ${SSPLAT_SLANG_VERSION}); fetching the pinned release")
else()
message(STATUS "slangc not found; fetching the pinned release "
"${SSPLAT_SLANG_VERSION}")
endif()
if(WIN32)
set(_slang_plat "windows-x86_64")
set(_slang_ext "zip")
set(_slang_exe "bin/slangc.exe")
else()
if(CMAKE_HOST_SYSTEM_PROCESSOR MATCHES "aarch64|arm64")
set(_slang_arch "aarch64")
else()
set(_slang_arch "x86_64")
endif()
if(CMAKE_HOST_SYSTEM_NAME STREQUAL "Darwin")
set(_slang_plat "macos-${_slang_arch}")
else()
set(_slang_plat "linux-${_slang_arch}")
endif()
set(_slang_ext "tar.gz")
set(_slang_exe "bin/slangc")
endif()
set(_slang_name "slang-${SSPLAT_SLANG_VERSION}-${_slang_plat}")
set(_slang_dir ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name})
set(_slangc ${_slang_dir}/${_slang_exe})
if(NOT EXISTS ${_slangc})
set(_slang_url "https://github.com/shader-slang/slang/releases/download/v${SSPLAT_SLANG_VERSION}/${_slang_name}.${_slang_ext}")
set(_slang_archive ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name}.${_slang_ext})
message(STATUS "Downloading ${_slang_url}")
file(DOWNLOAD ${_slang_url} ${_slang_archive}
SHOW_PROGRESS STATUS _dl_status)
list(GET _dl_status 0 _dl_code)
if(NOT _dl_code EQUAL 0)
message(FATAL_ERROR "Failed to download slang release: "
"${_dl_status}. Install slang-${SSPLAT_SLANG_VERSION} "
"and pass -DSSPLAT_SLANGC=/path/to/slangc.")
endif()
file(ARCHIVE_EXTRACT INPUT ${_slang_archive}
DESTINATION ${_slang_dir})
file(REMOVE ${_slang_archive})
endif()
_ssplat_slangc_version("${_slangc}" _slangc_ver)
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
message(FATAL_ERROR "Fetched slangc reports version "
"'${_slangc_ver}', expected ${SSPLAT_SLANG_VERSION}")
endif()
endif()
message(STATUS "Using slangc ${_slangc} (${_slangc_ver})")
set(${out_var} ${_slangc} PARENT_SCOPE)
endfunction()
+69
View File
@@ -0,0 +1,69 @@
# Collects the engine/kernel sources from cmake/sources.txt.
#
# The glob patterns live in a plain text file rather than here so that setup.py
# (which builds the same sources through torch.utils.cpp_extension, without
# CMake) reads the exact same list. See cmake/sources.txt.
# _ssplat_expand_section(<section> <out_var>)
#
# CONFIGURE_DEPENDS: kernel families are routinely split into several
# translation units, each named after what it does (see docs/codegen.md), and
# codegen adds/removes files under src/instantiations/. Without it a newly
# added source is silently not compiled and only surfaces as an undefined
# reference at link time.
function(_ssplat_expand_section section out_var)
set(spec ${SSPLAT_ROOT}/cmake/sources.txt)
file(STRINGS ${spec} lines)
set(patterns "")
set(in_section FALSE)
foreach(line ${lines})
string(STRIP "${line}" line)
if(line STREQUAL "" OR line MATCHES "^#")
continue()
endif()
if(line MATCHES "^\\[(.*)\\]$")
if(CMAKE_MATCH_1 STREQUAL section)
set(in_section TRUE)
else()
set(in_section FALSE)
endif()
continue()
endif()
if(in_section)
list(APPEND patterns ${SSPLAT_ROOT}/${line})
endif()
endforeach()
if(NOT patterns)
message(FATAL_ERROR "No [${section}] patterns found in ${spec}")
endif()
file(GLOB sources CONFIGURE_DEPENDS ${patterns})
# Drop ROCm sources (neither build compiles them).
list(FILTER sources EXCLUDE REGEX "hip")
if(NOT sources)
message(FATAL_ERROR "cmake/sources.txt [${section}] matched no files")
endif()
# Reconfigure when the spec itself changes.
set_property(DIRECTORY ${SSPLAT_ROOT} APPEND
PROPERTY CMAKE_CONFIGURE_DEPENDS ${spec})
set(${out_var} ${sources} PARENT_SCOPE)
endfunction()
# ssplat_collect_sources(<out_var>) -- everything the CUDA `csrc` lib compiles.
function(ssplat_collect_sources out_var)
_ssplat_expand_section(cuda result)
set(${out_var} ${result} PARENT_SCOPE)
endfunction()
# ssplat_collect_portable_sources(<out_var>) -- the torch-free, CUDA-free
# subset compiled as csrc_portable by the Vulkan build.
function(ssplat_collect_portable_sources out_var)
_ssplat_expand_section(portable result)
set(${out_var} ${result} PARENT_SCOPE)
endfunction()
+46
View File
@@ -0,0 +1,46 @@
# Engine/kernel source globs, relative to the repository root.
#
# SINGLE SOURCE OF TRUTH. Consumed by:
# - cmake/SsplatSources.cmake (the development build, both backends)
# - setup.py (the pip / torch.utils.cpp_extension build)
# Keep it a plain list of glob patterns, one per line: both readers parse it
# line-by-line, and setup.py must be able to do so without CMake installed.
# Blank lines and '#' comments are ignored. Globs are NOT recursive -- list each
# directory, so adding one is a deliberate act.
#
# Any path containing "hip" is dropped by both readers (ROCm sources are not
# part of either build).
#
# Not listed here, on purpose: src/app/** and src/backend/** (their targets own
# their sources -- see SsplatApps.cmake / the backend modules).
[cuda]
# Everything the CUDA `csrc` library compiles.
src/instantiations/*.cu
src/kernels/background/*.cu
src/kernels/bilagrid/*.cu
src/kernels/bilagrid/*.cpp
src/kernels/densify/*.cu
src/kernels/loss/*.cu
src/kernels/optim/*.cu
src/kernels/pixelwise/*.cu
src/kernels/ppisp/*.cu
src/kernels/projection/*.cu
src/kernels/raster/*.cu
src/kernels/tile/*.cu
src/kernels/visualize/*.cu
src/mesh/*.cu
src/mesh/*.cpp
src/engine/*.cpp
src/data/*.cpp
src/external/*.cpp
src/bindings/*.cpp
[portable]
# The torch-free subset: the same list minus the CUDA kernels and minus
# bindings/ext.cpp (Torch-only). Compiled as csrc_portable by the Vulkan build.
src/kernels/bilagrid/*.cpp
src/mesh/*.cpp
src/engine/*.cpp
src/data/*.cpp
src/external/*.cpp
+3 -3
View File
@@ -16,12 +16,12 @@ the detail.
Authoritative documents that live next to their code rather than here:
- `spirulae_splat/splat/cuda/csrc/backend/README.md` — the backend seam design.
- `spirulae_splat/splat/cuda/csrc/backend/vulkan/README.md` — the Vulkan
- `src/backend/README.md` — the backend seam design.
- `src/backend/vulkan/README.md` — the Vulkan
backend in depth: device baseline, capability variants, memory model,
atomics, Slang notes, kernel coverage. **The single most detailed document
in the repo**; read it before touching Vulkan code.
- `spirulae_splat/splat/cuda/csrc/app/README.md` — working notes on the
- `src/app/README.md` — working notes on the
standalone CLI and the native GUI.
- `viewer/README.md` — the standalone WebGL2/WASM viewer.
- `scripts/README.md` — dataset preprocessing tools.
+21 -21
View File
@@ -4,51 +4,51 @@
```
┌──────────────────────────────────────────────────────────┐
front ends │ ssplat-train (CLI) ssplat-gui (ImGui) spirulae-train │
│ csrc/app/main.cpp csrc/app/gui/ Python package │
front ends │ ssplat-train (CLI) ssplat-gui (ImGui) Python package │
│ src/app/cli/ src/app/gui/ spirulae_splat/│
└───────────┬────────────────┬──────────────────┬──────────┘
│ │ │
┌────────┴────────────────┴───┐ │ pybind11
│ TrainerCore (csrc/app/) │ │ (csrc/ext.cpp)
│ config → dataset → seeding │ │
│ TrainerCore (src/app/) │ │ src/bindings/
│ config → dataset → seeding │ │ ext.cpp
│ → step loop │ │
└──────────────┬──────────────┘ │
│ │
┌─────────────────┴─────────────────────────────┴──────────┐
engine │ Engine*.cpp — torch-free, CUDA-free, process-global │
│ state in EngineState.h; DataManager owns image I/O │
engine │ src/engine/*.cpp — torch-free, CUDA-free, process-global │
│ state in EngineState.h; src/data/ owns image I/O │
└──────────────────────────┬───────────────────────────────┘
│ launch declarations
┌──────────────────────────┴───────────────────────────────┐
backend seam │ csrc/backend/api/*.h + BackendRuntime.h │
backend seam │ src/backend/api/*.h + BackendRuntime.h │
└───────────┬──────────────────────────────┬───────────────┘
│ │
┌──────────────┴──────────┐ ┌────────────┴──────────────┐
backends │ CUDA: csrc/*.cu │ │ Vulkan: backend/vulkan/ │
│ + cuda/ins/*.cu │ │ SPIR-V from slang/vulkan │
└──────────────┬──────────┘ └────────────┬──────────────┘
┌──────────────┴───────────┐ ┌────────────┴──────────────┐
backends │ CUDA: src/kernels/**/*.cu│ │ Vulkan: src/backend/vulkan│
│ + src/instantiations/ │ │ launchers + shaders/ │
└──────────────┬───────────┘ └────────────┬──────────────┘
└──────────┬───────────────────┘
│
slang/*.slang — shared device math
backend/common/SortScan.h — CUB / onesweep radix
src/shaders/*.slang — shared device math
src/backend/common/SortScan.h — CUB / onesweep radix
```
The seam is the important part, and it is the best-documented part of the
repo: read `csrc/backend/README.md`, then `csrc/backend/vulkan/README.md`.
repo: read `src/backend/README.md`, then `src/backend/vulkan/README.md`.
## Where responsibilities live
| responsibility | code |
|---|---|
| training config (source of truth) | Python dataclasses in `spirulae_splat/modules/` → codegen → `app/generated/cli_config.h` |
| dataset parsing (native) | `csrc/app/{Colmap,Nerfstudio,Metashape}Parser.cpp`, `DatasetCommon.cpp`, `DatasetParser.h` |
| training config (source of truth) | Python dataclasses in `spirulae_splat/modules/` → codegen → `src/app/generated/cli_config.h` |
| dataset parsing (native) | `src/data/parsers/{Colmap,Nerfstudio,Metashape}Parser.cpp`, `DatasetCommon.cpp`, `DatasetParser.h` |
| dataset parsing (Python) | `spirulae_splat/modules/{dataparser,colmap_utils,metashape_utils,camera_utils}.py` — *duplicate; being retired* |
| image cache / prefetch / warp | `csrc/DataManager.cpp` |
| training session orchestration | `csrc/app/TrainerCore.cpp` and `spirulae_splat/modules/trainer.py` — *duplicate; being retired* |
| image cache / prefetch / warp | `src/data/DataManager.cpp` |
| training session orchestration | `src/app/TrainerCore.cpp` and `spirulae_splat/modules/trainer.py` — *duplicate; being retired* |
| the actual training step | `Engine*.cpp`, entered via `engine_train_step_managed` |
| kernels | `csrc/*.cu` (CUDA) + `slang/vulkan/*.slang` + `backend/vulkan/kernels/*.cpp` |
| kernels | `src/kernels/**/*.cu` (CUDA) + `src/backend/vulkan/shaders/*.slang` + `src/backend/vulkan/kernels/*.cpp` |
| web viewer client | `spirulae_splat/viewer/viewer.html` — single source; the C++ viewer embeds it at build time |
| web viewer server | `csrc/app/{Viewer,HttpServer,RenderWorker}.cpp` and `spirulae_splat/viewer/*.py` — *duplicate; being retired* |
| web viewer server | `src/app/webviewer/{Viewer,HttpServer,RenderWorker}.cpp` and `spirulae_splat/viewer/*.py` — *duplicate; being retired* |
| standalone WASM viewer | `viewer/` — independent, but compiles the C++ parsers in place |
Three subsystems are currently implemented twice, once in Python and once in
@@ -93,6 +93,6 @@ silently, so it is asserted only by parity tests.
3DGS, Mip (anti-aliased 3DGS), and 3DGUT are compile-time *types*, not runtime
branches: `Primitive.cuh`, `Primitive3DGS.cuh`, `Primitive3DGUT.cuh`,
`PrimitiveBase3DGS.cuh`. Camera model, SH degree and distortion are value
parameters. This is what makes `cuda/ins/` (111 generated instantiation TUs)
parameters. This is what makes `src/instantiations/` (111 generated instantiation TUs)
necessary on the CUDA side, and what maps to Slang specialization constants on
the Vulkan side.
+9 -9
View File
@@ -5,11 +5,11 @@ Two interchangeable compute backends behind one seam. **Both must work.**
The authoritative documents live next to the code and are more detailed than
this page — this is the orientation, they are the reference:
- **`spirulae_splat/splat/cuda/csrc/backend/README.md`** — the seam's design:
- **`src/backend/README.md`** — the seam's design:
what `api/` is, why the per-kernel `.cuh` headers are authoritative and the
module headers are forwarders, and how `BackendRuntime.h` abstracts
alloc/memcpy/memset/events.
- **`spirulae_splat/splat/cuda/csrc/backend/vulkan/README.md`** — the Vulkan
- **`src/backend/vulkan/README.md`** — the Vulkan
backend in depth: verified premises, device baseline and per-capability
pipeline variants, memory model, atomics strategy, sort/scan, Slang compiler
notes, and **the kernel coverage table**. Read it before touching anything
@@ -22,8 +22,8 @@ Engine*.cpp ──► backend/api/*.h (launch declarations)
backend/api/BackendRuntime.h (alloc / memcpy / memset / events)
│
┌───────────────┴───────────────┐
CUDA: csrc/*.cu Vulkan: backend/vulkan/
+ cuda/ins/*.cu SPIR-V pipelines from slang/vulkan/
CUDA: src/kernels/**/*.cu Vulkan: src/backend/vulkan/
+ src/instantiations/ runtime + kernels/ + shaders/
│ │
└────► backend/common/SortScan.h ◄────┘
CUB (CUDA) / onesweep radix + prefix sum (Vulkan)
@@ -39,10 +39,10 @@ under `-DSSPLAT_BACKEND_VULKAN` with no CUDA toolkit installed.
## Shared device math
`slang/*.slang` is compiled **twice**: to CUDA headers (`csrc/generated/*.cuh`)
`src/shaders/*.slang` is compiled **twice**: to CUDA headers (`src/generated/*.cuh`)
and to SPIR-V. The projection, SH, primitive and PPISP math therefore has one
source. `slang/vulkan/*.slang` adds the Vulkan-only compute entry points and
the compatibility layers (`int64_compat`, `int8_compat`, `atomic_float`).
source. `src/backend/vulkan/shaders/*.slang` adds the Vulkan-only compute entry
points and the compatibility layers (`int64_compat`, `int8_compat`, `atomic_float`).
Don't fork device math per backend. If a backend needs different code, it
belongs in a capability variant (below), not a copy.
@@ -74,12 +74,12 @@ two-field struct, because one Intel Windows driver segfaults in
Three deliverables, always:
1. **CUDA**: `csrc/<Kernel>.cu` (+ `_kernel.cuh` if the device body is shared
1. **CUDA**: `src/kernels/<family>/<Kernel>.cu` (+ `_kernel.cuh` if the device body is shared
between fwd/bwd or eval3d/non-eval3d), with
`/*[AutoHeaderGeneratorExport]*/` on the launcher. Rerun
`generate_headers.py`, and `generate_kernel_instantiation.py` if it is
templated over primitive / camera model / SH degree.
2. **Vulkan**: the Slang entry point in `slang/vulkan/*.slang` and its
2. **Vulkan**: the Slang entry point in `backend/vulkan/shaders/*.slang` and its
launcher in `backend/vulkan/kernels/*.cpp`.
3. **Parity test**: a tool in `backend/tests/` that builds under both backends
and does `dump` / `compare`. See [testing.md](testing.md).
+35 -11
View File
@@ -2,15 +2,38 @@
There are **two independent build systems**, by design and by history:
- `CMakeLists.txt` — every native artifact (engine, both backends, CLI, GUI,
meshing tool, parity tests, and optionally the Python extension).
- `CMakeLists.txt` + `cmake/` — every native artifact (engine, both backends,
CLI, GUI, meshing tool, parity tests, and optionally the Python extension).
- `setup.py` / `pyproject.toml` — the pip path, using torch's `CUDAExtension`.
It globs its own source list and carries its own flags. **It does not read
`CMakeLists.txt`**, so changes that touch compile flags or add source
directories must be applied to both.
It carries its own compile/link flags, deliberately: it targets a released
wheel (its own `WITH_SYMBOLS` / `LINE_INFO` env gates), CMake targets local
development. **Flag changes must be applied to both.**
They share exactly one thing: the source list, `cmake/sources.txt`. It is a
plain list of glob patterns relative to the repo root, read line-by-line by
`cmake/SsplatSources.cmake` and by `setup.py`'s `get_sources()` — plain text so
the pip build never needs CMake installed. Adding a source directory means
editing that file once.
For development, always use CMake via the dev scripts.
## CMake layout
`CMakeLists.txt` is just the running order; the logic lives in `cmake/`:
| file | what it does |
|---|---|
| `SsplatOptions.cmake` | options (`SSPLAT_BUILD_CLI/GUI`, `SSPLAT_NO_TORCH`, `SSPLAT_DEBUG_SYMBOLS`, …), backend selection, tree paths (`SSPLAT_ROOT`, `SSPLAT_CSRC`, …) |
| `SsplatSources.cmake` | `ssplat_collect_sources()` — expands `sources.txt` with `CONFIGURE_DEPENDS` |
| `SsplatBackendCuda.cmake` | Torch probe, CUDA arch detection, flags, the `csrc` library, CUDA-side parity tools |
| `SsplatBackendVulkan.cmake` | portable engine object lib, slangc + SPIR-V embed, `ssplat_backend_vulkan`, the Vulkan-side tests |
| `SsplatSlang.cmake` | `ssplat_find_slangc()` — pinned version, PATH lookup, fetch on miss |
| `SsplatEmbed.cmake` | `ssplat_embed_file()` — bake a file into a byte-array header |
| `SsplatApps.cmake` | `ssplat-train`, `ssplat-mesh`, `ssplat-gui` (backend-agnostic) |
Exactly one backend module runs. It leaves behind `SSPLAT_WITH_TORCH` and
`SSPLAT_APP_LIBS`, which is the whole contract `SsplatApps.cmake` depends on.
## Dev builds
```bash
@@ -72,9 +95,9 @@ exist *only* because of libtorch and are skipped in no-torch builds.
**Vulkan.** Needs the Vulkan SDK. No Python and no CUDA toolkit required.
Slang is pinned to a specific version (`SSPLAT_SLANG_VERSION` in
`CMakeLists.txt`); if the `slangc` on PATH doesn't match, CMake fetches the
`cmake/SsplatSlang.cmake`); if the `slangc` on PATH doesn't match, CMake fetches the
pinned release. SPIR-V blobs are **never committed** — they are compiled at
build time (one `slangc` edge per blob, see `slang/SpirvShaders.cmake`) and
build time (one `slangc` edge per blob, see `src/backend/vulkan/shaders/SpirvShaders.cmake`) and
embedded into the binary. On an offline machine, transfer a matching `slangc`
and point `-DSSPLAT_SLANGC=` at it.
@@ -89,7 +112,8 @@ suppress). Pass a trailing `-DSSPLAT_NO_TORCH=OFF` to try anyway.
## Build-time cost
The long poles are the biggest `.cu` translation units and the CUDA template
instantiation set in `cuda/ins/` (111 files). Splitting an oversized `.cu`
using the `Name.Part.cu` convention (see [codegen.md](codegen.md)) both
improves readability and parallelizes the build — no build-file edits needed,
since CMake and `setup.py` both glob.
instantiation set in `src/instantiations/` (111 files). Splitting an oversized `.cu`
into function-named parts (see [codegen.md](codegen.md)) both improves
readability and parallelizes the build — no build-file edits needed, since
both builds glob `cmake/sources.txt`. Only `HEADER_SOURCES` in
`generate_headers.py` has to learn the new file.
+23 -21
View File
@@ -8,19 +8,19 @@ available and fall back to the committed files when it isn't.
Generated locations:
```
spirulae_splat/splat/cuda/csrc/generated/ device-math headers from Slang
spirulae_splat/splat/cuda/csrc/app/generated/ cli_config.h, viewer_html.h
spirulae_splat/splat/cuda/csrc/backend/api/ backend module forwarders
spirulae_splat/splat/cuda/ins/ kernel instantiation TUs
src/generated/ device-math headers from Slang
src/app/generated/ cli_config.h, viewer_html.h
src/backend/api/ backend module forwarders
src/instantiations/ kernel instantiation TUs
```
All generators run from the **repo root**:
```bash
python3 spirulae_splat/generate_headers.py
python3 spirulae_splat/generate_kernel_instantiation.py
python3 spirulae_splat/generate_cli_config.py
python3 spirulae_splat/generate_backend_api.py
python3 tools/codegen/generate_headers.py
python3 tools/codegen/generate_kernel_instantiation.py
python3 tools/codegen/generate_cli_config.py
python3 tools/codegen/generate_backend_api.py
# generate_vulkan_stubs.py takes arguments; see below
```
@@ -28,14 +28,14 @@ python3 spirulae_splat/generate_backend_api.py
## `generate_headers.py` — `.cu` → `.cuh` declarations
Scans `csrc/*.cu` for functions preceded by the marker
Scans the `.cu` files listed in `HEADER_SOURCES` for functions preceded by the marker
```cpp
/*[AutoHeaderGeneratorExport]*/
void my_kernel_launcher(...) { ... }
```
and rewrites the declaration section of `csrc/<Name>.cuh`.
and rewrites the declaration section of the matching `<Name>.cuh`.
**Invariants:**
@@ -69,12 +69,13 @@ Because of invariant 3, splitting is nearly free:
# put the shared preamble in <Family>Common.cuh and shared device helpers in a
# descriptively named header (BilinearSample.cuh), then:
# - add the new files to HEADER_SOURCES['PixelWise']
python3 spirulae_splat/generate_headers.py # PixelWise.cuh's contents are unchanged
python3 tools/codegen/generate_headers.py # PixelWise.cuh's contents are unchanged
```
No build file changes: `CMakeLists.txt` globs `csrc/*.cu` with
`CONFIGURE_DEPENDS`, and `setup.py` globs at build time. No caller changes:
the header keeps its name and its declaration set.
No build file changes: both build systems glob the kernel directories through
`cmake/sources.txt` (CMake with `CONFIGURE_DEPENDS`, so a new file is picked up
without a manual reconfigure). No caller changes: the header keeps its name and
its declaration set.
Naming rules when splitting:
@@ -90,8 +91,9 @@ Naming rules when splitting:
## `generate_kernel_instantiation.py` — explicit template instantiations
Reads kernel declarations out of `csrc/*.cuh` and emits one small TU per
(primitive × camera model × SH degree × …) combination into `cuda/ins/`
Reads kernel declarations out of `src/kernels/**/*.cuh` and emits one small TU
per (primitive × camera model × SH degree × …) combination into
`src/instantiations/`
(currently 111 files). This keeps each nvcc edge small and parallelizable
rather than instantiating everything in one TU.
@@ -105,7 +107,7 @@ rather than instantiating everything in one TU.
`OptimizerConfig`, plus the tyro preset subclasses.
The script parses them with `ast` (no torch import — it runs on a fresh
checkout before `csrc.so` exists) and emits `csrc/app/generated/cli_config.h`:
checkout before `csrc.so` exists) and emits `src/app/generated/cli_config.h`:
- `struct SsplatConfig` — every field, flattened, with the Python defaults
baked in;
@@ -128,7 +130,7 @@ Docstrings on the dataclass fields become CLI help text and GUI tooltips.
## `generate_backend_api.py` — backend module headers
Emits `csrc/backend/api/*.h` (`Projection.h`, `Rasterization.h`,
Emits `src/backend/api/*.h` (`Projection.h`, `Rasterization.h`,
`PixelWise.h`, …) as **forwarders** that `#include` the relevant per-kernel
`.cuh` headers, grouped by subsystem.
@@ -141,7 +143,7 @@ and an individual `.cuh` without ODR violations.
## `generate_vulkan_stubs.py` — link-probe for unported kernels
```bash
python3 spirulae_splat/generate_vulkan_stubs.py <build-vulkan-dir> <output.cpp>
python3 tools/codegen/generate_vulkan_stubs.py <build-vulkan-dir> <output.cpp>
```
Link-probes `csrc_portable`'s objects against `libssplat_backend_vulkan.a`,
@@ -166,6 +168,6 @@ Rerun it whenever the engine gains a new kernel call, **or when a port lands**
Regenerated at configure time when the HTML changes. `Viewer.cpp` has a dev
override so HTML edits don't require a rebuild.
- Slang → SPIR-V blobs → `vk_shaders_embedded.cpp`
(`slang/SpirvShaders.cmake` + `slang/spirv_tool.cpp`). SPIR-V is **never
committed**; one `slangc` edge per blob, with capability variants
(`backend/vulkan/shaders/SpirvShaders.cmake` +
`backend/vulkan/shaders/spirv_tool.cpp`). SPIR-V is **never committed**; one `slangc` edge per blob, with capability variants
(`.atomicadd`, `.noint64`, `.int8`) compiled per entry point.
+2 -2
View File
@@ -3,7 +3,7 @@
## Supported layouts
Three formats, auto-detected in this order: **Nerfstudio**, then **COLMAP**,
then **Metashape**. Both the native parsers (`csrc/app/*Parser.cpp`) and the
then **Metashape**. Both the native parsers (`src/data/parsers/*Parser.cpp`) and the
Python parsers (`spirulae_splat/modules/dataparser.py`) probe in the same
order.
@@ -21,7 +21,7 @@ The COLMAP reconstruction directory is auto-detected over
Perspective (with full radial / tangential / thin-prism distortion),
equidistant and equisolid fisheye (including >180° FOV as produced by 360
cameras), and equirectangular/spherical. See `csrc/CameraModel.h` — it is
cameras), and equirectangular/spherical. See `src/core/CameraModel.h` — it is
plain C++17 with no CUDA dependency, which is why the WASM viewer can reuse
it.
+10 -10
View File
@@ -166,14 +166,14 @@ and the transform-backward math. The v2 revival is the moment to collapse this:
1. **Selector component** — `BilagridBwdSelector` (contextual Gaussian TS
ported from vksplat) + `SSPLAT_BILAGRID_BWD` override + unit test with
synthetic timings. No kernel changes. **DONE** — header-only
`csrc/BilagridBwdSelector.h`; host unit test
`csrc/backend/tests/bilagrid_selector_test.cpp` (auto-globbed; also builds
`kernels/bilagrid/BilagridBwdSelector.h`; host unit test
`src/backend/tests/bilagrid_selector_test.cpp` (auto-globbed; also builds
standalone) verifies forced-init coverage, contextual separation (two
resolutions → two optima), elasticity across a mid-run reward flip, and all
override forms. ~97% tail exploitation on the synthetic model.
2. **Async timing plumbing** — per-dispatch `backend::Event` timing, one-iter
readback, `(key, arm)` attribution at the backward hook. **DONE** —
`csrc/BilagridBwdTiming.h` (`BwdTimingRing`) wired into
`kernels/bilagrid/BilagridBwdTiming.h` (`BwdTimingRing`) wired into
`_engine_bilagrid_backward_hook` (harvest at hook entry; begin/end around the
RGB/depth/normal dispatches). Gated by `SSPLAT_BILAGRID_PROFILE=1` (logs
per-family GPU ms; zero overhead when unset). Validated on RTX 4080S / CUDA:
@@ -324,14 +324,14 @@ Two device-agnostic tunings support keeping v2 everywhere cheaply:
## 7. Files this touches
- `csrc/BilagridUniformSampleBwdV2_kernel.cuh` (+ per-family v2 headers) —
- `kernels/bilagrid/BilagridUniformSampleBwdV2_kernel.cuh` (+ per-family v2 headers) —
revive/generalize.
- `csrc/Bilagrid*UniformSample.cu`, `csrc/EngineBilagrid.cpp` (backward hook) —
- `kernels/bilagrid/Bilagrid*UniformSample.cu`, `engine/EngineBilagrid.cpp` (backward hook) —
launch + selection wiring.
- `csrc/BilagridConfig.cuh` — shared block/atomic helpers.
- `slang/vulkan/bilagrid_{common,affine,ppisp,loglinear,depth,normal}.slang` —
- `kernels/bilagrid/BilagridConfig.cuh` — shared block/atomic helpers.
- `backend/vulkan/shaders/bilagrid_{common,affine,ppisp,loglinear,depth,normal}.slang` —
add v2 entries.
- `csrc/backend/vulkan/kernels/Bilagrid.cpp` — v2 dispatch in `launch_family_*`.
- `csrc/backend/tests/bilagrid_parity.cpp` — per-arm parity sweep.
- **New:** `csrc/BilagridBwdSelector.h` (the selector).
- `src/backend/vulkan/kernels/Bilagrid.cpp` — v2 dispatch in `launch_family_*`.
- `src/backend/tests/bilagrid_parity.cpp` — per-arm parity sweep.
- **New:** `kernels/bilagrid/BilagridBwdSelector.h` (the selector).
```
+101 -2
View File
@@ -1,6 +1,6 @@
# Repository restructure proposal
Status: **proposal**, nothing here has been applied yet.
Status: phases 0-4 applied (see §8); §4 and phases 5-7 still proposal.
Goal: make the tree navigable for humans and agents, remove duplicated
Python/C++ subsystems, split oversized files, and land durable documentation
@@ -549,4 +549,103 @@ from drifting, and is the one I'd want a real regression run behind.
`architecture`, `build`, `backends`, `codegen`, `datasets`, `testing`,
`README` index, `notes/`. Written against the *current* layout, so they are
correct today and get path updates in phase 4.
- **Phase 2 — in progress.** File splits inside existing directories.
- **Phase 2 — done.** File splits inside existing directories: `PixelWise.cu`
(3668) → 7 function-named `.cu` + `PixelWiseCommon.cuh` + `BilinearSample.cuh`;
`Optimizer.cu` (2655) → 8 `.cu` + `OptimizerCommon.cuh`; `Densify.cu` (2471)
→ 5 `.cu` + 3 `.cuh`. `generate_headers.py` grew an explicit `HEADER_SOURCES`
map so a source file's name is free of the header it feeds. Declaration sets
of `PixelWise.cuh` / `Optimizer.cuh` / `Densify.cuh` verified unchanged
(43 / 26 / 20).
- **Phase 3 — done.** `CMakeLists.txt` (723 lines) → a 30-line running order
plus `cmake/`: `SsplatOptions`, `SsplatSources`, `SsplatBackendCuda`,
`SsplatBackendVulkan`, `SsplatSlang`, `SsplatEmbed`, `SsplatApps`. The
backend `if/else` is now `include(SsplatBackendCuda|Vulkan)`, each leaving
`SSPLAT_WITH_TORCH` + `SSPLAT_APP_LIBS` for the shared app targets; the
duplicated per-app boilerplate (libpython, `-static-libstdc++`, `BUILD_RPATH`,
include dirs) collapsed into `ssplat_configure_app()`, and the two
copy-pasted hex-embed blocks into `ssplat_embed_file()`.
Source lists are no longer duplicated: `cmake/sources.txt` holds the glob
patterns and is read by both `cmake/SsplatSources.cmake` and `setup.py`'s new
`get_sources()`. Plain text rather than CMake so the pip build needs no CMake
— full scikit-build-core delegation was rejected as too disruptive for users
still on the pip path (§7.1).
Verified by diffing the generated `build.ninja` old-vs-new for three configs
(torch CUDA / no-torch CUDA+GUI+tests / Vulkan): identical target sets,
**zero** flag, define, or link-library differences; the only deltas are
source ordering and one extra harmless `-I<builddir>` on `ssplat-mesh`.
Both backends rebuild clean and the parity suite is 17/17.
- **Phase 4 — done (native half).** The directory move. All native code is now
under a root `src/`, subdivided per §2: `core/ primitives/ kernels/{projection,
raster,tile,pixelwise,ppisp,bilagrid,optim,densify,loss,background,visualize}/
engine/ data/{,parsers/} mesh/ backend/ shaders/ app/{cli,gui,webviewer}/
bindings/ generated/ instantiations/ external/`. `spirulae_splat/splat/cuda/`
no longer holds native code. Also moved: the five codegen tools →
`tools/codegen/`, `tests/*.py` + the stray `csrc/tests/test_delaunay3d.py` →
`tests/python/`, `csrc/tests/delaunay3d_bench.cpp` → `tests/native/`.
**`src/` is now the include root** and every local include is path-qualified
against it (`#include "core/Common.cuh"`) — 716 quoted plus 145 angle-bracket
directives rewritten mechanically, none left relative. A file's include list
now names the subsystems it depends on.
Build/codegen plumbing updated in the same pass: `cmake/sources.txt` gained
`[cuda]` / `[portable]` sections (the Vulkan build's old flat `csrc/*.cpp`
glob no longer expresses "the torch-free subset"), `SpirvShaders.cmake` is
parameterised on the shaders dir, all five generators emit and read
src-relative paths, `.gitattributes`, `viewer/CMakeLists.txt` (which compiles
the parsers in place) and `build_develop.{bash,bat}` follow.
`generate_kernel_instantiation.py` now slugs its output filenames from
include *basenames*, so path-qualifying did not rename all 111 generated TUs.
Verified: every moved file's content diff is include lines only (checked
blob-by-blob against HEAD, 52 residual lines, all include directives or the
one comment naming a header); CUDA build, Vulkan build and Torch-extension
build all green; parity 17/17 plus the three Vulkan smoke tests; all five
generators re-run with zero drift.
**Deliberately deferred:** §2's reorganisation *inside* the Python package
(`spirulae_splat/{cli,config,training,data,metrics,utils}`) and the
`viewer/` → `web/` rename. The former changes public import paths
(`spirulae_splat.splat.cuda`, `spirulae_splat.modules.*`) that `scripts/` and
external users depend on, which per §7.1 should be paced; the latter belongs
with the viewer unification in phase 6. So `spirulae_splat/splat/` still
exists, holding only Python.
- **Phase 4b — backend symmetry, partial (done).** Reviewed whether the CUDA
and Vulkan halves should mirror each other. Conclusion: make each backend
*self-contained*, not make the two *look* like peers.
Done: `src/shaders/vulkan/` (39 files) → `src/backend/vulkan/shaders/`, and
`SpirvShaders.cmake` + `spirv_tool.cpp` with it. The Vulkan backend was the
only subsystem split across two top-level directories (host code under
`backend/`, device code under `shaders/`); it is now one directory —
runtime + kernels + shaders + its own SPIR-V build. `src/shaders/` now means
exactly what it says: the 8 Slang files compiled *twice*, to
`src/generated/*.cuh` and to SPIR-V. The Vulkan shaders reach the shared math
by src-relative path (`#include "shaders/densify.slang"`, with `-I<src>`),
matching the C++ convention and disambiguating `densify.slang`, which exists
in both directories.
Rejected: mirroring the CUDA corpus under `src/backend/cuda/kernels/`. The
`<Name>.cuh` declaration headers are the *backend-neutral contract* (Vulkan
TUs include them; they parse with no CUDA toolkit) and they are generated
from the `<Name>.cu` sitting next to them. A full mirror forces a choice
between putting the neutral contract inside `backend/cuda/` (so Vulkan
includes headers out of the CUDA backend — incoherent) and separating each
`.cuh` from the `.cu` it is generated from (breaking the repo's most
load-bearing codegen invariant). Symmetry is not worth either. The two sides
are also not peers in fact: CUDA is the reference implementation the contract
is *extracted from*, and a symmetric tree would assert a peerhood the codegen
does not have.
Two bugs surfaced and were fixed on the way:
- `backend/api/*.h` forwarders were emitting `#include "../../<Name>.cuh"`,
stale since the phase-4 move. Nothing includes them yet, so nothing caught
it. They are now src-relative and each one is compile-checked to parse
under `-DSSPLAT_BACKEND_VULKAN` with no CUDA toolkit.
- `build_develop.bash` always exited 0: the trailing `libcsrc.so` move
masked the build's status, so a failed build reported green. Now
propagated.
+4 -4
View File
@@ -4,7 +4,7 @@ Three suites, in descending order of how much you should trust them.
## 1. Native cross-backend parity tests (the important ones)
`csrc/backend/tests/*.cpp` — currently 17 tools covering projection (fwd, bwd,
`src/backend/tests/*.cpp` — currently 17 tools covering projection (fwd, bwd,
quant-grad), rasterization bwd, tile intersect, warp, FPBO, optimizer (general
+ geometry), densify, per-pixel train, PPISP, bilagrid, multi-scale loss, plus
`backend/tests/engine/` which drives the *real* engine end to end
@@ -68,7 +68,7 @@ If you script a Python training run that serves the viewer, pass
## 3. Python tests (`tests/`)
```bash
pytest tests/
pytest tests/python/
```
`test_per_pixel.py`, `test_ppisp.py`, `test_rasterization.py`,
@@ -80,8 +80,8 @@ have not been run recently. Treat a failure here as "needs investigation",
not as a release blocker — but do investigate, since they are the only
remaining check on the `_torch_impl.py` reference math.
There is also `csrc/tests/test_delaunay3d.py` (a Python test living in the C++
tree) and `csrc/tests/delaunay3d_bench.cpp`.
There is also `tests/python/test_delaunay3d.py` and the standalone benchmark
`tests/native/delaunay3d_bench.cpp` (not built by CMake; compile it directly).
## What to run before calling a change done
+38 -11
View File
@@ -38,6 +38,40 @@ def get_ext():
return CustomBuildExtention.with_options(no_python_abi_suffix=True, use_ninja=True)
def get_sources():
"""Engine/kernel sources, from the list CMake also builds from.
cmake/sources.txt is the single source of truth: [section] headers, then
one glob pattern per line relative to the repo root. This build wants the
[cuda] section (everything the csrc library compiles). Keeping it plain
text means the pip build does not need CMake installed. See
cmake/SsplatSources.cmake for the other consumer.
"""
root = Path(__file__).parent
spec = root / "cmake" / "sources.txt"
sources = []
in_section = False
for line in spec.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#"):
continue
if line.startswith("[") and line.endswith("]"):
in_section = line == "[cuda]"
continue
if in_section:
sources += glob.glob(str(root / line))
if not sources:
raise RuntimeError(f"{spec} [cuda] matched no files")
# Relative to the repo root, as setuptools expects (it mirrors the source
# path into build/temp.*/).
sources = [os.path.relpath(s, root) for s in sources]
# Drop ROCm sources (neither build compiles them).
return [s for s in sources if "hip" not in s]
def get_extensions():
import torch
from torch.utils.cpp_extension import CUDAExtension
@@ -50,13 +84,7 @@ def get_extensions():
else:
raise RuntimeError("CUDA is required for this extension.")
extensions_dir = Path("spirulae_splat/splat/cuda")
sources = (
glob.glob(str(extensions_dir / "ins" / "*.cu")) +
glob.glob(str(extensions_dir / "csrc" / "*.cu")) +
glob.glob(str(extensions_dir / "csrc" / "*.cpp"))
)
sources = [s for s in sources if "hip" not in s]
sources = get_sources()
undef_macros = []
define_macros = []
@@ -135,10 +163,9 @@ def get_extensions():
extension = CUDAExtension(
"spirulae_splat.csrc",
sources,
include_dirs=[
str((extensions_dir / "csrc").absolute()),
str((extensions_dir / "csrc" / "glm").absolute()),
],
# src/ is the include root: local includes are path-qualified
# relative to it (e.g. #include "core/Common.cuh").
include_dirs=[str((Path(__file__).parent / "src").absolute())],
define_macros=define_macros,
undef_macros=undef_macros,
extra_compile_args=extra_compile_args,
-149
View File
@@ -1,149 +0,0 @@
import re
import os
from pathlib import Path
def strip_if_zero_blocks(code):
lines = code.split('\n')
result = []
depth = 0 # 0 = active, >0 = inside a #if 0 dead block
for line in lines:
stripped = line.strip()
if depth == 0:
if re.match(r'#\s*if\s+0\b', stripped):
depth = 1
else:
result.append(line)
else:
if re.match(r'#\s*if', stripped):
depth += 1
elif re.match(r'#\s*endif', stripped):
depth -= 1
return '\n'.join(result)
def extract_function_declarations(code):
# Regex to match non-inline function declarations
function_decl_pattern = re.compile(r"""
\/\*\[AutoHeaderGeneratorExport\]\*\/\s*
(.*?\))\s*\{
""", re.MULTILINE | re.VERBOSE | re.DOTALL)
matches = function_decl_pattern.findall(code)
decls = []
for m in matches:
decls.append(m.strip()+';')
return decls
def write_if_changed(path, new_text):
old = None
if os.path.exists(path):
with open(path, "r") as f:
old = f.read()
if old != new_text:
with open(path, "w") as f:
f.write(new_text)
return True
return False
def generate_header(filename, sources):
path = "spirulae_splat/splat/cuda/csrc/"
# `sources` is listed in a fixed order, so the emitted declaration order is
# stable across machines and the committed header does not churn.
code = ""
for source_filename in sources:
full = path + source_filename
if not os.path.exists(full):
raise FileNotFoundError(
f"{filename}.cuh lists a missing source: {source_filename}. "
"Update HEADER_SOURCES in generate_headers.py."
)
code += open(full).read()
decls = extract_function_declarations(strip_if_zero_blocks(code))
splitter = "/* == AUTO HEADER GENERATOR - DO NOT EDIT THIS LINE OR ANYTHING BELOW THIS LINE == */\n"
include = open(path + f"{filename}.cuh").read()
include = include.split(splitter)[0].strip()
header = '\n\n\n'.join([include, splitter]+decls) + "\n"
cuh_path = path + f"{filename}.cuh"
if write_if_changed(cuh_path, header):
print("Generated", cuh_path)
return True
return False
# <header stem> -> the .cu translation units whose /*[AutoHeaderGeneratorExport]*/
# functions it declares, in emission order.
#
# A family may be spread over several TUs, each named after what it *does*
# rather than after the header it feeds (splitting oversized files is
# encouraged -- see docs/codegen.md). The mapping is explicit rather than
# inferred from filenames so that a source file's name is free, and so a
# typo'd or moved file fails loudly instead of silently dropping declarations.
HEADER_SOURCES = {
'IntersectTile': ['IntersectTile.cu'],
'BackgroundSphericalHarmonics': ['BackgroundSphericalHarmonics.cu'],
'PerSplatLoss': ['PerSplatLoss.cu'],
'PerPixelLoss': ['PerPixelLoss.cu'],
# Image-space per-pixel operations, split by function.
'PixelWise': [
'ImageConvert.cu', # uint8/uint16 -> float; rendered -> expected depth
'ImageColorOps.cu', # background blending, log map, overexposure reg
'DepthGeometry.cu', # depth -> points/normal, depth-normal loss,
# ray <-> linear depth
'ImageDistort.cu', # distort / undistort
'ImageWarp.cu', # wide <-> pinhole warps, incl. byte-fused
'GtDepthNormalWarp.cu', # GT depth/normal wide -> pinhole warps
'Ppisp.cu', # per-pixel image signal processing
],
'Projection': [], # types only; kernels live in the Fwd/Bwd headers
'ProjectionFwd': ['ProjectionFwd.cu'],
'ProjectionBwd': ['ProjectionBwd.cu'],
'ProjectionBwdQuantGrad': ['ProjectionBwdQuantGrad.cu'],
'ProjectionPackedFwd': ['ProjectionPackedFwd.cu'],
'RasterizationFwd': ['RasterizationFwd.cu'],
'RasterizationBwd': ['RasterizationBwd.cu'],
'RasterizationEval3DFwd': ['RasterizationEval3DFwd.cu'],
'RasterizationEval3DBwd': ['RasterizationEval3DBwd.cu'],
# Optimizer variants, split by function.
'Optimizer': [
'TensorSetZero.cu', # set_zero_tensor
'AdamOptim.cu', # Adam (incl. multi/stepped/8-bit/quat)
'NewtonOptim.cu',
'ScaleAgnosticMeanOptim.cu',
'FusedGeometryOptim.cu',
'FusedAppearanceOptim.cu',
'TrustRegion3DGS2Optim.cu', # 3DGS^2-TR family
'ColorOptim.cu', # linear-RGB / trust-region RGB + SH
],
'FusedProjectionBwdOptim': ['FusedProjectionBwdOptim.cu'],
# Densification, split by function.
'Densify': [
'DensifySampling.cu', # quantile/median, indexing, scatter, sampling
'DensifyScoring.cu', # covariance scale init, param update
'Relocation.cu',
'McmcRelocation.cu', # MCMC relocation + noise
'DensifySplitFilter.cu', # long-axis split, image edge filters
],
'BilagridUtils': ['BilagridUtils.cu'],
'Visualizer': ['Visualizer.cu'],
'FusedSSIM': ['FusedSSIM.cu'],
}
def generate_headers():
"""Regenerate headers only if needed."""
num_generated = 0
for name, sources in HEADER_SOURCES.items():
if generate_header(name, sources):
num_generated += 1
print(f"Generated {num_generated}/{len(HEADER_SOURCES)} new headers")
generate_headers()
+1 -1
View File
@@ -1,5 +1,5 @@
"""NumPy host-side mirrors of the block quantization codecs in
spirulae_splat/splat/cuda/csrc/Tensor.h, for adapting a checkpoint's quantized
src/core/Tensor.h, for adapting a checkpoint's quantized
buffers when resuming with a different layout (fewer Gaussians / different SH
degree). Decode -> transform (gather / band-resample) -> re-encode with fresh
per-block bounds. All host / CPU, vectorized (no per-element Python loops).
@@ -1,2 +0,0 @@
#define IS_EVAL3D 0
#include "RasterizationEval3DBwd_kernel.cuh"
@@ -1,2 +0,0 @@
#define IS_EVAL3D 0
#include "RasterizationEval3DFwd_kernel.cuh"
@@ -1,17 +0,0 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
//
// Splat projection forward/backward, packed variants, quantized-gradient backward, and the fused projection-backward + optimizer path.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
#include "../../ProjectionFwd.cuh"
#include "../../ProjectionBwd.cuh"
#include "../../ProjectionPackedFwd.cuh"
#include "../../ProjectionBwdQuantGrad.cuh"
#include "../../FusedProjectionBwdOptim.cuh"
@@ -1,17 +0,0 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
//
// Tile-based rasterization forward/backward (2D and eval-3D/3DGUT).
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
#include "../../RasterizationFwd.cuh"
#include "../../RasterizationBwd.cuh"
#include "../../RasterizationEval3DFwd.cuh"
#include "../../RasterizationEval3DBwd.cuh"
#include "../../RasterizationMomentsFwd.cuh"
@@ -1,6 +1,6 @@
# Standalone CLI Trainer (`ssplat-train`) + Native GUI (`ssplat-gui`)
Working notes for the no-Python trainer under `csrc/app/`. Written 2026-07-10,
Working notes for the no-Python trainer under `src/app/`. Written 2026-07-10,
GUI added 2026-07-12; covers code structure, verification status, and
remaining ports. Long-term context: Phase 1 (standalone C++ trainer) and
Phase 2 (Dear ImGui/GLFW GUI) of the GUI-app plan are done; Phase 3 is
@@ -37,7 +37,7 @@ system libs), `SSPLAT_BUILD_CLI` is forced ON, CUDA archs come from
`nvidia-smi --query-gpu=compute_cap` (override with `-DTORCH_CUDA_ARCH_LIST`),
and the libpython link + static-libstdc++/nftw interposition workarounds are
skipped (they exist only because of libtorch). Generated headers
(`csrc/generated/`, `app/generated/`, `cuda/ins/`) are committed, so a fresh
(`src/generated/`, `app/generated/`, `src/instantiations/`) are committed, so a fresh
checkout builds with no Python at all. With Torch present the extension build
is unchanged (shared `libcsrc`, same flags as before).
@@ -166,7 +166,7 @@ HTTP/viewer.
## Config codegen (source of truth = Python dataclasses)
`spirulae_splat/generate_cli_config.py` parses `TrainerConfig` (+ preset
`tools/codegen/generate_cli_config.py` parses `TrainerConfig` (+ preset
subclasses) in `modules/trainer.py` and the nested
`SpirulaeSplatDataParserConfig` / `SpirulaeSplatDataManagerConfig` /
`SpirulaeSplatModelConfig` / `OptimizerConfig`. Field docstrings become help
@@ -3,8 +3,8 @@
// in app/README.md) and is numerically verified against it -- keep edits
// in sync with the Python side.
#include "TrainerCore.h"
#include "Knn.h"
#include "app/TrainerCore.h"
#include "data/Knn.h"
#ifndef _WIN32
#include <ftw.h>
@@ -23,10 +23,10 @@
// setup_engine() calls engine_reset(), so a fresh session can follow a
// finished one in the same process (the GUI's "train again" path).
#include "../Engine.h"
#include "DatasetParser.h"
#include "RenderWorker.h"
#include "generated/cli_config.h"
#include "engine/Engine.h"
#include "data/DatasetParser.h"
#include "app/webviewer/RenderWorker.h"
#include "app/generated/cli_config.h"
#include <array>
#include <atomic>
@@ -15,8 +15,8 @@
// adds only CLI parsing, --help, stdout progress printing, and the web
// viewer wiring.
#include "TrainerCore.h"
#include "Viewer.h"
#include "app/TrainerCore.h"
#include "app/webviewer/Viewer.h"
#include <algorithm>
#include <chrono>
@@ -13,10 +13,10 @@
// the same C++ parsers the CLI trainer uses, and calls meshing::generate_mesh
// (Meshing.h) for the heavy lifting.
#include "../Meshing.h"
#include "../Camera.h"
#include "DatasetParser.h"
#include "Json.h"
#include "mesh/Meshing.h"
#include "core/Camera.h"
#include "data/DatasetParser.h"
#include "data/Json.h"
#include <algorithm>
#include <cmath>
@@ -1,6 +1,6 @@
#pragma once
// AUTO-GENERATED by spirulae_splat/generate_cli_config.py -- DO NOT EDIT.
// AUTO-GENERATED by tools/codegen/generate_cli_config.py -- DO NOT EDIT.
// Source of truth: the Python training config dataclasses
// (TrainerConfig + presets, SpirulaeSplatDataParserConfig,
// SpirulaeSplatDataManagerConfig, SpirulaeSplatModelConfig,
@@ -1,13 +1,13 @@
// ColmapRunner.cpp -- see ColmapRunner.h. CLI flags mirror
// scripts/run_colmap.bash (COLMAP >= 4.x; use_gpu-style flags are gone).
#include "ColmapRunner.h"
#include "FrameSelect.h"
#include "Subprocess.h"
#include "app/gui/ColmapRunner.h"
#include "app/gui/FrameSelect.h"
#include "app/gui/Subprocess.h"
#include "app_generated/mask_py.h" // kMaskPy[] (CMake-embedded scripts/mask.py)
#include "../../external/stb_image.h" // stbi_info (image size probe)
#include "external/stb_image.h" // stbi_info (image size probe)
#ifndef _WIN32
#include <ftw.h>
@@ -3,7 +3,7 @@
// special cases (the point is that new Python config fields show up here
// with zero GUI work).
#include "ConfigUI.h"
#include "app/gui/ConfigUI.h"
#include "imgui.h"
#include "imgui_stdlib.h"
@@ -13,7 +13,7 @@
// New config fields added on the Python side appear here automatically after
// re-running generate_cli_config.py -- no GUI change needed.
#include "../generated/cli_config.h"
#include "app/generated/cli_config.h"
namespace gui {
@@ -1,6 +1,6 @@
// FileDialog.cpp -- see FileDialog.h.
#include "FileDialog.h"
#include "app/gui/FileDialog.h"
#include "imgui.h"
#include "imgui_stdlib.h"
@@ -1,8 +1,8 @@
// FrameSelect.cpp -- see FrameSelect.h.
#include "FrameSelect.h"
#include "app/gui/FrameSelect.h"
#include "../../external/stb_image.h"
#include "external/stb_image.h"
#include <algorithm>
#include <chrono>
@@ -1,6 +1,6 @@
// GlLoader.cpp -- see GlLoader.h.
#include "GlLoader.h"
#include "app/gui/GlLoader.h"
#include <GLFW/glfw3.h>
@@ -1,9 +1,9 @@
// GuiApp.cpp -- see GuiApp.h.
#include "GuiApp.h"
#include "app/gui/GuiApp.h"
#include "../Json.h"
#include "Subprocess.h"
#include "data/Json.h"
#include "app/gui/Subprocess.h"
#include "imgui.h"
#include "imgui_stdlib.h"
@@ -4,13 +4,13 @@
// between the config editor, COLMAP runner, train runner, and the native
// viewport.
#include "../../backend/api/BackendRuntime.h"
#include "../generated/cli_config.h"
#include "ColmapRunner.h"
#include "ConfigUI.h"
#include "FileDialog.h"
#include "TrainRunner.h"
#include "ViewportPanel.h"
#include "backend/api/BackendRuntime.h"
#include "app/generated/cli_config.h"
#include "app/gui/ColmapRunner.h"
#include "app/gui/ConfigUI.h"
#include "app/gui/FileDialog.h"
#include "app/gui/TrainRunner.h"
#include "app/gui/ViewportPanel.h"
#include <deque>
#include <string>
@@ -2,7 +2,7 @@
// core context (lowest version that also satisfies macOS) + Dear ImGui
// bootstrap, then hands every frame to gui::GuiApp.
#include "GuiApp.h"
#include "app/gui/GuiApp.h"
#include "imgui.h"
#include "imgui_impl_glfw.h"
@@ -1,7 +1,7 @@
// NavCamera.cpp -- see NavCamera.h. Line-for-line port of viewer.html's
// v3 / quat helpers, `cam`, and `Nav`.
#include "NavCamera.h"
#include "app/gui/NavCamera.h"
#include <GLFW/glfw3.h>
@@ -1,9 +1,9 @@
// PreviewRenderer.cpp -- see PreviewRenderer.h.
#include "PreviewRenderer.h"
#include "app/gui/PreviewRenderer.h"
#include "GlLoader.h"
#include "../TrainerCore.h"
#include "app/gui/GlLoader.h"
#include "app/TrainerCore.h"
#include <algorithm>
#include <cmath>
@@ -1,6 +1,6 @@
// Subprocess.cpp -- see Subprocess.h.
#include "Subprocess.h"
#include "app/gui/Subprocess.h"
#include <algorithm>
#include <cstring>
@@ -1,6 +1,6 @@
// TrainRunner.cpp -- see TrainRunner.h.
#include "TrainRunner.h"
#include "app/gui/TrainRunner.h"
#include <algorithm>
@@ -13,8 +13,8 @@
// - engine_ready() flips true after engine setup; only then may a
// RenderWorker attach.
#include "../TrainerCore.h"
#include "../Viewer.h"
#include "app/TrainerCore.h"
#include "app/webviewer/Viewer.h"
#include <atomic>
#include <deque>
@@ -1,10 +1,10 @@
// ViewportPanel.cpp -- see ViewportPanel.h.
#include "ViewportPanel.h"
#include "app/gui/ViewportPanel.h"
#include "../TrainerCore.h"
#include "ConfigUI.h" // help_tooltip_on_hover
#include "GlLoader.h" // GL types + 1.1 entry points
#include "app/TrainerCore.h"
#include "app/gui/ConfigUI.h" // help_tooltip_on_hover
#include "app/gui/GlLoader.h" // GL types + 1.1 entry points
#include "imgui.h"
@@ -14,9 +14,9 @@
// GUI thread only. attach*() after the corresponding phase; detach() BEFORE
// the session it renders from is destroyed.
#include "../RenderWorker.h"
#include "NavCamera.h"
#include "PreviewRenderer.h"
#include "app/webviewer/RenderWorker.h"
#include "app/gui/NavCamera.h"
#include "app/gui/PreviewRenderer.h"
#include <cstdint>
#include <string>
@@ -1,6 +1,6 @@
// HttpServer.cpp -- see HttpServer.h.
#include "HttpServer.h"
#include "app/webviewer/HttpServer.h"
#include <cstdio>
#include <cstring>
@@ -2,10 +2,10 @@
// render math is a verified port of trainer.py _render/render + model.py
// get_outputs (viewer subset) + viewer/annotation.py.
#include "RenderWorker.h"
#include "app/webviewer/RenderWorker.h"
#include "../Engine.h" // engine render + viewer entry points
#include "../Camera.h" // camera_model_from_name
#include "engine/Engine.h" // engine render + viewer entry points
#include "core/Camera.h" // camera_model_from_name
#include <algorithm>
#include <atomic>
@@ -19,7 +19,7 @@
// callers either block (wait_result, HTTP thread) or poll (try_get_result,
// GUI frame loop).
#include "DatasetParser.h"
#include "data/DatasetParser.h"
#include <array>
#include <atomic>
@@ -1,16 +1,16 @@
// Viewer.cpp -- see Viewer.h. HTTP endpoint handlers are ports of
// viewer/http_server.py; the render path itself lives in RenderWorker.cpp.
#include "Viewer.h"
#include "HttpServer.h"
#include "app/webviewer/Viewer.h"
#include "app/webviewer/HttpServer.h"
// Implementation lives in csrc/stb_image_write_impl.cpp (shared with the
// Implementation lives in external/stb_image_write_impl.cpp (shared with the
// mesh texture export); declarations only here.
#include "../external/stb_image_write.h"
#include "external/stb_image_write.h"
#include "app_generated/viewer_html.h" // kViewerHtml[] (CMake-embedded)
#include "../Camera.h" // camera_model_from_name
#include "core/Camera.h" // camera_model_from_name
#include <algorithm>
#include <cstdio>
@@ -17,8 +17,8 @@
// annotation blit) lives in RenderWorker.h -- shared with the native GUI
// viewport. This file adds only the HTTP layer + JPEG encode.
#include "DatasetParser.h"
#include "RenderWorker.h"
#include "data/DatasetParser.h"
#include "app/webviewer/RenderWorker.h"
#include <memory>
#include <string>
@@ -12,7 +12,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
│
┌───────────────┴───────────────┐
backend/cuda/ backend/vulkan/
(existing csrc/*.cu, (SPIR-V pipeline dispatch over a
(existing src/kernels/**, (SPIR-V pipeline dispatch over a
unchanged, stays canonical) VulkanGSPipeline-style device layer)
│ │
└────────► backend/common/SortScan.h ◄────────┘
@@ -25,7 +25,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
- `Projection.h`, `Rasterization.h`, `TileIntersect.h`, `PixelWise.h`,
`Optimizer.h`, `Densify.h`, `Loss.h`, `Background.h`, `Bilagrid.h`,
`Visualizer.h` — **generated forwarders**
(`spirulae_splat/generate_backend_api.py`) that group the per-kernel
(`tools/codegen/generate_backend_api.py`) that group the per-kernel
launch headers per subsystem. The AUTHORITATIVE declarations and types
live in the per-kernel `.cuh` headers (declaration sections generated by
`generate_headers.py`), which are CUDA-include-free and parse under
@@ -88,7 +88,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
executable. It becomes a real build as backend/vulkan/ fills in (step 5).
5. **backend/vulkan/** implements `BackendRuntime.h`, `SortScan.h`, then the
api modules kernel-by-kernel (Slang `-target spirv` entry points reusing
the existing `slang/*.slang` device functions), in the phase order:
the existing `src/shaders/*.slang` device functions), in the phase order:
projection → SH/background → tile intersect + sort → rasterization fwd →
visualizer, then training (losses, backward, optimizer, densify).
@@ -170,5 +170,5 @@ float event_elapsed_ms(Event* start, Event* end); // requires enable_timing
} // namespace backend
#ifndef SSPLAT_BACKEND_VULKAN
#include "../cuda/BackendRuntimeCuda.h"
#include "backend/cuda/BackendRuntimeCuda.h"
#endif
@@ -1,13 +1,13 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Background spherical-harmonics render forward/backward.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../BackgroundSphericalHarmonics.cuh"
#include "kernels/background/BackgroundSphericalHarmonics.cuh"
@@ -1,14 +1,14 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Bilateral-grid utilities and the sample/TV-loss/fused-Adam launcher family.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../BilagridUtils.cuh"
#include "../../BilagridBindings.h"
#include "kernels/bilagrid/BilagridUtils.cuh"
#include "kernels/bilagrid/BilagridBindings.h"
@@ -1,13 +1,13 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Densification: MCMC and long-axis-split add/relocate, edge filters, scatter utilities.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../Densify.cuh"
#include "kernels/densify/Densify.cuh"
@@ -1,15 +1,15 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Fused SSIM, multi-scale per-pixel losses, per-splat losses.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../FusedSSIM.cuh"
#include "../../PerPixelLoss.cuh"
#include "../../PerSplatLoss.cuh"
#include "kernels/loss/FusedSSIM.cuh"
#include "kernels/loss/PerPixelLoss.cuh"
#include "kernels/loss/PerSplatLoss.cuh"
@@ -1,13 +1,13 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Fused Adam / AdaGrad / Newton optimizer steps, quantized and offloaded variants.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../Optimizer.cuh"
#include "kernels/optim/Optimizer.cuh"
@@ -1,13 +1,13 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Per-pixel image ops: background blend, depth/normal transforms, warps, color-space, PPISP.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../PixelWise.cuh"
#include "kernels/pixelwise/PixelWise.cuh"
+17
View File
@@ -0,0 +1,17 @@
#pragma once
// Backend API module — GENERATED forwarder by
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Splat projection forward/backward, packed variants, quantized-gradient backward, and the fused projection-backward + optimizer path.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "kernels/projection/ProjectionFwd.cuh"
#include "kernels/projection/ProjectionBwd.cuh"
#include "kernels/projection/ProjectionPackedFwd.cuh"
#include "kernels/projection/ProjectionBwdQuantGrad.cuh"
#include "kernels/optim/FusedProjectionBwdOptim.cuh"
+17
View File
@@ -0,0 +1,17 @@
#pragma once
// Backend API module — GENERATED forwarder by
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Tile-based rasterization forward/backward (2D and eval-3D/3DGUT).
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "kernels/raster/RasterizationFwd.cuh"
#include "kernels/raster/RasterizationBwd.cuh"
#include "kernels/raster/RasterizationEval3DFwd.cuh"
#include "kernels/raster/RasterizationEval3DBwd.cuh"
#include "kernels/raster/RasterizationMomentsFwd.cuh"
@@ -11,6 +11,6 @@
// BackendTypes.h / BackendRuntime.h, after which this include chain is valid
// in non-CUDA builds unchanged.
#include "BackendTypes.h"
#include "backend/api/BackendTypes.h"
#include "../../Primitive.cuh"
#include "primitives/Primitive.cuh"
@@ -1,14 +1,14 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Splat-to-tile intersection: key generation, radix sort, tile ranges.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../IntersectTile.cuh"
#include "../../SplatTileIntersector.cuh"
#include "kernels/tile/IntersectTile.cuh"
#include "kernels/tile/SplatTileIntersector.cuh"
@@ -1,13 +1,13 @@
#pragma once
// Backend API module — GENERATED forwarder by
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
//
// Interactive viewer render / blit entry points.
//
// Authoritative declarations live in the per-kernel headers included below
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
// Implementations: the .cu files next to each header (CUDA) and
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
#include "../../Visualizer.cuh"
#include "kernels/visualize/Visualizer.cuh"
@@ -21,7 +21,7 @@
//
// Zero overhead and zero behavior change when SSPLAT_PROFILE is unset.
#include "../api/BackendRuntime.h"
#include "backend/api/BackendRuntime.h"
#include <atomic>
#include <chrono>
@@ -34,7 +34,7 @@
#include <cstdint>
#include "../api/BackendRuntime.h"
#include "backend/api/BackendRuntime.h"
namespace backend {
@@ -10,7 +10,7 @@
#include <mutex>
#include <unordered_map>
#include "../common/Profiler.h"
#include "backend/common/Profiler.h"
namespace backend {
@@ -21,9 +21,9 @@
// contribution far below the loose tolerance. Quantized packed cells
// compare as integer codes with a +-1 allowance.
#include <BilagridBindings.h>
#include <EngineInternal.h>
#include <Tensor.h>
#include <kernels/bilagrid/BilagridBindings.h>
#include <engine/EngineInternal.h>
#include <core/Tensor.h>
#include <cmath>
#include <cstdio>
@@ -6,7 +6,7 @@
//
// Exit code 0 = all checks pass, 1 = a check failed.
#include "../../BilagridBwdSelector.h"
#include "kernels/bilagrid/BilagridBwdSelector.h"
#include <cstdio>
#include <random>
@@ -26,7 +26,7 @@
// kernel, dst = cur + i deterministic) covers many concurrent dst splats,
// including neighbor-word masked atomic stores in the quantized layouts.
#include <Densify.cuh>
#include <kernels/densify/Densify.cuh>
#include <cmath>
#include <cstdio>
@@ -12,10 +12,10 @@
// pixels). The blit's uint8 output allows |d| <= 2 per byte (fp rounding at
// the byte quantization boundary) with the same violation cap.
#include <Engine.h>
#include <EngineState.h>
#include <PixelWise.cuh>
#include <Visualizer.cuh>
#include <engine/Engine.h>
#include <engine/EngineState.h>
#include <kernels/pixelwise/PixelWise.cuh>
#include <kernels/visualize/Visualizer.cuh>
#include <cmath>
#include <cstdio>
@@ -20,9 +20,9 @@
// all optimizer steps (fed by atomically accumulated gradients).
// Densification is disabled so the splat count stays fixed.
#include <Engine.h>
#include <EngineState.h>
#include <PixelWise.cuh>
#include <engine/Engine.h>
#include <engine/EngineState.h>
#include <kernels/pixelwise/PixelWise.cuh>
#include <cmath>
#include <cstdio>
@@ -12,9 +12,9 @@
// N is a multiple of the 256-thread block (see projqgrad_parity note on the
// CUDA kernel's tail-thread reads).
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <FusedProjectionBwdOptim.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <kernels/optim/FusedProjectionBwdOptim.cuh>
#include <cmath>
#include <cstdio>
@@ -37,8 +37,8 @@
// readouts on both backends, so every config runs the identical computation
// twice and compares the second return value.
#include <PerPixelLoss.cuh>
#include <Tensor.h>
#include <kernels/loss/PerPixelLoss.cuh>
#include <core/Tensor.h>
#include <array>
#include <cmath>
@@ -16,8 +16,8 @@
// arithmetic is contraction-sensitive (can differ CUDA vs Vulkan), so for
// that mode only the deterministic count channel (accum.y) is compared.
#include <Optimizer.cuh>
#include <Densify.cuh>
#include <kernels/optim/Optimizer.cuh>
#include <kernels/densify/Densify.cuh>
#include <cmath>
#include <cstdio>
@@ -18,7 +18,7 @@
// quant combination is excluded, as in fpbo_parity (u = g1/(sqrt_g2 + eps)
// is unbounded for radii == 0 splats — true on CUDA too).
#include <Optimizer.cuh>
#include <kernels/optim/Optimizer.cuh>
#include <cmath>
#include <cstdio>
@@ -1,5 +1,5 @@
// Backend parity tool for the PPISP image-transform + regularization launch
// APIs (csrc/PixelWise.cuh PPISP subset), across all three param types
// APIs (kernels/pixelwise/PixelWise.cuh PPISP subset), across all three param types
// (original / rqs / no_crf) and cam_indices on/off. The SAME source builds
// under both backends:
//
@@ -8,7 +8,7 @@
//
// Ref format: [nf tight floats] [nl loose floats].
//
// Both backends run the same canonical slang/ppisp.slang math (CUDA via the
// Both backends run the same canonical shaders/ppisp.slang math (CUDA via the
// 2026.2.1 emission, Vulkan via SPIR-V), so per-pixel outputs and per-image
// gradients differ only by fast-math rounding -> tight channel. Atomically
// accumulated buffers (v_ppisp_params from the pixel backward, the summed
@@ -16,9 +16,9 @@
// regularization BACKWARD is fed synthetic host-side raw_losses/v_losses so
// its outputs stay deterministic (tight).
#include <PixelWise.cuh>
#include <EngineInternal.h>
#include <Tensor.h>
#include <kernels/pixelwise/PixelWise.cuh>
#include <engine/EngineInternal.h>
#include <core/Tensor.h>
#include <cmath>
#include <cstdio>
@@ -11,9 +11,9 @@
// whose order differs between backends, so the comparison is tolerance-based
// with a small violation-fraction cap.
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <ProjectionBwd.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <kernels/projection/ProjectionBwd.cuh>
#include <cmath>
#include <cstdio>
@@ -10,8 +10,8 @@
// based (fast-math exp/sqrt chains differ across compilers) with a small
// allowance for borderline-cull flips, which change entire rows.
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <cmath>
#include <cstdio>
@@ -13,9 +13,9 @@
// The register camera-loop accumulation is deterministic; codes compare with
// a +-1 quantum tolerance (codec rounding + last-ulp bound differences).
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <ProjectionBwdQuantGrad.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <kernels/projection/ProjectionBwdQuantGrad.cuh>
#include <cmath>
#include <cstdio>
@@ -11,9 +11,9 @@
// contributions in a different order than CUDA's block-atomic scheme, so
// comparison uses the usual tolerance + small violation-fraction cap.
#include <PixelWise.cuh>
#include <BackgroundSphericalHarmonics.cuh>
#include <EngineInternal.h>
#include <kernels/pixelwise/PixelWise.cuh>
#include <kernels/background/BackgroundSphericalHarmonics.cuh>
#include <engine/EngineInternal.h>
#include <cmath>
#include <cstdio>
@@ -11,13 +11,13 @@
// whose order differs between backends (and between runs on CUDA), so the
// comparison is tolerance-based with a small violation-fraction cap.
#include <IntersectTile.cuh>
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <RasterizationEval3DFwd.cuh>
#include <RasterizationFwd.cuh>
#include <RasterizationBwd.cuh>
#include <RasterizationEval3DBwd.cuh>
#include <kernels/tile/IntersectTile.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <kernels/raster/RasterizationEval3DFwd.cuh>
#include <kernels/raster/RasterizationFwd.cuh>
#include <kernels/raster/RasterizationBwd.cuh>
#include <kernels/raster/RasterizationEval3DBwd.cuh>
#include <cmath>
#include <cstdio>
@@ -13,12 +13,12 @@
// depth keys may sort differently across backends (last-ulp projection
// divergence), so a small violation fraction is allowed.
#include <BackgroundSphericalHarmonics.cuh>
#include <IntersectTile.cuh>
#include <ProjectionFwd.cuh>
#include <ProjectionPackedFwd.cuh>
#include <RasterizationEval3DFwd.cuh>
#include <RasterizationFwd.cuh>
#include <kernels/background/BackgroundSphericalHarmonics.cuh>
#include <kernels/tile/IntersectTile.cuh>
#include <kernels/projection/ProjectionFwd.cuh>
#include <kernels/projection/ProjectionPackedFwd.cuh>
#include <kernels/raster/RasterizationEval3DFwd.cuh>
#include <kernels/raster/RasterizationFwd.cuh>
#include <cmath>
#include <cstdio>
@@ -1,5 +1,5 @@
// Backend parity tool for the dataset GT warp + byte->float conversion
// launch APIs (csrc/PixelWise.cuh launch_warp_* + EngineInternal.h raw
// launch APIs (kernels/pixelwise/PixelWise.cuh launch_warp_* + EngineInternal.h raw
// converters). The SAME source builds under both backends:
//
// CUDA build: ./warp_parity dump ref.bin
@@ -24,9 +24,9 @@
// pixels), the equirectangular variants (x-wrap sampling), and the four
// raw byte->float converters with odd element counts (word-tail handling).
#include <PixelWise.cuh>
#include <EngineInternal.h>
#include <Tensor.h>
#include <kernels/pixelwise/PixelWise.cuh>
#include <engine/EngineInternal.h>
#include <core/Tensor.h>
#include <cmath>
#include <cstdio>
@@ -12,7 +12,7 @@ reproduced.
## Verified premises (probed 2026-07 with slang-2026.2.1, SDK 1.3.296;
## Vulkan backend now pins slang-2026.12.0.1 — see "Slang compiler notes")
- The existing `slang/*.slang` device functions compile to SPIR-V unchanged:
- The existing `src/shaders/*.slang` device functions compile to SPIR-V unchanged:
a compute entry point that `#include`s `primitive_3dgs.slang` and calls
`projection_3dgs<cam, ...>` validates for Vulkan 1.2 requiring only
`PhysicalStorageBufferAddresses` + `Int64` (+ `Shader`).
@@ -35,14 +35,14 @@ for A/B testing):
- `VK_EXT_shader_atomic_float` (`shaderBufferFloat32AtomicAdd`) — training
backward passes (~1,069 float atomicAdd sites). Every entry that calls
`atomic_add_f32` is compiled twice by `slang/spirv_tool.cpp`: the base
`atomic_add_f32` is compiled twice by `shaders/spirv_tool.cpp`: the base
blob uses a CAS-loop emulation (MoltenVK, llvmpipe), and a
`.atomicadd`-suffixed blob uses native `OpAtomicFAddEXT`; the pipeline
layer picks per device at module load (no in-shader branch). Unlike
VkSplat, both variants are always built.
- `shaderInt64` — 64-bit sort keys, morton codes, large-buffer indexing.
Entries in an int64_compat-including source get a `.noint64` variant
compiled with `-DSSPLAT_EMULATE_INT64` (`slang/vulkan/int64_compat.slang`),
compiled with `-DSSPLAT_EMULATE_INT64` (`backend/vulkan/shaders/int64_compat.slang`),
and the embed step keeps only those whose base blob actually declares
`OpCapability Int64` (the rest compile to a base-identical blob and are
dropped). With emulation: index arithmetic narrows to int32
@@ -61,7 +61,7 @@ for A/B testing):
safe shape (`acc_add`/`acc_sub` in sort_scan.slang carry explicitly).
- `shaderInt8` + `storageBuffer8BitAccess` — byte-buffer access. The
baseline reads/writes uint8 buffers as packed u32 words
(`slang/vulkan/int8_compat.slang`: `u8_load`/`s8_load`/`u8_store`; word
(`backend/vulkan/shaders/int8_compat.slang`: `u8_load`/`s8_load`/`u8_store`; word
reads past `n` stay in bounds because allocations are 16-byte rounded,
and byte stores are an InterlockedAnd+Or pair). Entries that call those
helpers get an `.int8` variant with native byte loads and plain byte
@@ -162,12 +162,12 @@ Variant axes follow the CUDA instantiation structure
## Shaders: build and shipping
Entry points live in `slang/vulkan/*.slang` and `#include` the existing
`slang/*.slang` module files (same trick as the CUDA side's
Entry points live in `backend/vulkan/shaders/*.slang` and `#include` the existing
`src/shaders/*.slang` module files (same trick as the CUDA side's
`#define CudaDeviceExport ForceInline` include). SPIR-V is NEVER committed
(unlike VkSplat): CMake compiles it at build time. The build is fully native
(no Python): `slang/spirv_tool.cpp` — a self-contained C++17 host tool
compiled once via `try_compile`, driven by `slang/SpirvShaders.cmake` —
(no Python): `shaders/spirv_tool.cpp` — a self-contained C++17 host tool
compiled once via `try_compile`, driven by `shaders/SpirvShaders.cmake` —
`discover`s every blob (scanning entries + feature variants over the
transitive #include closure), and CMake emits one slangc `-target spirv`
custom command per blob so the build's `-j` bounds how many run at once
@@ -270,7 +270,7 @@ Mesa ANV, llvmpipe — all tests + validation layers clean on all three):
Phase-2 state (projection forward): `kernels/ProjectionFwd.cpp` +
`kernels/ProjectionPackedFwd.cpp` implement the full `ProjectionFwd.cuh` /
`ProjectionPackedFwd.cuh` APIs (fp32 SH) for all three primitives via
`slang/vulkan/projection_fwd.slang` — the same `projection_3dgs<cam,...>` /
`backend/vulkan/shaders/projection_fwd.slang` — the same `projection_3dgs<cam,...>` /
`sh{N}_to_color` slang device functions the CUDA kernels use. Spec-constant
axes: camera model, SH degree, antialiased (MIP), eval3d (3DGUT; its
"conic" output lands in the screen scale slot, used by the eval3d
@@ -287,7 +287,7 @@ configs, ~7.2M floats, zero tolerance violations on all three local devices
(packed nnz and compaction order match exactly).
**SH value-quant (q8/q16)** is implemented WITHOUT any 8/16-bit device
features: `slang/vulkan/sh_quant.slang` mirrors harmonics.slang's
features: `backend/vulkan/shaders/sh_quant.slang` mirrors harmonics.slang's
`_sh_load_q{8,16}` + `sh_coeffs_to_color_q{8,16}` math but reads the packed
byte/short buffer as u32 words (the no-8-bit-types rule above), so quantized
checkpoints render on the plain baseline. `kShValueBits` (0/8/16) is a new
@@ -303,19 +303,19 @@ Phase-3 + raster-fwd state (full forward render): the whole
projection -> tile-intersect -> rasterize -> background chain now runs on
Vulkan.
- `kernels/BackgroundShFwd.cpp` + `slang/vulkan/background_sh.slang`:
- `kernels/BackgroundShFwd.cpp` + `backend/vulkan/shaders/background_sh.slang`:
per-pixel skybox via the same `generate_ray` / `transform_ray_d` /
`sh{N}_to_color_dir` exports the CUDA kernel wraps. SH degree is a spec
constant; camera model is a runtime int (as in CUDA). The backward throws
until the training phase (`TODO` in the file).
- `kernels/IntersectTile.cpp` + `slang/vulkan/intersect_tile.slang`:
- `kernels/IntersectTile.cpp` + `backend/vulkan/shaders/intersect_tile.slang`:
`do_intersect_tile_generic` = count -> `backend::inclusive_sum<int64>` ->
8-byte n_isects readback -> key write -> `backend::sort_pairs<int64,int32>`
(begin 0, end 32 + tile bits, exactly the CUB call) -> offset kernel.
Ellipse-vs-AABB mode is a runtime null-check on proj_conic (the CUDA
template bool only folds that check). `do_intersect_tile_post` has no
callers and is not ported.
- `kernels/RasterizeFwd.cpp` + `slang/vulkan/rasterize_fwd.slang`: closely
- `kernels/RasterizeFwd.cpp` + `backend/vulkan/shaders/rasterize_fwd.slang`: closely
follows RasterizationEval3DFwd_kernel.cuh (the CUDA source of both the 2D
and eval3D kernels). Two entries (2D shared by 3DGS/MIP; 3DGUT eval3d
reusing `evaluate_alpha_3dgs` / `evaluate_color_3dgs` /
@@ -330,7 +330,7 @@ Vulkan.
violations — near-tie depth keys sort differently when the projected
depths differ in the last ulp across compilers, reordering the blend for
a handful of pixels; the violation-fraction cap absorbs this.
- the SPIR-V entry scanner (now `slang/spirv_tool.cpp`) tolerates `//`
- the SPIR-V entry scanner (now `shaders/spirv_tool.cpp`) tolerates `//`
comments between the [shader]/[numthreads] attributes and `void`.
Phase-4 completion + engine-level end-to-end (2026-07): the REAL engine
@@ -338,12 +338,12 @@ render path (`forward_3dgs` -> background blend -> color space ->
depth-normal -> `engine_blit_view`) runs on Vulkan, verified against CUDA at
the engine level.
- `kernels/PixelWiseRender.cpp` + `slang/vulkan/pixel_wise_render.slang`:
- `kernels/PixelWiseRender.cpp` + `backend/vulkan/shaders/pixel_wise_render.slang`:
the render-path PixelWise subset — `blend_background_forward`,
`blend_background_noise_forward` (hash_uint3 ported bit-exact),
`rgb_to_srgb_forward`, `depth_to_normal_forward{,_tv}` (16x16 shared-mem
apron tile as a flat 256-thread workgroup per the X-dim subgroup rule).
- `kernels/Visualizer.cpp` + `slang/vulkan/visualizer.slang`: full
- `kernels/Visualizer.cpp` + `backend/vulkan/shaders/visualizer.slang`: full
visualizer — frustum geometry, LBVH build (morton keys over
`backend::sort_pairs<int64,int32>` bits 0..63, Karras internal nodes,
bottom-up AABBs), depth min/max, the annotate/colormap/overlay blit, the
@@ -388,12 +388,12 @@ the engine level.
(throwing, debug-sync hook), or_fallback, resolve_sh_quant. New kernel TUs
should not hand-roll these.
- **Optimizer port (phase 5, first slice)**: `kernels/Optimizer.cpp` +
`slang/vulkan/optimizer.slang` implement fused_adam_step (fp32),
`backend/vulkan/shaders/optimizer.slang` implement fused_adam_step (fp32),
fused_adam_step_quantized (4/8-bit state), fused_adam_step_quantized_value
(state x value x optional gradq int8 grad), fused_adagrad_step,
float_add_into, increment_int32_inplace; `kernels/Densify.cpp` +
`densify.slang` implement densify_update_weight. The quantized-state
codecs live in `slang/vulkan/optim_quant.slang` (QuantizedAdamState /
codecs live in `backend/vulkan/shaders/optim_quant.slang` (QuantizedAdamState /
QuantizedTensor / gradq::Codec<8> mirrors): packed cells are read as u32
words and written by cooperative word assembly in groupshared (a
256-cell quant block is always word-aligned; CUDA leaves final-word tail
@@ -500,12 +500,12 @@ the engine level.
- **Densify/MCMC + remaining optimizer variants (phase 5, fifth slice)**:
`kernels/Densify.cpp` implements the full densification API
(densify_clip_scale, MCMC/revised noise, long-axis-split relocate/add,
MCMC relocate/add) over `slang/vulkan/densify.slang`, which includes the
canonical `slang/densify.slang` device code (long_axis_split_3dgs, the
MCMC relocate/add) over `backend/vulkan/shaders/densify.slang`, which includes the
canonical `shaders/densify.slang` device code (long_axis_split_3dgs, the
randn3 noise kernels) and mirrors the McmcRelocation.cu-only mcmc_relocation
math; `kernels/Optimizer.cpp` gains the adamtr trust-region color
variants (`slang/vulkan/optim_color.slang`) and the fused 3DGS geometry
step (`slang/vulkan/optim_geometry.slang`, sharing the FPBO helpers
variants (`backend/vulkan/shaders/optim_color.slang`) and the fused 3DGS geometry
step (`backend/vulkan/shaders/optim_geometry.slang`, sharing the FPBO helpers
refactored into optim_quant.slang: oq_state_read/oq_accum_us/
oq_state_encode/oq_sqrt3/oq_sqrt4). Design notes:
- **Weighted sampling without replacement**: the efraimidis-spirakis
@@ -548,14 +548,14 @@ the engine level.
sh_value_bits 32/16/8 x bounds layouts x non-SH quant x degree
warmup). All three devices pass; validation clean.
- **Bilagrid/PPISP color pipeline (phase 5, sixth slice)**:
`kernels/Bilagrid.cpp` implements the full csrc/BilagridBindings.h raw
`kernels/Bilagrid.cpp` implements the full kernels/bilagrid/BilagridBindings.h raw
launcher family (affine / PPISP / log-linear / depth / normal samplers:
uniform + patched + per-pixel-coords + packed, forward + backward v1/v2)
plus TV loss / channel mean and the EngineInternal.h fused
TV-Adam/AdaGrad, q16 encode, identity init and scatter;
`kernels/PpispImage.cpp` implements the PixelWise.cuh PPISP image
transform + regularization APIs and the PpispInit.cu helpers. Device
work: `slang/vulkan/bilagrid_{common,affine,ppisp,loglinear,depth,
work: `backend/vulkan/shaders/bilagrid_{common,affine,ppisp,loglinear,depth,
normal,tv,optim}.slang` + `ppisp_image.slang` (37 new entries). Design
notes:
- **BilagridReader** maps to three never-null pointer fields
@@ -563,7 +563,7 @@ the engine level.
`grid_indices == nullptr -> identity` semantic becomes a
has_grid_indices flag field (llvmpipe speculation rule). Uniform vs
patched layouts share one entry via a `kPatched` spec axis.
- **PPISP math is the canonical slang** (`slang/ppisp.slang`) on both
- **PPISP math is the canonical slang** (`shaders/ppisp.slang`) on both
backends: the image transform + reg losses call the same
apply_ppisp* / compute_*_loss* exports the CUDA build gets through
generated/ppisp.cuh, so ppisp_parity's per-pixel values AND gradients
@@ -620,7 +620,7 @@ the engine level.
Gaussian-Thompson-sampling stats but making them per-key, so it needs no
offline calibration pass (unlike fused-bilagrid) and does not average over
resolutions (unlike vksplat's context-free bandit). Full design +
phasing: `csrc/BilagridBackwardSelection.md`. When porting v2, the
phasing: `docs/notes/bilagrid-backward-selection.md`. When porting v2, the
`needs_image_grad` axis is a Slang spec-constant (false for depth/normal,
which skip the image-grad scatter — the v2 analogue of the existing
`v_in == nullptr` skip), and the grid-fold tail rule + null-`v_in` guard
@@ -634,7 +634,7 @@ the engine level.
for real (was a manual stub; the `<true>` instantiation is used by
BilagridBindings). Device work in 4 new modules:
- `multi_scale_loss.slang` includes the CANONICAL
`slang/per_pixel_losses.slang` (CudaDeviceExport -> ForceInline, like
`shaders/per_pixel_losses.slang` (CudaDeviceExport -> ForceInline, like
ppisp), so the per-pixel loss math and its autodiff backward are the
same functions CUDA emits -> per-pixel grads land tight. Bilinear
GT-resolution sampling/scatter re-derived from Interpolation.cuh.
@@ -771,7 +771,7 @@ the engine level.
`--device <index|name substring>`; `ssplat-gui` has a Device combo in
the train panel that locks once training starts.
- **Training stubs**: `kernels/TrainingStubs.gen.cpp` (generated by
`spirulae_splat/generate_vulkan_stubs.py` from a link probe of
`tools/codegen/generate_vulkan_stubs.py` from a link probe of
csrc_portable vs the backend lib) + `TrainingStubsManual.cpp`
(meshing::OccupancyEvaluator methods) provide throwing
definitions for every still-CUDA-only launch function, so the FULL
@@ -1,12 +1,12 @@
// Vulkan implementation of backend/common/SortScan.h over the kernels in
// slang/vulkan/sort_scan.slang (embedded SPIR-V). Design notes in README.md.
// shaders/sort_scan.slang (embedded SPIR-V). Design notes in README.md.
//
// Scratch is a global grow-only pool (mirroring DeviceScratch's single-
// stream assumption): concurrent sort/scan calls on different streams are
// not supported, matching how the engine uses these primitives today.
#include "../common/SortScan.h"
#include "VulkanPipelines.h"
#include "backend/common/SortScan.h"
#include "backend/vulkan/VulkanPipelines.h"
#include <algorithm>
#include <cstring>
@@ -18,7 +18,7 @@ namespace {
using vk::SpecList;
// Must mirror slang/vulkan/sort_scan.slang.
// Must mirror shaders/sort_scan.slang.
constexpr uint32_t kSortPart = 2048;
constexpr uint32_t kRadix = 256;
constexpr uint32_t kScanBlock = 2048;

Some files were not shown because too many files have changed in this diff Show More