mirror of
https://github.com/harry7557558/spirula-studio.git
synced 2026-10-02 02:44:54 +08:00
code restructure
This commit is contained in:
+4
-4
@@ -1,8 +1,8 @@
|
||||
# Third-party dependencies (miniz, stb, npy, ...)
|
||||
spirulae_splat/splat/cuda/csrc/external/** linguist-vendored
|
||||
src/external/** linguist-vendored
|
||||
|
||||
# Auto-generated code (generate_headers.py, generate_kernel_instantiation.py,
|
||||
# generate_cli_config.py)
|
||||
spirulae_splat/splat/cuda/csrc/generated/** linguist-vendored linguist-generated
|
||||
spirulae_splat/splat/cuda/csrc/app/generated/** linguist-vendored linguist-generated
|
||||
spirulae_splat/splat/cuda/ins/** linguist-vendored linguist-generated
|
||||
src/generated/** linguist-vendored linguist-generated
|
||||
src/app/generated/** linguist-vendored linguist-generated
|
||||
src/instantiations/** linguist-vendored linguist-generated
|
||||
|
||||
@@ -24,53 +24,67 @@ Direction of travel, so you don't push the wrong way:
|
||||
## Repo map
|
||||
|
||||
```
|
||||
CMakeLists.txt all native builds (723 lines, 4 modes — see docs/build.md)
|
||||
setup.py / pyproject.toml the pip/torch-extension build (independent of CMake)
|
||||
CMakeLists.txt running order only; the logic is in cmake/
|
||||
cmake/ build modules (options, backends, apps) + sources.txt,
|
||||
the source list BOTH build systems read
|
||||
setup.py / pyproject.toml the pip/torch-extension build (own flags; see docs/build.md)
|
||||
build_develop.bash/.bat the dev build entry points — USE THESE, not pip install
|
||||
AGENTS.md CLAUDE.md this file (CLAUDE.md is a pointer to it)
|
||||
docs/ architecture, build, backends, codegen, testing, notes
|
||||
src/ ALL native code (see below)
|
||||
tools/codegen/ the five codegen tools (see "Codegen" below)
|
||||
tests/python/ Python tests (may be stale; see docs/testing.md)
|
||||
tests/native/ standalone native benchmarks
|
||||
scripts/ dataset preprocessing CLI tools (Python, standalone)
|
||||
tests/ Python tests (may be stale; see docs/testing.md)
|
||||
viewer/ standalone WebGL2 + WASM viewer, independent of training
|
||||
(build-time exception: compiles csrc/app/*Parser.cpp in place)
|
||||
spirulae_splat/ Python package
|
||||
(build-time exception: compiles src/data/parsers/*.cpp
|
||||
in place)
|
||||
spirulae_splat/ Python package — a *client* of the engine, on its way out
|
||||
├── ss_{trainer,benchmark,viewer,meshing}.py console-script entry points
|
||||
├── generate_*.py the five codegen tools (see "Codegen" below)
|
||||
├── modules/ config dataclasses, training driver, eval metrics, resume
|
||||
├── viewer/ Python HTTP viewer server + viewer.html (html is shared
|
||||
│ with the C++ viewer, which embeds it at build time)
|
||||
└── splat/cuda/
|
||||
├── _backend.py _wrapper*.py the extension import + lazy function wrappers
|
||||
├── slang/ Slang device math, shared by CUDA and Vulkan
|
||||
│ └── vulkan/ Vulkan-only compute entry points
|
||||
├── ins/ GENERATED kernel instantiations (111 files)
|
||||
└── csrc/ ALL native code (see below)
|
||||
└── splat/cuda/*.py the extension import + lazy function wrappers
|
||||
```
|
||||
|
||||
`spirulae_splat/splat/cuda/csrc/` is where nearly everything lives. The path
|
||||
is vestigial — `splat/` is a gsplat leftover, `cuda/` also holds the Slang
|
||||
shaders and the Vulkan backend, and `csrc/` holds the engine, the dataset
|
||||
parsers, the CLI, the GUI and the web viewer. Inside it:
|
||||
`src/` is the include root: every local include is path-qualified relative to
|
||||
it (`#include "core/Common.cuh"`, `#include "backend/api/BackendTypes.h"`), so
|
||||
a file's include lines tell you which subsystems it depends on.
|
||||
|
||||
```
|
||||
csrc/
|
||||
├── Engine*.cpp/.h the torch-free training engine (process-global singleton)
|
||||
├── <Kernel>.cu kernel launchers + __global__ kernels
|
||||
├── <Kernel>.cuh GENERATED declaration section (see "Codegen")
|
||||
├── <Kernel>_kernel.cuh the device-side kernel body, split out for reuse
|
||||
├── Bilagrid* 43 files — bilateral grid; see docs/notes/
|
||||
├── Primitive*.cuh 3DGS / Mip / 3DGUT primitive traits
|
||||
├── DataManager.{cpp,h} image cache / prefetch / warp pipeline
|
||||
├── Tensor.h Camera.h Common.cuh core types and device helpers
|
||||
├── Mesh* Delaunay3D.* meshing pipeline
|
||||
├── ext.cpp the pybind11 module (144 m.def's)
|
||||
├── app/ ssplat-train CLI, dataset parsers, web viewer, gui/
|
||||
src/
|
||||
├── core/ Tensor.h, Camera.h, Common.cuh, GradQuant.cuh, …
|
||||
│ the types and device helpers everything uses
|
||||
├── primitives/ Primitive*.cuh — 3DGS / Mip / 3DGUT traits
|
||||
│ (compile-time types, not runtime branches)
|
||||
├── kernels/ one directory per family; each holds the launchers
|
||||
│ │ (<Name>.cu), the GENERATED declaration header
|
||||
│ │ (<Name>.cuh) and the device body (<Name>_kernel.cuh)
|
||||
│ ├── projection/ raster/ tile/
|
||||
│ ├── pixelwise/ ppisp/ bilagrid/ (bilagrid is 42 files, see docs/notes/)
|
||||
│ ├── optim/ densify/ loss/ background/ visualize/
|
||||
├── engine/ Engine*.cpp/.h — the torch-free training engine
|
||||
│ (process-global singleton)
|
||||
├── data/ DataManager (image cache / prefetch / warp) and
|
||||
│ └── parsers/ COLMAP / Nerfstudio / Metashape readers
|
||||
├── mesh/ meshing pipeline, Delaunay3D, UV, export
|
||||
├── backend/ the backend seam — READ backend/README.md
|
||||
│ ├── api/ backend-neutral launch declarations (GENERATED forwarders)
|
||||
│ ├── cuda/ common/ CUDA runtime shim, SortScan, Profiler
|
||||
│ └── vulkan/ Vulkan runtime + kernels/ — READ backend/vulkan/README.md
|
||||
├── generated/ external/ do not hand-edit / vendored
|
||||
└── tests/ backend/tests/ native parity tests
|
||||
│ ├── vulkan/ the whole Vulkan backend: runtime + kernels/ +
|
||||
│ │ shaders/ (its own entry points and SPIR-V build)
|
||||
│ │ — READ backend/vulkan/README.md
|
||||
│ └── tests/ native cross-backend parity tests
|
||||
├── shaders/ Slang device math SHARED by both backends — compiled
|
||||
│ twice, to src/generated/*.cuh and to SPIR-V
|
||||
├── app/ the native applications
|
||||
│ ├── cli/ main.cpp (ssplat-train), mesh_main.cpp (ssplat-mesh)
|
||||
│ ├── gui/ Dear ImGui desktop app (ssplat-gui)
|
||||
│ ├── webviewer/ HTTP server + render worker for the embedded viewer
|
||||
│ └── TrainerCore.{h,cpp} the CLI/GUI training loop
|
||||
├── bindings/ext.cpp the pybind11 module (144 m.def's)
|
||||
├── generated/ app/generated/ instantiations/ GENERATED — do not hand-edit
|
||||
└── external/ vendored (miniz, stb, npy)
|
||||
```
|
||||
|
||||
## Building
|
||||
@@ -97,17 +111,17 @@ Full matrix and per-platform notes: `docs/build.md`.
|
||||
|
||||
## Codegen — the invariants that bite
|
||||
|
||||
Generated trees are marked in `.gitattributes` (`csrc/generated/`,
|
||||
`csrc/app/generated/`, `cuda/ins/`) and are **committed**, so a fresh checkout
|
||||
Generated trees are marked in `.gitattributes` (`src/generated/`,
|
||||
`src/app/generated/`, `src/instantiations/`) and are **committed**, so a fresh checkout
|
||||
builds with no Python at all. Five generators, all run from the repo root:
|
||||
|
||||
| generator | reads | writes |
|
||||
|---|---|---|
|
||||
| `generate_headers.py` | `/*[AutoHeaderGeneratorExport]*/` markers in `csrc/*.cu` | the declaration section of `csrc/<Name>.cuh` |
|
||||
| `generate_kernel_instantiation.py` | kernel decls in `csrc/*.cuh` | `cuda/ins/*.cu` |
|
||||
| `generate_cli_config.py` | the Python config dataclasses (`ast`-parsed, no torch import) | `csrc/app/generated/cli_config.h` |
|
||||
| `generate_backend_api.py` | per-kernel `.cuh` headers | `csrc/backend/api/*.h` forwarders |
|
||||
| `generate_vulkan_stubs.py` | link-probes the Vulkan build | throwing stubs for unported kernels |
|
||||
| `tools/codegen/generate_headers.py` | `/*[AutoHeaderGeneratorExport]*/` markers in `src/kernels/**/*.cu` | the declaration section of the matching `<Name>.cuh` |
|
||||
| `tools/codegen/generate_kernel_instantiation.py` | kernel decls in `src/kernels/**/*_kernel.cuh` | `src/instantiations/*.cu` |
|
||||
| `tools/codegen/generate_cli_config.py` | the Python config dataclasses (`ast`-parsed, no torch import) | `src/app/generated/cli_config.h` |
|
||||
| `tools/codegen/generate_backend_api.py` | per-kernel `.cuh` headers | `src/backend/api/*.h` forwarders |
|
||||
| `tools/codegen/generate_vulkan_stubs.py` | link-probes the Vulkan build | throwing stubs for unported kernels |
|
||||
|
||||
Rules:
|
||||
|
||||
@@ -116,13 +130,13 @@ Rules:
|
||||
regenerated from the `.cu`.
|
||||
2. To export a launch function to the engine, put
|
||||
`/*[AutoHeaderGeneratorExport]*/` immediately above its definition and
|
||||
rerun `generate_headers.py`.
|
||||
rerun `tools/codegen/generate_headers.py`.
|
||||
3. **A header can be fed by several `.cu` files**, listed explicitly in
|
||||
`HEADER_SOURCES` in `generate_headers.py`. This is the sanctioned way to
|
||||
`HEADER_SOURCES` in `tools/codegen/generate_headers.py`. This is the sanctioned way to
|
||||
split a large `.cu`: name each part after **what it does**, not after the
|
||||
header it feeds (e.g. `ImageWarp.cu`, `DepthGeometry.cu` → `PixelWise.cuh`),
|
||||
and add it to the list. Both CMake and `setup.py` glob `*.cu`, so no build
|
||||
file changes are needed. A listed file that doesn't exist is a hard error,
|
||||
and add it to the list. Both build systems glob via `cmake/sources.txt`, so
|
||||
no build file changes are needed. A listed file that doesn't exist is a hard error,
|
||||
so a rename can't silently drop declarations.
|
||||
4. The Python config dataclasses are the **single source of truth** for the
|
||||
training config. Adding a field there makes it appear in the native CLI,
|
||||
@@ -136,14 +150,14 @@ Rules:
|
||||
Every kernel-level change needs **three** things, or the Vulkan build breaks
|
||||
or silently diverges:
|
||||
|
||||
1. the CUDA implementation (`csrc/<Kernel>.cu` + `_kernel.cuh`),
|
||||
1. the CUDA implementation (`src/kernels/<family>/<Kernel>.cu` + `_kernel.cuh`),
|
||||
2. the Slang implementation (`cuda/slang/vulkan/*.slang`) and its launcher
|
||||
(`csrc/backend/vulkan/kernels/*.cpp`),
|
||||
3. a parity test in `csrc/backend/tests/` that runs both and compares.
|
||||
(`src/backend/vulkan/kernels/*.cpp`),
|
||||
3. a parity test in `src/backend/tests/` that runs both and compares.
|
||||
|
||||
If the Vulkan side isn't ready, `generate_vulkan_stubs.py` emits a throwing
|
||||
If the Vulkan side isn't ready, `tools/codegen/generate_vulkan_stubs.py` emits a throwing
|
||||
stub so the portable engine still links — that's a deliberate TODO marker, not
|
||||
a finished state. Coverage is tracked in `csrc/backend/vulkan/README.md`,
|
||||
a finished state. Coverage is tracked in `src/backend/vulkan/README.md`,
|
||||
which is the authoritative and unusually detailed document on the Vulkan
|
||||
design (device baseline, capability variants, memory model, atomics). Read it
|
||||
before touching anything under `backend/vulkan/`.
|
||||
@@ -154,10 +168,10 @@ before touching anything under `backend/vulkan/`.
|
||||
# native parity tests (CUDA build)
|
||||
bash build_develop.bash -DSSPLAT_BUILD_BACKEND_TESTS=ON && ./build/<test_name>
|
||||
# the Vulkan build produces the same test binaries unconditionally
|
||||
pytest tests/ # Python tests — currently unmaintained, may fail
|
||||
pytest tests/python/ # Python tests — currently unmaintained, may fail
|
||||
```
|
||||
|
||||
Each `csrc/backend/tests/*.cpp` becomes an executable of the same name.
|
||||
Each `src/backend/tests/*.cpp` becomes an executable of the same name.
|
||||
`backend/tests/engine/*` drive the real engine end to end. Details and the
|
||||
CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
|
||||
|
||||
@@ -166,7 +180,7 @@ CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
|
||||
- Files are named after **what they do**, not after the file they were split
|
||||
out of, and not after the header they declare into.
|
||||
- `<Name>_kernel.cuh` is reserved: it means "the device body that
|
||||
`generate_kernel_instantiation.py` instantiates". Don't use the `_kernel`
|
||||
`tools/codegen/generate_kernel_instantiation.py` instantiates". Don't use the `_kernel`
|
||||
suffix for an ordinary shared-helper header — give it a descriptive name
|
||||
(e.g. `BilinearSample.cuh`).
|
||||
- A `<Family>Common.cuh` holds the shared preamble when a family spans several
|
||||
@@ -179,7 +193,7 @@ CUDA-vs-Vulkan reference-dump workflow: `docs/testing.md`.
|
||||
```
|
||||
These are load-bearing for the split workflow — keep them meaningful.
|
||||
- Slang device math is shared by both backends; don't fork it per backend.
|
||||
- `csrc/*.cpp` (as opposed to `.cu`) means "portable, compiles for Vulkan
|
||||
- A `.cpp` (as opposed to `.cu`) under `src/` means "portable, compiles for Vulkan
|
||||
builds too". Keep the engine layer in `.cpp` and CUDA-free.
|
||||
|
||||
## Gotchas worth knowing before you hit them
|
||||
|
||||
+21
-716
@@ -1,727 +1,32 @@
|
||||
cmake_minimum_required(VERSION 3.18 FATAL_ERROR)
|
||||
project(spirulae_splat LANGUAGES C CXX) # CUDA enabled below unless SSPLAT_BACKEND=vulkan
|
||||
project(spirulae_splat LANGUAGES C CXX) # CUDA enabled by the cuda backend module
|
||||
|
||||
# This file is intended to be used with `build_develop.bash` for development purpose.
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Find Torch (optional)
|
||||
# Development build (see build_develop.bash / build_develop.bat and
|
||||
# docs/build.md). The pip build is separate -- setup.py -- and shares only the
|
||||
# source list, cmake/sources.txt.
|
||||
#
|
||||
# Torch + Python are only needed for the Python extension module (ext.cpp).
|
||||
# When either is missing -- or SSPLAT_NO_TORCH=ON -- the extension is skipped
|
||||
# and only the standalone ssplat-train CLI is built (engine is torch-free).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
option(SSPLAT_BUILD_CLI "Build the standalone ssplat-train CLI" OFF)
|
||||
option(SSPLAT_BUILD_GUI "Build the native GUI app ssplat-gui (fetches GLFW + Dear ImGui)" OFF)
|
||||
option(SSPLAT_NO_TORCH "Skip Torch/Python even if present; build only ssplat-train" OFF)
|
||||
|
||||
# Debug symbols / line info are OFF by default: they bloat the binaries
|
||||
# massively (nvcc host -g, CUDA cubin lineinfo/source-in-ptx, and slangc -g2
|
||||
# in the embedded SPIR-V). Turn on for profiling/debugging builds. The pip
|
||||
# build (setup.py) is independent and unaffected (it has its own WITH_SYMBOLS
|
||||
# and LINE_INFO env gates).
|
||||
option(SSPLAT_DEBUG_SYMBOLS "Emit debug symbols / line info (host -g, CUDA cubin lineinfo, SPIR-V -g2)" OFF)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Compute backend selection
|
||||
# The logic lives in cmake/; this file is just the running order:
|
||||
#
|
||||
# cuda (default): the full build, exactly as before.
|
||||
# vulkan: the portable engine layer (Engine*.cpp + host support) against the
|
||||
# Vulkan compute runtime (csrc/backend/vulkan/, see its README.md), built
|
||||
# WITHOUT the CUDA toolkit. Produces the backend tests, the ssplat-train CLI
|
||||
# (always) and ssplat-gui (with SSPLAT_BUILD_GUI=ON); no Python extension.
|
||||
# ---------------------------------------------------------------------------
|
||||
set(SSPLAT_BACKEND "cuda" CACHE STRING "Compute backend: cuda | vulkan")
|
||||
set_property(CACHE SSPLAT_BACKEND PROPERTY STRINGS cuda vulkan)
|
||||
# SsplatOptions.cmake options, backend selection, tree paths
|
||||
# SsplatSources.cmake the engine/kernel source list
|
||||
# SsplatBackendCuda.cmake CUDA kernels + optional Torch extension } one
|
||||
# SsplatBackendVulkan.cmake portable engine + Vulkan compute runtime } of
|
||||
# SsplatApps.cmake ssplat-train / ssplat-mesh / ssplat-gui
|
||||
#
|
||||
# Each backend module leaves SSPLAT_WITH_TORCH and SSPLAT_APP_LIBS behind for
|
||||
# the shared app targets.
|
||||
|
||||
list(APPEND CMAKE_MODULE_PATH ${CMAKE_CURRENT_SOURCE_DIR}/cmake)
|
||||
|
||||
include(SsplatOptions)
|
||||
include(SsplatSources)
|
||||
|
||||
if(SSPLAT_BACKEND STREQUAL "vulkan")
|
||||
message(STATUS "SSPLAT_BACKEND=vulkan: portable engine layer + "
|
||||
"backend/vulkan runtime; kernel coverage is tracked in "
|
||||
"csrc/backend/vulkan/README.md.")
|
||||
# The Python extension is CUDA-only, so the CLI trainer is this build's
|
||||
# primary artifact (same forcing as the CUDA no-torch build).
|
||||
set(SSPLAT_BUILD_CLI ON)
|
||||
|
||||
set(CMAKE_CXX_STANDARD 17)
|
||||
set(CMAKE_CXX_STANDARD_REQUIRED ON)
|
||||
|
||||
set(SSPLAT_CSRC ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc)
|
||||
|
||||
find_package(Vulkan REQUIRED)
|
||||
|
||||
# Portable engine layer (compile-checked; links against the Vulkan
|
||||
# runtime once kernel launchers exist for the symbols it needs).
|
||||
file(GLOB SSPLAT_PORTABLE_SOURCES CONFIGURE_DEPENDS ${SSPLAT_CSRC}/*.cpp)
|
||||
list(REMOVE_ITEM SSPLAT_PORTABLE_SOURCES ${SSPLAT_CSRC}/ext.cpp)
|
||||
|
||||
add_library(csrc_portable OBJECT ${SSPLAT_PORTABLE_SOURCES})
|
||||
target_compile_definitions(csrc_portable PUBLIC SSPLAT_BACKEND_VULKAN)
|
||||
target_include_directories(csrc_portable PUBLIC ${SSPLAT_CSRC})
|
||||
|
||||
find_package(OpenMP)
|
||||
if(OpenMP_CXX_FOUND)
|
||||
target_link_libraries(csrc_portable PUBLIC OpenMP::OpenMP_CXX)
|
||||
endif()
|
||||
|
||||
# ---- slangc: locate a matching compiler or fetch the pinned release ----
|
||||
# SPIR-V blobs are never committed; they are compiled at build time (one
|
||||
# slangc edge per blob, see slang/SpirvShaders.cmake) and embedded below.
|
||||
# No Python is required for the Vulkan build.
|
||||
set(SSPLAT_SLANG_VERSION "2026.12.0.1")
|
||||
set(SSPLAT_SLANGC "" CACHE FILEPATH
|
||||
"slangc to use (empty = find in PATH, fetch on miss/mismatch)")
|
||||
|
||||
function(_ssplat_slangc_version exe out_var)
|
||||
set(${out_var} "" PARENT_SCOPE)
|
||||
if(NOT exe OR NOT EXISTS ${exe})
|
||||
return()
|
||||
endif()
|
||||
execute_process(COMMAND ${exe} -v
|
||||
OUTPUT_VARIABLE _v ERROR_VARIABLE _v
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE ERROR_STRIP_TRAILING_WHITESPACE
|
||||
RESULT_VARIABLE _r)
|
||||
if(_r EQUAL 0)
|
||||
string(STRIP "${_v}" _v)
|
||||
set(${out_var} "${_v}" PARENT_SCOPE)
|
||||
endif()
|
||||
endfunction()
|
||||
|
||||
set(_slangc "${SSPLAT_SLANGC}")
|
||||
if(NOT _slangc)
|
||||
find_program(_slangc_in_path slangc)
|
||||
if(_slangc_in_path)
|
||||
set(_slangc ${_slangc_in_path})
|
||||
endif()
|
||||
endif()
|
||||
_ssplat_slangc_version("${_slangc}" _slangc_ver)
|
||||
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
|
||||
if(_slangc)
|
||||
message(STATUS "slangc at ${_slangc} is version '${_slangc_ver}' "
|
||||
"(need ${SSPLAT_SLANG_VERSION}); fetching the pinned release")
|
||||
else()
|
||||
message(STATUS "slangc not found; fetching the pinned release "
|
||||
"${SSPLAT_SLANG_VERSION}")
|
||||
endif()
|
||||
if(WIN32)
|
||||
set(_slang_plat "windows-x86_64")
|
||||
set(_slang_ext "zip")
|
||||
set(_slang_exe "bin/slangc.exe")
|
||||
else()
|
||||
if(CMAKE_HOST_SYSTEM_PROCESSOR MATCHES "aarch64|arm64")
|
||||
set(_slang_arch "aarch64")
|
||||
else()
|
||||
set(_slang_arch "x86_64")
|
||||
endif()
|
||||
if(CMAKE_HOST_SYSTEM_NAME STREQUAL "Darwin")
|
||||
set(_slang_plat "macos-${_slang_arch}")
|
||||
else()
|
||||
set(_slang_plat "linux-${_slang_arch}")
|
||||
endif()
|
||||
set(_slang_ext "tar.gz")
|
||||
set(_slang_exe "bin/slangc")
|
||||
endif()
|
||||
set(_slang_name "slang-${SSPLAT_SLANG_VERSION}-${_slang_plat}")
|
||||
set(_slang_dir ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name})
|
||||
set(_slangc ${_slang_dir}/${_slang_exe})
|
||||
if(NOT EXISTS ${_slangc})
|
||||
set(_slang_url "https://github.com/shader-slang/slang/releases/download/v${SSPLAT_SLANG_VERSION}/${_slang_name}.${_slang_ext}")
|
||||
set(_slang_archive ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name}.${_slang_ext})
|
||||
message(STATUS "Downloading ${_slang_url}")
|
||||
file(DOWNLOAD ${_slang_url} ${_slang_archive}
|
||||
SHOW_PROGRESS STATUS _dl_status)
|
||||
list(GET _dl_status 0 _dl_code)
|
||||
if(NOT _dl_code EQUAL 0)
|
||||
message(FATAL_ERROR "Failed to download slang release: "
|
||||
"${_dl_status}. Install slang-${SSPLAT_SLANG_VERSION} "
|
||||
"and pass -DSSPLAT_SLANGC=/path/to/slangc.")
|
||||
endif()
|
||||
file(ARCHIVE_EXTRACT INPUT ${_slang_archive}
|
||||
DESTINATION ${_slang_dir})
|
||||
file(REMOVE ${_slang_archive})
|
||||
endif()
|
||||
_ssplat_slangc_version("${_slangc}" _slangc_ver)
|
||||
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
|
||||
message(FATAL_ERROR "Fetched slangc reports version "
|
||||
"'${_slangc_ver}', expected ${SSPLAT_SLANG_VERSION}")
|
||||
endif()
|
||||
endif()
|
||||
message(STATUS "Using slangc ${_slangc} (${_slangc_ver})")
|
||||
|
||||
# ---- compile Slang -> SPIR-V (build tree) and embed ----
|
||||
# Native (no Python): slang/spirv_tool.cpp is compiled by the module below,
|
||||
# which enumerates one slangc custom command per blob (so -j bounds RAM) and
|
||||
# embeds the result. SPIR-V debug info (slangc -g2) is opt-in via
|
||||
# SSPLAT_DEBUG_SYMBOLS; it otherwise dominates the embedded blob size.
|
||||
set(SSPLAT_CUDA_DIR ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda)
|
||||
set(SSPLAT_SPIRV_DIR ${CMAKE_CURRENT_BINARY_DIR}/spirv)
|
||||
set(SSPLAT_SPIRV_EMBED ${CMAKE_CURRENT_BINARY_DIR}/vk_shaders_embedded.cpp)
|
||||
include(${SSPLAT_CUDA_DIR}/slang/SpirvShaders.cmake)
|
||||
ssplat_setup_spirv(${SSPLAT_CUDA_DIR} ${_slangc} ${SSPLAT_SPIRV_DIR}
|
||||
${SSPLAT_SPIRV_EMBED} ${SSPLAT_DEBUG_SYMBOLS})
|
||||
|
||||
# Vulkan backend runtime (device context, memory, streams, events) and
|
||||
# kernel launchers.
|
||||
file(GLOB SSPLAT_VULKAN_SOURCES CONFIGURE_DEPENDS ${SSPLAT_CSRC}/backend/vulkan/*.cpp
|
||||
${SSPLAT_CSRC}/backend/vulkan/kernels/*.cpp)
|
||||
list(APPEND SSPLAT_VULKAN_SOURCES ${SSPLAT_SPIRV_EMBED})
|
||||
add_library(ssplat_backend_vulkan STATIC ${SSPLAT_VULKAN_SOURCES})
|
||||
target_compile_definitions(ssplat_backend_vulkan PUBLIC SSPLAT_BACKEND_VULKAN)
|
||||
target_include_directories(ssplat_backend_vulkan PUBLIC ${SSPLAT_CSRC})
|
||||
target_link_libraries(ssplat_backend_vulkan PUBLIC Vulkan::Vulkan)
|
||||
|
||||
# Backend tests (run manually; see backend/vulkan/README.md).
|
||||
file(GLOB SSPLAT_VULKAN_TESTS CONFIGURE_DEPENDS ${SSPLAT_CSRC}/backend/vulkan/tests/*.cpp)
|
||||
foreach(test_src ${SSPLAT_VULKAN_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
|
||||
endforeach()
|
||||
|
||||
# Cross-backend parity tools (same sources build in the CUDA branch).
|
||||
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
|
||||
${SSPLAT_CSRC}/backend/tests/*.cpp)
|
||||
foreach(test_src ${SSPLAT_PARITY_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
|
||||
endforeach()
|
||||
|
||||
# Engine-level parity tools: drive the real engine (forward_3dgs +
|
||||
# background + blit) through the portable layer. The backend lib
|
||||
# provides the ported kernels plus throwing stubs
|
||||
# (kernels/TrainingStubs*.cpp) for everything still CUDA-only.
|
||||
find_package(Threads REQUIRED)
|
||||
file(GLOB SSPLAT_ENGINE_TESTS CONFIGURE_DEPENDS
|
||||
${SSPLAT_CSRC}/backend/tests/engine/*.cpp)
|
||||
foreach(test_src ${SSPLAT_ENGINE_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE csrc_portable
|
||||
ssplat_backend_vulkan Threads::Threads)
|
||||
endforeach()
|
||||
|
||||
# Torch/Python extension is CUDA-only; the app targets below build
|
||||
# against the portable engine + Vulkan backend instead.
|
||||
set(SSPLAT_WITH_TORCH OFF)
|
||||
set(SSPLAT_APP_LIBS csrc_portable ssplat_backend_vulkan Threads::Threads)
|
||||
include(SsplatBackendVulkan)
|
||||
else()
|
||||
|
||||
enable_language(CUDA)
|
||||
|
||||
set(SSPLAT_WITH_TORCH OFF)
|
||||
if(NOT SSPLAT_NO_TORCH)
|
||||
find_package(Python3 QUIET COMPONENTS Interpreter)
|
||||
if(Python3_Interpreter_FOUND)
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(torch.utils.cmake_prefix_path)"
|
||||
OUTPUT_VARIABLE PYTORCH_CMAKE_PREFIX_PATH
|
||||
RESULT_VARIABLE SSPLAT_TORCH_PROBE_RESULT
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
message(STATUS "PyTorch CMake prefix path: ${PYTORCH_CMAKE_PREFIX_PATH}")
|
||||
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import sysconfig; print(sysconfig.get_paths()['include'])"
|
||||
OUTPUT_VARIABLE PYTHON_ADDITION_INCLUDE_PATH
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
message(STATUS "Python additional include path: ${PYTHON_ADDITION_INCLUDE_PATH}")
|
||||
|
||||
if(SSPLAT_TORCH_PROBE_RESULT EQUAL 0)
|
||||
set(CMAKE_PREFIX_PATH ${PYTORCH_CMAKE_PREFIX_PATH})
|
||||
find_package(Torch QUIET)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(Torch_FOUND)
|
||||
message(STATUS "Found Torch: ${TORCH_VERSION}")
|
||||
find_library(TORCH_PYTHON_LIBRARY torch_python PATHS "${TORCH_INSTALL_PREFIX}/lib")
|
||||
set(SSPLAT_WITH_TORCH ON)
|
||||
endif()
|
||||
include(SsplatBackendCuda)
|
||||
endif()
|
||||
|
||||
if(NOT SSPLAT_WITH_TORCH)
|
||||
message(STATUS "Torch/Python3 not available (or SSPLAT_NO_TORCH=ON) -- "
|
||||
"skipping the Python extension; building the standalone ssplat-train CLI only.")
|
||||
set(SSPLAT_BUILD_CLI ON)
|
||||
endif()
|
||||
|
||||
find_package(CUDAToolkit REQUIRED)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Find all sources
|
||||
# ---------------------------------------------------------------------------
|
||||
# CONFIGURE_DEPENDS: kernel families are routinely split into several TUs
|
||||
# (Name.Part.cu -- see docs/codegen.md), and codegen adds/removes files under
|
||||
# ins/. Without it a new source is silently not compiled and only surfaces as
|
||||
# an undefined reference at link time.
|
||||
file(GLOB SPLAT_SOURCES CONFIGURE_DEPENDS
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/*.cpp
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/*.cu
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/ins/*.cu
|
||||
)
|
||||
|
||||
foreach(src ${SPLAT_SOURCES})
|
||||
if(src MATCHES "hip")
|
||||
list(REMOVE_ITEM SPLAT_SOURCES ${src})
|
||||
endif()
|
||||
endforeach()
|
||||
|
||||
if(NOT SSPLAT_WITH_TORCH)
|
||||
# ext.cpp (the pybind11 module) is the only Torch-dependent source.
|
||||
list(REMOVE_ITEM SPLAT_SOURCES
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/ext.cpp)
|
||||
endif()
|
||||
|
||||
# message(STATUS "Extension sources:")
|
||||
# foreach(src ${SPLAT_SOURCES})
|
||||
# message(STATUS " ${src}")
|
||||
# endforeach()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Detect CUDA architectures
|
||||
# ---------------------------------------------------------------------------
|
||||
if(NOT DEFINED TORCH_CUDA_ARCH_LIST)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(' '.join(f'{a[0]}{a[1]}' for a in [torch.cuda.get_device_capability(i) for i in range(torch.cuda.device_count())]))"
|
||||
OUTPUT_VARIABLE TORCH_CUDA_ARCH_LIST
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
endif()
|
||||
|
||||
if(NOT TORCH_CUDA_ARCH_LIST)
|
||||
# No Torch (or Torch saw no device): ask the driver directly.
|
||||
execute_process(
|
||||
COMMAND nvidia-smi --query-gpu=compute_cap --format=csv,noheader
|
||||
OUTPUT_VARIABLE SSPLAT_SMI_CAPS
|
||||
RESULT_VARIABLE SSPLAT_SMI_RESULT
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
if(SSPLAT_SMI_RESULT EQUAL 0)
|
||||
string(REPLACE "\r" "" SSPLAT_SMI_CAPS "${SSPLAT_SMI_CAPS}")
|
||||
string(REPLACE "\n" " " TORCH_CUDA_ARCH_LIST "${SSPLAT_SMI_CAPS}")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(NOT TORCH_CUDA_ARCH_LIST)
|
||||
message(FATAL_ERROR "Could not detect a CUDA GPU. "
|
||||
"Pass -DTORCH_CUDA_ARCH_LIST=\"8.6\" (or similar) to set the target architecture.")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# normalize: "8.6" / "8.6 7.5" / "86;75" all become a deduped list like "86;75"
|
||||
string(REGEX REPLACE "\\." "" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
|
||||
string(REPLACE " " ";" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
|
||||
list(REMOVE_DUPLICATES TORCH_CUDA_ARCH_LIST)
|
||||
|
||||
message(STATUS "CUDA architecture(s): ${TORCH_CUDA_ARCH_LIST}")
|
||||
|
||||
set(CMAKE_CUDA_ARCHITECTURES "")
|
||||
foreach(arch ${TORCH_CUDA_ARCH_LIST})
|
||||
list(APPEND CMAKE_CUDA_ARCHITECTURES "${arch}")
|
||||
endforeach()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Compiler Flags
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# CXX flags
|
||||
set(CMAKE_CXX_STANDARD 17)
|
||||
set(CMAKE_CXX_STANDARD_REQUIRED ON)
|
||||
|
||||
if(MSVC)
|
||||
set(SPLAT_CXX_FLAGS "/O2")
|
||||
else()
|
||||
set(SPLAT_CXX_FLAGS "-O3")
|
||||
# set(SPLAT_CXX_FLAGS "-g")
|
||||
list(APPEND SPLAT_CXX_FLAGS "-Wno-sign-compare")
|
||||
endif()
|
||||
|
||||
if(WIN32)
|
||||
add_compile_definitions(_USE_MATH_DEFINES NOMINMAX _CRT_SECURE_NO_WARNINGS)
|
||||
endif()
|
||||
|
||||
if (MSVC)
|
||||
add_compile_options(
|
||||
$<$<COMPILE_LANGUAGE:CXX>:/Zc:preprocessor>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:-Xcompiler=/Zc:preprocessor>
|
||||
)
|
||||
endif()
|
||||
|
||||
# Detect OpenMP
|
||||
find_package(OpenMP)
|
||||
if(OpenMP_CXX_FOUND AND NOT APPLE)
|
||||
message(STATUS "Compiling with OpenMP")
|
||||
list(APPEND SPLAT_CXX_FLAGS "-DAT_PARALLEL_OPENMP")
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# Compile flag only: the OpenMP runtime comes from torch's bundled
|
||||
# libgomp. Linking the system one too would shadow it (CMake warns
|
||||
# about the unsafe rpath). The no-torch build instead links the
|
||||
# OpenMP::OpenMP_CXX target, whose flags are language-guarded so
|
||||
# nvcc never sees them.
|
||||
list(APPEND SPLAT_CXX_FLAGS ${OpenMP_CXX_FLAGS})
|
||||
endif()
|
||||
else()
|
||||
message(STATUS "Compiling without OpenMP...")
|
||||
endif()
|
||||
|
||||
# macOS ARM64 special case
|
||||
if(APPLE AND CMAKE_SYSTEM_PROCESSOR MATCHES "arm64")
|
||||
message(STATUS "Enabling macOS ARM64 flags")
|
||||
add_compile_options("-arch" "arm64")
|
||||
add_link_options("-arch" "arm64")
|
||||
endif()
|
||||
|
||||
# nvcc flags (mirrors setup.py)
|
||||
set(SPLAT_NVCC_FLAGS
|
||||
"-O3" "--use_fast_math"
|
||||
"--expt-relaxed-constexpr"
|
||||
"-Xcudafe=--diag_suppress=20012"
|
||||
"-Xcudafe=--diag_suppress=550"
|
||||
"--threads" "0"
|
||||
"-Xptxas" "--warn-on-double-precision-use"
|
||||
)
|
||||
|
||||
# cubin/PTX line info (adds embedded source; needed for profiling but bloats
|
||||
# the fatbin). Off unless SSPLAT_DEBUG_SYMBOLS.
|
||||
if(SSPLAT_DEBUG_SYMBOLS)
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-lineinfo" "--generate-line-info" "--source-in-ptx")
|
||||
endif()
|
||||
|
||||
# Add gencode flags
|
||||
foreach(arch ${TORCH_CUDA_ARCH_LIST})
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-gencode" "arch=compute_${arch},code=sm_${arch}")
|
||||
endforeach()
|
||||
|
||||
# Host side optimizations
|
||||
if(NOT WIN32)
|
||||
# list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-O3,-march=native")
|
||||
# Host debug symbols (nvcc -Xcompiler=-g). Off unless SSPLAT_DEBUG_SYMBOLS.
|
||||
if(SSPLAT_DEBUG_SYMBOLS)
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-g")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Define the extension module
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
add_library(csrc SHARED ${SPLAT_SOURCES})
|
||||
else()
|
||||
# The engine API has no dllexport annotations (Linux .so exports
|
||||
# everything by default); a static lib sidesteps that on Windows and
|
||||
# makes ssplat-train a self-contained executable.
|
||||
add_library(csrc STATIC ${SPLAT_SOURCES})
|
||||
endif()
|
||||
|
||||
# Include directories
|
||||
target_include_directories(csrc PRIVATE
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
|
||||
# ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc/glm
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
|
||||
# Definitions
|
||||
if(WIN32)
|
||||
target_compile_definitions(csrc PRIVATE spirulae_splat_EXPORTS)
|
||||
endif()
|
||||
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
target_include_directories(csrc PRIVATE
|
||||
${Python3_INCLUDE_DIRS}
|
||||
${PYTHON_ADDITION_INCLUDE_PATH}
|
||||
)
|
||||
|
||||
target_compile_definitions(csrc PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:TORCH_EXTENSION_NAME=csrc>
|
||||
)
|
||||
|
||||
# Link Torch
|
||||
target_link_libraries(csrc
|
||||
${TORCH_LIBRARIES}
|
||||
${TORCH_PYTHON_LIBRARY}
|
||||
${Python3_LIBRARIES}
|
||||
)
|
||||
else()
|
||||
# Torch normally drags these in for us
|
||||
find_package(Threads REQUIRED)
|
||||
target_link_libraries(csrc CUDA::cudart Threads::Threads)
|
||||
if(OpenMP_CXX_FOUND AND NOT APPLE)
|
||||
target_link_libraries(csrc OpenMP::OpenMP_CXX)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# Apply flags
|
||||
target_compile_options(csrc PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:${SPLAT_NVCC_FLAGS}>
|
||||
)
|
||||
|
||||
# Optional: strip symbols (setup.py uses '-s' if WITH_SYMBOLS==False)
|
||||
if(DEFINED WITH_SYMBOLS AND NOT WITH_SYMBOLS)
|
||||
if(NOT WIN32)
|
||||
target_link_options(csrc PRIVATE "-s")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# Required by PyTorch C++ extensions
|
||||
set_property(TARGET csrc PROPERTY CXX_STANDARD 17)
|
||||
set_property(TARGET csrc PROPERTY CUDA_STANDARD 17)
|
||||
|
||||
set(SSPLAT_APP_LIBS csrc CUDA::cudart)
|
||||
endif() # SSPLAT_BACKEND (vulkan | cuda core; app targets below are shared)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Standalone CLI trainer -- no Python at runtime; links the csrc shared
|
||||
# library for the engine. Off by default (unless Torch is unavailable, which
|
||||
# forces it on) so the normal extension build is unaffected.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
if(SSPLAT_BUILD_CLI OR SSPLAT_BUILD_GUI)
|
||||
# Embed the (unchanged) web viewer client into the binary so the exe is
|
||||
# self-contained. Regenerated at configure time when viewer.html changes.
|
||||
set(SSPLAT_VIEWER_HTML ${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/viewer/viewer.html)
|
||||
set(SSPLAT_VIEWER_HTML_HDR ${CMAKE_BINARY_DIR}/app_generated/viewer_html.h)
|
||||
file(READ ${SSPLAT_VIEWER_HTML} _ssplat_vh HEX)
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _ssplat_vh_bytes ${_ssplat_vh})
|
||||
file(WRITE ${SSPLAT_VIEWER_HTML_HDR}
|
||||
"#pragma once\n"
|
||||
"// AUTO-GENERATED from spirulae_splat/viewer/viewer.html -- do not edit.\n"
|
||||
"#include <cstddef>\n"
|
||||
"inline const unsigned char kViewerHtml[] = {${_ssplat_vh_bytes}};\n"
|
||||
"inline const size_t kViewerHtmlSize = sizeof(kViewerHtml);\n")
|
||||
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS ${SSPLAT_VIEWER_HTML})
|
||||
endif()
|
||||
|
||||
# Cross-backend parity tools (dump reference outputs; the Vulkan build's
|
||||
# twin executable compares against them). See csrc/backend/vulkan/README.md.
|
||||
# The Vulkan branch builds its comparing twins unconditionally above.
|
||||
option(SSPLAT_BUILD_BACKEND_TESTS "Build backend parity test tools" OFF)
|
||||
if(SSPLAT_BUILD_BACKEND_TESTS AND NOT SSPLAT_BACKEND STREQUAL "vulkan")
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# Same libpython + static-libstdc++ notes as ssplat-train below.
|
||||
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
|
||||
endif()
|
||||
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
|
||||
spirulae_splat/splat/cuda/csrc/backend/tests/*.cpp
|
||||
spirulae_splat/splat/cuda/csrc/backend/tests/engine/*.cpp)
|
||||
foreach(test_src ${SSPLAT_PARITY_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_include_directories(${test_name} PRIVATE
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
|
||||
${CUDAToolkit_INCLUDE_DIRS})
|
||||
target_link_libraries(${test_name} PRIVATE csrc CUDA::cudart)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
target_link_libraries(${test_name} PRIVATE Python3::Python)
|
||||
if(NOT WIN32)
|
||||
target_link_options(${test_name} PRIVATE
|
||||
-static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
endif()
|
||||
endforeach()
|
||||
endif()
|
||||
|
||||
if(SSPLAT_BUILD_CLI)
|
||||
add_executable(ssplat-train
|
||||
spirulae_splat/splat/cuda/csrc/app/main.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/TrainerCore.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/HttpServer.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/RenderWorker.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/Viewer.cpp
|
||||
spirulae_splat/splat/cuda/csrc/external/miniz.c
|
||||
)
|
||||
target_include_directories(ssplat-train PRIVATE
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
|
||||
${CMAKE_BINARY_DIR}
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
target_link_libraries(ssplat-train PRIVATE ${SSPLAT_APP_LIBS})
|
||||
if (WIN32)
|
||||
target_link_libraries(ssplat-train PRIVATE ws2_32)
|
||||
endif()
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# libpython is only needed because csrc's pybind/libtorch_python layer
|
||||
# leaves CPython symbols undefined (a Python process normally provides
|
||||
# them at import time). Drops out with the no-torch build.
|
||||
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
|
||||
target_link_libraries(ssplat-train PRIVATE Python3::Python)
|
||||
endif()
|
||||
target_compile_options(ssplat-train PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
)
|
||||
# Statically link libstdc++ into the exe so its C++ runtime calls bind at
|
||||
# link time. Without this, libtorch.so (pulled in via csrc) interposes
|
||||
# its own bundled std::filesystem symbols at load time, and e.g.
|
||||
# remove_all resolves to an ABI-incompatible copy that jumps to null.
|
||||
if(SSPLAT_WITH_TORCH AND NOT WIN32)
|
||||
target_link_options(ssplat-train PRIVATE -static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
# libcsrc.so is renamed to spirulae_splat/csrc.so by build_develop.bash;
|
||||
# keep the exe's lookup pointing at the build tree copy.
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
set_property(TARGET ssplat-train PROPERTY BUILD_RPATH
|
||||
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
|
||||
else()
|
||||
set_property(TARGET ssplat-train PROPERTY BUILD_RPATH "$ORIGIN")
|
||||
endif()
|
||||
set_property(TARGET ssplat-train PROPERTY CXX_STANDARD 17)
|
||||
|
||||
# ---- standalone mesh-extraction CLI (shares the dataset parsers) ----
|
||||
# CUDA-only: the Vulkan backend stubs the meshing kernels
|
||||
# (meshing::OccupancyEvaluator), so the tool would throw at startup.
|
||||
if(NOT SSPLAT_BACKEND STREQUAL "vulkan")
|
||||
add_executable(ssplat-mesh
|
||||
spirulae_splat/splat/cuda/csrc/app/mesh_main.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
|
||||
spirulae_splat/splat/cuda/csrc/external/miniz.c
|
||||
)
|
||||
target_include_directories(ssplat-mesh PRIVATE
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
target_link_libraries(ssplat-mesh PRIVATE csrc CUDA::cudart)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
target_link_libraries(ssplat-mesh PRIVATE Python3::Python)
|
||||
endif()
|
||||
target_compile_options(ssplat-mesh PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
)
|
||||
if(SSPLAT_WITH_TORCH AND NOT WIN32)
|
||||
target_link_options(ssplat-mesh PRIVATE -static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
set_property(TARGET ssplat-mesh PROPERTY BUILD_RPATH
|
||||
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
|
||||
else()
|
||||
set_property(TARGET ssplat-mesh PROPERTY BUILD_RPATH "$ORIGIN")
|
||||
endif()
|
||||
set_property(TARGET ssplat-mesh PROPERTY CXX_STANDARD 17)
|
||||
endif() # NOT vulkan (ssplat-mesh)
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Native GUI (ssplat-gui) -- optional, off by default. Dear ImGui + GLFW +
|
||||
# OpenGL 3 (imgui's embedded GL loader; no GLEW/GLAD). Dependencies are
|
||||
# fetched at configure time via FetchContent and built statically, so the
|
||||
# only system requirements are OpenGL dev headers (Linux: libgl1-mesa-dev +
|
||||
# xorg-dev for GLFW; Windows/macOS: nothing extra).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
if(SSPLAT_BUILD_GUI)
|
||||
include(FetchContent)
|
||||
|
||||
set(GLFW_BUILD_DOCS OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_BUILD_TESTS OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_INSTALL OFF CACHE BOOL "" FORCE)
|
||||
FetchContent_Declare(glfw
|
||||
GIT_REPOSITORY https://github.com/glfw/glfw.git
|
||||
GIT_TAG 3.4
|
||||
GIT_SHALLOW TRUE)
|
||||
FetchContent_Declare(imgui
|
||||
GIT_REPOSITORY https://github.com/ocornut/imgui.git
|
||||
GIT_TAG v1.92.8
|
||||
GIT_SHALLOW TRUE)
|
||||
FetchContent_MakeAvailable(glfw imgui)
|
||||
|
||||
find_package(OpenGL REQUIRED)
|
||||
|
||||
# imgui ships no CMakeLists; build the core + glfw/gl3 backends +
|
||||
# std::string InputText helper as a small static lib.
|
||||
add_library(imgui_glfw STATIC
|
||||
${imgui_SOURCE_DIR}/imgui.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_draw.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_tables.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_widgets.cpp
|
||||
${imgui_SOURCE_DIR}/backends/imgui_impl_glfw.cpp
|
||||
${imgui_SOURCE_DIR}/backends/imgui_impl_opengl3.cpp
|
||||
${imgui_SOURCE_DIR}/misc/cpp/imgui_stdlib.cpp
|
||||
)
|
||||
target_include_directories(imgui_glfw PUBLIC
|
||||
${imgui_SOURCE_DIR}
|
||||
${imgui_SOURCE_DIR}/backends
|
||||
${imgui_SOURCE_DIR}/misc/cpp
|
||||
)
|
||||
target_link_libraries(imgui_glfw PUBLIC glfw)
|
||||
set_property(TARGET imgui_glfw PROPERTY CXX_STANDARD 17)
|
||||
|
||||
# Embed scripts/mask.py (AI masking helper, run via external Python) so
|
||||
# the exe is self-contained. Same mechanism as the viewer.html embed.
|
||||
set(SSPLAT_MASK_PY ${CMAKE_CURRENT_SOURCE_DIR}/scripts/mask.py)
|
||||
set(SSPLAT_MASK_PY_HDR ${CMAKE_BINARY_DIR}/app_generated/mask_py.h)
|
||||
file(READ ${SSPLAT_MASK_PY} _ssplat_mp HEX)
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _ssplat_mp_bytes ${_ssplat_mp})
|
||||
file(WRITE ${SSPLAT_MASK_PY_HDR}
|
||||
"#pragma once\n"
|
||||
"// AUTO-GENERATED from scripts/mask.py -- do not edit.\n"
|
||||
"#include <cstddef>\n"
|
||||
"inline const unsigned char kMaskPy[] = {${_ssplat_mp_bytes}};\n"
|
||||
"inline const size_t kMaskPySize = sizeof(kMaskPy);\n")
|
||||
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS ${SSPLAT_MASK_PY})
|
||||
|
||||
add_executable(ssplat-gui
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/GuiMain.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/GuiApp.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/ConfigUI.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/FileDialog.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/Subprocess.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/ColmapRunner.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/FrameSelect.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/TrainRunner.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/GlLoader.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/NavCamera.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/PreviewRenderer.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/gui/ViewportPanel.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/TrainerCore.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/RenderWorker.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/HttpServer.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/Viewer.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/ColmapParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/NerfstudioParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/MetashapeParser.cpp
|
||||
spirulae_splat/splat/cuda/csrc/app/DatasetCommon.cpp
|
||||
spirulae_splat/splat/cuda/csrc/external/miniz.c
|
||||
)
|
||||
target_include_directories(ssplat-gui PRIVATE
|
||||
${CMAKE_CURRENT_SOURCE_DIR}/spirulae_splat/splat/cuda/csrc
|
||||
${CMAKE_BINARY_DIR}
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
target_link_libraries(ssplat-gui PRIVATE
|
||||
${SSPLAT_APP_LIBS} imgui_glfw OpenGL::GL)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
|
||||
target_link_libraries(ssplat-gui PRIVATE Python3::Python)
|
||||
endif()
|
||||
if (WIN32)
|
||||
target_link_libraries(ssplat-gui PRIVATE ws2_32)
|
||||
endif()
|
||||
target_compile_options(ssplat-gui PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
)
|
||||
# Same libtorch std::filesystem interposition workaround as ssplat-train.
|
||||
if(SSPLAT_WITH_TORCH AND NOT WIN32)
|
||||
target_link_options(ssplat-gui PRIVATE -static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
set_property(TARGET ssplat-gui PROPERTY BUILD_RPATH
|
||||
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
|
||||
else()
|
||||
set_property(TARGET ssplat-gui PROPERTY BUILD_RPATH "$ORIGIN")
|
||||
endif()
|
||||
set_property(TARGET ssplat-gui PROPERTY CXX_STANDARD 17)
|
||||
endif()
|
||||
include(SsplatApps)
|
||||
|
||||
message(STATUS "spirulae_splat configured (backend: ${SSPLAT_BACKEND}).")
|
||||
|
||||
+11
-5
@@ -7,14 +7,14 @@
|
||||
# Regenerate headers/config. Skipped when python3 is unavailable -- the
|
||||
# generated files are committed, so the build still works without it.
|
||||
if command -v python3 >/dev/null 2>&1; then
|
||||
python3 spirulae_splat/generate_headers.py
|
||||
python3 spirulae_splat/generate_kernel_instantiation.py
|
||||
python3 spirulae_splat/generate_cli_config.py
|
||||
python3 tools/codegen/generate_headers.py
|
||||
python3 tools/codegen/generate_kernel_instantiation.py
|
||||
python3 tools/codegen/generate_cli_config.py
|
||||
else
|
||||
echo "python3 not found -- skipping codegen (using committed generated files)"
|
||||
fi
|
||||
|
||||
cmake -G Ninja -B build "$@"
|
||||
cmake -G Ninja -B build "$@" || exit $?
|
||||
|
||||
echo ""
|
||||
|
||||
@@ -29,7 +29,13 @@ echo "Available RAM : ${AVAILABLE_MB} MB"
|
||||
echo "CPU cores : ${CPU_CORES}"
|
||||
echo "Using jobs : ${JOBS}"
|
||||
echo ""
|
||||
cmake --build build --verbose -j"${JOBS}"
|
||||
# Propagate the build's exit status: without this the script always exits 0
|
||||
# (the trailing libcsrc.so move succeeds regardless) and a failed build reads
|
||||
# as a green one.
|
||||
if ! cmake --build build --verbose -j"${JOBS}"; then
|
||||
echo "BUILD FAILED" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Torch build only: expose the extension as spirulae_splat/csrc.so, with a
|
||||
# symlink back so ssplat-train's $ORIGIN lookup keeps working. A no-torch
|
||||
|
||||
+3
-3
@@ -16,9 +16,9 @@ rem ---------------------------------------------------------------------------
|
||||
rem Codegen (optional -- the generated files are committed)
|
||||
rem ---------------------------------------------------------------------------
|
||||
where python >nul 2>&1 || goto :skip_codegen
|
||||
python spirulae_splat\generate_headers.py || goto :codegen_warn
|
||||
python spirulae_splat\generate_kernel_instantiation.py || goto :codegen_warn
|
||||
python spirulae_splat\generate_cli_config.py || goto :codegen_warn
|
||||
python tools\codegen\generate_headers.py || goto :codegen_warn
|
||||
python tools\codegen\generate_kernel_instantiation.py || goto :codegen_warn
|
||||
python tools\codegen\generate_cli_config.py || goto :codegen_warn
|
||||
goto :msvc_env
|
||||
:codegen_warn
|
||||
echo codegen failed -- continuing with committed generated files
|
||||
|
||||
@@ -0,0 +1,177 @@
|
||||
# The native applications: ssplat-train (CLI), ssplat-mesh, ssplat-gui.
|
||||
#
|
||||
# Backend-agnostic: whichever backend module ran first left the engine to link
|
||||
# in SSPLAT_APP_LIBS and set SSPLAT_WITH_TORCH. ssplat-mesh is the exception --
|
||||
# the Vulkan backend stubs the meshing kernels, so it is CUDA-only.
|
||||
|
||||
include(SsplatEmbed)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Source groups shared by several app targets
|
||||
# ---------------------------------------------------------------------------
|
||||
# Dataset parsing (COLMAP / Nerfstudio / Metashape) + miniz for the zipped
|
||||
# Metashape .psx payloads.
|
||||
set(SSPLAT_DATASET_SOURCES
|
||||
${SSPLAT_SRC}/data/parsers/ColmapParser.cpp
|
||||
${SSPLAT_SRC}/data/parsers/NerfstudioParser.cpp
|
||||
${SSPLAT_SRC}/data/parsers/MetashapeParser.cpp
|
||||
${SSPLAT_SRC}/data/parsers/DatasetCommon.cpp
|
||||
${SSPLAT_SRC}/external/miniz.c
|
||||
)
|
||||
|
||||
# The embedded web viewer: HTTP server + the render worker feeding it.
|
||||
set(SSPLAT_VIEWER_SOURCES
|
||||
${SSPLAT_SRC}/app/webviewer/HttpServer.cpp
|
||||
${SSPLAT_SRC}/app/webviewer/RenderWorker.cpp
|
||||
${SSPLAT_SRC}/app/webviewer/Viewer.cpp
|
||||
)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# ssplat_configure_app(<target>)
|
||||
#
|
||||
# The settings every app target needs: include paths, the engine libraries,
|
||||
# host flags, and the two link-time workarounds a Torch-enabled build requires.
|
||||
# ---------------------------------------------------------------------------
|
||||
function(ssplat_configure_app target)
|
||||
target_include_directories(${target} PRIVATE
|
||||
${SSPLAT_SRC}
|
||||
${CMAKE_BINARY_DIR} # app_generated/{viewer_html,mask_py}.h
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
target_link_libraries(${target} PRIVATE ${SSPLAT_APP_LIBS})
|
||||
target_compile_options(${target} PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
)
|
||||
set_property(TARGET ${target} PROPERTY CXX_STANDARD 17)
|
||||
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# libpython is only needed because csrc's pybind/libtorch_python layer
|
||||
# leaves CPython symbols undefined (a Python process normally provides
|
||||
# them at import time). Drops out with the no-torch build.
|
||||
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
|
||||
target_link_libraries(${target} PRIVATE Python3::Python)
|
||||
|
||||
# Statically link libstdc++ into the exe so its C++ runtime calls bind
|
||||
# at link time. Without this, libtorch.so (pulled in via csrc)
|
||||
# interposes its own bundled std::filesystem symbols at load time, and
|
||||
# e.g. remove_all resolves to an ABI-incompatible copy that jumps to
|
||||
# null.
|
||||
if(NOT WIN32)
|
||||
target_link_options(${target} PRIVATE -static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
|
||||
# libcsrc.so is renamed to spirulae_splat/csrc.so by
|
||||
# build_develop.bash; keep the exe's lookup pointing at the build tree
|
||||
# copy (and at libtorch).
|
||||
set_property(TARGET ${target} PROPERTY BUILD_RPATH
|
||||
"$ORIGIN;${TORCH_INSTALL_PREFIX}/lib")
|
||||
else()
|
||||
set_property(TARGET ${target} PROPERTY BUILD_RPATH "$ORIGIN")
|
||||
endif()
|
||||
endfunction()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Embedded web viewer client
|
||||
# ---------------------------------------------------------------------------
|
||||
if(SSPLAT_BUILD_CLI OR SSPLAT_BUILD_GUI)
|
||||
# Embed the (unchanged) web viewer client into the binary so the exe is
|
||||
# self-contained. Regenerated at configure time when viewer.html changes.
|
||||
ssplat_embed_file(
|
||||
${SSPLAT_ROOT}/spirulae_splat/viewer/viewer.html
|
||||
${CMAKE_BINARY_DIR}/app_generated/viewer_html.h
|
||||
ViewerHtml)
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# ssplat-train -- standalone CLI trainer, no Python at runtime. Off by default
|
||||
# (unless Torch is unavailable, or the backend is Vulkan, which force it on).
|
||||
# ---------------------------------------------------------------------------
|
||||
if(SSPLAT_BUILD_CLI)
|
||||
add_executable(ssplat-train
|
||||
${SSPLAT_SRC}/app/cli/main.cpp
|
||||
${SSPLAT_SRC}/app/TrainerCore.cpp
|
||||
${SSPLAT_DATASET_SOURCES}
|
||||
${SSPLAT_VIEWER_SOURCES}
|
||||
)
|
||||
ssplat_configure_app(ssplat-train)
|
||||
if(WIN32)
|
||||
target_link_libraries(ssplat-train PRIVATE ws2_32)
|
||||
endif()
|
||||
|
||||
# ---- standalone mesh-extraction CLI (shares the dataset parsers) ----
|
||||
# CUDA-only: the Vulkan backend stubs the meshing kernels
|
||||
# (meshing::OccupancyEvaluator), so the tool would throw at startup.
|
||||
if(NOT SSPLAT_BACKEND STREQUAL "vulkan")
|
||||
add_executable(ssplat-mesh
|
||||
${SSPLAT_SRC}/app/cli/mesh_main.cpp
|
||||
${SSPLAT_DATASET_SOURCES}
|
||||
)
|
||||
ssplat_configure_app(ssplat-mesh)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# ssplat-gui -- optional, off by default. Dear ImGui + GLFW + OpenGL 3
|
||||
# (imgui's embedded GL loader; no GLEW/GLAD). Dependencies are fetched at
|
||||
# configure time via FetchContent and built statically, so the only system
|
||||
# requirements are OpenGL dev headers (Linux: libgl1-mesa-dev + xorg-dev for
|
||||
# GLFW; Windows/macOS: nothing extra).
|
||||
# ---------------------------------------------------------------------------
|
||||
if(SSPLAT_BUILD_GUI)
|
||||
include(FetchContent)
|
||||
|
||||
set(GLFW_BUILD_DOCS OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_BUILD_TESTS OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
|
||||
set(GLFW_INSTALL OFF CACHE BOOL "" FORCE)
|
||||
FetchContent_Declare(glfw
|
||||
GIT_REPOSITORY https://github.com/glfw/glfw.git
|
||||
GIT_TAG 3.4
|
||||
GIT_SHALLOW TRUE)
|
||||
FetchContent_Declare(imgui
|
||||
GIT_REPOSITORY https://github.com/ocornut/imgui.git
|
||||
GIT_TAG v1.92.8
|
||||
GIT_SHALLOW TRUE)
|
||||
FetchContent_MakeAvailable(glfw imgui)
|
||||
|
||||
find_package(OpenGL REQUIRED)
|
||||
|
||||
# imgui ships no CMakeLists; build the core + glfw/gl3 backends +
|
||||
# std::string InputText helper as a small static lib.
|
||||
add_library(imgui_glfw STATIC
|
||||
${imgui_SOURCE_DIR}/imgui.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_draw.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_tables.cpp
|
||||
${imgui_SOURCE_DIR}/imgui_widgets.cpp
|
||||
${imgui_SOURCE_DIR}/backends/imgui_impl_glfw.cpp
|
||||
${imgui_SOURCE_DIR}/backends/imgui_impl_opengl3.cpp
|
||||
${imgui_SOURCE_DIR}/misc/cpp/imgui_stdlib.cpp
|
||||
)
|
||||
target_include_directories(imgui_glfw PUBLIC
|
||||
${imgui_SOURCE_DIR}
|
||||
${imgui_SOURCE_DIR}/backends
|
||||
${imgui_SOURCE_DIR}/misc/cpp
|
||||
)
|
||||
target_link_libraries(imgui_glfw PUBLIC glfw)
|
||||
set_property(TARGET imgui_glfw PROPERTY CXX_STANDARD 17)
|
||||
|
||||
# Embed scripts/mask.py (AI masking helper, run via external Python) so
|
||||
# the exe is self-contained. Same mechanism as the viewer.html embed.
|
||||
ssplat_embed_file(
|
||||
${SSPLAT_ROOT}/scripts/mask.py
|
||||
${CMAKE_BINARY_DIR}/app_generated/mask_py.h
|
||||
MaskPy)
|
||||
|
||||
file(GLOB SSPLAT_GUI_SOURCES CONFIGURE_DEPENDS ${SSPLAT_SRC}/app/gui/*.cpp)
|
||||
add_executable(ssplat-gui
|
||||
${SSPLAT_GUI_SOURCES}
|
||||
${SSPLAT_SRC}/app/TrainerCore.cpp
|
||||
${SSPLAT_DATASET_SOURCES}
|
||||
${SSPLAT_VIEWER_SOURCES}
|
||||
)
|
||||
ssplat_configure_app(ssplat-gui)
|
||||
target_link_libraries(ssplat-gui PRIVATE imgui_glfw OpenGL::GL)
|
||||
if(WIN32)
|
||||
target_link_libraries(ssplat-gui PRIVATE ws2_32)
|
||||
endif()
|
||||
endif()
|
||||
@@ -0,0 +1,275 @@
|
||||
# SSPLAT_BACKEND=cuda (default): the CUDA kernels plus, when Torch and Python
|
||||
# are available, the `csrc` Python extension module.
|
||||
#
|
||||
# Defines, for the shared app targets in SsplatApps.cmake:
|
||||
# SSPLAT_WITH_TORCH ON when the Python extension is being built
|
||||
# SSPLAT_APP_LIBS the engine library the apps link
|
||||
# SPLAT_CXX_FLAGS host flags the app targets reuse
|
||||
|
||||
enable_language(CUDA)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Find Torch (optional)
|
||||
# ---------------------------------------------------------------------------
|
||||
set(SSPLAT_WITH_TORCH OFF)
|
||||
if(NOT SSPLAT_NO_TORCH)
|
||||
find_package(Python3 QUIET COMPONENTS Interpreter)
|
||||
if(Python3_Interpreter_FOUND)
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(torch.utils.cmake_prefix_path)"
|
||||
OUTPUT_VARIABLE PYTORCH_CMAKE_PREFIX_PATH
|
||||
RESULT_VARIABLE SSPLAT_TORCH_PROBE_RESULT
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
message(STATUS "PyTorch CMake prefix path: ${PYTORCH_CMAKE_PREFIX_PATH}")
|
||||
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import sysconfig; print(sysconfig.get_paths()['include'])"
|
||||
OUTPUT_VARIABLE PYTHON_ADDITION_INCLUDE_PATH
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
message(STATUS "Python additional include path: ${PYTHON_ADDITION_INCLUDE_PATH}")
|
||||
|
||||
if(SSPLAT_TORCH_PROBE_RESULT EQUAL 0)
|
||||
set(CMAKE_PREFIX_PATH ${PYTORCH_CMAKE_PREFIX_PATH})
|
||||
find_package(Torch QUIET)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(Torch_FOUND)
|
||||
message(STATUS "Found Torch: ${TORCH_VERSION}")
|
||||
find_library(TORCH_PYTHON_LIBRARY torch_python PATHS "${TORCH_INSTALL_PREFIX}/lib")
|
||||
set(SSPLAT_WITH_TORCH ON)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(NOT SSPLAT_WITH_TORCH)
|
||||
message(STATUS "Torch/Python3 not available (or SSPLAT_NO_TORCH=ON) -- "
|
||||
"skipping the Python extension; building the standalone ssplat-train CLI only.")
|
||||
set(SSPLAT_BUILD_CLI ON)
|
||||
endif()
|
||||
|
||||
find_package(CUDAToolkit REQUIRED)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Sources
|
||||
# ---------------------------------------------------------------------------
|
||||
ssplat_collect_sources(SPLAT_SOURCES)
|
||||
|
||||
if(NOT SSPLAT_WITH_TORCH)
|
||||
# bindings/ext.cpp (the pybind11 module) is the only Torch-dependent source.
|
||||
list(REMOVE_ITEM SPLAT_SOURCES ${SSPLAT_SRC}/bindings/ext.cpp)
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Detect CUDA architectures
|
||||
# ---------------------------------------------------------------------------
|
||||
if(NOT DEFINED TORCH_CUDA_ARCH_LIST)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
execute_process(
|
||||
COMMAND ${Python3_EXECUTABLE} -c "import torch; print(' '.join(f'{a[0]}{a[1]}' for a in [torch.cuda.get_device_capability(i) for i in range(torch.cuda.device_count())]))"
|
||||
OUTPUT_VARIABLE TORCH_CUDA_ARCH_LIST
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
endif()
|
||||
|
||||
if(NOT TORCH_CUDA_ARCH_LIST)
|
||||
# No Torch (or Torch saw no device): ask the driver directly.
|
||||
execute_process(
|
||||
COMMAND nvidia-smi --query-gpu=compute_cap --format=csv,noheader
|
||||
OUTPUT_VARIABLE SSPLAT_SMI_CAPS
|
||||
RESULT_VARIABLE SSPLAT_SMI_RESULT
|
||||
ERROR_QUIET
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE)
|
||||
if(SSPLAT_SMI_RESULT EQUAL 0)
|
||||
string(REPLACE "\r" "" SSPLAT_SMI_CAPS "${SSPLAT_SMI_CAPS}")
|
||||
string(REPLACE "\n" " " TORCH_CUDA_ARCH_LIST "${SSPLAT_SMI_CAPS}")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(NOT TORCH_CUDA_ARCH_LIST)
|
||||
message(FATAL_ERROR "Could not detect a CUDA GPU. "
|
||||
"Pass -DTORCH_CUDA_ARCH_LIST=\"8.6\" (or similar) to set the target architecture.")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# normalize: "8.6" / "8.6 7.5" / "86;75" all become a deduped list like "86;75"
|
||||
string(REGEX REPLACE "\\." "" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
|
||||
string(REPLACE " " ";" TORCH_CUDA_ARCH_LIST "${TORCH_CUDA_ARCH_LIST}")
|
||||
list(REMOVE_DUPLICATES TORCH_CUDA_ARCH_LIST)
|
||||
|
||||
message(STATUS "CUDA architecture(s): ${TORCH_CUDA_ARCH_LIST}")
|
||||
|
||||
set(CMAKE_CUDA_ARCHITECTURES "")
|
||||
foreach(arch ${TORCH_CUDA_ARCH_LIST})
|
||||
list(APPEND CMAKE_CUDA_ARCHITECTURES "${arch}")
|
||||
endforeach()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Compiler flags
|
||||
#
|
||||
# These intentionally mirror, but do not share, setup.py's flags: the pip build
|
||||
# targets a released wheel (its own WITH_SYMBOLS / LINE_INFO env gates) while
|
||||
# this one targets local development. Only the *source list* is shared, via
|
||||
# cmake/sources.txt.
|
||||
# ---------------------------------------------------------------------------
|
||||
if(MSVC)
|
||||
set(SPLAT_CXX_FLAGS "/O2")
|
||||
else()
|
||||
set(SPLAT_CXX_FLAGS "-O3")
|
||||
list(APPEND SPLAT_CXX_FLAGS "-Wno-sign-compare")
|
||||
endif()
|
||||
|
||||
if(MSVC)
|
||||
add_compile_options(
|
||||
$<$<COMPILE_LANGUAGE:CXX>:/Zc:preprocessor>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:-Xcompiler=/Zc:preprocessor>
|
||||
)
|
||||
endif()
|
||||
|
||||
# Detect OpenMP
|
||||
find_package(OpenMP)
|
||||
if(OpenMP_CXX_FOUND AND NOT APPLE)
|
||||
message(STATUS "Compiling with OpenMP")
|
||||
list(APPEND SPLAT_CXX_FLAGS "-DAT_PARALLEL_OPENMP")
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# Compile flag only: the OpenMP runtime comes from torch's bundled
|
||||
# libgomp. Linking the system one too would shadow it (CMake warns
|
||||
# about the unsafe rpath). The no-torch build instead links the
|
||||
# OpenMP::OpenMP_CXX target, whose flags are language-guarded so
|
||||
# nvcc never sees them.
|
||||
list(APPEND SPLAT_CXX_FLAGS ${OpenMP_CXX_FLAGS})
|
||||
endif()
|
||||
else()
|
||||
message(STATUS "Compiling without OpenMP...")
|
||||
endif()
|
||||
|
||||
# macOS ARM64 special case
|
||||
if(APPLE AND CMAKE_SYSTEM_PROCESSOR MATCHES "arm64")
|
||||
message(STATUS "Enabling macOS ARM64 flags")
|
||||
add_compile_options("-arch" "arm64")
|
||||
add_link_options("-arch" "arm64")
|
||||
endif()
|
||||
|
||||
# nvcc flags
|
||||
set(SPLAT_NVCC_FLAGS
|
||||
"-O3" "--use_fast_math"
|
||||
"--expt-relaxed-constexpr"
|
||||
"-Xcudafe=--diag_suppress=20012"
|
||||
"-Xcudafe=--diag_suppress=550"
|
||||
"--threads" "0"
|
||||
"-Xptxas" "--warn-on-double-precision-use"
|
||||
)
|
||||
|
||||
# cubin/PTX line info (adds embedded source; needed for profiling but bloats
|
||||
# the fatbin). Off unless SSPLAT_DEBUG_SYMBOLS.
|
||||
if(SSPLAT_DEBUG_SYMBOLS)
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-lineinfo" "--generate-line-info" "--source-in-ptx")
|
||||
endif()
|
||||
|
||||
# Add gencode flags
|
||||
foreach(arch ${TORCH_CUDA_ARCH_LIST})
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-gencode" "arch=compute_${arch},code=sm_${arch}")
|
||||
endforeach()
|
||||
|
||||
# Host side optimizations
|
||||
if(NOT WIN32)
|
||||
# list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-O3,-march=native")
|
||||
# Host debug symbols (nvcc -Xcompiler=-g). Off unless SSPLAT_DEBUG_SYMBOLS.
|
||||
if(SSPLAT_DEBUG_SYMBOLS)
|
||||
list(APPEND SPLAT_NVCC_FLAGS "-Xcompiler=-g")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The engine library / Python extension module
|
||||
# ---------------------------------------------------------------------------
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
add_library(csrc SHARED ${SPLAT_SOURCES})
|
||||
else()
|
||||
# The engine API has no dllexport annotations (Linux .so exports
|
||||
# everything by default); a static lib sidesteps that on Windows and
|
||||
# makes ssplat-train a self-contained executable.
|
||||
add_library(csrc STATIC ${SPLAT_SOURCES})
|
||||
endif()
|
||||
|
||||
target_include_directories(csrc PRIVATE
|
||||
${SSPLAT_SRC}
|
||||
${CUDAToolkit_INCLUDE_DIRS}
|
||||
)
|
||||
|
||||
if(WIN32)
|
||||
target_compile_definitions(csrc PRIVATE spirulae_splat_EXPORTS)
|
||||
endif()
|
||||
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
target_include_directories(csrc PRIVATE
|
||||
${Python3_INCLUDE_DIRS}
|
||||
${PYTHON_ADDITION_INCLUDE_PATH}
|
||||
)
|
||||
|
||||
target_compile_definitions(csrc PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:TORCH_EXTENSION_NAME=csrc>
|
||||
)
|
||||
|
||||
target_link_libraries(csrc
|
||||
${TORCH_LIBRARIES}
|
||||
${TORCH_PYTHON_LIBRARY}
|
||||
${Python3_LIBRARIES}
|
||||
)
|
||||
else()
|
||||
# Torch normally drags these in for us
|
||||
find_package(Threads REQUIRED)
|
||||
target_link_libraries(csrc CUDA::cudart Threads::Threads)
|
||||
if(OpenMP_CXX_FOUND AND NOT APPLE)
|
||||
target_link_libraries(csrc OpenMP::OpenMP_CXX)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
target_compile_options(csrc PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:${SPLAT_CXX_FLAGS}>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:${SPLAT_NVCC_FLAGS}>
|
||||
)
|
||||
|
||||
# Optional: strip symbols (setup.py uses '-s' if WITH_SYMBOLS==False)
|
||||
if(DEFINED WITH_SYMBOLS AND NOT WITH_SYMBOLS)
|
||||
if(NOT WIN32)
|
||||
target_link_options(csrc PRIVATE "-s")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# Required by PyTorch C++ extensions
|
||||
set_property(TARGET csrc PROPERTY CXX_STANDARD 17)
|
||||
set_property(TARGET csrc PROPERTY CUDA_STANDARD 17)
|
||||
|
||||
set(SSPLAT_APP_LIBS csrc CUDA::cudart)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Cross-backend parity tools -- CUDA side (dump reference outputs; the Vulkan
|
||||
# build's twin executables compare against them). See backend/vulkan/README.md.
|
||||
# The Vulkan branch builds its comparing twins unconditionally.
|
||||
# ---------------------------------------------------------------------------
|
||||
if(SSPLAT_BUILD_BACKEND_TESTS)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
# Same libpython + static-libstdc++ notes as ssplat-train.
|
||||
find_package(Python3 REQUIRED COMPONENTS Development.Embed)
|
||||
endif()
|
||||
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS
|
||||
${SSPLAT_SRC}/backend/tests/*.cpp
|
||||
${SSPLAT_SRC}/backend/tests/engine/*.cpp)
|
||||
foreach(test_src ${SSPLAT_PARITY_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_include_directories(${test_name} PRIVATE
|
||||
${SSPLAT_SRC}
|
||||
${CUDAToolkit_INCLUDE_DIRS})
|
||||
target_link_libraries(${test_name} PRIVATE csrc CUDA::cudart)
|
||||
if(SSPLAT_WITH_TORCH)
|
||||
target_link_libraries(${test_name} PRIVATE Python3::Python)
|
||||
if(NOT WIN32)
|
||||
target_link_options(${test_name} PRIVATE
|
||||
-static-libgcc -static-libstdc++)
|
||||
endif()
|
||||
endif()
|
||||
endforeach()
|
||||
endif()
|
||||
@@ -0,0 +1,101 @@
|
||||
# SSPLAT_BACKEND=vulkan: the portable engine layer built against the Vulkan
|
||||
# compute runtime, without the CUDA toolkit. Kernel coverage is tracked in
|
||||
# src/backend/vulkan/README.md.
|
||||
#
|
||||
# Defines, for the shared app targets in SsplatApps.cmake:
|
||||
# SSPLAT_WITH_TORCH = OFF (the Python extension is CUDA-only)
|
||||
# SSPLAT_APP_LIBS (portable engine + Vulkan runtime + threads)
|
||||
|
||||
message(STATUS "SSPLAT_BACKEND=vulkan: portable engine layer + "
|
||||
"backend/vulkan runtime; kernel coverage is tracked in "
|
||||
"src/backend/vulkan/README.md.")
|
||||
|
||||
# The Python extension is CUDA-only, so the CLI trainer is this build's
|
||||
# primary artifact (same forcing as the CUDA no-torch build).
|
||||
set(SSPLAT_BUILD_CLI ON)
|
||||
|
||||
find_package(Vulkan REQUIRED)
|
||||
find_package(Threads REQUIRED)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Portable engine layer
|
||||
# ---------------------------------------------------------------------------
|
||||
# Compile-checked here; links against the Vulkan runtime once kernel launchers
|
||||
# exist for the symbols it needs.
|
||||
ssplat_collect_portable_sources(SSPLAT_PORTABLE_SOURCES)
|
||||
|
||||
add_library(csrc_portable OBJECT ${SSPLAT_PORTABLE_SOURCES})
|
||||
target_compile_definitions(csrc_portable PUBLIC SSPLAT_BACKEND_VULKAN)
|
||||
target_include_directories(csrc_portable PUBLIC ${SSPLAT_SRC})
|
||||
|
||||
find_package(OpenMP)
|
||||
if(OpenMP_CXX_FOUND)
|
||||
target_link_libraries(csrc_portable PUBLIC OpenMP::OpenMP_CXX)
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Slang -> SPIR-V (build tree), then embed
|
||||
# ---------------------------------------------------------------------------
|
||||
# Native (no Python): backend/vulkan/shaders/spirv_tool.cpp is compiled by the
|
||||
# module below,
|
||||
# which enumerates one slangc custom command per blob (so -j bounds RAM) and
|
||||
# embeds the result. SPIR-V debug info (slangc -g2) is opt-in via
|
||||
# SSPLAT_DEBUG_SYMBOLS; it otherwise dominates the embedded blob size.
|
||||
include(SsplatSlang)
|
||||
ssplat_find_slangc(SSPLAT_SLANGC_EXE)
|
||||
|
||||
set(SSPLAT_SPIRV_DIR ${CMAKE_CURRENT_BINARY_DIR}/spirv)
|
||||
set(SSPLAT_SPIRV_EMBED ${CMAKE_CURRENT_BINARY_DIR}/vk_shaders_embedded.cpp)
|
||||
include(${SSPLAT_VK_SHADERS}/SpirvShaders.cmake)
|
||||
ssplat_setup_spirv(${SSPLAT_SHADERS} ${SSPLAT_VK_SHADERS} ${SSPLAT_SLANGC_EXE}
|
||||
${SSPLAT_SPIRV_DIR} ${SSPLAT_SPIRV_EMBED} ${SSPLAT_DEBUG_SYMBOLS})
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Vulkan backend runtime (device context, memory, streams, events) + kernels
|
||||
# ---------------------------------------------------------------------------
|
||||
file(GLOB SSPLAT_VULKAN_SOURCES CONFIGURE_DEPENDS
|
||||
${SSPLAT_SRC}/backend/vulkan/*.cpp
|
||||
${SSPLAT_SRC}/backend/vulkan/kernels/*.cpp)
|
||||
list(APPEND SSPLAT_VULKAN_SOURCES ${SSPLAT_SPIRV_EMBED})
|
||||
|
||||
add_library(ssplat_backend_vulkan STATIC ${SSPLAT_VULKAN_SOURCES})
|
||||
target_compile_definitions(ssplat_backend_vulkan PUBLIC SSPLAT_BACKEND_VULKAN)
|
||||
target_include_directories(ssplat_backend_vulkan PUBLIC ${SSPLAT_SRC})
|
||||
target_link_libraries(ssplat_backend_vulkan PUBLIC Vulkan::Vulkan)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Tests (run manually; see backend/vulkan/README.md)
|
||||
# ---------------------------------------------------------------------------
|
||||
# Vulkan-only unit tests.
|
||||
file(GLOB SSPLAT_VULKAN_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/vulkan/tests/*.cpp)
|
||||
foreach(test_src ${SSPLAT_VULKAN_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
|
||||
endforeach()
|
||||
|
||||
# Cross-backend parity tools (the same sources build in the CUDA branch, where
|
||||
# they dump the reference outputs these compare against).
|
||||
file(GLOB SSPLAT_PARITY_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/tests/*.cpp)
|
||||
foreach(test_src ${SSPLAT_PARITY_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE ssplat_backend_vulkan)
|
||||
endforeach()
|
||||
|
||||
# Engine-level parity tools: drive the real engine (forward_3dgs + background
|
||||
# + blit) through the portable layer. The backend lib provides the ported
|
||||
# kernels plus throwing stubs (kernels/TrainingStubs*.cpp) for everything still
|
||||
# CUDA-only.
|
||||
file(GLOB SSPLAT_ENGINE_TESTS CONFIGURE_DEPENDS ${SSPLAT_SRC}/backend/tests/engine/*.cpp)
|
||||
foreach(test_src ${SSPLAT_ENGINE_TESTS})
|
||||
get_filename_component(test_name ${test_src} NAME_WE)
|
||||
add_executable(${test_name} ${test_src})
|
||||
target_link_libraries(${test_name} PRIVATE csrc_portable
|
||||
ssplat_backend_vulkan Threads::Threads)
|
||||
endforeach()
|
||||
|
||||
# Torch/Python extension is CUDA-only; the app targets build against the
|
||||
# portable engine + Vulkan backend instead.
|
||||
set(SSPLAT_WITH_TORCH OFF)
|
||||
set(SSPLAT_APP_LIBS csrc_portable ssplat_backend_vulkan Threads::Threads)
|
||||
@@ -0,0 +1,20 @@
|
||||
# Embedding data files into the executables as byte arrays, so the apps are
|
||||
# self-contained (no runtime lookup of viewer.html or scripts/mask.py).
|
||||
|
||||
# ssplat_embed_file(<input> <output_header> <symbol>)
|
||||
#
|
||||
# Writes a header defining `k<symbol>[]` / `k<symbol>Size` holding the bytes of
|
||||
# <input>. Regenerated at configure time whenever <input> changes.
|
||||
function(ssplat_embed_file input output_header symbol)
|
||||
file(READ ${input} _hex HEX)
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _bytes ${_hex})
|
||||
file(RELATIVE_PATH _rel ${SSPLAT_ROOT} ${input})
|
||||
file(WRITE ${output_header}
|
||||
"#pragma once\n"
|
||||
"// AUTO-GENERATED from ${_rel} -- do not edit.\n"
|
||||
"#include <cstddef>\n"
|
||||
"inline const unsigned char k${symbol}[] = {${_bytes}};\n"
|
||||
"inline const size_t k${symbol}Size = sizeof(k${symbol});\n")
|
||||
set_property(DIRECTORY ${SSPLAT_ROOT} APPEND
|
||||
PROPERTY CMAKE_CONFIGURE_DEPENDS ${input})
|
||||
endfunction()
|
||||
@@ -0,0 +1,61 @@
|
||||
# Build options, backend selection, and the source-tree paths every other
|
||||
# module builds on. Included first by the top-level CMakeLists.txt.
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Paths
|
||||
# ---------------------------------------------------------------------------
|
||||
# SSPLAT_ROOT is the repository root (the directory holding CMakeLists.txt).
|
||||
# Everything else is expressed relative to it so the modules never care where
|
||||
# they were included from.
|
||||
# SSPLAT_SRC is also the include root: every local #include in the native
|
||||
# tree is written relative to it (e.g. `#include "core/Common.cuh"`).
|
||||
set(SSPLAT_ROOT ${CMAKE_CURRENT_SOURCE_DIR})
|
||||
set(SSPLAT_SRC ${SSPLAT_ROOT}/src)
|
||||
set(SSPLAT_SHADERS ${SSPLAT_SRC}/shaders) # shared Slang device math
|
||||
set(SSPLAT_VK_SHADERS ${SSPLAT_SRC}/backend/vulkan/shaders) # Vulkan-only entry points
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Options
|
||||
#
|
||||
# Torch + Python are only needed for the Python extension module (ext.cpp).
|
||||
# When either is missing -- or SSPLAT_NO_TORCH=ON -- the extension is skipped
|
||||
# and only the standalone ssplat-train CLI is built (engine is torch-free).
|
||||
# ---------------------------------------------------------------------------
|
||||
option(SSPLAT_BUILD_CLI "Build the standalone ssplat-train CLI" OFF)
|
||||
option(SSPLAT_BUILD_GUI "Build the native GUI app ssplat-gui (fetches GLFW + Dear ImGui)" OFF)
|
||||
option(SSPLAT_NO_TORCH "Skip Torch/Python even if present; build only ssplat-train" OFF)
|
||||
option(SSPLAT_BUILD_BACKEND_TESTS "Build backend parity test tools" OFF)
|
||||
|
||||
# Debug symbols / line info are OFF by default: they bloat the binaries
|
||||
# massively (nvcc host -g, CUDA cubin lineinfo/source-in-ptx, and slangc -g2
|
||||
# in the embedded SPIR-V). Turn on for profiling/debugging builds. The pip
|
||||
# build (setup.py) is independent and unaffected (it has its own WITH_SYMBOLS
|
||||
# and LINE_INFO env gates).
|
||||
option(SSPLAT_DEBUG_SYMBOLS "Emit debug symbols / line info (host -g, CUDA cubin lineinfo, SPIR-V -g2)" OFF)
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Compute backend selection
|
||||
#
|
||||
# cuda (default): the full build -- CUDA kernels, optional Torch extension,
|
||||
# and the app targets.
|
||||
# vulkan: the portable engine layer (Engine*.cpp + host support) against the
|
||||
# Vulkan compute runtime (src/backend/vulkan/, see its README.md), built
|
||||
# WITHOUT the CUDA toolkit. Produces the backend tests, the ssplat-train CLI
|
||||
# (always) and ssplat-gui (with SSPLAT_BUILD_GUI=ON); no Python extension.
|
||||
# ---------------------------------------------------------------------------
|
||||
set(SSPLAT_BACKEND "cuda" CACHE STRING "Compute backend: cuda | vulkan")
|
||||
set_property(CACHE SSPLAT_BACKEND PROPERTY STRINGS cuda vulkan)
|
||||
|
||||
if(NOT SSPLAT_BACKEND STREQUAL "cuda" AND NOT SSPLAT_BACKEND STREQUAL "vulkan")
|
||||
message(FATAL_ERROR "SSPLAT_BACKEND must be 'cuda' or 'vulkan', got '${SSPLAT_BACKEND}'")
|
||||
endif()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Shared C++ settings
|
||||
# ---------------------------------------------------------------------------
|
||||
set(CMAKE_CXX_STANDARD 17)
|
||||
set(CMAKE_CXX_STANDARD_REQUIRED ON)
|
||||
|
||||
if(WIN32)
|
||||
add_compile_definitions(_USE_MATH_DEFINES NOMINMAX _CRT_SECURE_NO_WARNINGS)
|
||||
endif()
|
||||
@@ -0,0 +1,95 @@
|
||||
# Locating slangc (the Slang -> SPIR-V compiler) for the Vulkan backend.
|
||||
#
|
||||
# SPIR-V blobs are never committed; they are compiled at build time (one slangc
|
||||
# edge per blob, see src/backend/vulkan/shaders/SpirvShaders.cmake) and embedded into the
|
||||
# binary. No Python is required for the Vulkan build.
|
||||
|
||||
set(SSPLAT_SLANG_VERSION "2026.12.0.1")
|
||||
set(SSPLAT_SLANGC "" CACHE FILEPATH
|
||||
"slangc to use (empty = find in PATH, fetch on miss/mismatch)")
|
||||
|
||||
function(_ssplat_slangc_version exe out_var)
|
||||
set(${out_var} "" PARENT_SCOPE)
|
||||
if(NOT exe OR NOT EXISTS ${exe})
|
||||
return()
|
||||
endif()
|
||||
execute_process(COMMAND ${exe} -v
|
||||
OUTPUT_VARIABLE _v ERROR_VARIABLE _v
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE ERROR_STRIP_TRAILING_WHITESPACE
|
||||
RESULT_VARIABLE _r)
|
||||
if(_r EQUAL 0)
|
||||
string(STRIP "${_v}" _v)
|
||||
set(${out_var} "${_v}" PARENT_SCOPE)
|
||||
endif()
|
||||
endfunction()
|
||||
|
||||
# ssplat_find_slangc(<out_var>)
|
||||
#
|
||||
# Returns a slangc executable matching SSPLAT_SLANG_VERSION: the one named by
|
||||
# -DSSPLAT_SLANGC, else one in PATH, else the pinned upstream release
|
||||
# downloaded into the build tree.
|
||||
function(ssplat_find_slangc out_var)
|
||||
set(_slangc "${SSPLAT_SLANGC}")
|
||||
if(NOT _slangc)
|
||||
find_program(_slangc_in_path slangc)
|
||||
if(_slangc_in_path)
|
||||
set(_slangc ${_slangc_in_path})
|
||||
endif()
|
||||
endif()
|
||||
_ssplat_slangc_version("${_slangc}" _slangc_ver)
|
||||
|
||||
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
|
||||
if(_slangc)
|
||||
message(STATUS "slangc at ${_slangc} is version '${_slangc_ver}' "
|
||||
"(need ${SSPLAT_SLANG_VERSION}); fetching the pinned release")
|
||||
else()
|
||||
message(STATUS "slangc not found; fetching the pinned release "
|
||||
"${SSPLAT_SLANG_VERSION}")
|
||||
endif()
|
||||
if(WIN32)
|
||||
set(_slang_plat "windows-x86_64")
|
||||
set(_slang_ext "zip")
|
||||
set(_slang_exe "bin/slangc.exe")
|
||||
else()
|
||||
if(CMAKE_HOST_SYSTEM_PROCESSOR MATCHES "aarch64|arm64")
|
||||
set(_slang_arch "aarch64")
|
||||
else()
|
||||
set(_slang_arch "x86_64")
|
||||
endif()
|
||||
if(CMAKE_HOST_SYSTEM_NAME STREQUAL "Darwin")
|
||||
set(_slang_plat "macos-${_slang_arch}")
|
||||
else()
|
||||
set(_slang_plat "linux-${_slang_arch}")
|
||||
endif()
|
||||
set(_slang_ext "tar.gz")
|
||||
set(_slang_exe "bin/slangc")
|
||||
endif()
|
||||
set(_slang_name "slang-${SSPLAT_SLANG_VERSION}-${_slang_plat}")
|
||||
set(_slang_dir ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name})
|
||||
set(_slangc ${_slang_dir}/${_slang_exe})
|
||||
if(NOT EXISTS ${_slangc})
|
||||
set(_slang_url "https://github.com/shader-slang/slang/releases/download/v${SSPLAT_SLANG_VERSION}/${_slang_name}.${_slang_ext}")
|
||||
set(_slang_archive ${CMAKE_CURRENT_BINARY_DIR}/${_slang_name}.${_slang_ext})
|
||||
message(STATUS "Downloading ${_slang_url}")
|
||||
file(DOWNLOAD ${_slang_url} ${_slang_archive}
|
||||
SHOW_PROGRESS STATUS _dl_status)
|
||||
list(GET _dl_status 0 _dl_code)
|
||||
if(NOT _dl_code EQUAL 0)
|
||||
message(FATAL_ERROR "Failed to download slang release: "
|
||||
"${_dl_status}. Install slang-${SSPLAT_SLANG_VERSION} "
|
||||
"and pass -DSSPLAT_SLANGC=/path/to/slangc.")
|
||||
endif()
|
||||
file(ARCHIVE_EXTRACT INPUT ${_slang_archive}
|
||||
DESTINATION ${_slang_dir})
|
||||
file(REMOVE ${_slang_archive})
|
||||
endif()
|
||||
_ssplat_slangc_version("${_slangc}" _slangc_ver)
|
||||
if(NOT _slangc_ver STREQUAL SSPLAT_SLANG_VERSION)
|
||||
message(FATAL_ERROR "Fetched slangc reports version "
|
||||
"'${_slangc_ver}', expected ${SSPLAT_SLANG_VERSION}")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
message(STATUS "Using slangc ${_slangc} (${_slangc_ver})")
|
||||
set(${out_var} ${_slangc} PARENT_SCOPE)
|
||||
endfunction()
|
||||
@@ -0,0 +1,69 @@
|
||||
# Collects the engine/kernel sources from cmake/sources.txt.
|
||||
#
|
||||
# The glob patterns live in a plain text file rather than here so that setup.py
|
||||
# (which builds the same sources through torch.utils.cpp_extension, without
|
||||
# CMake) reads the exact same list. See cmake/sources.txt.
|
||||
|
||||
# _ssplat_expand_section(<section> <out_var>)
|
||||
#
|
||||
# CONFIGURE_DEPENDS: kernel families are routinely split into several
|
||||
# translation units, each named after what it does (see docs/codegen.md), and
|
||||
# codegen adds/removes files under src/instantiations/. Without it a newly
|
||||
# added source is silently not compiled and only surfaces as an undefined
|
||||
# reference at link time.
|
||||
function(_ssplat_expand_section section out_var)
|
||||
set(spec ${SSPLAT_ROOT}/cmake/sources.txt)
|
||||
file(STRINGS ${spec} lines)
|
||||
|
||||
set(patterns "")
|
||||
set(in_section FALSE)
|
||||
foreach(line ${lines})
|
||||
string(STRIP "${line}" line)
|
||||
if(line STREQUAL "" OR line MATCHES "^#")
|
||||
continue()
|
||||
endif()
|
||||
if(line MATCHES "^\\[(.*)\\]$")
|
||||
if(CMAKE_MATCH_1 STREQUAL section)
|
||||
set(in_section TRUE)
|
||||
else()
|
||||
set(in_section FALSE)
|
||||
endif()
|
||||
continue()
|
||||
endif()
|
||||
if(in_section)
|
||||
list(APPEND patterns ${SSPLAT_ROOT}/${line})
|
||||
endif()
|
||||
endforeach()
|
||||
|
||||
if(NOT patterns)
|
||||
message(FATAL_ERROR "No [${section}] patterns found in ${spec}")
|
||||
endif()
|
||||
|
||||
file(GLOB sources CONFIGURE_DEPENDS ${patterns})
|
||||
|
||||
# Drop ROCm sources (neither build compiles them).
|
||||
list(FILTER sources EXCLUDE REGEX "hip")
|
||||
|
||||
if(NOT sources)
|
||||
message(FATAL_ERROR "cmake/sources.txt [${section}] matched no files")
|
||||
endif()
|
||||
|
||||
# Reconfigure when the spec itself changes.
|
||||
set_property(DIRECTORY ${SSPLAT_ROOT} APPEND
|
||||
PROPERTY CMAKE_CONFIGURE_DEPENDS ${spec})
|
||||
|
||||
set(${out_var} ${sources} PARENT_SCOPE)
|
||||
endfunction()
|
||||
|
||||
# ssplat_collect_sources(<out_var>) -- everything the CUDA `csrc` lib compiles.
|
||||
function(ssplat_collect_sources out_var)
|
||||
_ssplat_expand_section(cuda result)
|
||||
set(${out_var} ${result} PARENT_SCOPE)
|
||||
endfunction()
|
||||
|
||||
# ssplat_collect_portable_sources(<out_var>) -- the torch-free, CUDA-free
|
||||
# subset compiled as csrc_portable by the Vulkan build.
|
||||
function(ssplat_collect_portable_sources out_var)
|
||||
_ssplat_expand_section(portable result)
|
||||
set(${out_var} ${result} PARENT_SCOPE)
|
||||
endfunction()
|
||||
@@ -0,0 +1,46 @@
|
||||
# Engine/kernel source globs, relative to the repository root.
|
||||
#
|
||||
# SINGLE SOURCE OF TRUTH. Consumed by:
|
||||
# - cmake/SsplatSources.cmake (the development build, both backends)
|
||||
# - setup.py (the pip / torch.utils.cpp_extension build)
|
||||
# Keep it a plain list of glob patterns, one per line: both readers parse it
|
||||
# line-by-line, and setup.py must be able to do so without CMake installed.
|
||||
# Blank lines and '#' comments are ignored. Globs are NOT recursive -- list each
|
||||
# directory, so adding one is a deliberate act.
|
||||
#
|
||||
# Any path containing "hip" is dropped by both readers (ROCm sources are not
|
||||
# part of either build).
|
||||
#
|
||||
# Not listed here, on purpose: src/app/** and src/backend/** (their targets own
|
||||
# their sources -- see SsplatApps.cmake / the backend modules).
|
||||
|
||||
[cuda]
|
||||
# Everything the CUDA `csrc` library compiles.
|
||||
src/instantiations/*.cu
|
||||
src/kernels/background/*.cu
|
||||
src/kernels/bilagrid/*.cu
|
||||
src/kernels/bilagrid/*.cpp
|
||||
src/kernels/densify/*.cu
|
||||
src/kernels/loss/*.cu
|
||||
src/kernels/optim/*.cu
|
||||
src/kernels/pixelwise/*.cu
|
||||
src/kernels/ppisp/*.cu
|
||||
src/kernels/projection/*.cu
|
||||
src/kernels/raster/*.cu
|
||||
src/kernels/tile/*.cu
|
||||
src/kernels/visualize/*.cu
|
||||
src/mesh/*.cu
|
||||
src/mesh/*.cpp
|
||||
src/engine/*.cpp
|
||||
src/data/*.cpp
|
||||
src/external/*.cpp
|
||||
src/bindings/*.cpp
|
||||
|
||||
[portable]
|
||||
# The torch-free subset: the same list minus the CUDA kernels and minus
|
||||
# bindings/ext.cpp (Torch-only). Compiled as csrc_portable by the Vulkan build.
|
||||
src/kernels/bilagrid/*.cpp
|
||||
src/mesh/*.cpp
|
||||
src/engine/*.cpp
|
||||
src/data/*.cpp
|
||||
src/external/*.cpp
|
||||
+3
-3
@@ -16,12 +16,12 @@ the detail.
|
||||
|
||||
Authoritative documents that live next to their code rather than here:
|
||||
|
||||
- `spirulae_splat/splat/cuda/csrc/backend/README.md` — the backend seam design.
|
||||
- `spirulae_splat/splat/cuda/csrc/backend/vulkan/README.md` — the Vulkan
|
||||
- `src/backend/README.md` — the backend seam design.
|
||||
- `src/backend/vulkan/README.md` — the Vulkan
|
||||
backend in depth: device baseline, capability variants, memory model,
|
||||
atomics, Slang notes, kernel coverage. **The single most detailed document
|
||||
in the repo**; read it before touching Vulkan code.
|
||||
- `spirulae_splat/splat/cuda/csrc/app/README.md` — working notes on the
|
||||
- `src/app/README.md` — working notes on the
|
||||
standalone CLI and the native GUI.
|
||||
- `viewer/README.md` — the standalone WebGL2/WASM viewer.
|
||||
- `scripts/README.md` — dataset preprocessing tools.
|
||||
|
||||
+21
-21
@@ -4,51 +4,51 @@
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
front ends │ ssplat-train (CLI) ssplat-gui (ImGui) spirulae-train │
|
||||
│ csrc/app/main.cpp csrc/app/gui/ Python package │
|
||||
front ends │ ssplat-train (CLI) ssplat-gui (ImGui) Python package │
|
||||
│ src/app/cli/ src/app/gui/ spirulae_splat/│
|
||||
└───────────┬────────────────┬──────────────────┬──────────┘
|
||||
│ │ │
|
||||
┌────────┴────────────────┴───┐ │ pybind11
|
||||
│ TrainerCore (csrc/app/) │ │ (csrc/ext.cpp)
|
||||
│ config → dataset → seeding │ │
|
||||
│ TrainerCore (src/app/) │ │ src/bindings/
|
||||
│ config → dataset → seeding │ │ ext.cpp
|
||||
│ → step loop │ │
|
||||
└──────────────┬──────────────┘ │
|
||||
│ │
|
||||
┌─────────────────┴─────────────────────────────┴──────────┐
|
||||
engine │ Engine*.cpp — torch-free, CUDA-free, process-global │
|
||||
│ state in EngineState.h; DataManager owns image I/O │
|
||||
engine │ src/engine/*.cpp — torch-free, CUDA-free, process-global │
|
||||
│ state in EngineState.h; src/data/ owns image I/O │
|
||||
└──────────────────────────┬───────────────────────────────┘
|
||||
│ launch declarations
|
||||
┌──────────────────────────┴───────────────────────────────┐
|
||||
backend seam │ csrc/backend/api/*.h + BackendRuntime.h │
|
||||
backend seam │ src/backend/api/*.h + BackendRuntime.h │
|
||||
└───────────┬──────────────────────────────┬───────────────┘
|
||||
│ │
|
||||
┌──────────────┴──────────┐ ┌────────────┴──────────────┐
|
||||
backends │ CUDA: csrc/*.cu │ │ Vulkan: backend/vulkan/ │
|
||||
│ + cuda/ins/*.cu │ │ SPIR-V from slang/vulkan │
|
||||
└──────────────┬──────────┘ └────────────┬──────────────┘
|
||||
┌──────────────┴───────────┐ ┌────────────┴──────────────┐
|
||||
backends │ CUDA: src/kernels/**/*.cu│ │ Vulkan: src/backend/vulkan│
|
||||
│ + src/instantiations/ │ │ launchers + shaders/ │
|
||||
└──────────────┬───────────┘ └────────────┬──────────────┘
|
||||
└──────────┬───────────────────┘
|
||||
│
|
||||
slang/*.slang — shared device math
|
||||
backend/common/SortScan.h — CUB / onesweep radix
|
||||
src/shaders/*.slang — shared device math
|
||||
src/backend/common/SortScan.h — CUB / onesweep radix
|
||||
```
|
||||
|
||||
The seam is the important part, and it is the best-documented part of the
|
||||
repo: read `csrc/backend/README.md`, then `csrc/backend/vulkan/README.md`.
|
||||
repo: read `src/backend/README.md`, then `src/backend/vulkan/README.md`.
|
||||
|
||||
## Where responsibilities live
|
||||
|
||||
| responsibility | code |
|
||||
|---|---|
|
||||
| training config (source of truth) | Python dataclasses in `spirulae_splat/modules/` → codegen → `app/generated/cli_config.h` |
|
||||
| dataset parsing (native) | `csrc/app/{Colmap,Nerfstudio,Metashape}Parser.cpp`, `DatasetCommon.cpp`, `DatasetParser.h` |
|
||||
| training config (source of truth) | Python dataclasses in `spirulae_splat/modules/` → codegen → `src/app/generated/cli_config.h` |
|
||||
| dataset parsing (native) | `src/data/parsers/{Colmap,Nerfstudio,Metashape}Parser.cpp`, `DatasetCommon.cpp`, `DatasetParser.h` |
|
||||
| dataset parsing (Python) | `spirulae_splat/modules/{dataparser,colmap_utils,metashape_utils,camera_utils}.py` — *duplicate; being retired* |
|
||||
| image cache / prefetch / warp | `csrc/DataManager.cpp` |
|
||||
| training session orchestration | `csrc/app/TrainerCore.cpp` and `spirulae_splat/modules/trainer.py` — *duplicate; being retired* |
|
||||
| image cache / prefetch / warp | `src/data/DataManager.cpp` |
|
||||
| training session orchestration | `src/app/TrainerCore.cpp` and `spirulae_splat/modules/trainer.py` — *duplicate; being retired* |
|
||||
| the actual training step | `Engine*.cpp`, entered via `engine_train_step_managed` |
|
||||
| kernels | `csrc/*.cu` (CUDA) + `slang/vulkan/*.slang` + `backend/vulkan/kernels/*.cpp` |
|
||||
| kernels | `src/kernels/**/*.cu` (CUDA) + `src/backend/vulkan/shaders/*.slang` + `src/backend/vulkan/kernels/*.cpp` |
|
||||
| web viewer client | `spirulae_splat/viewer/viewer.html` — single source; the C++ viewer embeds it at build time |
|
||||
| web viewer server | `csrc/app/{Viewer,HttpServer,RenderWorker}.cpp` and `spirulae_splat/viewer/*.py` — *duplicate; being retired* |
|
||||
| web viewer server | `src/app/webviewer/{Viewer,HttpServer,RenderWorker}.cpp` and `spirulae_splat/viewer/*.py` — *duplicate; being retired* |
|
||||
| standalone WASM viewer | `viewer/` — independent, but compiles the C++ parsers in place |
|
||||
|
||||
Three subsystems are currently implemented twice, once in Python and once in
|
||||
@@ -93,6 +93,6 @@ silently, so it is asserted only by parity tests.
|
||||
3DGS, Mip (anti-aliased 3DGS), and 3DGUT are compile-time *types*, not runtime
|
||||
branches: `Primitive.cuh`, `Primitive3DGS.cuh`, `Primitive3DGUT.cuh`,
|
||||
`PrimitiveBase3DGS.cuh`. Camera model, SH degree and distortion are value
|
||||
parameters. This is what makes `cuda/ins/` (111 generated instantiation TUs)
|
||||
parameters. This is what makes `src/instantiations/` (111 generated instantiation TUs)
|
||||
necessary on the CUDA side, and what maps to Slang specialization constants on
|
||||
the Vulkan side.
|
||||
|
||||
+9
-9
@@ -5,11 +5,11 @@ Two interchangeable compute backends behind one seam. **Both must work.**
|
||||
The authoritative documents live next to the code and are more detailed than
|
||||
this page — this is the orientation, they are the reference:
|
||||
|
||||
- **`spirulae_splat/splat/cuda/csrc/backend/README.md`** — the seam's design:
|
||||
- **`src/backend/README.md`** — the seam's design:
|
||||
what `api/` is, why the per-kernel `.cuh` headers are authoritative and the
|
||||
module headers are forwarders, and how `BackendRuntime.h` abstracts
|
||||
alloc/memcpy/memset/events.
|
||||
- **`spirulae_splat/splat/cuda/csrc/backend/vulkan/README.md`** — the Vulkan
|
||||
- **`src/backend/vulkan/README.md`** — the Vulkan
|
||||
backend in depth: verified premises, device baseline and per-capability
|
||||
pipeline variants, memory model, atomics strategy, sort/scan, Slang compiler
|
||||
notes, and **the kernel coverage table**. Read it before touching anything
|
||||
@@ -22,8 +22,8 @@ Engine*.cpp ──► backend/api/*.h (launch declarations)
|
||||
backend/api/BackendRuntime.h (alloc / memcpy / memset / events)
|
||||
│
|
||||
┌───────────────┴───────────────┐
|
||||
CUDA: csrc/*.cu Vulkan: backend/vulkan/
|
||||
+ cuda/ins/*.cu SPIR-V pipelines from slang/vulkan/
|
||||
CUDA: src/kernels/**/*.cu Vulkan: src/backend/vulkan/
|
||||
+ src/instantiations/ runtime + kernels/ + shaders/
|
||||
│ │
|
||||
└────► backend/common/SortScan.h ◄────┘
|
||||
CUB (CUDA) / onesweep radix + prefix sum (Vulkan)
|
||||
@@ -39,10 +39,10 @@ under `-DSSPLAT_BACKEND_VULKAN` with no CUDA toolkit installed.
|
||||
|
||||
## Shared device math
|
||||
|
||||
`slang/*.slang` is compiled **twice**: to CUDA headers (`csrc/generated/*.cuh`)
|
||||
`src/shaders/*.slang` is compiled **twice**: to CUDA headers (`src/generated/*.cuh`)
|
||||
and to SPIR-V. The projection, SH, primitive and PPISP math therefore has one
|
||||
source. `slang/vulkan/*.slang` adds the Vulkan-only compute entry points and
|
||||
the compatibility layers (`int64_compat`, `int8_compat`, `atomic_float`).
|
||||
source. `src/backend/vulkan/shaders/*.slang` adds the Vulkan-only compute entry
|
||||
points and the compatibility layers (`int64_compat`, `int8_compat`, `atomic_float`).
|
||||
|
||||
Don't fork device math per backend. If a backend needs different code, it
|
||||
belongs in a capability variant (below), not a copy.
|
||||
@@ -74,12 +74,12 @@ two-field struct, because one Intel Windows driver segfaults in
|
||||
|
||||
Three deliverables, always:
|
||||
|
||||
1. **CUDA**: `csrc/<Kernel>.cu` (+ `_kernel.cuh` if the device body is shared
|
||||
1. **CUDA**: `src/kernels/<family>/<Kernel>.cu` (+ `_kernel.cuh` if the device body is shared
|
||||
between fwd/bwd or eval3d/non-eval3d), with
|
||||
`/*[AutoHeaderGeneratorExport]*/` on the launcher. Rerun
|
||||
`generate_headers.py`, and `generate_kernel_instantiation.py` if it is
|
||||
templated over primitive / camera model / SH degree.
|
||||
2. **Vulkan**: the Slang entry point in `slang/vulkan/*.slang` and its
|
||||
2. **Vulkan**: the Slang entry point in `backend/vulkan/shaders/*.slang` and its
|
||||
launcher in `backend/vulkan/kernels/*.cpp`.
|
||||
3. **Parity test**: a tool in `backend/tests/` that builds under both backends
|
||||
and does `dump` / `compare`. See [testing.md](testing.md).
|
||||
|
||||
+35
-11
@@ -2,15 +2,38 @@
|
||||
|
||||
There are **two independent build systems**, by design and by history:
|
||||
|
||||
- `CMakeLists.txt` — every native artifact (engine, both backends, CLI, GUI,
|
||||
meshing tool, parity tests, and optionally the Python extension).
|
||||
- `CMakeLists.txt` + `cmake/` — every native artifact (engine, both backends,
|
||||
CLI, GUI, meshing tool, parity tests, and optionally the Python extension).
|
||||
- `setup.py` / `pyproject.toml` — the pip path, using torch's `CUDAExtension`.
|
||||
It globs its own source list and carries its own flags. **It does not read
|
||||
`CMakeLists.txt`**, so changes that touch compile flags or add source
|
||||
directories must be applied to both.
|
||||
It carries its own compile/link flags, deliberately: it targets a released
|
||||
wheel (its own `WITH_SYMBOLS` / `LINE_INFO` env gates), CMake targets local
|
||||
development. **Flag changes must be applied to both.**
|
||||
|
||||
They share exactly one thing: the source list, `cmake/sources.txt`. It is a
|
||||
plain list of glob patterns relative to the repo root, read line-by-line by
|
||||
`cmake/SsplatSources.cmake` and by `setup.py`'s `get_sources()` — plain text so
|
||||
the pip build never needs CMake installed. Adding a source directory means
|
||||
editing that file once.
|
||||
|
||||
For development, always use CMake via the dev scripts.
|
||||
|
||||
## CMake layout
|
||||
|
||||
`CMakeLists.txt` is just the running order; the logic lives in `cmake/`:
|
||||
|
||||
| file | what it does |
|
||||
|---|---|
|
||||
| `SsplatOptions.cmake` | options (`SSPLAT_BUILD_CLI/GUI`, `SSPLAT_NO_TORCH`, `SSPLAT_DEBUG_SYMBOLS`, …), backend selection, tree paths (`SSPLAT_ROOT`, `SSPLAT_CSRC`, …) |
|
||||
| `SsplatSources.cmake` | `ssplat_collect_sources()` — expands `sources.txt` with `CONFIGURE_DEPENDS` |
|
||||
| `SsplatBackendCuda.cmake` | Torch probe, CUDA arch detection, flags, the `csrc` library, CUDA-side parity tools |
|
||||
| `SsplatBackendVulkan.cmake` | portable engine object lib, slangc + SPIR-V embed, `ssplat_backend_vulkan`, the Vulkan-side tests |
|
||||
| `SsplatSlang.cmake` | `ssplat_find_slangc()` — pinned version, PATH lookup, fetch on miss |
|
||||
| `SsplatEmbed.cmake` | `ssplat_embed_file()` — bake a file into a byte-array header |
|
||||
| `SsplatApps.cmake` | `ssplat-train`, `ssplat-mesh`, `ssplat-gui` (backend-agnostic) |
|
||||
|
||||
Exactly one backend module runs. It leaves behind `SSPLAT_WITH_TORCH` and
|
||||
`SSPLAT_APP_LIBS`, which is the whole contract `SsplatApps.cmake` depends on.
|
||||
|
||||
## Dev builds
|
||||
|
||||
```bash
|
||||
@@ -72,9 +95,9 @@ exist *only* because of libtorch and are skipped in no-torch builds.
|
||||
|
||||
**Vulkan.** Needs the Vulkan SDK. No Python and no CUDA toolkit required.
|
||||
Slang is pinned to a specific version (`SSPLAT_SLANG_VERSION` in
|
||||
`CMakeLists.txt`); if the `slangc` on PATH doesn't match, CMake fetches the
|
||||
`cmake/SsplatSlang.cmake`); if the `slangc` on PATH doesn't match, CMake fetches the
|
||||
pinned release. SPIR-V blobs are **never committed** — they are compiled at
|
||||
build time (one `slangc` edge per blob, see `slang/SpirvShaders.cmake`) and
|
||||
build time (one `slangc` edge per blob, see `src/backend/vulkan/shaders/SpirvShaders.cmake`) and
|
||||
embedded into the binary. On an offline machine, transfer a matching `slangc`
|
||||
and point `-DSSPLAT_SLANGC=` at it.
|
||||
|
||||
@@ -89,7 +112,8 @@ suppress). Pass a trailing `-DSSPLAT_NO_TORCH=OFF` to try anyway.
|
||||
## Build-time cost
|
||||
|
||||
The long poles are the biggest `.cu` translation units and the CUDA template
|
||||
instantiation set in `cuda/ins/` (111 files). Splitting an oversized `.cu`
|
||||
using the `Name.Part.cu` convention (see [codegen.md](codegen.md)) both
|
||||
improves readability and parallelizes the build — no build-file edits needed,
|
||||
since CMake and `setup.py` both glob.
|
||||
instantiation set in `src/instantiations/` (111 files). Splitting an oversized `.cu`
|
||||
into function-named parts (see [codegen.md](codegen.md)) both improves
|
||||
readability and parallelizes the build — no build-file edits needed, since
|
||||
both builds glob `cmake/sources.txt`. Only `HEADER_SOURCES` in
|
||||
`generate_headers.py` has to learn the new file.
|
||||
|
||||
+23
-21
@@ -8,19 +8,19 @@ available and fall back to the committed files when it isn't.
|
||||
Generated locations:
|
||||
|
||||
```
|
||||
spirulae_splat/splat/cuda/csrc/generated/ device-math headers from Slang
|
||||
spirulae_splat/splat/cuda/csrc/app/generated/ cli_config.h, viewer_html.h
|
||||
spirulae_splat/splat/cuda/csrc/backend/api/ backend module forwarders
|
||||
spirulae_splat/splat/cuda/ins/ kernel instantiation TUs
|
||||
src/generated/ device-math headers from Slang
|
||||
src/app/generated/ cli_config.h, viewer_html.h
|
||||
src/backend/api/ backend module forwarders
|
||||
src/instantiations/ kernel instantiation TUs
|
||||
```
|
||||
|
||||
All generators run from the **repo root**:
|
||||
|
||||
```bash
|
||||
python3 spirulae_splat/generate_headers.py
|
||||
python3 spirulae_splat/generate_kernel_instantiation.py
|
||||
python3 spirulae_splat/generate_cli_config.py
|
||||
python3 spirulae_splat/generate_backend_api.py
|
||||
python3 tools/codegen/generate_headers.py
|
||||
python3 tools/codegen/generate_kernel_instantiation.py
|
||||
python3 tools/codegen/generate_cli_config.py
|
||||
python3 tools/codegen/generate_backend_api.py
|
||||
# generate_vulkan_stubs.py takes arguments; see below
|
||||
```
|
||||
|
||||
@@ -28,14 +28,14 @@ python3 spirulae_splat/generate_backend_api.py
|
||||
|
||||
## `generate_headers.py` — `.cu` → `.cuh` declarations
|
||||
|
||||
Scans `csrc/*.cu` for functions preceded by the marker
|
||||
Scans the `.cu` files listed in `HEADER_SOURCES` for functions preceded by the marker
|
||||
|
||||
```cpp
|
||||
/*[AutoHeaderGeneratorExport]*/
|
||||
void my_kernel_launcher(...) { ... }
|
||||
```
|
||||
|
||||
and rewrites the declaration section of `csrc/<Name>.cuh`.
|
||||
and rewrites the declaration section of the matching `<Name>.cuh`.
|
||||
|
||||
**Invariants:**
|
||||
|
||||
@@ -69,12 +69,13 @@ Because of invariant 3, splitting is nearly free:
|
||||
# put the shared preamble in <Family>Common.cuh and shared device helpers in a
|
||||
# descriptively named header (BilinearSample.cuh), then:
|
||||
# - add the new files to HEADER_SOURCES['PixelWise']
|
||||
python3 spirulae_splat/generate_headers.py # PixelWise.cuh's contents are unchanged
|
||||
python3 tools/codegen/generate_headers.py # PixelWise.cuh's contents are unchanged
|
||||
```
|
||||
|
||||
No build file changes: `CMakeLists.txt` globs `csrc/*.cu` with
|
||||
`CONFIGURE_DEPENDS`, and `setup.py` globs at build time. No caller changes:
|
||||
the header keeps its name and its declaration set.
|
||||
No build file changes: both build systems glob the kernel directories through
|
||||
`cmake/sources.txt` (CMake with `CONFIGURE_DEPENDS`, so a new file is picked up
|
||||
without a manual reconfigure). No caller changes: the header keeps its name and
|
||||
its declaration set.
|
||||
|
||||
Naming rules when splitting:
|
||||
|
||||
@@ -90,8 +91,9 @@ Naming rules when splitting:
|
||||
|
||||
## `generate_kernel_instantiation.py` — explicit template instantiations
|
||||
|
||||
Reads kernel declarations out of `csrc/*.cuh` and emits one small TU per
|
||||
(primitive × camera model × SH degree × …) combination into `cuda/ins/`
|
||||
Reads kernel declarations out of `src/kernels/**/*.cuh` and emits one small TU
|
||||
per (primitive × camera model × SH degree × …) combination into
|
||||
`src/instantiations/`
|
||||
(currently 111 files). This keeps each nvcc edge small and parallelizable
|
||||
rather than instantiating everything in one TU.
|
||||
|
||||
@@ -105,7 +107,7 @@ rather than instantiating everything in one TU.
|
||||
`OptimizerConfig`, plus the tyro preset subclasses.
|
||||
|
||||
The script parses them with `ast` (no torch import — it runs on a fresh
|
||||
checkout before `csrc.so` exists) and emits `csrc/app/generated/cli_config.h`:
|
||||
checkout before `csrc.so` exists) and emits `src/app/generated/cli_config.h`:
|
||||
|
||||
- `struct SsplatConfig` — every field, flattened, with the Python defaults
|
||||
baked in;
|
||||
@@ -128,7 +130,7 @@ Docstrings on the dataclass fields become CLI help text and GUI tooltips.
|
||||
|
||||
## `generate_backend_api.py` — backend module headers
|
||||
|
||||
Emits `csrc/backend/api/*.h` (`Projection.h`, `Rasterization.h`,
|
||||
Emits `src/backend/api/*.h` (`Projection.h`, `Rasterization.h`,
|
||||
`PixelWise.h`, …) as **forwarders** that `#include` the relevant per-kernel
|
||||
`.cuh` headers, grouped by subsystem.
|
||||
|
||||
@@ -141,7 +143,7 @@ and an individual `.cuh` without ODR violations.
|
||||
## `generate_vulkan_stubs.py` — link-probe for unported kernels
|
||||
|
||||
```bash
|
||||
python3 spirulae_splat/generate_vulkan_stubs.py <build-vulkan-dir> <output.cpp>
|
||||
python3 tools/codegen/generate_vulkan_stubs.py <build-vulkan-dir> <output.cpp>
|
||||
```
|
||||
|
||||
Link-probes `csrc_portable`'s objects against `libssplat_backend_vulkan.a`,
|
||||
@@ -166,6 +168,6 @@ Rerun it whenever the engine gains a new kernel call, **or when a port lands**
|
||||
Regenerated at configure time when the HTML changes. `Viewer.cpp` has a dev
|
||||
override so HTML edits don't require a rebuild.
|
||||
- Slang → SPIR-V blobs → `vk_shaders_embedded.cpp`
|
||||
(`slang/SpirvShaders.cmake` + `slang/spirv_tool.cpp`). SPIR-V is **never
|
||||
committed**; one `slangc` edge per blob, with capability variants
|
||||
(`backend/vulkan/shaders/SpirvShaders.cmake` +
|
||||
`backend/vulkan/shaders/spirv_tool.cpp`). SPIR-V is **never committed**; one `slangc` edge per blob, with capability variants
|
||||
(`.atomicadd`, `.noint64`, `.int8`) compiled per entry point.
|
||||
|
||||
+2
-2
@@ -3,7 +3,7 @@
|
||||
## Supported layouts
|
||||
|
||||
Three formats, auto-detected in this order: **Nerfstudio**, then **COLMAP**,
|
||||
then **Metashape**. Both the native parsers (`csrc/app/*Parser.cpp`) and the
|
||||
then **Metashape**. Both the native parsers (`src/data/parsers/*Parser.cpp`) and the
|
||||
Python parsers (`spirulae_splat/modules/dataparser.py`) probe in the same
|
||||
order.
|
||||
|
||||
@@ -21,7 +21,7 @@ The COLMAP reconstruction directory is auto-detected over
|
||||
|
||||
Perspective (with full radial / tangential / thin-prism distortion),
|
||||
equidistant and equisolid fisheye (including >180° FOV as produced by 360
|
||||
cameras), and equirectangular/spherical. See `csrc/CameraModel.h` — it is
|
||||
cameras), and equirectangular/spherical. See `src/core/CameraModel.h` — it is
|
||||
plain C++17 with no CUDA dependency, which is why the WASM viewer can reuse
|
||||
it.
|
||||
|
||||
|
||||
@@ -166,14 +166,14 @@ and the transform-backward math. The v2 revival is the moment to collapse this:
|
||||
1. **Selector component** — `BilagridBwdSelector` (contextual Gaussian TS
|
||||
ported from vksplat) + `SSPLAT_BILAGRID_BWD` override + unit test with
|
||||
synthetic timings. No kernel changes. **DONE** — header-only
|
||||
`csrc/BilagridBwdSelector.h`; host unit test
|
||||
`csrc/backend/tests/bilagrid_selector_test.cpp` (auto-globbed; also builds
|
||||
`kernels/bilagrid/BilagridBwdSelector.h`; host unit test
|
||||
`src/backend/tests/bilagrid_selector_test.cpp` (auto-globbed; also builds
|
||||
standalone) verifies forced-init coverage, contextual separation (two
|
||||
resolutions → two optima), elasticity across a mid-run reward flip, and all
|
||||
override forms. ~97% tail exploitation on the synthetic model.
|
||||
2. **Async timing plumbing** — per-dispatch `backend::Event` timing, one-iter
|
||||
readback, `(key, arm)` attribution at the backward hook. **DONE** —
|
||||
`csrc/BilagridBwdTiming.h` (`BwdTimingRing`) wired into
|
||||
`kernels/bilagrid/BilagridBwdTiming.h` (`BwdTimingRing`) wired into
|
||||
`_engine_bilagrid_backward_hook` (harvest at hook entry; begin/end around the
|
||||
RGB/depth/normal dispatches). Gated by `SSPLAT_BILAGRID_PROFILE=1` (logs
|
||||
per-family GPU ms; zero overhead when unset). Validated on RTX 4080S / CUDA:
|
||||
@@ -324,14 +324,14 @@ Two device-agnostic tunings support keeping v2 everywhere cheaply:
|
||||
|
||||
## 7. Files this touches
|
||||
|
||||
- `csrc/BilagridUniformSampleBwdV2_kernel.cuh` (+ per-family v2 headers) —
|
||||
- `kernels/bilagrid/BilagridUniformSampleBwdV2_kernel.cuh` (+ per-family v2 headers) —
|
||||
revive/generalize.
|
||||
- `csrc/Bilagrid*UniformSample.cu`, `csrc/EngineBilagrid.cpp` (backward hook) —
|
||||
- `kernels/bilagrid/Bilagrid*UniformSample.cu`, `engine/EngineBilagrid.cpp` (backward hook) —
|
||||
launch + selection wiring.
|
||||
- `csrc/BilagridConfig.cuh` — shared block/atomic helpers.
|
||||
- `slang/vulkan/bilagrid_{common,affine,ppisp,loglinear,depth,normal}.slang` —
|
||||
- `kernels/bilagrid/BilagridConfig.cuh` — shared block/atomic helpers.
|
||||
- `backend/vulkan/shaders/bilagrid_{common,affine,ppisp,loglinear,depth,normal}.slang` —
|
||||
add v2 entries.
|
||||
- `csrc/backend/vulkan/kernels/Bilagrid.cpp` — v2 dispatch in `launch_family_*`.
|
||||
- `csrc/backend/tests/bilagrid_parity.cpp` — per-arm parity sweep.
|
||||
- **New:** `csrc/BilagridBwdSelector.h` (the selector).
|
||||
- `src/backend/vulkan/kernels/Bilagrid.cpp` — v2 dispatch in `launch_family_*`.
|
||||
- `src/backend/tests/bilagrid_parity.cpp` — per-arm parity sweep.
|
||||
- **New:** `kernels/bilagrid/BilagridBwdSelector.h` (the selector).
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Repository restructure proposal
|
||||
|
||||
Status: **proposal**, nothing here has been applied yet.
|
||||
Status: phases 0-4 applied (see §8); §4 and phases 5-7 still proposal.
|
||||
|
||||
Goal: make the tree navigable for humans and agents, remove duplicated
|
||||
Python/C++ subsystems, split oversized files, and land durable documentation
|
||||
@@ -549,4 +549,103 @@ from drifting, and is the one I'd want a real regression run behind.
|
||||
`architecture`, `build`, `backends`, `codegen`, `datasets`, `testing`,
|
||||
`README` index, `notes/`. Written against the *current* layout, so they are
|
||||
correct today and get path updates in phase 4.
|
||||
- **Phase 2 — in progress.** File splits inside existing directories.
|
||||
- **Phase 2 — done.** File splits inside existing directories: `PixelWise.cu`
|
||||
(3668) → 7 function-named `.cu` + `PixelWiseCommon.cuh` + `BilinearSample.cuh`;
|
||||
`Optimizer.cu` (2655) → 8 `.cu` + `OptimizerCommon.cuh`; `Densify.cu` (2471)
|
||||
→ 5 `.cu` + 3 `.cuh`. `generate_headers.py` grew an explicit `HEADER_SOURCES`
|
||||
map so a source file's name is free of the header it feeds. Declaration sets
|
||||
of `PixelWise.cuh` / `Optimizer.cuh` / `Densify.cuh` verified unchanged
|
||||
(43 / 26 / 20).
|
||||
- **Phase 3 — done.** `CMakeLists.txt` (723 lines) → a 30-line running order
|
||||
plus `cmake/`: `SsplatOptions`, `SsplatSources`, `SsplatBackendCuda`,
|
||||
`SsplatBackendVulkan`, `SsplatSlang`, `SsplatEmbed`, `SsplatApps`. The
|
||||
backend `if/else` is now `include(SsplatBackendCuda|Vulkan)`, each leaving
|
||||
`SSPLAT_WITH_TORCH` + `SSPLAT_APP_LIBS` for the shared app targets; the
|
||||
duplicated per-app boilerplate (libpython, `-static-libstdc++`, `BUILD_RPATH`,
|
||||
include dirs) collapsed into `ssplat_configure_app()`, and the two
|
||||
copy-pasted hex-embed blocks into `ssplat_embed_file()`.
|
||||
|
||||
Source lists are no longer duplicated: `cmake/sources.txt` holds the glob
|
||||
patterns and is read by both `cmake/SsplatSources.cmake` and `setup.py`'s new
|
||||
`get_sources()`. Plain text rather than CMake so the pip build needs no CMake
|
||||
— full scikit-build-core delegation was rejected as too disruptive for users
|
||||
still on the pip path (§7.1).
|
||||
|
||||
Verified by diffing the generated `build.ninja` old-vs-new for three configs
|
||||
(torch CUDA / no-torch CUDA+GUI+tests / Vulkan): identical target sets,
|
||||
**zero** flag, define, or link-library differences; the only deltas are
|
||||
source ordering and one extra harmless `-I<builddir>` on `ssplat-mesh`.
|
||||
Both backends rebuild clean and the parity suite is 17/17.
|
||||
|
||||
- **Phase 4 — done (native half).** The directory move. All native code is now
|
||||
under a root `src/`, subdivided per §2: `core/ primitives/ kernels/{projection,
|
||||
raster,tile,pixelwise,ppisp,bilagrid,optim,densify,loss,background,visualize}/
|
||||
engine/ data/{,parsers/} mesh/ backend/ shaders/ app/{cli,gui,webviewer}/
|
||||
bindings/ generated/ instantiations/ external/`. `spirulae_splat/splat/cuda/`
|
||||
no longer holds native code. Also moved: the five codegen tools →
|
||||
`tools/codegen/`, `tests/*.py` + the stray `csrc/tests/test_delaunay3d.py` →
|
||||
`tests/python/`, `csrc/tests/delaunay3d_bench.cpp` → `tests/native/`.
|
||||
|
||||
**`src/` is now the include root** and every local include is path-qualified
|
||||
against it (`#include "core/Common.cuh"`) — 716 quoted plus 145 angle-bracket
|
||||
directives rewritten mechanically, none left relative. A file's include list
|
||||
now names the subsystems it depends on.
|
||||
|
||||
Build/codegen plumbing updated in the same pass: `cmake/sources.txt` gained
|
||||
`[cuda]` / `[portable]` sections (the Vulkan build's old flat `csrc/*.cpp`
|
||||
glob no longer expresses "the torch-free subset"), `SpirvShaders.cmake` is
|
||||
parameterised on the shaders dir, all five generators emit and read
|
||||
src-relative paths, `.gitattributes`, `viewer/CMakeLists.txt` (which compiles
|
||||
the parsers in place) and `build_develop.{bash,bat}` follow.
|
||||
`generate_kernel_instantiation.py` now slugs its output filenames from
|
||||
include *basenames*, so path-qualifying did not rename all 111 generated TUs.
|
||||
|
||||
Verified: every moved file's content diff is include lines only (checked
|
||||
blob-by-blob against HEAD, 52 residual lines, all include directives or the
|
||||
one comment naming a header); CUDA build, Vulkan build and Torch-extension
|
||||
build all green; parity 17/17 plus the three Vulkan smoke tests; all five
|
||||
generators re-run with zero drift.
|
||||
|
||||
**Deliberately deferred:** §2's reorganisation *inside* the Python package
|
||||
(`spirulae_splat/{cli,config,training,data,metrics,utils}`) and the
|
||||
`viewer/` → `web/` rename. The former changes public import paths
|
||||
(`spirulae_splat.splat.cuda`, `spirulae_splat.modules.*`) that `scripts/` and
|
||||
external users depend on, which per §7.1 should be paced; the latter belongs
|
||||
with the viewer unification in phase 6. So `spirulae_splat/splat/` still
|
||||
exists, holding only Python.
|
||||
|
||||
- **Phase 4b — backend symmetry, partial (done).** Reviewed whether the CUDA
|
||||
and Vulkan halves should mirror each other. Conclusion: make each backend
|
||||
*self-contained*, not make the two *look* like peers.
|
||||
|
||||
Done: `src/shaders/vulkan/` (39 files) → `src/backend/vulkan/shaders/`, and
|
||||
`SpirvShaders.cmake` + `spirv_tool.cpp` with it. The Vulkan backend was the
|
||||
only subsystem split across two top-level directories (host code under
|
||||
`backend/`, device code under `shaders/`); it is now one directory —
|
||||
runtime + kernels + shaders + its own SPIR-V build. `src/shaders/` now means
|
||||
exactly what it says: the 8 Slang files compiled *twice*, to
|
||||
`src/generated/*.cuh` and to SPIR-V. The Vulkan shaders reach the shared math
|
||||
by src-relative path (`#include "shaders/densify.slang"`, with `-I<src>`),
|
||||
matching the C++ convention and disambiguating `densify.slang`, which exists
|
||||
in both directories.
|
||||
|
||||
Rejected: mirroring the CUDA corpus under `src/backend/cuda/kernels/`. The
|
||||
`<Name>.cuh` declaration headers are the *backend-neutral contract* (Vulkan
|
||||
TUs include them; they parse with no CUDA toolkit) and they are generated
|
||||
from the `<Name>.cu` sitting next to them. A full mirror forces a choice
|
||||
between putting the neutral contract inside `backend/cuda/` (so Vulkan
|
||||
includes headers out of the CUDA backend — incoherent) and separating each
|
||||
`.cuh` from the `.cu` it is generated from (breaking the repo's most
|
||||
load-bearing codegen invariant). Symmetry is not worth either. The two sides
|
||||
are also not peers in fact: CUDA is the reference implementation the contract
|
||||
is *extracted from*, and a symmetric tree would assert a peerhood the codegen
|
||||
does not have.
|
||||
|
||||
Two bugs surfaced and were fixed on the way:
|
||||
- `backend/api/*.h` forwarders were emitting `#include "../../<Name>.cuh"`,
|
||||
stale since the phase-4 move. Nothing includes them yet, so nothing caught
|
||||
it. They are now src-relative and each one is compile-checked to parse
|
||||
under `-DSSPLAT_BACKEND_VULKAN` with no CUDA toolkit.
|
||||
- `build_develop.bash` always exited 0: the trailing `libcsrc.so` move
|
||||
masked the build's status, so a failed build reported green. Now
|
||||
propagated.
|
||||
|
||||
+4
-4
@@ -4,7 +4,7 @@ Three suites, in descending order of how much you should trust them.
|
||||
|
||||
## 1. Native cross-backend parity tests (the important ones)
|
||||
|
||||
`csrc/backend/tests/*.cpp` — currently 17 tools covering projection (fwd, bwd,
|
||||
`src/backend/tests/*.cpp` — currently 17 tools covering projection (fwd, bwd,
|
||||
quant-grad), rasterization bwd, tile intersect, warp, FPBO, optimizer (general
|
||||
+ geometry), densify, per-pixel train, PPISP, bilagrid, multi-scale loss, plus
|
||||
`backend/tests/engine/` which drives the *real* engine end to end
|
||||
@@ -68,7 +68,7 @@ If you script a Python training run that serves the viewer, pass
|
||||
## 3. Python tests (`tests/`)
|
||||
|
||||
```bash
|
||||
pytest tests/
|
||||
pytest tests/python/
|
||||
```
|
||||
|
||||
`test_per_pixel.py`, `test_ppisp.py`, `test_rasterization.py`,
|
||||
@@ -80,8 +80,8 @@ have not been run recently. Treat a failure here as "needs investigation",
|
||||
not as a release blocker — but do investigate, since they are the only
|
||||
remaining check on the `_torch_impl.py` reference math.
|
||||
|
||||
There is also `csrc/tests/test_delaunay3d.py` (a Python test living in the C++
|
||||
tree) and `csrc/tests/delaunay3d_bench.cpp`.
|
||||
There is also `tests/python/test_delaunay3d.py` and the standalone benchmark
|
||||
`tests/native/delaunay3d_bench.cpp` (not built by CMake; compile it directly).
|
||||
|
||||
## What to run before calling a change done
|
||||
|
||||
|
||||
@@ -38,6 +38,40 @@ def get_ext():
|
||||
return CustomBuildExtention.with_options(no_python_abi_suffix=True, use_ninja=True)
|
||||
|
||||
|
||||
def get_sources():
|
||||
"""Engine/kernel sources, from the list CMake also builds from.
|
||||
|
||||
cmake/sources.txt is the single source of truth: [section] headers, then
|
||||
one glob pattern per line relative to the repo root. This build wants the
|
||||
[cuda] section (everything the csrc library compiles). Keeping it plain
|
||||
text means the pip build does not need CMake installed. See
|
||||
cmake/SsplatSources.cmake for the other consumer.
|
||||
"""
|
||||
root = Path(__file__).parent
|
||||
spec = root / "cmake" / "sources.txt"
|
||||
|
||||
sources = []
|
||||
in_section = False
|
||||
for line in spec.read_text().splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#"):
|
||||
continue
|
||||
if line.startswith("[") and line.endswith("]"):
|
||||
in_section = line == "[cuda]"
|
||||
continue
|
||||
if in_section:
|
||||
sources += glob.glob(str(root / line))
|
||||
if not sources:
|
||||
raise RuntimeError(f"{spec} [cuda] matched no files")
|
||||
|
||||
# Relative to the repo root, as setuptools expects (it mirrors the source
|
||||
# path into build/temp.*/).
|
||||
sources = [os.path.relpath(s, root) for s in sources]
|
||||
|
||||
# Drop ROCm sources (neither build compiles them).
|
||||
return [s for s in sources if "hip" not in s]
|
||||
|
||||
|
||||
def get_extensions():
|
||||
import torch
|
||||
from torch.utils.cpp_extension import CUDAExtension
|
||||
@@ -50,13 +84,7 @@ def get_extensions():
|
||||
else:
|
||||
raise RuntimeError("CUDA is required for this extension.")
|
||||
|
||||
extensions_dir = Path("spirulae_splat/splat/cuda")
|
||||
sources = (
|
||||
glob.glob(str(extensions_dir / "ins" / "*.cu")) +
|
||||
glob.glob(str(extensions_dir / "csrc" / "*.cu")) +
|
||||
glob.glob(str(extensions_dir / "csrc" / "*.cpp"))
|
||||
)
|
||||
sources = [s for s in sources if "hip" not in s]
|
||||
sources = get_sources()
|
||||
|
||||
undef_macros = []
|
||||
define_macros = []
|
||||
@@ -135,10 +163,9 @@ def get_extensions():
|
||||
extension = CUDAExtension(
|
||||
"spirulae_splat.csrc",
|
||||
sources,
|
||||
include_dirs=[
|
||||
str((extensions_dir / "csrc").absolute()),
|
||||
str((extensions_dir / "csrc" / "glm").absolute()),
|
||||
],
|
||||
# src/ is the include root: local includes are path-qualified
|
||||
# relative to it (e.g. #include "core/Common.cuh").
|
||||
include_dirs=[str((Path(__file__).parent / "src").absolute())],
|
||||
define_macros=define_macros,
|
||||
undef_macros=undef_macros,
|
||||
extra_compile_args=extra_compile_args,
|
||||
|
||||
@@ -1,149 +0,0 @@
|
||||
import re
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def strip_if_zero_blocks(code):
|
||||
lines = code.split('\n')
|
||||
result = []
|
||||
depth = 0 # 0 = active, >0 = inside a #if 0 dead block
|
||||
for line in lines:
|
||||
stripped = line.strip()
|
||||
if depth == 0:
|
||||
if re.match(r'#\s*if\s+0\b', stripped):
|
||||
depth = 1
|
||||
else:
|
||||
result.append(line)
|
||||
else:
|
||||
if re.match(r'#\s*if', stripped):
|
||||
depth += 1
|
||||
elif re.match(r'#\s*endif', stripped):
|
||||
depth -= 1
|
||||
return '\n'.join(result)
|
||||
|
||||
|
||||
def extract_function_declarations(code):
|
||||
# Regex to match non-inline function declarations
|
||||
function_decl_pattern = re.compile(r"""
|
||||
\/\*\[AutoHeaderGeneratorExport\]\*\/\s*
|
||||
(.*?\))\s*\{
|
||||
""", re.MULTILINE | re.VERBOSE | re.DOTALL)
|
||||
|
||||
matches = function_decl_pattern.findall(code)
|
||||
decls = []
|
||||
|
||||
for m in matches:
|
||||
decls.append(m.strip()+';')
|
||||
|
||||
return decls
|
||||
|
||||
|
||||
def write_if_changed(path, new_text):
|
||||
old = None
|
||||
if os.path.exists(path):
|
||||
with open(path, "r") as f:
|
||||
old = f.read()
|
||||
if old != new_text:
|
||||
with open(path, "w") as f:
|
||||
f.write(new_text)
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def generate_header(filename, sources):
|
||||
path = "spirulae_splat/splat/cuda/csrc/"
|
||||
|
||||
# `sources` is listed in a fixed order, so the emitted declaration order is
|
||||
# stable across machines and the committed header does not churn.
|
||||
code = ""
|
||||
for source_filename in sources:
|
||||
full = path + source_filename
|
||||
if not os.path.exists(full):
|
||||
raise FileNotFoundError(
|
||||
f"{filename}.cuh lists a missing source: {source_filename}. "
|
||||
"Update HEADER_SOURCES in generate_headers.py."
|
||||
)
|
||||
code += open(full).read()
|
||||
decls = extract_function_declarations(strip_if_zero_blocks(code))
|
||||
|
||||
splitter = "/* == AUTO HEADER GENERATOR - DO NOT EDIT THIS LINE OR ANYTHING BELOW THIS LINE == */\n"
|
||||
include = open(path + f"{filename}.cuh").read()
|
||||
include = include.split(splitter)[0].strip()
|
||||
|
||||
header = '\n\n\n'.join([include, splitter]+decls) + "\n"
|
||||
cuh_path = path + f"{filename}.cuh"
|
||||
if write_if_changed(cuh_path, header):
|
||||
print("Generated", cuh_path)
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
# <header stem> -> the .cu translation units whose /*[AutoHeaderGeneratorExport]*/
|
||||
# functions it declares, in emission order.
|
||||
#
|
||||
# A family may be spread over several TUs, each named after what it *does*
|
||||
# rather than after the header it feeds (splitting oversized files is
|
||||
# encouraged -- see docs/codegen.md). The mapping is explicit rather than
|
||||
# inferred from filenames so that a source file's name is free, and so a
|
||||
# typo'd or moved file fails loudly instead of silently dropping declarations.
|
||||
HEADER_SOURCES = {
|
||||
'IntersectTile': ['IntersectTile.cu'],
|
||||
'BackgroundSphericalHarmonics': ['BackgroundSphericalHarmonics.cu'],
|
||||
'PerSplatLoss': ['PerSplatLoss.cu'],
|
||||
'PerPixelLoss': ['PerPixelLoss.cu'],
|
||||
# Image-space per-pixel operations, split by function.
|
||||
'PixelWise': [
|
||||
'ImageConvert.cu', # uint8/uint16 -> float; rendered -> expected depth
|
||||
'ImageColorOps.cu', # background blending, log map, overexposure reg
|
||||
'DepthGeometry.cu', # depth -> points/normal, depth-normal loss,
|
||||
# ray <-> linear depth
|
||||
'ImageDistort.cu', # distort / undistort
|
||||
'ImageWarp.cu', # wide <-> pinhole warps, incl. byte-fused
|
||||
'GtDepthNormalWarp.cu', # GT depth/normal wide -> pinhole warps
|
||||
'Ppisp.cu', # per-pixel image signal processing
|
||||
],
|
||||
'Projection': [], # types only; kernels live in the Fwd/Bwd headers
|
||||
'ProjectionFwd': ['ProjectionFwd.cu'],
|
||||
'ProjectionBwd': ['ProjectionBwd.cu'],
|
||||
'ProjectionBwdQuantGrad': ['ProjectionBwdQuantGrad.cu'],
|
||||
'ProjectionPackedFwd': ['ProjectionPackedFwd.cu'],
|
||||
'RasterizationFwd': ['RasterizationFwd.cu'],
|
||||
'RasterizationBwd': ['RasterizationBwd.cu'],
|
||||
'RasterizationEval3DFwd': ['RasterizationEval3DFwd.cu'],
|
||||
'RasterizationEval3DBwd': ['RasterizationEval3DBwd.cu'],
|
||||
# Optimizer variants, split by function.
|
||||
'Optimizer': [
|
||||
'TensorSetZero.cu', # set_zero_tensor
|
||||
'AdamOptim.cu', # Adam (incl. multi/stepped/8-bit/quat)
|
||||
'NewtonOptim.cu',
|
||||
'ScaleAgnosticMeanOptim.cu',
|
||||
'FusedGeometryOptim.cu',
|
||||
'FusedAppearanceOptim.cu',
|
||||
'TrustRegion3DGS2Optim.cu', # 3DGS^2-TR family
|
||||
'ColorOptim.cu', # linear-RGB / trust-region RGB + SH
|
||||
],
|
||||
'FusedProjectionBwdOptim': ['FusedProjectionBwdOptim.cu'],
|
||||
# Densification, split by function.
|
||||
'Densify': [
|
||||
'DensifySampling.cu', # quantile/median, indexing, scatter, sampling
|
||||
'DensifyScoring.cu', # covariance scale init, param update
|
||||
'Relocation.cu',
|
||||
'McmcRelocation.cu', # MCMC relocation + noise
|
||||
'DensifySplitFilter.cu', # long-axis split, image edge filters
|
||||
],
|
||||
'BilagridUtils': ['BilagridUtils.cu'],
|
||||
'Visualizer': ['Visualizer.cu'],
|
||||
'FusedSSIM': ['FusedSSIM.cu'],
|
||||
}
|
||||
|
||||
|
||||
def generate_headers():
|
||||
"""Regenerate headers only if needed."""
|
||||
num_generated = 0
|
||||
for name, sources in HEADER_SOURCES.items():
|
||||
if generate_header(name, sources):
|
||||
num_generated += 1
|
||||
print(f"Generated {num_generated}/{len(HEADER_SOURCES)} new headers")
|
||||
|
||||
|
||||
generate_headers()
|
||||
@@ -1,5 +1,5 @@
|
||||
"""NumPy host-side mirrors of the block quantization codecs in
|
||||
spirulae_splat/splat/cuda/csrc/Tensor.h, for adapting a checkpoint's quantized
|
||||
src/core/Tensor.h, for adapting a checkpoint's quantized
|
||||
buffers when resuming with a different layout (fewer Gaussians / different SH
|
||||
degree). Decode -> transform (gather / band-resample) -> re-encode with fresh
|
||||
per-block bounds. All host / CPU, vectorized (no per-element Python loops).
|
||||
|
||||
@@ -1,2 +0,0 @@
|
||||
#define IS_EVAL3D 0
|
||||
#include "RasterizationEval3DBwd_kernel.cuh"
|
||||
@@ -1,2 +0,0 @@
|
||||
#define IS_EVAL3D 0
|
||||
#include "RasterizationEval3DFwd_kernel.cuh"
|
||||
@@ -1,17 +0,0 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Splat projection forward/backward, packed variants, quantized-gradient backward, and the fused projection-backward + optimizer path.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
|
||||
#include "../../ProjectionFwd.cuh"
|
||||
#include "../../ProjectionBwd.cuh"
|
||||
#include "../../ProjectionPackedFwd.cuh"
|
||||
#include "../../ProjectionBwdQuantGrad.cuh"
|
||||
#include "../../FusedProjectionBwdOptim.cuh"
|
||||
@@ -1,17 +0,0 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Tile-based rasterization forward/backward (2D and eval-3D/3DGUT).
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
|
||||
#include "../../RasterizationFwd.cuh"
|
||||
#include "../../RasterizationBwd.cuh"
|
||||
#include "../../RasterizationEval3DFwd.cuh"
|
||||
#include "../../RasterizationEval3DBwd.cuh"
|
||||
#include "../../RasterizationMomentsFwd.cuh"
|
||||
@@ -1,6 +1,6 @@
|
||||
# Standalone CLI Trainer (`ssplat-train`) + Native GUI (`ssplat-gui`)
|
||||
|
||||
Working notes for the no-Python trainer under `csrc/app/`. Written 2026-07-10,
|
||||
Working notes for the no-Python trainer under `src/app/`. Written 2026-07-10,
|
||||
GUI added 2026-07-12; covers code structure, verification status, and
|
||||
remaining ports. Long-term context: Phase 1 (standalone C++ trainer) and
|
||||
Phase 2 (Dear ImGui/GLFW GUI) of the GUI-app plan are done; Phase 3 is
|
||||
@@ -37,7 +37,7 @@ system libs), `SSPLAT_BUILD_CLI` is forced ON, CUDA archs come from
|
||||
`nvidia-smi --query-gpu=compute_cap` (override with `-DTORCH_CUDA_ARCH_LIST`),
|
||||
and the libpython link + static-libstdc++/nftw interposition workarounds are
|
||||
skipped (they exist only because of libtorch). Generated headers
|
||||
(`csrc/generated/`, `app/generated/`, `cuda/ins/`) are committed, so a fresh
|
||||
(`src/generated/`, `app/generated/`, `src/instantiations/`) are committed, so a fresh
|
||||
checkout builds with no Python at all. With Torch present the extension build
|
||||
is unchanged (shared `libcsrc`, same flags as before).
|
||||
|
||||
@@ -166,7 +166,7 @@ HTTP/viewer.
|
||||
|
||||
## Config codegen (source of truth = Python dataclasses)
|
||||
|
||||
`spirulae_splat/generate_cli_config.py` parses `TrainerConfig` (+ preset
|
||||
`tools/codegen/generate_cli_config.py` parses `TrainerConfig` (+ preset
|
||||
subclasses) in `modules/trainer.py` and the nested
|
||||
`SpirulaeSplatDataParserConfig` / `SpirulaeSplatDataManagerConfig` /
|
||||
`SpirulaeSplatModelConfig` / `OptimizerConfig`. Field docstrings become help
|
||||
@@ -3,8 +3,8 @@
|
||||
// in app/README.md) and is numerically verified against it -- keep edits
|
||||
// in sync with the Python side.
|
||||
|
||||
#include "TrainerCore.h"
|
||||
#include "Knn.h"
|
||||
#include "app/TrainerCore.h"
|
||||
#include "data/Knn.h"
|
||||
|
||||
#ifndef _WIN32
|
||||
#include <ftw.h>
|
||||
@@ -23,10 +23,10 @@
|
||||
// setup_engine() calls engine_reset(), so a fresh session can follow a
|
||||
// finished one in the same process (the GUI's "train again" path).
|
||||
|
||||
#include "../Engine.h"
|
||||
#include "DatasetParser.h"
|
||||
#include "RenderWorker.h"
|
||||
#include "generated/cli_config.h"
|
||||
#include "engine/Engine.h"
|
||||
#include "data/DatasetParser.h"
|
||||
#include "app/webviewer/RenderWorker.h"
|
||||
#include "app/generated/cli_config.h"
|
||||
|
||||
#include <array>
|
||||
#include <atomic>
|
||||
@@ -15,8 +15,8 @@
|
||||
// adds only CLI parsing, --help, stdout progress printing, and the web
|
||||
// viewer wiring.
|
||||
|
||||
#include "TrainerCore.h"
|
||||
#include "Viewer.h"
|
||||
#include "app/TrainerCore.h"
|
||||
#include "app/webviewer/Viewer.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
@@ -13,10 +13,10 @@
|
||||
// the same C++ parsers the CLI trainer uses, and calls meshing::generate_mesh
|
||||
// (Meshing.h) for the heavy lifting.
|
||||
|
||||
#include "../Meshing.h"
|
||||
#include "../Camera.h"
|
||||
#include "DatasetParser.h"
|
||||
#include "Json.h"
|
||||
#include "mesh/Meshing.h"
|
||||
#include "core/Camera.h"
|
||||
#include "data/DatasetParser.h"
|
||||
#include "data/Json.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <cmath>
|
||||
Generated
Vendored
+1
-1
@@ -1,6 +1,6 @@
|
||||
#pragma once
|
||||
|
||||
// AUTO-GENERATED by spirulae_splat/generate_cli_config.py -- DO NOT EDIT.
|
||||
// AUTO-GENERATED by tools/codegen/generate_cli_config.py -- DO NOT EDIT.
|
||||
// Source of truth: the Python training config dataclasses
|
||||
// (TrainerConfig + presets, SpirulaeSplatDataParserConfig,
|
||||
// SpirulaeSplatDataManagerConfig, SpirulaeSplatModelConfig,
|
||||
+4
-4
@@ -1,13 +1,13 @@
|
||||
// ColmapRunner.cpp -- see ColmapRunner.h. CLI flags mirror
|
||||
// scripts/run_colmap.bash (COLMAP >= 4.x; use_gpu-style flags are gone).
|
||||
|
||||
#include "ColmapRunner.h"
|
||||
#include "FrameSelect.h"
|
||||
#include "Subprocess.h"
|
||||
#include "app/gui/ColmapRunner.h"
|
||||
#include "app/gui/FrameSelect.h"
|
||||
#include "app/gui/Subprocess.h"
|
||||
|
||||
#include "app_generated/mask_py.h" // kMaskPy[] (CMake-embedded scripts/mask.py)
|
||||
|
||||
#include "../../external/stb_image.h" // stbi_info (image size probe)
|
||||
#include "external/stb_image.h" // stbi_info (image size probe)
|
||||
|
||||
#ifndef _WIN32
|
||||
#include <ftw.h>
|
||||
@@ -3,7 +3,7 @@
|
||||
// special cases (the point is that new Python config fields show up here
|
||||
// with zero GUI work).
|
||||
|
||||
#include "ConfigUI.h"
|
||||
#include "app/gui/ConfigUI.h"
|
||||
|
||||
#include "imgui.h"
|
||||
#include "imgui_stdlib.h"
|
||||
@@ -13,7 +13,7 @@
|
||||
// New config fields added on the Python side appear here automatically after
|
||||
// re-running generate_cli_config.py -- no GUI change needed.
|
||||
|
||||
#include "../generated/cli_config.h"
|
||||
#include "app/generated/cli_config.h"
|
||||
|
||||
namespace gui {
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
// FileDialog.cpp -- see FileDialog.h.
|
||||
|
||||
#include "FileDialog.h"
|
||||
#include "app/gui/FileDialog.h"
|
||||
|
||||
#include "imgui.h"
|
||||
#include "imgui_stdlib.h"
|
||||
+2
-2
@@ -1,8 +1,8 @@
|
||||
// FrameSelect.cpp -- see FrameSelect.h.
|
||||
|
||||
#include "FrameSelect.h"
|
||||
#include "app/gui/FrameSelect.h"
|
||||
|
||||
#include "../../external/stb_image.h"
|
||||
#include "external/stb_image.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <chrono>
|
||||
@@ -1,6 +1,6 @@
|
||||
// GlLoader.cpp -- see GlLoader.h.
|
||||
|
||||
#include "GlLoader.h"
|
||||
#include "app/gui/GlLoader.h"
|
||||
|
||||
#include <GLFW/glfw3.h>
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
// GuiApp.cpp -- see GuiApp.h.
|
||||
|
||||
#include "GuiApp.h"
|
||||
#include "app/gui/GuiApp.h"
|
||||
|
||||
#include "../Json.h"
|
||||
#include "Subprocess.h"
|
||||
#include "data/Json.h"
|
||||
#include "app/gui/Subprocess.h"
|
||||
|
||||
#include "imgui.h"
|
||||
#include "imgui_stdlib.h"
|
||||
@@ -4,13 +4,13 @@
|
||||
// between the config editor, COLMAP runner, train runner, and the native
|
||||
// viewport.
|
||||
|
||||
#include "../../backend/api/BackendRuntime.h"
|
||||
#include "../generated/cli_config.h"
|
||||
#include "ColmapRunner.h"
|
||||
#include "ConfigUI.h"
|
||||
#include "FileDialog.h"
|
||||
#include "TrainRunner.h"
|
||||
#include "ViewportPanel.h"
|
||||
#include "backend/api/BackendRuntime.h"
|
||||
#include "app/generated/cli_config.h"
|
||||
#include "app/gui/ColmapRunner.h"
|
||||
#include "app/gui/ConfigUI.h"
|
||||
#include "app/gui/FileDialog.h"
|
||||
#include "app/gui/TrainRunner.h"
|
||||
#include "app/gui/ViewportPanel.h"
|
||||
|
||||
#include <deque>
|
||||
#include <string>
|
||||
@@ -2,7 +2,7 @@
|
||||
// core context (lowest version that also satisfies macOS) + Dear ImGui
|
||||
// bootstrap, then hands every frame to gui::GuiApp.
|
||||
|
||||
#include "GuiApp.h"
|
||||
#include "app/gui/GuiApp.h"
|
||||
|
||||
#include "imgui.h"
|
||||
#include "imgui_impl_glfw.h"
|
||||
@@ -1,7 +1,7 @@
|
||||
// NavCamera.cpp -- see NavCamera.h. Line-for-line port of viewer.html's
|
||||
// v3 / quat helpers, `cam`, and `Nav`.
|
||||
|
||||
#include "NavCamera.h"
|
||||
#include "app/gui/NavCamera.h"
|
||||
|
||||
#include <GLFW/glfw3.h>
|
||||
|
||||
+3
-3
@@ -1,9 +1,9 @@
|
||||
// PreviewRenderer.cpp -- see PreviewRenderer.h.
|
||||
|
||||
#include "PreviewRenderer.h"
|
||||
#include "app/gui/PreviewRenderer.h"
|
||||
|
||||
#include "GlLoader.h"
|
||||
#include "../TrainerCore.h"
|
||||
#include "app/gui/GlLoader.h"
|
||||
#include "app/TrainerCore.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <cmath>
|
||||
@@ -1,6 +1,6 @@
|
||||
// Subprocess.cpp -- see Subprocess.h.
|
||||
|
||||
#include "Subprocess.h"
|
||||
#include "app/gui/Subprocess.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstring>
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
// TrainRunner.cpp -- see TrainRunner.h.
|
||||
|
||||
#include "TrainRunner.h"
|
||||
#include "app/gui/TrainRunner.h"
|
||||
|
||||
#include <algorithm>
|
||||
|
||||
@@ -13,8 +13,8 @@
|
||||
// - engine_ready() flips true after engine setup; only then may a
|
||||
// RenderWorker attach.
|
||||
|
||||
#include "../TrainerCore.h"
|
||||
#include "../Viewer.h"
|
||||
#include "app/TrainerCore.h"
|
||||
#include "app/webviewer/Viewer.h"
|
||||
|
||||
#include <atomic>
|
||||
#include <deque>
|
||||
+4
-4
@@ -1,10 +1,10 @@
|
||||
// ViewportPanel.cpp -- see ViewportPanel.h.
|
||||
|
||||
#include "ViewportPanel.h"
|
||||
#include "app/gui/ViewportPanel.h"
|
||||
|
||||
#include "../TrainerCore.h"
|
||||
#include "ConfigUI.h" // help_tooltip_on_hover
|
||||
#include "GlLoader.h" // GL types + 1.1 entry points
|
||||
#include "app/TrainerCore.h"
|
||||
#include "app/gui/ConfigUI.h" // help_tooltip_on_hover
|
||||
#include "app/gui/GlLoader.h" // GL types + 1.1 entry points
|
||||
|
||||
#include "imgui.h"
|
||||
|
||||
+3
-3
@@ -14,9 +14,9 @@
|
||||
// GUI thread only. attach*() after the corresponding phase; detach() BEFORE
|
||||
// the session it renders from is destroyed.
|
||||
|
||||
#include "../RenderWorker.h"
|
||||
#include "NavCamera.h"
|
||||
#include "PreviewRenderer.h"
|
||||
#include "app/webviewer/RenderWorker.h"
|
||||
#include "app/gui/NavCamera.h"
|
||||
#include "app/gui/PreviewRenderer.h"
|
||||
|
||||
#include <cstdint>
|
||||
#include <string>
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
// HttpServer.cpp -- see HttpServer.h.
|
||||
|
||||
#include "HttpServer.h"
|
||||
#include "app/webviewer/HttpServer.h"
|
||||
|
||||
#include <cstdio>
|
||||
#include <cstring>
|
||||
+3
-3
@@ -2,10 +2,10 @@
|
||||
// render math is a verified port of trainer.py _render/render + model.py
|
||||
// get_outputs (viewer subset) + viewer/annotation.py.
|
||||
|
||||
#include "RenderWorker.h"
|
||||
#include "app/webviewer/RenderWorker.h"
|
||||
|
||||
#include "../Engine.h" // engine render + viewer entry points
|
||||
#include "../Camera.h" // camera_model_from_name
|
||||
#include "engine/Engine.h" // engine render + viewer entry points
|
||||
#include "core/Camera.h" // camera_model_from_name
|
||||
|
||||
#include <algorithm>
|
||||
#include <atomic>
|
||||
+1
-1
@@ -19,7 +19,7 @@
|
||||
// callers either block (wait_result, HTTP thread) or poll (try_get_result,
|
||||
// GUI frame loop).
|
||||
|
||||
#include "DatasetParser.h"
|
||||
#include "data/DatasetParser.h"
|
||||
|
||||
#include <array>
|
||||
#include <atomic>
|
||||
@@ -1,16 +1,16 @@
|
||||
// Viewer.cpp -- see Viewer.h. HTTP endpoint handlers are ports of
|
||||
// viewer/http_server.py; the render path itself lives in RenderWorker.cpp.
|
||||
|
||||
#include "Viewer.h"
|
||||
#include "HttpServer.h"
|
||||
#include "app/webviewer/Viewer.h"
|
||||
#include "app/webviewer/HttpServer.h"
|
||||
|
||||
// Implementation lives in csrc/stb_image_write_impl.cpp (shared with the
|
||||
// Implementation lives in external/stb_image_write_impl.cpp (shared with the
|
||||
// mesh texture export); declarations only here.
|
||||
#include "../external/stb_image_write.h"
|
||||
#include "external/stb_image_write.h"
|
||||
|
||||
#include "app_generated/viewer_html.h" // kViewerHtml[] (CMake-embedded)
|
||||
|
||||
#include "../Camera.h" // camera_model_from_name
|
||||
#include "core/Camera.h" // camera_model_from_name
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdio>
|
||||
@@ -17,8 +17,8 @@
|
||||
// annotation blit) lives in RenderWorker.h -- shared with the native GUI
|
||||
// viewport. This file adds only the HTTP layer + JPEG encode.
|
||||
|
||||
#include "DatasetParser.h"
|
||||
#include "RenderWorker.h"
|
||||
#include "data/DatasetParser.h"
|
||||
#include "app/webviewer/RenderWorker.h"
|
||||
|
||||
#include <memory>
|
||||
#include <string>
|
||||
@@ -12,7 +12,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
|
||||
│
|
||||
┌───────────────┴───────────────┐
|
||||
backend/cuda/ backend/vulkan/
|
||||
(existing csrc/*.cu, (SPIR-V pipeline dispatch over a
|
||||
(existing src/kernels/**, (SPIR-V pipeline dispatch over a
|
||||
unchanged, stays canonical) VulkanGSPipeline-style device layer)
|
||||
│ │
|
||||
└────────► backend/common/SortScan.h ◄────────┘
|
||||
@@ -25,7 +25,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
|
||||
- `Projection.h`, `Rasterization.h`, `TileIntersect.h`, `PixelWise.h`,
|
||||
`Optimizer.h`, `Densify.h`, `Loss.h`, `Background.h`, `Bilagrid.h`,
|
||||
`Visualizer.h` — **generated forwarders**
|
||||
(`spirulae_splat/generate_backend_api.py`) that group the per-kernel
|
||||
(`tools/codegen/generate_backend_api.py`) that group the per-kernel
|
||||
launch headers per subsystem. The AUTHORITATIVE declarations and types
|
||||
live in the per-kernel `.cuh` headers (declaration sections generated by
|
||||
`generate_headers.py`), which are CUDA-include-free and parse under
|
||||
@@ -88,7 +88,7 @@ Engine*.cpp ──► backend/api/*.h (launch declarations, generated
|
||||
executable. It becomes a real build as backend/vulkan/ fills in (step 5).
|
||||
5. **backend/vulkan/** implements `BackendRuntime.h`, `SortScan.h`, then the
|
||||
api modules kernel-by-kernel (Slang `-target spirv` entry points reusing
|
||||
the existing `slang/*.slang` device functions), in the phase order:
|
||||
the existing `src/shaders/*.slang` device functions), in the phase order:
|
||||
projection → SH/background → tile intersect + sort → rasterization fwd →
|
||||
visualizer, then training (losses, backward, optimizer, densify).
|
||||
|
||||
+1
-1
@@ -170,5 +170,5 @@ float event_elapsed_ms(Event* start, Event* end); // requires enable_timing
|
||||
} // namespace backend
|
||||
|
||||
#ifndef SSPLAT_BACKEND_VULKAN
|
||||
#include "../cuda/BackendRuntimeCuda.h"
|
||||
#include "backend/cuda/BackendRuntimeCuda.h"
|
||||
#endif
|
||||
+4
-4
@@ -1,13 +1,13 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Background spherical-harmonics render forward/backward.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../BackgroundSphericalHarmonics.cuh"
|
||||
#include "kernels/background/BackgroundSphericalHarmonics.cuh"
|
||||
@@ -1,14 +1,14 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Bilateral-grid utilities and the sample/TV-loss/fused-Adam launcher family.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../BilagridUtils.cuh"
|
||||
#include "../../BilagridBindings.h"
|
||||
#include "kernels/bilagrid/BilagridUtils.cuh"
|
||||
#include "kernels/bilagrid/BilagridBindings.h"
|
||||
@@ -1,13 +1,13 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Densification: MCMC and long-axis-split add/relocate, edge filters, scatter utilities.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../Densify.cuh"
|
||||
#include "kernels/densify/Densify.cuh"
|
||||
@@ -1,15 +1,15 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Fused SSIM, multi-scale per-pixel losses, per-splat losses.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../FusedSSIM.cuh"
|
||||
#include "../../PerPixelLoss.cuh"
|
||||
#include "../../PerSplatLoss.cuh"
|
||||
#include "kernels/loss/FusedSSIM.cuh"
|
||||
#include "kernels/loss/PerPixelLoss.cuh"
|
||||
#include "kernels/loss/PerSplatLoss.cuh"
|
||||
+4
-4
@@ -1,13 +1,13 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Fused Adam / AdaGrad / Newton optimizer steps, quantized and offloaded variants.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../Optimizer.cuh"
|
||||
#include "kernels/optim/Optimizer.cuh"
|
||||
+4
-4
@@ -1,13 +1,13 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Per-pixel image ops: background blend, depth/normal transforms, warps, color-space, PPISP.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../PixelWise.cuh"
|
||||
#include "kernels/pixelwise/PixelWise.cuh"
|
||||
@@ -0,0 +1,17 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Splat projection forward/backward, packed variants, quantized-gradient backward, and the fused projection-backward + optimizer path.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "kernels/projection/ProjectionFwd.cuh"
|
||||
#include "kernels/projection/ProjectionBwd.cuh"
|
||||
#include "kernels/projection/ProjectionPackedFwd.cuh"
|
||||
#include "kernels/projection/ProjectionBwdQuantGrad.cuh"
|
||||
#include "kernels/optim/FusedProjectionBwdOptim.cuh"
|
||||
@@ -0,0 +1,17 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Tile-based rasterization forward/backward (2D and eval-3D/3DGUT).
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "kernels/raster/RasterizationFwd.cuh"
|
||||
#include "kernels/raster/RasterizationBwd.cuh"
|
||||
#include "kernels/raster/RasterizationEval3DFwd.cuh"
|
||||
#include "kernels/raster/RasterizationEval3DBwd.cuh"
|
||||
#include "kernels/raster/RasterizationMomentsFwd.cuh"
|
||||
+2
-2
@@ -11,6 +11,6 @@
|
||||
// BackendTypes.h / BackendRuntime.h, after which this include chain is valid
|
||||
// in non-CUDA builds unchanged.
|
||||
|
||||
#include "BackendTypes.h"
|
||||
#include "backend/api/BackendTypes.h"
|
||||
|
||||
#include "../../Primitive.cuh"
|
||||
#include "primitives/Primitive.cuh"
|
||||
+5
-5
@@ -1,14 +1,14 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Splat-to-tile intersection: key generation, radix sort, tile ranges.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../IntersectTile.cuh"
|
||||
#include "../../SplatTileIntersector.cuh"
|
||||
#include "kernels/tile/IntersectTile.cuh"
|
||||
#include "kernels/tile/SplatTileIntersector.cuh"
|
||||
+4
-4
@@ -1,13 +1,13 @@
|
||||
#pragma once
|
||||
// Backend API module — GENERATED forwarder by
|
||||
// spirulae_splat/generate_backend_api.py. DO NOT EDIT.
|
||||
// tools/codegen/generate_backend_api.py. DO NOT EDIT.
|
||||
//
|
||||
// Interactive viewer render / blit entry points.
|
||||
//
|
||||
// Authoritative declarations live in the per-kernel headers included below
|
||||
// (declaration sections generated from /*[AutoHeaderGeneratorExport]*/
|
||||
// markers by generate_headers.py; CUDA-include-free via BackendTypes.h).
|
||||
// Implementations: backend/cuda (the existing .cu files), backend/vulkan
|
||||
// (SPIR-V pipeline dispatch). See csrc/backend/README.md.
|
||||
// Implementations: the .cu files next to each header (CUDA) and
|
||||
// backend/vulkan/kernels/ (SPIR-V pipeline dispatch). See backend/README.md.
|
||||
|
||||
#include "../../Visualizer.cuh"
|
||||
#include "kernels/visualize/Visualizer.cuh"
|
||||
+1
-1
@@ -21,7 +21,7 @@
|
||||
//
|
||||
// Zero overhead and zero behavior change when SSPLAT_PROFILE is unset.
|
||||
|
||||
#include "../api/BackendRuntime.h"
|
||||
#include "backend/api/BackendRuntime.h"
|
||||
|
||||
#include <atomic>
|
||||
#include <chrono>
|
||||
+1
-1
@@ -34,7 +34,7 @@
|
||||
|
||||
#include <cstdint>
|
||||
|
||||
#include "../api/BackendRuntime.h"
|
||||
#include "backend/api/BackendRuntime.h"
|
||||
|
||||
namespace backend {
|
||||
|
||||
+1
-1
@@ -10,7 +10,7 @@
|
||||
#include <mutex>
|
||||
#include <unordered_map>
|
||||
|
||||
#include "../common/Profiler.h"
|
||||
#include "backend/common/Profiler.h"
|
||||
|
||||
namespace backend {
|
||||
|
||||
+3
-3
@@ -21,9 +21,9 @@
|
||||
// contribution far below the loose tolerance. Quantized packed cells
|
||||
// compare as integer codes with a +-1 allowance.
|
||||
|
||||
#include <BilagridBindings.h>
|
||||
#include <EngineInternal.h>
|
||||
#include <Tensor.h>
|
||||
#include <kernels/bilagrid/BilagridBindings.h>
|
||||
#include <engine/EngineInternal.h>
|
||||
#include <core/Tensor.h>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+1
-1
@@ -6,7 +6,7 @@
|
||||
//
|
||||
// Exit code 0 = all checks pass, 1 = a check failed.
|
||||
|
||||
#include "../../BilagridBwdSelector.h"
|
||||
#include "kernels/bilagrid/BilagridBwdSelector.h"
|
||||
|
||||
#include <cstdio>
|
||||
#include <random>
|
||||
+1
-1
@@ -26,7 +26,7 @@
|
||||
// kernel, dst = cur + i deterministic) covers many concurrent dst splats,
|
||||
// including neighbor-word masked atomic stores in the quantized layouts.
|
||||
|
||||
#include <Densify.cuh>
|
||||
#include <kernels/densify/Densify.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+4
-4
@@ -12,10 +12,10 @@
|
||||
// pixels). The blit's uint8 output allows |d| <= 2 per byte (fp rounding at
|
||||
// the byte quantization boundary) with the same violation cap.
|
||||
|
||||
#include <Engine.h>
|
||||
#include <EngineState.h>
|
||||
#include <PixelWise.cuh>
|
||||
#include <Visualizer.cuh>
|
||||
#include <engine/Engine.h>
|
||||
#include <engine/EngineState.h>
|
||||
#include <kernels/pixelwise/PixelWise.cuh>
|
||||
#include <kernels/visualize/Visualizer.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+3
-3
@@ -20,9 +20,9 @@
|
||||
// all optimizer steps (fed by atomically accumulated gradients).
|
||||
// Densification is disabled so the splat count stays fixed.
|
||||
|
||||
#include <Engine.h>
|
||||
#include <EngineState.h>
|
||||
#include <PixelWise.cuh>
|
||||
#include <engine/Engine.h>
|
||||
#include <engine/EngineState.h>
|
||||
#include <kernels/pixelwise/PixelWise.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+3
-3
@@ -12,9 +12,9 @@
|
||||
// N is a multiple of the 256-thread block (see projqgrad_parity note on the
|
||||
// CUDA kernel's tail-thread reads).
|
||||
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <FusedProjectionBwdOptim.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
#include <kernels/optim/FusedProjectionBwdOptim.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+2
-2
@@ -37,8 +37,8 @@
|
||||
// readouts on both backends, so every config runs the identical computation
|
||||
// twice and compares the second return value.
|
||||
|
||||
#include <PerPixelLoss.cuh>
|
||||
#include <Tensor.h>
|
||||
#include <kernels/loss/PerPixelLoss.cuh>
|
||||
#include <core/Tensor.h>
|
||||
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
+2
-2
@@ -16,8 +16,8 @@
|
||||
// arithmetic is contraction-sensitive (can differ CUDA vs Vulkan), so for
|
||||
// that mode only the deterministic count channel (accum.y) is compared.
|
||||
|
||||
#include <Optimizer.cuh>
|
||||
#include <Densify.cuh>
|
||||
#include <kernels/optim/Optimizer.cuh>
|
||||
#include <kernels/densify/Densify.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+1
-1
@@ -18,7 +18,7 @@
|
||||
// quant combination is excluded, as in fpbo_parity (u = g1/(sqrt_g2 + eps)
|
||||
// is unbounded for radii == 0 splats — true on CUDA too).
|
||||
|
||||
#include <Optimizer.cuh>
|
||||
#include <kernels/optim/Optimizer.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+5
-5
@@ -1,5 +1,5 @@
|
||||
// Backend parity tool for the PPISP image-transform + regularization launch
|
||||
// APIs (csrc/PixelWise.cuh PPISP subset), across all three param types
|
||||
// APIs (kernels/pixelwise/PixelWise.cuh PPISP subset), across all three param types
|
||||
// (original / rqs / no_crf) and cam_indices on/off. The SAME source builds
|
||||
// under both backends:
|
||||
//
|
||||
@@ -8,7 +8,7 @@
|
||||
//
|
||||
// Ref format: [nf tight floats] [nl loose floats].
|
||||
//
|
||||
// Both backends run the same canonical slang/ppisp.slang math (CUDA via the
|
||||
// Both backends run the same canonical shaders/ppisp.slang math (CUDA via the
|
||||
// 2026.2.1 emission, Vulkan via SPIR-V), so per-pixel outputs and per-image
|
||||
// gradients differ only by fast-math rounding -> tight channel. Atomically
|
||||
// accumulated buffers (v_ppisp_params from the pixel backward, the summed
|
||||
@@ -16,9 +16,9 @@
|
||||
// regularization BACKWARD is fed synthetic host-side raw_losses/v_losses so
|
||||
// its outputs stay deterministic (tight).
|
||||
|
||||
#include <PixelWise.cuh>
|
||||
#include <EngineInternal.h>
|
||||
#include <Tensor.h>
|
||||
#include <kernels/pixelwise/PixelWise.cuh>
|
||||
#include <engine/EngineInternal.h>
|
||||
#include <core/Tensor.h>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+3
-3
@@ -11,9 +11,9 @@
|
||||
// whose order differs between backends, so the comparison is tolerance-based
|
||||
// with a small violation-fraction cap.
|
||||
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <ProjectionBwd.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
#include <kernels/projection/ProjectionBwd.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+2
-2
@@ -10,8 +10,8 @@
|
||||
// based (fast-math exp/sqrt chains differ across compilers) with a small
|
||||
// allowance for borderline-cull flips, which change entire rows.
|
||||
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+3
-3
@@ -13,9 +13,9 @@
|
||||
// The register camera-loop accumulation is deterministic; codes compare with
|
||||
// a +-1 quantum tolerance (codec rounding + last-ulp bound differences).
|
||||
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <ProjectionBwdQuantGrad.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
#include <kernels/projection/ProjectionBwdQuantGrad.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+3
-3
@@ -11,9 +11,9 @@
|
||||
// contributions in a different order than CUDA's block-atomic scheme, so
|
||||
// comparison uses the usual tolerance + small violation-fraction cap.
|
||||
|
||||
#include <PixelWise.cuh>
|
||||
#include <BackgroundSphericalHarmonics.cuh>
|
||||
#include <EngineInternal.h>
|
||||
#include <kernels/pixelwise/PixelWise.cuh>
|
||||
#include <kernels/background/BackgroundSphericalHarmonics.cuh>
|
||||
#include <engine/EngineInternal.h>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+7
-7
@@ -11,13 +11,13 @@
|
||||
// whose order differs between backends (and between runs on CUDA), so the
|
||||
// comparison is tolerance-based with a small violation-fraction cap.
|
||||
|
||||
#include <IntersectTile.cuh>
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <RasterizationEval3DFwd.cuh>
|
||||
#include <RasterizationFwd.cuh>
|
||||
#include <RasterizationBwd.cuh>
|
||||
#include <RasterizationEval3DBwd.cuh>
|
||||
#include <kernels/tile/IntersectTile.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
#include <kernels/raster/RasterizationEval3DFwd.cuh>
|
||||
#include <kernels/raster/RasterizationFwd.cuh>
|
||||
#include <kernels/raster/RasterizationBwd.cuh>
|
||||
#include <kernels/raster/RasterizationEval3DBwd.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+6
-6
@@ -13,12 +13,12 @@
|
||||
// depth keys may sort differently across backends (last-ulp projection
|
||||
// divergence), so a small violation fraction is allowed.
|
||||
|
||||
#include <BackgroundSphericalHarmonics.cuh>
|
||||
#include <IntersectTile.cuh>
|
||||
#include <ProjectionFwd.cuh>
|
||||
#include <ProjectionPackedFwd.cuh>
|
||||
#include <RasterizationEval3DFwd.cuh>
|
||||
#include <RasterizationFwd.cuh>
|
||||
#include <kernels/background/BackgroundSphericalHarmonics.cuh>
|
||||
#include <kernels/tile/IntersectTile.cuh>
|
||||
#include <kernels/projection/ProjectionFwd.cuh>
|
||||
#include <kernels/projection/ProjectionPackedFwd.cuh>
|
||||
#include <kernels/raster/RasterizationEval3DFwd.cuh>
|
||||
#include <kernels/raster/RasterizationFwd.cuh>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+4
-4
@@ -1,5 +1,5 @@
|
||||
// Backend parity tool for the dataset GT warp + byte->float conversion
|
||||
// launch APIs (csrc/PixelWise.cuh launch_warp_* + EngineInternal.h raw
|
||||
// launch APIs (kernels/pixelwise/PixelWise.cuh launch_warp_* + EngineInternal.h raw
|
||||
// converters). The SAME source builds under both backends:
|
||||
//
|
||||
// CUDA build: ./warp_parity dump ref.bin
|
||||
@@ -24,9 +24,9 @@
|
||||
// pixels), the equirectangular variants (x-wrap sampling), and the four
|
||||
// raw byte->float converters with odd element counts (word-tail handling).
|
||||
|
||||
#include <PixelWise.cuh>
|
||||
#include <EngineInternal.h>
|
||||
#include <Tensor.h>
|
||||
#include <kernels/pixelwise/PixelWise.cuh>
|
||||
#include <engine/EngineInternal.h>
|
||||
#include <core/Tensor.h>
|
||||
|
||||
#include <cmath>
|
||||
#include <cstdio>
|
||||
+28
-28
@@ -12,7 +12,7 @@ reproduced.
|
||||
## Verified premises (probed 2026-07 with slang-2026.2.1, SDK 1.3.296;
|
||||
## Vulkan backend now pins slang-2026.12.0.1 — see "Slang compiler notes")
|
||||
|
||||
- The existing `slang/*.slang` device functions compile to SPIR-V unchanged:
|
||||
- The existing `src/shaders/*.slang` device functions compile to SPIR-V unchanged:
|
||||
a compute entry point that `#include`s `primitive_3dgs.slang` and calls
|
||||
`projection_3dgs<cam, ...>` validates for Vulkan 1.2 requiring only
|
||||
`PhysicalStorageBufferAddresses` + `Int64` (+ `Shader`).
|
||||
@@ -35,14 +35,14 @@ for A/B testing):
|
||||
|
||||
- `VK_EXT_shader_atomic_float` (`shaderBufferFloat32AtomicAdd`) — training
|
||||
backward passes (~1,069 float atomicAdd sites). Every entry that calls
|
||||
`atomic_add_f32` is compiled twice by `slang/spirv_tool.cpp`: the base
|
||||
`atomic_add_f32` is compiled twice by `shaders/spirv_tool.cpp`: the base
|
||||
blob uses a CAS-loop emulation (MoltenVK, llvmpipe), and a
|
||||
`.atomicadd`-suffixed blob uses native `OpAtomicFAddEXT`; the pipeline
|
||||
layer picks per device at module load (no in-shader branch). Unlike
|
||||
VkSplat, both variants are always built.
|
||||
- `shaderInt64` — 64-bit sort keys, morton codes, large-buffer indexing.
|
||||
Entries in an int64_compat-including source get a `.noint64` variant
|
||||
compiled with `-DSSPLAT_EMULATE_INT64` (`slang/vulkan/int64_compat.slang`),
|
||||
compiled with `-DSSPLAT_EMULATE_INT64` (`backend/vulkan/shaders/int64_compat.slang`),
|
||||
and the embed step keeps only those whose base blob actually declares
|
||||
`OpCapability Int64` (the rest compile to a base-identical blob and are
|
||||
dropped). With emulation: index arithmetic narrows to int32
|
||||
@@ -61,7 +61,7 @@ for A/B testing):
|
||||
safe shape (`acc_add`/`acc_sub` in sort_scan.slang carry explicitly).
|
||||
- `shaderInt8` + `storageBuffer8BitAccess` — byte-buffer access. The
|
||||
baseline reads/writes uint8 buffers as packed u32 words
|
||||
(`slang/vulkan/int8_compat.slang`: `u8_load`/`s8_load`/`u8_store`; word
|
||||
(`backend/vulkan/shaders/int8_compat.slang`: `u8_load`/`s8_load`/`u8_store`; word
|
||||
reads past `n` stay in bounds because allocations are 16-byte rounded,
|
||||
and byte stores are an InterlockedAnd+Or pair). Entries that call those
|
||||
helpers get an `.int8` variant with native byte loads and plain byte
|
||||
@@ -162,12 +162,12 @@ Variant axes follow the CUDA instantiation structure
|
||||
|
||||
## Shaders: build and shipping
|
||||
|
||||
Entry points live in `slang/vulkan/*.slang` and `#include` the existing
|
||||
`slang/*.slang` module files (same trick as the CUDA side's
|
||||
Entry points live in `backend/vulkan/shaders/*.slang` and `#include` the existing
|
||||
`src/shaders/*.slang` module files (same trick as the CUDA side's
|
||||
`#define CudaDeviceExport ForceInline` include). SPIR-V is NEVER committed
|
||||
(unlike VkSplat): CMake compiles it at build time. The build is fully native
|
||||
(no Python): `slang/spirv_tool.cpp` — a self-contained C++17 host tool
|
||||
compiled once via `try_compile`, driven by `slang/SpirvShaders.cmake` —
|
||||
(no Python): `shaders/spirv_tool.cpp` — a self-contained C++17 host tool
|
||||
compiled once via `try_compile`, driven by `shaders/SpirvShaders.cmake` —
|
||||
`discover`s every blob (scanning entries + feature variants over the
|
||||
transitive #include closure), and CMake emits one slangc `-target spirv`
|
||||
custom command per blob so the build's `-j` bounds how many run at once
|
||||
@@ -270,7 +270,7 @@ Mesa ANV, llvmpipe — all tests + validation layers clean on all three):
|
||||
Phase-2 state (projection forward): `kernels/ProjectionFwd.cpp` +
|
||||
`kernels/ProjectionPackedFwd.cpp` implement the full `ProjectionFwd.cuh` /
|
||||
`ProjectionPackedFwd.cuh` APIs (fp32 SH) for all three primitives via
|
||||
`slang/vulkan/projection_fwd.slang` — the same `projection_3dgs<cam,...>` /
|
||||
`backend/vulkan/shaders/projection_fwd.slang` — the same `projection_3dgs<cam,...>` /
|
||||
`sh{N}_to_color` slang device functions the CUDA kernels use. Spec-constant
|
||||
axes: camera model, SH degree, antialiased (MIP), eval3d (3DGUT; its
|
||||
"conic" output lands in the screen scale slot, used by the eval3d
|
||||
@@ -287,7 +287,7 @@ configs, ~7.2M floats, zero tolerance violations on all three local devices
|
||||
(packed nnz and compaction order match exactly).
|
||||
|
||||
**SH value-quant (q8/q16)** is implemented WITHOUT any 8/16-bit device
|
||||
features: `slang/vulkan/sh_quant.slang` mirrors harmonics.slang's
|
||||
features: `backend/vulkan/shaders/sh_quant.slang` mirrors harmonics.slang's
|
||||
`_sh_load_q{8,16}` + `sh_coeffs_to_color_q{8,16}` math but reads the packed
|
||||
byte/short buffer as u32 words (the no-8-bit-types rule above), so quantized
|
||||
checkpoints render on the plain baseline. `kShValueBits` (0/8/16) is a new
|
||||
@@ -303,19 +303,19 @@ Phase-3 + raster-fwd state (full forward render): the whole
|
||||
projection -> tile-intersect -> rasterize -> background chain now runs on
|
||||
Vulkan.
|
||||
|
||||
- `kernels/BackgroundShFwd.cpp` + `slang/vulkan/background_sh.slang`:
|
||||
- `kernels/BackgroundShFwd.cpp` + `backend/vulkan/shaders/background_sh.slang`:
|
||||
per-pixel skybox via the same `generate_ray` / `transform_ray_d` /
|
||||
`sh{N}_to_color_dir` exports the CUDA kernel wraps. SH degree is a spec
|
||||
constant; camera model is a runtime int (as in CUDA). The backward throws
|
||||
until the training phase (`TODO` in the file).
|
||||
- `kernels/IntersectTile.cpp` + `slang/vulkan/intersect_tile.slang`:
|
||||
- `kernels/IntersectTile.cpp` + `backend/vulkan/shaders/intersect_tile.slang`:
|
||||
`do_intersect_tile_generic` = count -> `backend::inclusive_sum<int64>` ->
|
||||
8-byte n_isects readback -> key write -> `backend::sort_pairs<int64,int32>`
|
||||
(begin 0, end 32 + tile bits, exactly the CUB call) -> offset kernel.
|
||||
Ellipse-vs-AABB mode is a runtime null-check on proj_conic (the CUDA
|
||||
template bool only folds that check). `do_intersect_tile_post` has no
|
||||
callers and is not ported.
|
||||
- `kernels/RasterizeFwd.cpp` + `slang/vulkan/rasterize_fwd.slang`: closely
|
||||
- `kernels/RasterizeFwd.cpp` + `backend/vulkan/shaders/rasterize_fwd.slang`: closely
|
||||
follows RasterizationEval3DFwd_kernel.cuh (the CUDA source of both the 2D
|
||||
and eval3D kernels). Two entries (2D shared by 3DGS/MIP; 3DGUT eval3d
|
||||
reusing `evaluate_alpha_3dgs` / `evaluate_color_3dgs` /
|
||||
@@ -330,7 +330,7 @@ Vulkan.
|
||||
violations — near-tie depth keys sort differently when the projected
|
||||
depths differ in the last ulp across compilers, reordering the blend for
|
||||
a handful of pixels; the violation-fraction cap absorbs this.
|
||||
- the SPIR-V entry scanner (now `slang/spirv_tool.cpp`) tolerates `//`
|
||||
- the SPIR-V entry scanner (now `shaders/spirv_tool.cpp`) tolerates `//`
|
||||
comments between the [shader]/[numthreads] attributes and `void`.
|
||||
|
||||
Phase-4 completion + engine-level end-to-end (2026-07): the REAL engine
|
||||
@@ -338,12 +338,12 @@ render path (`forward_3dgs` -> background blend -> color space ->
|
||||
depth-normal -> `engine_blit_view`) runs on Vulkan, verified against CUDA at
|
||||
the engine level.
|
||||
|
||||
- `kernels/PixelWiseRender.cpp` + `slang/vulkan/pixel_wise_render.slang`:
|
||||
- `kernels/PixelWiseRender.cpp` + `backend/vulkan/shaders/pixel_wise_render.slang`:
|
||||
the render-path PixelWise subset — `blend_background_forward`,
|
||||
`blend_background_noise_forward` (hash_uint3 ported bit-exact),
|
||||
`rgb_to_srgb_forward`, `depth_to_normal_forward{,_tv}` (16x16 shared-mem
|
||||
apron tile as a flat 256-thread workgroup per the X-dim subgroup rule).
|
||||
- `kernels/Visualizer.cpp` + `slang/vulkan/visualizer.slang`: full
|
||||
- `kernels/Visualizer.cpp` + `backend/vulkan/shaders/visualizer.slang`: full
|
||||
visualizer — frustum geometry, LBVH build (morton keys over
|
||||
`backend::sort_pairs<int64,int32>` bits 0..63, Karras internal nodes,
|
||||
bottom-up AABBs), depth min/max, the annotate/colormap/overlay blit, the
|
||||
@@ -388,12 +388,12 @@ the engine level.
|
||||
(throwing, debug-sync hook), or_fallback, resolve_sh_quant. New kernel TUs
|
||||
should not hand-roll these.
|
||||
- **Optimizer port (phase 5, first slice)**: `kernels/Optimizer.cpp` +
|
||||
`slang/vulkan/optimizer.slang` implement fused_adam_step (fp32),
|
||||
`backend/vulkan/shaders/optimizer.slang` implement fused_adam_step (fp32),
|
||||
fused_adam_step_quantized (4/8-bit state), fused_adam_step_quantized_value
|
||||
(state x value x optional gradq int8 grad), fused_adagrad_step,
|
||||
float_add_into, increment_int32_inplace; `kernels/Densify.cpp` +
|
||||
`densify.slang` implement densify_update_weight. The quantized-state
|
||||
codecs live in `slang/vulkan/optim_quant.slang` (QuantizedAdamState /
|
||||
codecs live in `backend/vulkan/shaders/optim_quant.slang` (QuantizedAdamState /
|
||||
QuantizedTensor / gradq::Codec<8> mirrors): packed cells are read as u32
|
||||
words and written by cooperative word assembly in groupshared (a
|
||||
256-cell quant block is always word-aligned; CUDA leaves final-word tail
|
||||
@@ -500,12 +500,12 @@ the engine level.
|
||||
- **Densify/MCMC + remaining optimizer variants (phase 5, fifth slice)**:
|
||||
`kernels/Densify.cpp` implements the full densification API
|
||||
(densify_clip_scale, MCMC/revised noise, long-axis-split relocate/add,
|
||||
MCMC relocate/add) over `slang/vulkan/densify.slang`, which includes the
|
||||
canonical `slang/densify.slang` device code (long_axis_split_3dgs, the
|
||||
MCMC relocate/add) over `backend/vulkan/shaders/densify.slang`, which includes the
|
||||
canonical `shaders/densify.slang` device code (long_axis_split_3dgs, the
|
||||
randn3 noise kernels) and mirrors the McmcRelocation.cu-only mcmc_relocation
|
||||
math; `kernels/Optimizer.cpp` gains the adamtr trust-region color
|
||||
variants (`slang/vulkan/optim_color.slang`) and the fused 3DGS geometry
|
||||
step (`slang/vulkan/optim_geometry.slang`, sharing the FPBO helpers
|
||||
variants (`backend/vulkan/shaders/optim_color.slang`) and the fused 3DGS geometry
|
||||
step (`backend/vulkan/shaders/optim_geometry.slang`, sharing the FPBO helpers
|
||||
refactored into optim_quant.slang: oq_state_read/oq_accum_us/
|
||||
oq_state_encode/oq_sqrt3/oq_sqrt4). Design notes:
|
||||
- **Weighted sampling without replacement**: the efraimidis-spirakis
|
||||
@@ -548,14 +548,14 @@ the engine level.
|
||||
sh_value_bits 32/16/8 x bounds layouts x non-SH quant x degree
|
||||
warmup). All three devices pass; validation clean.
|
||||
- **Bilagrid/PPISP color pipeline (phase 5, sixth slice)**:
|
||||
`kernels/Bilagrid.cpp` implements the full csrc/BilagridBindings.h raw
|
||||
`kernels/Bilagrid.cpp` implements the full kernels/bilagrid/BilagridBindings.h raw
|
||||
launcher family (affine / PPISP / log-linear / depth / normal samplers:
|
||||
uniform + patched + per-pixel-coords + packed, forward + backward v1/v2)
|
||||
plus TV loss / channel mean and the EngineInternal.h fused
|
||||
TV-Adam/AdaGrad, q16 encode, identity init and scatter;
|
||||
`kernels/PpispImage.cpp` implements the PixelWise.cuh PPISP image
|
||||
transform + regularization APIs and the PpispInit.cu helpers. Device
|
||||
work: `slang/vulkan/bilagrid_{common,affine,ppisp,loglinear,depth,
|
||||
work: `backend/vulkan/shaders/bilagrid_{common,affine,ppisp,loglinear,depth,
|
||||
normal,tv,optim}.slang` + `ppisp_image.slang` (37 new entries). Design
|
||||
notes:
|
||||
- **BilagridReader** maps to three never-null pointer fields
|
||||
@@ -563,7 +563,7 @@ the engine level.
|
||||
`grid_indices == nullptr -> identity` semantic becomes a
|
||||
has_grid_indices flag field (llvmpipe speculation rule). Uniform vs
|
||||
patched layouts share one entry via a `kPatched` spec axis.
|
||||
- **PPISP math is the canonical slang** (`slang/ppisp.slang`) on both
|
||||
- **PPISP math is the canonical slang** (`shaders/ppisp.slang`) on both
|
||||
backends: the image transform + reg losses call the same
|
||||
apply_ppisp* / compute_*_loss* exports the CUDA build gets through
|
||||
generated/ppisp.cuh, so ppisp_parity's per-pixel values AND gradients
|
||||
@@ -620,7 +620,7 @@ the engine level.
|
||||
Gaussian-Thompson-sampling stats but making them per-key, so it needs no
|
||||
offline calibration pass (unlike fused-bilagrid) and does not average over
|
||||
resolutions (unlike vksplat's context-free bandit). Full design +
|
||||
phasing: `csrc/BilagridBackwardSelection.md`. When porting v2, the
|
||||
phasing: `docs/notes/bilagrid-backward-selection.md`. When porting v2, the
|
||||
`needs_image_grad` axis is a Slang spec-constant (false for depth/normal,
|
||||
which skip the image-grad scatter — the v2 analogue of the existing
|
||||
`v_in == nullptr` skip), and the grid-fold tail rule + null-`v_in` guard
|
||||
@@ -634,7 +634,7 @@ the engine level.
|
||||
for real (was a manual stub; the `<true>` instantiation is used by
|
||||
BilagridBindings). Device work in 4 new modules:
|
||||
- `multi_scale_loss.slang` includes the CANONICAL
|
||||
`slang/per_pixel_losses.slang` (CudaDeviceExport -> ForceInline, like
|
||||
`shaders/per_pixel_losses.slang` (CudaDeviceExport -> ForceInline, like
|
||||
ppisp), so the per-pixel loss math and its autodiff backward are the
|
||||
same functions CUDA emits -> per-pixel grads land tight. Bilinear
|
||||
GT-resolution sampling/scatter re-derived from Interpolation.cuh.
|
||||
@@ -771,7 +771,7 @@ the engine level.
|
||||
`--device <index|name substring>`; `ssplat-gui` has a Device combo in
|
||||
the train panel that locks once training starts.
|
||||
- **Training stubs**: `kernels/TrainingStubs.gen.cpp` (generated by
|
||||
`spirulae_splat/generate_vulkan_stubs.py` from a link probe of
|
||||
`tools/codegen/generate_vulkan_stubs.py` from a link probe of
|
||||
csrc_portable vs the backend lib) + `TrainingStubsManual.cpp`
|
||||
(meshing::OccupancyEvaluator methods) provide throwing
|
||||
definitions for every still-CUDA-only launch function, so the FULL
|
||||
+4
-4
@@ -1,12 +1,12 @@
|
||||
// Vulkan implementation of backend/common/SortScan.h over the kernels in
|
||||
// slang/vulkan/sort_scan.slang (embedded SPIR-V). Design notes in README.md.
|
||||
// shaders/sort_scan.slang (embedded SPIR-V). Design notes in README.md.
|
||||
//
|
||||
// Scratch is a global grow-only pool (mirroring DeviceScratch's single-
|
||||
// stream assumption): concurrent sort/scan calls on different streams are
|
||||
// not supported, matching how the engine uses these primitives today.
|
||||
|
||||
#include "../common/SortScan.h"
|
||||
#include "VulkanPipelines.h"
|
||||
#include "backend/common/SortScan.h"
|
||||
#include "backend/vulkan/VulkanPipelines.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstring>
|
||||
@@ -18,7 +18,7 @@ namespace {
|
||||
|
||||
using vk::SpecList;
|
||||
|
||||
// Must mirror slang/vulkan/sort_scan.slang.
|
||||
// Must mirror shaders/sort_scan.slang.
|
||||
constexpr uint32_t kSortPart = 2048;
|
||||
constexpr uint32_t kRadix = 256;
|
||||
constexpr uint32_t kScanBlock = 2048;
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user