vram: alias memory in forward and backward

This commit is contained in:
Harry Chen
2026-08-31 00:01:06 -04:00
parent 42761fdb85
commit 7f85d049c9
12 changed files with 393 additions and 66 deletions
+8
View File
@@ -527,6 +527,14 @@ no ceremony — do not ask, do not leave a note saying you removed it.
`kernels/bilagrid/BilagridConfig.cuh` for the 3D case) and pass the fold
factor to the kernel. Related: the bilagrid samplers index cells with
int32, which `engine_init_bilagrid_*` enforces up front.
- **Some pool buffers do not own their memory.** A slot listed in
`POOL_ALIAS_TABLE` (`core/PoolSlots.h`) is a slice of one arena that every
`PoolPhase` reuses, so the arena costs the largest phase rather than the sum
(626 MiB on a 15M-splat step). Acquiring one outside its phase throws;
`SS_POOL_ALIAS_POISON=1` fills the arena at every phase switch so a read
that outlives its phase becomes NaNs a parity test catches, and
`SS_POOL_ALIAS=0` turns the whole thing off. Read
`docs/notes/vram-splat-x-img.md` before adding a row.
- **`SS_PROFILE=1`** enables the per-stage backend timing breakdown
(H2D / D2H / D2D / memset / device / host), header-only, both backends, plus
a per-category VRAM breakdown after any run that trained. What the biggest
+68 -13
View File
@@ -101,30 +101,85 @@ Worth revisiting only together with:
cost scales with `n_isects` while the gather cost scales with `nnz`, so the
same change should come out ahead. Measure before believing it.
## The phase arena
`isect.ids_a/b` and the tile counts are allocated, used and dead inside
`do_intersect_tile_generic` -- no caller reads the key array it returns.
`raster_bwd.v_screen` does not exist until the backward, and is gone by the
next forward. They can never be live at once, so they share one allocation.
`POOL_ALIAS_TABLE` (`core/PoolSlots.h`) is the whole declaration: one row per
buffer, naming the `PoolPhase` its bytes are valid in. A phase-tagged slot
owns nothing; `DevicePool` carves it out of a single arena with a bump
allocator, and `pool_begin_phase()` resets the cursor. The arena therefore
costs the LARGEST phase instead of the sum:
Three paired art_gallery runs at `--cap-max 15000000` (the arena and the
intersection buffers both follow `n_isects`, which moves with the cameras a run
happens to sample, so pair the measurements):
| | alias off | alias on |
|---|---|---|
| pool + scratch, MiB | 6594.65 / 6537.27 / 6544.67 | 5971.80 / 6078.39 / 6021.85 |
| arena, MiB | - | 653.17 / 724.61 / 686.69 |
| GPU-exec, ms per 10 iters | 6256 / 5762 / 5923 | 5611 / 5942 / 6005 |
535 MiB off the pool on average (8.2%), for no measurable time: a bump
allocation is a pointer add, and the arena is reallocated only when a phase
asks for more than it holds. The saving is `raster_bwd.v_screen` in full --
the intersect phase is the larger of the two, so it sets the arena and the
gradient buffer rides along free.
### Adding a buffer to a phase
1. Add its row to `POOL_ALIAS_TABLE`. If the phase is new, add the enumerator
and the one `pool_begin_phase()` call that opens it -- put that call in the
launcher that allocates the phase's first buffer, not in the engine layer.
2. Run the parity suite with `SS_POOL_ALIAS_POISON=1`, which fills the arena
with 0xff at every phase switch. A buffer that is still being read turns
into NaNs the tests catch instead of the plausible stale data they would
not. A deliberately wrong row (tagging `isect.offsets`, which the raster
backward reads) fails `raster_bwd_parity` even without the poison, and
hangs with it -- that is the shape of the failure to expect.
3. `SS_POOL_ALIAS=0` turns the whole mechanism off (every slot owns its
allocation again). It changes memory layout and nothing else, so it is both
the A/B lever and the escape hatch if a lifetime claim turns out wrong.
### What the mechanism checks for you
- Acquiring a phase-tagged slot while a different phase is open throws, naming
the buffer and both phases. That is the ordering bug caught for free.
- A request the arena cannot fit yet takes a private allocation for that call
and raises the arena's target, so a growing `n_isects` degrades to today's
behavior for one step rather than overflowing.
- The VRAM report marks aliased rows `arena` in the cap column (they own
nothing) and prints the arena's one allocation as `of which alias arena`, so
the category caps and the pool total still add up.
### Candidates not yet taken
`raster_bwd.accum_weight` (114 MiB) would fit the `RasterBwd` phase and balance
it against the intersect phase, for a net ~53 MiB. It is not tagged because
the non-split training path can leave `engine().fwd.accum_weight` pointing at
the previous step's buffer when a step produces none -- harmless today (stale
but plausible numbers), garbage once the bytes are reused. Fix that stale view
the way `_engine_train_step_split_batch` already does, then tag it.
## Remaining opportunities, largest first
1. **Lifetime aliasing for phase-local scratch** (~550-1200 MiB). `isect.ids_a/b`
are allocated, sorted and dead inside `do_intersect_tile_generic` -- no
caller reads the returned key array. `raster_bwd.v_screen` is not allocated
until the backward. They can never be live at once, so one allocation could
serve both. The pool is one `device_malloc` per slot with no notion of
lifetime, so this needs a real change: an arena with explicit scoping, and a
fail-fast in-use flag rather than a comment, because getting a lifetime
wrong here corrupts silently. The viewer/eval render path shares these
slots, so it has to be part of the analysis.
2. **`proj.aabb` as 4 x int16** (16 -> 8 bytes per `nnz`, 115 MiB here). The
1. **`proj.aabb` as 4 x int16** (16 -> 8 bytes per `nnz`, 115 MiB here). The
projection already clamps the box to `[0, W-1] x [0, H-1]`, and image sizes
are now checked against 32767, so the values fit exactly; round outward
(floor the min, ceil the max) to stay conservative. Wide but mechanical:
the type is threaded through both projections, the intersector, three
backward launchers and the meshing raster in both backends, and the
backward passes only use it as an "is this pair valid" predicate.
3. **Packed screen rows** (up to 230 MiB). `opac` and `rgb` would survive fp16
2. **Packed screen rows** (up to 230 MiB). `opac` and `rgb` would survive fp16
comfortably; `xy`, `conic` and `depth` would not. The blocker is
`raster_bwd.v_screen`, which is accumulated with float atomics -- fp16
atomic add is not portable to the Vulkan baseline, so only the forward row
could shrink, and then the two buffers stop sharing a layout.
4. **Splitting one training image into tile bands** -- low priority. It bounds
3. **Splitting one training image into tile bands** -- low priority. It bounds
`n_isects` per launch rather than per image, which is the only lever that
helps the 4K case without touching precision. It does not help the
`nnz`-sized buffers at all, which is why it is below the three above.
`nnz`-sized buffers at all, which is why it is below the two above.
@@ -97,6 +97,11 @@ std::tuple<
const uint32_t total_count = (uint32_t)depths.numel();
const uint32_t N = packed ? total_count : total_count / I;
// Opens the tile-intersect phase: the key and count buffers below are
// scratch that dies here, so they share the arena with the raster
// backward (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::TileIsect);
intersect_check_image_size(image_width, image_height);
uint32_t tile_width = _CEIL_DIV(image_width, TILE_SIZE_IX);
uint32_t tile_height = _CEIL_DIV(image_height, TILE_SIZE_IY);
@@ -95,6 +95,9 @@ std::tuple<
const bool md = v_median.data_ptr() != nullptr;
const bool packed = gaussian_ids.data_ptr() != nullptr;
// Opens the raster-backward phase: the screen-gradient buffer below shares
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::RasterBwd);
if (!v_splats_w.has_value())
v_splats_w = Vanilla3DGS<0>::WorldBuffer::zeros_pool(
splats_w, PoolSlot::RasterBwdVWorld);
@@ -221,6 +224,9 @@ std::tuple<
const bool md = v_median.data_ptr() != nullptr;
const bool packed = gaussian_ids.data_ptr() != nullptr;
// Opens the raster-backward phase: the screen-gradient buffer below shares
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::RasterBwd);
if (!v_splats_w.has_value())
v_splats_w = Vanilla3DGUT<0>::WorldBuffer::zeros_pool(
splats_w, PoolSlot::RasterBwdVWorld);
+58
View File
@@ -453,6 +453,64 @@ constexpr const char* to_string(SaveClass s) {
}
// ---- Aliasing: phases that share one arena --------------------------------
// A slot listed in POOL_ALIAS_TABLE below owns no storage: every phase carves
// out of ONE arena, so it costs the largest phase and not the sum. Adding a
// row is a lifetime claim -- docs/notes/vram-splat-x-img.md says how to test it.
enum class PoolPhase : uint8_t {
None = 0, // owns its allocation -- every slot not listed below
TileIsect, // scratch that dies inside do_intersect_tile_generic
RasterBwd, // raster backward -> projection backward / fused optim step
Count
};
// X(PoolSlot enumerator, PoolPhase enumerator)
#define POOL_ALIAS_TABLE(X) \
/* The sort keys and tile counts: no caller reads the key array the \
intersector returns, and the counts never leave it. */ \
X(IsectIdsA , TileIsect) \
X(IsectIdsB , TileIsect) \
X(IsectTilesPerSplat , TileIsect) \
X(IsectCumTiles , TileIsect) \
/* Screen-space gradients: written by the raster backward, consumed by the \
projection backward, or stashed for the fused optim step -- all before \
the next forward's intersection. forward_3dgs drops the stashed view. */ \
X(RasterBwdVScreen , RasterBwd)
struct AliasRow { PoolSlot slot; PoolPhase phase; };
inline constexpr AliasRow kAliasRows[] = {
#define X(name, ph) AliasRow{ PoolSlot::name, PoolPhase::ph },
POOL_ALIAS_TABLE(X)
#undef X
};
constexpr PoolPhase slot_phase(PoolSlot s) {
for (const AliasRow& r : kAliasRows)
if (r.slot == s) return r.phase;
return PoolPhase::None;
}
constexpr bool ce_alias_rows_unique() {
const size_t n = sizeof(kAliasRows) / sizeof(kAliasRows[0]);
for (size_t i = 0; i < n; ++i)
for (size_t j = i + 1; j < n; ++j)
if (kAliasRows[i].slot == kAliasRows[j].slot) return false;
return true;
}
static_assert(ce_alias_rows_unique(),
"POOL_ALIAS_TABLE: a slot is listed twice");
constexpr const char* to_string(PoolPhase p) {
switch (p) {
case PoolPhase::None: return "none";
case PoolPhase::TileIsect: return "tile-isect";
case PoolPhase::RasterBwd: return "raster-bwd";
default: return "?";
}
}
// ---- Sub-index scheme -----------------------------------------------------
// A single logical slot may back a bounded number of physical allocations
// (a QuantizedTensor* needs packed bytes + per-block bounds). The pool is keyed
+196 -32
View File
@@ -153,19 +153,37 @@ private:
class DevicePool {
struct Slot {
void* ptr = nullptr;
size_t cap_bytes = 0; // allocated
size_t cap_bytes = 0; // allocated (0 when the arena owns it)
size_t used_bytes = 0; // logical (last requested)
uint8_t dtype_tag = (uint8_t)NpyScalar::u1; // scalar type of last acquire<T>
bool owns = true; // false: ptr is a slice of _arena
uint32_t epoch = 0; // phase instance this slice came from
};
struct DynSlot {
Slot slot;
VramCategory cat = VramCategory::Other;
PoolPhase phase = PoolPhase::None;
};
// The one allocation every PoolPhase carves its buffers out of.
// `want` is the largest cursor any phase has asked for since the last
// grow; growing happens at a phase switch, when nothing holds a slice.
struct Arena {
void* ptr = nullptr;
size_t cap = 0;
size_t cursor = 0;
size_t want = 0;
};
static constexpr size_t kArenaAlign = 256;
std::vector<Slot> _slots; // indexed by PoolKey
std::unordered_map<std::string, DynSlot> _dyn; // acquire_dynamic scratch
// Guards _slots / _dyn. Multiple threads (e.g., trainer + viewer
// post-processor) call acquire concurrently.
Arena _arena;
PoolPhase _phase = PoolPhase::None;
uint32_t _epoch = 0; // bumped per begin_phase; 0 never matches a Slot
// Guards _slots / _dyn / _arena against concurrent acquire. It does NOT
// make the phase protocol thread-safe: two threads in different phases
// would fight over the arena, which the engine mutex is what prevents.
mutable std::mutex _mu;
DevicePool() : _slots(kNumPoolKeys) {}
@@ -180,12 +198,34 @@ class DevicePool {
return on;
}
// SS_POOL_ALIAS=0 gives every slot its own allocation again; the kill
// switch for the arena, and the A/B lever for measuring what it saves.
static bool _alias_enabled() {
static const bool on = [] {
const char* e = spirula::env("POOL_ALIAS");
return !(e && e[0] == '0');
}();
return on;
}
// SS_POOL_ALIAS_POISON=1 fills the arena on every phase switch, so a read
// that outlives its phase produces NaNs instead of plausible stale data.
static bool _poison_arena() {
static const bool on = [] {
const char* e = spirula::env("POOL_ALIAS_POISON");
return e && e[0] && e[0] != '0';
}();
return on;
}
// Grow-if-needed on a Slot already selected by the caller (mutex held).
template<typename T>
static T* _acquire_into(Slot& slot, size_t n) {
size_t bytes = n * sizeof(T);
// An arena slice is not capacity this slot may reuse; drop it first.
if (!slot.owns) { slot.ptr = nullptr; slot.cap_bytes = 0; slot.owns = true; }
if (bytes > slot.cap_bytes) {
if (slot.ptr) backend::device_free(slot.ptr);
if (slot.ptr && slot.owns) backend::device_free(slot.ptr);
// Reset before the (throwing) allocation so an OOM leaves the slot
// in a consistent empty state rather than {null ptr, stale cap}.
slot.ptr = nullptr;
@@ -197,23 +237,104 @@ class DevicePool {
// zero and catastrophic where it is not. NaN makes it show.
if (_poison_pool()) backend::memset_sync(slot.ptr, 0xff, bytes);
}
slot.owns = true;
slot.used_bytes = bytes;
slot.dtype_tag = (uint8_t)npy_scalar_tag<T>();
return static_cast<T*>(slot.ptr);
}
// Bump-allocate out of the arena (mutex held). What does not fit yet gets
// a private allocation for this call and raises `want`, so the next phase
// switch grows the arena and the call after that is aliased again.
template<typename T>
T* _acquire_arena(Slot& slot, size_t n) {
// Each acquire gets a FRESH slice, so a second one in the same phase
// would silently move the buffer -- say so rather than hand it back.
if (slot.epoch == _epoch)
throw std::runtime_error(
"DevicePool: arena-backed buffer acquired twice in one phase; "
"resize it once, or split the phase");
slot.epoch = _epoch;
const size_t bytes = n * sizeof(T);
const size_t off = (_arena.cursor + kArenaAlign - 1) & ~(kArenaAlign - 1);
// The cursor advances even when the slice does not fit, or `want`
// would come out one buffer wide instead of one phase wide.
_arena.cursor = off + bytes;
if (_arena.want < _arena.cursor) _arena.want = _arena.cursor;
if (!_arena.ptr || off + bytes > _arena.cap)
return _acquire_into<T>(slot, n);
if (slot.ptr && slot.owns) backend::device_free(slot.ptr);
slot.ptr = static_cast<char*>(_arena.ptr) + off;
slot.cap_bytes = 0;
slot.used_bytes = bytes;
slot.dtype_tag = (uint8_t)npy_scalar_tag<T>();
slot.owns = false;
return static_cast<T*>(slot.ptr);
}
// The one place the phase rule is enforced (mutex held). `what` names the
// buffer for the error; enum slots and acquire_dynamic both come through.
template<typename T>
T* _acquire_checked(Slot& slot, size_t n, PoolPhase ph, const char* what) {
if (ph == PoolPhase::None || !_alias_enabled())
return _acquire_into<T>(slot, n);
if (ph != _phase)
throw std::runtime_error(
std::string("DevicePool: ") + what + " is arena-backed in phase "
+ to_string(ph) + " but the open phase is " + to_string(_phase)
+ " -- see POOL_ALIAS_TABLE in core/PoolSlots.h");
return _acquire_arena<T>(slot, n);
}
public:
static DevicePool& global() {
static DevicePool pool;
return pool;
}
// Open `p`. Every slice the previous phase carved out of the arena dies
// here -- POOL_ALIAS_TABLE's rows are the claim that nothing still reads
// one. Grows the arena first, which is safe for the same reason.
void begin_phase(PoolPhase p) {
std::lock_guard<std::mutex> lock(_mu);
_phase = p;
++_epoch;
if (!_alias_enabled() || p == PoolPhase::None) return;
if (_arena.want > _arena.cap) {
for (Slot& s : _slots)
if (!s.owns) { s.ptr = nullptr; s.owns = true; }
for (auto& kv : _dyn) {
Slot& s = kv.second.slot;
if (!s.owns) { s.ptr = nullptr; s.owns = true; }
}
if (_arena.ptr) backend::device_free(_arena.ptr);
_arena.ptr = nullptr;
_arena.cap = 0;
_arena.ptr = backend::device_malloc_checked(_arena.want);
_arena.cap = _arena.want;
}
_arena.cursor = 0;
if (_poison_arena() && _arena.ptr)
backend::memset_sync(_arena.ptr, 0xff, _arena.cap);
}
struct ArenaStats { size_t cap_bytes; size_t want_bytes; size_t slots; };
ArenaStats arenaStats() const {
std::lock_guard<std::mutex> lock(_mu);
size_t n = 0;
for (const Slot& s : _slots) if (!s.owns) ++n;
for (const auto& kv : _dyn) if (!kv.second.slot.owns) ++n;
return { _arena.cap, _arena.want, n };
}
// Return a pointer to at least `n` elements of type T for the given key.
// Reallocates only when current capacity is exceeded.
template<typename T>
T* acquire(PoolKey key, size_t n) {
std::lock_guard<std::mutex> lock(_mu);
return _acquire_into<T>(_slots[key], n);
const PoolSlot slot = pool_key_slot(key);
return _acquire_checked<T>(_slots[key], n, slot_phase(slot),
slot_name(slot));
}
// Convenience: acquire a (slot, sub) buffer. sub defaults to the main slot.
@@ -225,11 +346,13 @@ public:
// Escape hatch for Never-saved scratch whose key is built at runtime and so
// has no compile-time PoolSlot. `cat` places it in the VRAM report; `name`
// is the display/debug key. Returns raw bytes (caller casts).
void* acquire_dynamic(VramCategory cat, const std::string& name, size_t bytes) {
void* acquire_dynamic(VramCategory cat, const std::string& name,
size_t bytes, PoolPhase phase = PoolPhase::None) {
std::lock_guard<std::mutex> lock(_mu);
DynSlot& d = _dyn[name];
d.cat = cat;
return _acquire_into<uint8_t>(d.slot, bytes);
d.phase = phase;
return _acquire_checked<uint8_t>(d.slot, bytes, phase, name.c_str());
}
// Free a single logical buffer (all of its sub-allocations).
@@ -237,7 +360,8 @@ public:
std::lock_guard<std::mutex> lock(_mu);
for (uint32_t sub = 0; sub <= kSubMask; ++sub) {
Slot& s = _slots[pool_key(slot, sub)];
if (s.ptr) { backend::device_free(s.ptr); s = Slot{}; }
if (s.ptr && s.owns) backend::device_free(s.ptr);
s = Slot{};
}
}
@@ -254,7 +378,8 @@ public:
std::lock_guard<std::mutex> lock(_mu);
for (auto it = _dyn.begin(); it != _dyn.end();) {
if (it->first.rfind(prefix, 0) == 0) {
if (it->second.slot.ptr) backend::device_free(it->second.slot.ptr);
if (it->second.slot.ptr && it->second.slot.owns)
backend::device_free(it->second.slot.ptr);
it = _dyn.erase(it);
} else {
++it;
@@ -265,15 +390,23 @@ public:
// Free all managed device memory.
void freeAll() {
std::lock_guard<std::mutex> lock(_mu);
for (auto& s : _slots) if (s.ptr) { backend::device_free(s.ptr); s = Slot{}; }
for (auto& kv : _dyn) if (kv.second.slot.ptr) backend::device_free(kv.second.slot.ptr);
for (auto& s : _slots) {
if (s.ptr && s.owns) backend::device_free(s.ptr);
s = Slot{};
}
for (auto& kv : _dyn)
if (kv.second.slot.ptr && kv.second.slot.owns)
backend::device_free(kv.second.slot.ptr);
_dyn.clear();
if (_arena.ptr) backend::device_free(_arena.ptr);
_arena = Arena{};
_phase = PoolPhase::None;
}
// Total bytes allocated (capacity, not logical size).
size_t totalAllocBytes() const {
std::lock_guard<std::mutex> lock(_mu);
size_t total = 0;
size_t total = _arena.cap;
for (auto& s : _slots) total += s.cap_bytes;
for (auto& kv : _dyn) total += kv.second.slot.cap_bytes;
return total;
@@ -287,7 +420,7 @@ public:
std::vector<std::tuple<std::string, size_t, size_t>> result;
for (uint32_t i = 0; i < _slots.size(); ++i) {
const Slot& s = _slots[i];
if (!s.ptr && s.cap_bytes == 0) continue;
if (!s.ptr && s.cap_bytes == 0 && s.used_bytes == 0) continue;
std::string name = std::string(slot_name(pool_key_slot(i)))
+ sub_suffix(pool_key_sub(i));
result.emplace_back(std::move(name), s.used_bytes, s.cap_bytes);
@@ -298,24 +431,44 @@ public:
return result;
}
// One row per live buffer. `arena` means the bytes are a slice of the
// shared arena, so `cap` is 0 and the allocation is counted once in
// arenaStats() instead.
struct BreakdownRow {
std::string name;
VramCategory cat;
size_t used_bytes;
size_t cap_bytes;
bool arena;
};
std::vector<BreakdownRow> breakdown() const {
std::lock_guard<std::mutex> lock(_mu);
std::vector<BreakdownRow> result;
for (uint32_t i = 0; i < _slots.size(); ++i) {
const Slot& s = _slots[i];
if (!s.ptr && s.cap_bytes == 0 && s.used_bytes == 0) continue;
PoolSlot slot = pool_key_slot(i);
result.push_back({ std::string(slot_name(slot)) +
sub_suffix(pool_key_sub(i)),
slot_category(slot), s.used_bytes, s.cap_bytes,
!s.owns });
}
for (auto& kv : _dyn)
result.push_back({ kv.first, kv.second.cat,
kv.second.slot.used_bytes,
kv.second.slot.cap_bytes,
!kv.second.slot.owns });
return result;
}
// Categorized breakdown: [(name, category, used_bytes, cap_bytes), ...].
// category is (int)VramCategory. Used by the C++-driven VRAM report.
std::vector<std::tuple<std::string, int, size_t, size_t>>
getBreakdownCategorized() const {
std::lock_guard<std::mutex> lock(_mu);
std::vector<std::tuple<std::string, int, size_t, size_t>> result;
for (uint32_t i = 0; i < _slots.size(); ++i) {
const Slot& s = _slots[i];
if (!s.ptr && s.cap_bytes == 0) continue;
PoolSlot slot = pool_key_slot(i);
std::string name = std::string(slot_name(slot))
+ sub_suffix(pool_key_sub(i));
result.emplace_back(std::move(name), (int)slot_category(slot),
s.used_bytes, s.cap_bytes);
}
for (auto& kv : _dyn)
result.emplace_back(kv.first, (int)kv.second.cat,
kv.second.slot.used_bytes, kv.second.slot.cap_bytes);
for (auto& r : breakdown())
result.emplace_back(r.name, (int)r.cat, r.used_bytes, r.cap_bytes);
return result;
}
@@ -449,6 +602,14 @@ public:
};
// Open a pool phase; the arena's slices from the previous one die here.
// Called by the launcher that owns a phase, not by the engine layer -- see
// POOL_ALIAS_TABLE in core/PoolSlots.h.
inline void pool_begin_phase(PoolPhase p) {
DevicePool::global().begin_phase(p);
}
// Free all device memory managed by the pool and scratch buffer.
// Safe to call at VRAM pressure or before process exit.
inline void freeAllDeviceMemory() {
@@ -504,9 +665,10 @@ public:
// Dynamic-keyed resize for Never-class scratch whose name is built at
// runtime (prefix + index / suffix). Routes to the pool's side map with an
// explicit category. See DevicePool::acquire_dynamic.
void resize_dynamic(VramCategory cat, const std::string& name, int64_t numel) {
void resize_dynamic(VramCategory cat, const std::string& name, int64_t numel,
PoolPhase phase = PoolPhase::None) {
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
cat, name, (size_t)numel * sizeof(T));
cat, name, (size_t)numel * sizeof(T), phase);
_numel = numel;
}
@@ -591,9 +753,10 @@ public:
// Dynamic-keyed resize (Never-class scratch, runtime-built name).
void resize_dynamic(VramCategory cat, const std::string& name,
int64_t numel_0, int64_t numel_1) {
int64_t numel_0, int64_t numel_1,
PoolPhase phase = PoolPhase::None) {
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
cat, name, (size_t)(numel_0 * numel_1) * sizeof(T));
cat, name, (size_t)(numel_0 * numel_1) * sizeof(T), phase);
_numel_0 = numel_0;
_numel_1 = numel_1;
}
@@ -694,9 +857,10 @@ public:
// Dynamic-keyed resize (Never-class scratch, runtime-built name).
void resize_dynamic(VramCategory cat, const std::string& name,
int64_t numel_0, int64_t numel_1, int64_t numel_2) {
int64_t numel_0, int64_t numel_1, int64_t numel_2,
PoolPhase phase = PoolPhase::None) {
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
cat, name, (size_t)(numel_0 * numel_1 * numel_2) * sizeof(T));
cat, name, (size_t)(numel_0 * numel_1 * numel_2) * sizeof(T), phase);
_numel_0 = numel_0;
_numel_1 = numel_1;
_numel_2 = numel_2;
+5
View File
@@ -33,6 +33,11 @@ void forward_3dgs(
engine().sh_degree = sh_degree;
engine().packed = packed;
// The stashed screen gradients live in the arena the intersect below
// reuses, so this view is already dead. Dropping it turns a stale
// fused-optim read into that step's "v_splats_s not stashed" error.
engine().fwd.v_splats_s.clear();
// Build splats as DeviceTensorFloatND from typed device buffers.
// For DeviceVector<float> (opacities), build a [N, 1] shape FND manually.
DeviceTensorFloatND fnd_opac;
+29 -15
View File
@@ -260,17 +260,20 @@ size_t engine_get_scratch_bytes() {
// ===========================================================================
std::string engine_vram_report() {
auto rows = DevicePool::global().getBreakdownCategorized();
// Arena-backed rows own nothing, so their cap is 0 and the arena's one
// allocation is a line of its own; the totals still add up.
auto rows = DevicePool::global().breakdown();
auto arena = DevicePool::global().arenaStats();
const int kNCat = (int)VramCategory::Count;
size_t used[kNCat] = {}, cap[kNCat] = {}, count[kNCat] = {};
size_t used_total = 0, cap_total = 0;
for (const auto& r : rows) {
int c = std::get<1>(r);
int c = (int)r.cat;
count[c]++;
used[c] += std::get<2>(r);
cap[c] += std::get<3>(r);
used_total += std::get<2>(r);
cap_total += std::get<3>(r);
used[c] += r.used_bytes;
cap[c] += r.cap_bytes;
used_total += r.used_bytes;
cap_total += r.cap_bytes;
}
std::string out;
@@ -293,10 +296,13 @@ std::string engine_vram_report() {
}
size_t scratch = engine_get_scratch_bytes();
add("%-30s %11zu %11.2f %11.2f\n", " pool total", rows.size(),
mib(used_total), mib(cap_total));
mib(used_total), mib(cap_total + arena.cap_bytes));
if (arena.cap_bytes || arena.slots)
add("%-30s %11zu %11s %11.2f\n", " of which alias arena", arena.slots,
"-", mib(arena.cap_bytes));
add("%-30s %11s %11s %11.2f\n", " scratch buffer", "-", "-", mib(scratch));
add("%-30s %11s %11s %11.2f\n", " pool + scratch", "", "",
mib(cap_total + scratch));
mib(cap_total + arena.cap_bytes + scratch));
// The driver's figures place the pool against everything else the process
// holds: staging buffers, backend allocations, NN models, GUI surfaces.
@@ -308,16 +314,24 @@ std::string engine_vram_report() {
add("%-30s %11s %11s %11.2f / %.2f\n", " device in use / total", "",
"", mib(mem.used_bytes), mib(mem.total_bytes));
std::sort(rows.begin(), rows.end(), [](const auto& a, const auto& b) {
return std::get<3>(a) > std::get<3>(b);
// Aliased rows own nothing, so rank on whichever of the two is bigger.
auto weight = [](const auto& r) {
return std::max(r.used_bytes, r.cap_bytes);
};
std::sort(rows.begin(), rows.end(), [&](const auto& a, const auto& b) {
return weight(a) > weight(b);
});
add("%-30s %11s %11s %11s\n", "individual buffers", "category", "used_MiB",
"cap_MiB");
for (size_t i = 0; i < rows.size(); i++) {
if (std::get<3>(rows[i]) == 0) break;
add(" %-28s %11s %11.2f %11.2f\n", std::get<0>(rows[i]).c_str(),
to_string((VramCategory)std::get<1>(rows[i])),
mib(std::get<2>(rows[i])), mib(std::get<3>(rows[i])));
for (const auto& r : rows) {
if (weight(r) == 0) break;
const char* cat = to_string(r.cat);
if (r.arena)
add(" %-28s %11s %11.2f %11s\n", r.name.c_str(), cat,
mib(r.used_bytes), "arena");
else
add(" %-28s %11s %11.2f %11.2f\n", r.name.c_str(), cat,
mib(r.used_bytes), mib(r.cap_bytes));
}
add("[spirula-profile] --------------------------\n");
return out;
+3
View File
@@ -140,6 +140,9 @@ inline std::tuple<
std::optional<std::vector<DeviceTensorFloatND>> v_splats_w,
std::optional<std::vector<DeviceTensorFloatND>> v_splats_s
) {
// Opens the raster-backward phase: the screen-gradient buffer below shares
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::RasterBwd);
if (!v_splats_w.has_value())
v_splats_w = SplatPrimitive::WorldBuffer::zeros_pool(splats_w, PoolSlot::RasterBwdVWorld);
if (!v_splats_s.has_value())
@@ -188,6 +188,9 @@ inline std::tuple<
std::optional<std::vector<DeviceTensorFloatND>> v_splats_s,
bool need_viewmat_grad
) {
// Opens the raster-backward phase: the screen-gradient buffer below shares
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::RasterBwd);
if (!v_splats_w.has_value())
v_splats_w = SplatPrimitive::WorldBuffer::zeros_pool(splats_w, PoolSlot::RasterBwdVWorld);
if (!v_splats_s.has_value())
+5
View File
@@ -395,6 +395,11 @@ std::tuple<
const uint32_t total_count = (uint32_t)depths.numel();
const uint32_t N = packed ? total_count : total_count / I;
// Opens the tile-intersect phase: the key and count buffers below are
// scratch that dies here, so they share the arena with the raster
// backward (POOL_ALIAS_TABLE, core/PoolSlots.h).
pool_begin_phase(PoolPhase::TileIsect);
intersect_check_image_size(image_width, image_height);
uint32_t tile_width = _CEIL_DIV(image_width, TILE_SIZE_IX);
uint32_t tile_height = _CEIL_DIV(image_height, TILE_SIZE_IY);
+7 -6
View File
@@ -428,7 +428,7 @@ public:
DeviceTensor2D<float> b;
b.resize_dynamic(slot_category(key_prefix),
std::string(slot_name(key_prefix)) + ".0",
size, (int64_t)STRIDE);
size, (int64_t)STRIDE, slot_phase(key_prefix));
return { DeviceTensorFloatND(b) };
}
@@ -537,17 +537,18 @@ public:
int64_t size, std::array<int32_t, N> strides, PoolSlot key_prefix
) {
const VramCategory cat = slot_category(key_prefix);
const PoolPhase ph = slot_phase(key_prefix);
const std::string base = slot_name(key_prefix);
std::vector<DeviceTensorFloatND> res;
for (int i = 0; i < N; ++i) {
if (strides[i] <= 0) { res.push_back(DeviceTensorFloatND()); continue; }
std::string key = base + "." + std::to_string(i);
switch (strides[i]) {
case 1: { DeviceVector<float> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
case 2: { DeviceVector<float2> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
case 3: { DeviceVector<float3> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
case 4: { DeviceVector<float4> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
default: { DeviceTensor2D<float> b; b.resize_dynamic(cat, key, size, (int64_t)strides[i]); res.push_back(DeviceTensorFloatND(b)); break; }
case 1: { DeviceVector<float> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
case 2: { DeviceVector<float2> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
case 3: { DeviceVector<float3> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
case 4: { DeviceVector<float4> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
default: { DeviceTensor2D<float> b; b.resize_dynamic(cat, key, size, (int64_t)strides[i], ph); res.push_back(DeviceTensorFloatND(b)); break; }
}
}
return res;