mirror of
https://github.com/harry7557558/spirula-studio.git
synced 2026-10-02 02:44:54 +08:00
vram: alias memory in forward and backward
This commit is contained in:
@@ -527,6 +527,14 @@ no ceremony — do not ask, do not leave a note saying you removed it.
|
||||
`kernels/bilagrid/BilagridConfig.cuh` for the 3D case) and pass the fold
|
||||
factor to the kernel. Related: the bilagrid samplers index cells with
|
||||
int32, which `engine_init_bilagrid_*` enforces up front.
|
||||
- **Some pool buffers do not own their memory.** A slot listed in
|
||||
`POOL_ALIAS_TABLE` (`core/PoolSlots.h`) is a slice of one arena that every
|
||||
`PoolPhase` reuses, so the arena costs the largest phase rather than the sum
|
||||
(626 MiB on a 15M-splat step). Acquiring one outside its phase throws;
|
||||
`SS_POOL_ALIAS_POISON=1` fills the arena at every phase switch so a read
|
||||
that outlives its phase becomes NaNs a parity test catches, and
|
||||
`SS_POOL_ALIAS=0` turns the whole thing off. Read
|
||||
`docs/notes/vram-splat-x-img.md` before adding a row.
|
||||
- **`SS_PROFILE=1`** enables the per-stage backend timing breakdown
|
||||
(H2D / D2H / D2D / memset / device / host), header-only, both backends, plus
|
||||
a per-category VRAM breakdown after any run that trained. What the biggest
|
||||
|
||||
@@ -101,30 +101,85 @@ Worth revisiting only together with:
|
||||
cost scales with `n_isects` while the gather cost scales with `nnz`, so the
|
||||
same change should come out ahead. Measure before believing it.
|
||||
|
||||
## The phase arena
|
||||
|
||||
`isect.ids_a/b` and the tile counts are allocated, used and dead inside
|
||||
`do_intersect_tile_generic` -- no caller reads the key array it returns.
|
||||
`raster_bwd.v_screen` does not exist until the backward, and is gone by the
|
||||
next forward. They can never be live at once, so they share one allocation.
|
||||
|
||||
`POOL_ALIAS_TABLE` (`core/PoolSlots.h`) is the whole declaration: one row per
|
||||
buffer, naming the `PoolPhase` its bytes are valid in. A phase-tagged slot
|
||||
owns nothing; `DevicePool` carves it out of a single arena with a bump
|
||||
allocator, and `pool_begin_phase()` resets the cursor. The arena therefore
|
||||
costs the LARGEST phase instead of the sum:
|
||||
|
||||
Three paired art_gallery runs at `--cap-max 15000000` (the arena and the
|
||||
intersection buffers both follow `n_isects`, which moves with the cameras a run
|
||||
happens to sample, so pair the measurements):
|
||||
|
||||
| | alias off | alias on |
|
||||
|---|---|---|
|
||||
| pool + scratch, MiB | 6594.65 / 6537.27 / 6544.67 | 5971.80 / 6078.39 / 6021.85 |
|
||||
| arena, MiB | - | 653.17 / 724.61 / 686.69 |
|
||||
| GPU-exec, ms per 10 iters | 6256 / 5762 / 5923 | 5611 / 5942 / 6005 |
|
||||
|
||||
535 MiB off the pool on average (8.2%), for no measurable time: a bump
|
||||
allocation is a pointer add, and the arena is reallocated only when a phase
|
||||
asks for more than it holds. The saving is `raster_bwd.v_screen` in full --
|
||||
the intersect phase is the larger of the two, so it sets the arena and the
|
||||
gradient buffer rides along free.
|
||||
|
||||
### Adding a buffer to a phase
|
||||
|
||||
1. Add its row to `POOL_ALIAS_TABLE`. If the phase is new, add the enumerator
|
||||
and the one `pool_begin_phase()` call that opens it -- put that call in the
|
||||
launcher that allocates the phase's first buffer, not in the engine layer.
|
||||
2. Run the parity suite with `SS_POOL_ALIAS_POISON=1`, which fills the arena
|
||||
with 0xff at every phase switch. A buffer that is still being read turns
|
||||
into NaNs the tests catch instead of the plausible stale data they would
|
||||
not. A deliberately wrong row (tagging `isect.offsets`, which the raster
|
||||
backward reads) fails `raster_bwd_parity` even without the poison, and
|
||||
hangs with it -- that is the shape of the failure to expect.
|
||||
3. `SS_POOL_ALIAS=0` turns the whole mechanism off (every slot owns its
|
||||
allocation again). It changes memory layout and nothing else, so it is both
|
||||
the A/B lever and the escape hatch if a lifetime claim turns out wrong.
|
||||
|
||||
### What the mechanism checks for you
|
||||
|
||||
- Acquiring a phase-tagged slot while a different phase is open throws, naming
|
||||
the buffer and both phases. That is the ordering bug caught for free.
|
||||
- A request the arena cannot fit yet takes a private allocation for that call
|
||||
and raises the arena's target, so a growing `n_isects` degrades to today's
|
||||
behavior for one step rather than overflowing.
|
||||
- The VRAM report marks aliased rows `arena` in the cap column (they own
|
||||
nothing) and prints the arena's one allocation as `of which alias arena`, so
|
||||
the category caps and the pool total still add up.
|
||||
|
||||
### Candidates not yet taken
|
||||
|
||||
`raster_bwd.accum_weight` (114 MiB) would fit the `RasterBwd` phase and balance
|
||||
it against the intersect phase, for a net ~53 MiB. It is not tagged because
|
||||
the non-split training path can leave `engine().fwd.accum_weight` pointing at
|
||||
the previous step's buffer when a step produces none -- harmless today (stale
|
||||
but plausible numbers), garbage once the bytes are reused. Fix that stale view
|
||||
the way `_engine_train_step_split_batch` already does, then tag it.
|
||||
|
||||
## Remaining opportunities, largest first
|
||||
|
||||
1. **Lifetime aliasing for phase-local scratch** (~550-1200 MiB). `isect.ids_a/b`
|
||||
are allocated, sorted and dead inside `do_intersect_tile_generic` -- no
|
||||
caller reads the returned key array. `raster_bwd.v_screen` is not allocated
|
||||
until the backward. They can never be live at once, so one allocation could
|
||||
serve both. The pool is one `device_malloc` per slot with no notion of
|
||||
lifetime, so this needs a real change: an arena with explicit scoping, and a
|
||||
fail-fast in-use flag rather than a comment, because getting a lifetime
|
||||
wrong here corrupts silently. The viewer/eval render path shares these
|
||||
slots, so it has to be part of the analysis.
|
||||
2. **`proj.aabb` as 4 x int16** (16 -> 8 bytes per `nnz`, 115 MiB here). The
|
||||
1. **`proj.aabb` as 4 x int16** (16 -> 8 bytes per `nnz`, 115 MiB here). The
|
||||
projection already clamps the box to `[0, W-1] x [0, H-1]`, and image sizes
|
||||
are now checked against 32767, so the values fit exactly; round outward
|
||||
(floor the min, ceil the max) to stay conservative. Wide but mechanical:
|
||||
the type is threaded through both projections, the intersector, three
|
||||
backward launchers and the meshing raster in both backends, and the
|
||||
backward passes only use it as an "is this pair valid" predicate.
|
||||
3. **Packed screen rows** (up to 230 MiB). `opac` and `rgb` would survive fp16
|
||||
2. **Packed screen rows** (up to 230 MiB). `opac` and `rgb` would survive fp16
|
||||
comfortably; `xy`, `conic` and `depth` would not. The blocker is
|
||||
`raster_bwd.v_screen`, which is accumulated with float atomics -- fp16
|
||||
atomic add is not portable to the Vulkan baseline, so only the forward row
|
||||
could shrink, and then the two buffers stop sharing a layout.
|
||||
4. **Splitting one training image into tile bands** -- low priority. It bounds
|
||||
3. **Splitting one training image into tile bands** -- low priority. It bounds
|
||||
`n_isects` per launch rather than per image, which is the only lever that
|
||||
helps the 4K case without touching precision. It does not help the
|
||||
`nnz`-sized buffers at all, which is why it is below the three above.
|
||||
`nnz`-sized buffers at all, which is why it is below the two above.
|
||||
|
||||
@@ -97,6 +97,11 @@ std::tuple<
|
||||
const uint32_t total_count = (uint32_t)depths.numel();
|
||||
const uint32_t N = packed ? total_count : total_count / I;
|
||||
|
||||
// Opens the tile-intersect phase: the key and count buffers below are
|
||||
// scratch that dies here, so they share the arena with the raster
|
||||
// backward (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::TileIsect);
|
||||
|
||||
intersect_check_image_size(image_width, image_height);
|
||||
uint32_t tile_width = _CEIL_DIV(image_width, TILE_SIZE_IX);
|
||||
uint32_t tile_height = _CEIL_DIV(image_height, TILE_SIZE_IY);
|
||||
|
||||
@@ -95,6 +95,9 @@ std::tuple<
|
||||
const bool md = v_median.data_ptr() != nullptr;
|
||||
const bool packed = gaussian_ids.data_ptr() != nullptr;
|
||||
|
||||
// Opens the raster-backward phase: the screen-gradient buffer below shares
|
||||
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::RasterBwd);
|
||||
if (!v_splats_w.has_value())
|
||||
v_splats_w = Vanilla3DGS<0>::WorldBuffer::zeros_pool(
|
||||
splats_w, PoolSlot::RasterBwdVWorld);
|
||||
@@ -221,6 +224,9 @@ std::tuple<
|
||||
const bool md = v_median.data_ptr() != nullptr;
|
||||
const bool packed = gaussian_ids.data_ptr() != nullptr;
|
||||
|
||||
// Opens the raster-backward phase: the screen-gradient buffer below shares
|
||||
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::RasterBwd);
|
||||
if (!v_splats_w.has_value())
|
||||
v_splats_w = Vanilla3DGUT<0>::WorldBuffer::zeros_pool(
|
||||
splats_w, PoolSlot::RasterBwdVWorld);
|
||||
|
||||
@@ -453,6 +453,64 @@ constexpr const char* to_string(SaveClass s) {
|
||||
}
|
||||
|
||||
|
||||
// ---- Aliasing: phases that share one arena --------------------------------
|
||||
|
||||
// A slot listed in POOL_ALIAS_TABLE below owns no storage: every phase carves
|
||||
// out of ONE arena, so it costs the largest phase and not the sum. Adding a
|
||||
// row is a lifetime claim -- docs/notes/vram-splat-x-img.md says how to test it.
|
||||
enum class PoolPhase : uint8_t {
|
||||
None = 0, // owns its allocation -- every slot not listed below
|
||||
TileIsect, // scratch that dies inside do_intersect_tile_generic
|
||||
RasterBwd, // raster backward -> projection backward / fused optim step
|
||||
Count
|
||||
};
|
||||
|
||||
// X(PoolSlot enumerator, PoolPhase enumerator)
|
||||
#define POOL_ALIAS_TABLE(X) \
|
||||
/* The sort keys and tile counts: no caller reads the key array the \
|
||||
intersector returns, and the counts never leave it. */ \
|
||||
X(IsectIdsA , TileIsect) \
|
||||
X(IsectIdsB , TileIsect) \
|
||||
X(IsectTilesPerSplat , TileIsect) \
|
||||
X(IsectCumTiles , TileIsect) \
|
||||
/* Screen-space gradients: written by the raster backward, consumed by the \
|
||||
projection backward, or stashed for the fused optim step -- all before \
|
||||
the next forward's intersection. forward_3dgs drops the stashed view. */ \
|
||||
X(RasterBwdVScreen , RasterBwd)
|
||||
|
||||
struct AliasRow { PoolSlot slot; PoolPhase phase; };
|
||||
inline constexpr AliasRow kAliasRows[] = {
|
||||
#define X(name, ph) AliasRow{ PoolSlot::name, PoolPhase::ph },
|
||||
POOL_ALIAS_TABLE(X)
|
||||
#undef X
|
||||
};
|
||||
|
||||
constexpr PoolPhase slot_phase(PoolSlot s) {
|
||||
for (const AliasRow& r : kAliasRows)
|
||||
if (r.slot == s) return r.phase;
|
||||
return PoolPhase::None;
|
||||
}
|
||||
|
||||
constexpr bool ce_alias_rows_unique() {
|
||||
const size_t n = sizeof(kAliasRows) / sizeof(kAliasRows[0]);
|
||||
for (size_t i = 0; i < n; ++i)
|
||||
for (size_t j = i + 1; j < n; ++j)
|
||||
if (kAliasRows[i].slot == kAliasRows[j].slot) return false;
|
||||
return true;
|
||||
}
|
||||
static_assert(ce_alias_rows_unique(),
|
||||
"POOL_ALIAS_TABLE: a slot is listed twice");
|
||||
|
||||
constexpr const char* to_string(PoolPhase p) {
|
||||
switch (p) {
|
||||
case PoolPhase::None: return "none";
|
||||
case PoolPhase::TileIsect: return "tile-isect";
|
||||
case PoolPhase::RasterBwd: return "raster-bwd";
|
||||
default: return "?";
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
// ---- Sub-index scheme -----------------------------------------------------
|
||||
// A single logical slot may back a bounded number of physical allocations
|
||||
// (a QuantizedTensor* needs packed bytes + per-block bounds). The pool is keyed
|
||||
|
||||
+196
-32
@@ -153,19 +153,37 @@ private:
|
||||
class DevicePool {
|
||||
struct Slot {
|
||||
void* ptr = nullptr;
|
||||
size_t cap_bytes = 0; // allocated
|
||||
size_t cap_bytes = 0; // allocated (0 when the arena owns it)
|
||||
size_t used_bytes = 0; // logical (last requested)
|
||||
uint8_t dtype_tag = (uint8_t)NpyScalar::u1; // scalar type of last acquire<T>
|
||||
bool owns = true; // false: ptr is a slice of _arena
|
||||
uint32_t epoch = 0; // phase instance this slice came from
|
||||
};
|
||||
struct DynSlot {
|
||||
Slot slot;
|
||||
VramCategory cat = VramCategory::Other;
|
||||
PoolPhase phase = PoolPhase::None;
|
||||
};
|
||||
|
||||
// The one allocation every PoolPhase carves its buffers out of.
|
||||
// `want` is the largest cursor any phase has asked for since the last
|
||||
// grow; growing happens at a phase switch, when nothing holds a slice.
|
||||
struct Arena {
|
||||
void* ptr = nullptr;
|
||||
size_t cap = 0;
|
||||
size_t cursor = 0;
|
||||
size_t want = 0;
|
||||
};
|
||||
static constexpr size_t kArenaAlign = 256;
|
||||
|
||||
std::vector<Slot> _slots; // indexed by PoolKey
|
||||
std::unordered_map<std::string, DynSlot> _dyn; // acquire_dynamic scratch
|
||||
// Guards _slots / _dyn. Multiple threads (e.g., trainer + viewer
|
||||
// post-processor) call acquire concurrently.
|
||||
Arena _arena;
|
||||
PoolPhase _phase = PoolPhase::None;
|
||||
uint32_t _epoch = 0; // bumped per begin_phase; 0 never matches a Slot
|
||||
// Guards _slots / _dyn / _arena against concurrent acquire. It does NOT
|
||||
// make the phase protocol thread-safe: two threads in different phases
|
||||
// would fight over the arena, which the engine mutex is what prevents.
|
||||
mutable std::mutex _mu;
|
||||
|
||||
DevicePool() : _slots(kNumPoolKeys) {}
|
||||
@@ -180,12 +198,34 @@ class DevicePool {
|
||||
return on;
|
||||
}
|
||||
|
||||
// SS_POOL_ALIAS=0 gives every slot its own allocation again; the kill
|
||||
// switch for the arena, and the A/B lever for measuring what it saves.
|
||||
static bool _alias_enabled() {
|
||||
static const bool on = [] {
|
||||
const char* e = spirula::env("POOL_ALIAS");
|
||||
return !(e && e[0] == '0');
|
||||
}();
|
||||
return on;
|
||||
}
|
||||
|
||||
// SS_POOL_ALIAS_POISON=1 fills the arena on every phase switch, so a read
|
||||
// that outlives its phase produces NaNs instead of plausible stale data.
|
||||
static bool _poison_arena() {
|
||||
static const bool on = [] {
|
||||
const char* e = spirula::env("POOL_ALIAS_POISON");
|
||||
return e && e[0] && e[0] != '0';
|
||||
}();
|
||||
return on;
|
||||
}
|
||||
|
||||
// Grow-if-needed on a Slot already selected by the caller (mutex held).
|
||||
template<typename T>
|
||||
static T* _acquire_into(Slot& slot, size_t n) {
|
||||
size_t bytes = n * sizeof(T);
|
||||
// An arena slice is not capacity this slot may reuse; drop it first.
|
||||
if (!slot.owns) { slot.ptr = nullptr; slot.cap_bytes = 0; slot.owns = true; }
|
||||
if (bytes > slot.cap_bytes) {
|
||||
if (slot.ptr) backend::device_free(slot.ptr);
|
||||
if (slot.ptr && slot.owns) backend::device_free(slot.ptr);
|
||||
// Reset before the (throwing) allocation so an OOM leaves the slot
|
||||
// in a consistent empty state rather than {null ptr, stale cap}.
|
||||
slot.ptr = nullptr;
|
||||
@@ -197,23 +237,104 @@ class DevicePool {
|
||||
// zero and catastrophic where it is not. NaN makes it show.
|
||||
if (_poison_pool()) backend::memset_sync(slot.ptr, 0xff, bytes);
|
||||
}
|
||||
slot.owns = true;
|
||||
slot.used_bytes = bytes;
|
||||
slot.dtype_tag = (uint8_t)npy_scalar_tag<T>();
|
||||
return static_cast<T*>(slot.ptr);
|
||||
}
|
||||
|
||||
// Bump-allocate out of the arena (mutex held). What does not fit yet gets
|
||||
// a private allocation for this call and raises `want`, so the next phase
|
||||
// switch grows the arena and the call after that is aliased again.
|
||||
template<typename T>
|
||||
T* _acquire_arena(Slot& slot, size_t n) {
|
||||
// Each acquire gets a FRESH slice, so a second one in the same phase
|
||||
// would silently move the buffer -- say so rather than hand it back.
|
||||
if (slot.epoch == _epoch)
|
||||
throw std::runtime_error(
|
||||
"DevicePool: arena-backed buffer acquired twice in one phase; "
|
||||
"resize it once, or split the phase");
|
||||
slot.epoch = _epoch;
|
||||
const size_t bytes = n * sizeof(T);
|
||||
const size_t off = (_arena.cursor + kArenaAlign - 1) & ~(kArenaAlign - 1);
|
||||
// The cursor advances even when the slice does not fit, or `want`
|
||||
// would come out one buffer wide instead of one phase wide.
|
||||
_arena.cursor = off + bytes;
|
||||
if (_arena.want < _arena.cursor) _arena.want = _arena.cursor;
|
||||
if (!_arena.ptr || off + bytes > _arena.cap)
|
||||
return _acquire_into<T>(slot, n);
|
||||
if (slot.ptr && slot.owns) backend::device_free(slot.ptr);
|
||||
slot.ptr = static_cast<char*>(_arena.ptr) + off;
|
||||
slot.cap_bytes = 0;
|
||||
slot.used_bytes = bytes;
|
||||
slot.dtype_tag = (uint8_t)npy_scalar_tag<T>();
|
||||
slot.owns = false;
|
||||
return static_cast<T*>(slot.ptr);
|
||||
}
|
||||
|
||||
// The one place the phase rule is enforced (mutex held). `what` names the
|
||||
// buffer for the error; enum slots and acquire_dynamic both come through.
|
||||
template<typename T>
|
||||
T* _acquire_checked(Slot& slot, size_t n, PoolPhase ph, const char* what) {
|
||||
if (ph == PoolPhase::None || !_alias_enabled())
|
||||
return _acquire_into<T>(slot, n);
|
||||
if (ph != _phase)
|
||||
throw std::runtime_error(
|
||||
std::string("DevicePool: ") + what + " is arena-backed in phase "
|
||||
+ to_string(ph) + " but the open phase is " + to_string(_phase)
|
||||
+ " -- see POOL_ALIAS_TABLE in core/PoolSlots.h");
|
||||
return _acquire_arena<T>(slot, n);
|
||||
}
|
||||
|
||||
public:
|
||||
static DevicePool& global() {
|
||||
static DevicePool pool;
|
||||
return pool;
|
||||
}
|
||||
|
||||
// Open `p`. Every slice the previous phase carved out of the arena dies
|
||||
// here -- POOL_ALIAS_TABLE's rows are the claim that nothing still reads
|
||||
// one. Grows the arena first, which is safe for the same reason.
|
||||
void begin_phase(PoolPhase p) {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
_phase = p;
|
||||
++_epoch;
|
||||
if (!_alias_enabled() || p == PoolPhase::None) return;
|
||||
if (_arena.want > _arena.cap) {
|
||||
for (Slot& s : _slots)
|
||||
if (!s.owns) { s.ptr = nullptr; s.owns = true; }
|
||||
for (auto& kv : _dyn) {
|
||||
Slot& s = kv.second.slot;
|
||||
if (!s.owns) { s.ptr = nullptr; s.owns = true; }
|
||||
}
|
||||
if (_arena.ptr) backend::device_free(_arena.ptr);
|
||||
_arena.ptr = nullptr;
|
||||
_arena.cap = 0;
|
||||
_arena.ptr = backend::device_malloc_checked(_arena.want);
|
||||
_arena.cap = _arena.want;
|
||||
}
|
||||
_arena.cursor = 0;
|
||||
if (_poison_arena() && _arena.ptr)
|
||||
backend::memset_sync(_arena.ptr, 0xff, _arena.cap);
|
||||
}
|
||||
|
||||
struct ArenaStats { size_t cap_bytes; size_t want_bytes; size_t slots; };
|
||||
ArenaStats arenaStats() const {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
size_t n = 0;
|
||||
for (const Slot& s : _slots) if (!s.owns) ++n;
|
||||
for (const auto& kv : _dyn) if (!kv.second.slot.owns) ++n;
|
||||
return { _arena.cap, _arena.want, n };
|
||||
}
|
||||
|
||||
// Return a pointer to at least `n` elements of type T for the given key.
|
||||
// Reallocates only when current capacity is exceeded.
|
||||
template<typename T>
|
||||
T* acquire(PoolKey key, size_t n) {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
return _acquire_into<T>(_slots[key], n);
|
||||
const PoolSlot slot = pool_key_slot(key);
|
||||
return _acquire_checked<T>(_slots[key], n, slot_phase(slot),
|
||||
slot_name(slot));
|
||||
}
|
||||
|
||||
// Convenience: acquire a (slot, sub) buffer. sub defaults to the main slot.
|
||||
@@ -225,11 +346,13 @@ public:
|
||||
// Escape hatch for Never-saved scratch whose key is built at runtime and so
|
||||
// has no compile-time PoolSlot. `cat` places it in the VRAM report; `name`
|
||||
// is the display/debug key. Returns raw bytes (caller casts).
|
||||
void* acquire_dynamic(VramCategory cat, const std::string& name, size_t bytes) {
|
||||
void* acquire_dynamic(VramCategory cat, const std::string& name,
|
||||
size_t bytes, PoolPhase phase = PoolPhase::None) {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
DynSlot& d = _dyn[name];
|
||||
d.cat = cat;
|
||||
return _acquire_into<uint8_t>(d.slot, bytes);
|
||||
d.phase = phase;
|
||||
return _acquire_checked<uint8_t>(d.slot, bytes, phase, name.c_str());
|
||||
}
|
||||
|
||||
// Free a single logical buffer (all of its sub-allocations).
|
||||
@@ -237,7 +360,8 @@ public:
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
for (uint32_t sub = 0; sub <= kSubMask; ++sub) {
|
||||
Slot& s = _slots[pool_key(slot, sub)];
|
||||
if (s.ptr) { backend::device_free(s.ptr); s = Slot{}; }
|
||||
if (s.ptr && s.owns) backend::device_free(s.ptr);
|
||||
s = Slot{};
|
||||
}
|
||||
}
|
||||
|
||||
@@ -254,7 +378,8 @@ public:
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
for (auto it = _dyn.begin(); it != _dyn.end();) {
|
||||
if (it->first.rfind(prefix, 0) == 0) {
|
||||
if (it->second.slot.ptr) backend::device_free(it->second.slot.ptr);
|
||||
if (it->second.slot.ptr && it->second.slot.owns)
|
||||
backend::device_free(it->second.slot.ptr);
|
||||
it = _dyn.erase(it);
|
||||
} else {
|
||||
++it;
|
||||
@@ -265,15 +390,23 @@ public:
|
||||
// Free all managed device memory.
|
||||
void freeAll() {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
for (auto& s : _slots) if (s.ptr) { backend::device_free(s.ptr); s = Slot{}; }
|
||||
for (auto& kv : _dyn) if (kv.second.slot.ptr) backend::device_free(kv.second.slot.ptr);
|
||||
for (auto& s : _slots) {
|
||||
if (s.ptr && s.owns) backend::device_free(s.ptr);
|
||||
s = Slot{};
|
||||
}
|
||||
for (auto& kv : _dyn)
|
||||
if (kv.second.slot.ptr && kv.second.slot.owns)
|
||||
backend::device_free(kv.second.slot.ptr);
|
||||
_dyn.clear();
|
||||
if (_arena.ptr) backend::device_free(_arena.ptr);
|
||||
_arena = Arena{};
|
||||
_phase = PoolPhase::None;
|
||||
}
|
||||
|
||||
// Total bytes allocated (capacity, not logical size).
|
||||
size_t totalAllocBytes() const {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
size_t total = 0;
|
||||
size_t total = _arena.cap;
|
||||
for (auto& s : _slots) total += s.cap_bytes;
|
||||
for (auto& kv : _dyn) total += kv.second.slot.cap_bytes;
|
||||
return total;
|
||||
@@ -287,7 +420,7 @@ public:
|
||||
std::vector<std::tuple<std::string, size_t, size_t>> result;
|
||||
for (uint32_t i = 0; i < _slots.size(); ++i) {
|
||||
const Slot& s = _slots[i];
|
||||
if (!s.ptr && s.cap_bytes == 0) continue;
|
||||
if (!s.ptr && s.cap_bytes == 0 && s.used_bytes == 0) continue;
|
||||
std::string name = std::string(slot_name(pool_key_slot(i)))
|
||||
+ sub_suffix(pool_key_sub(i));
|
||||
result.emplace_back(std::move(name), s.used_bytes, s.cap_bytes);
|
||||
@@ -298,24 +431,44 @@ public:
|
||||
return result;
|
||||
}
|
||||
|
||||
// One row per live buffer. `arena` means the bytes are a slice of the
|
||||
// shared arena, so `cap` is 0 and the allocation is counted once in
|
||||
// arenaStats() instead.
|
||||
struct BreakdownRow {
|
||||
std::string name;
|
||||
VramCategory cat;
|
||||
size_t used_bytes;
|
||||
size_t cap_bytes;
|
||||
bool arena;
|
||||
};
|
||||
|
||||
std::vector<BreakdownRow> breakdown() const {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
std::vector<BreakdownRow> result;
|
||||
for (uint32_t i = 0; i < _slots.size(); ++i) {
|
||||
const Slot& s = _slots[i];
|
||||
if (!s.ptr && s.cap_bytes == 0 && s.used_bytes == 0) continue;
|
||||
PoolSlot slot = pool_key_slot(i);
|
||||
result.push_back({ std::string(slot_name(slot)) +
|
||||
sub_suffix(pool_key_sub(i)),
|
||||
slot_category(slot), s.used_bytes, s.cap_bytes,
|
||||
!s.owns });
|
||||
}
|
||||
for (auto& kv : _dyn)
|
||||
result.push_back({ kv.first, kv.second.cat,
|
||||
kv.second.slot.used_bytes,
|
||||
kv.second.slot.cap_bytes,
|
||||
!kv.second.slot.owns });
|
||||
return result;
|
||||
}
|
||||
|
||||
// Categorized breakdown: [(name, category, used_bytes, cap_bytes), ...].
|
||||
// category is (int)VramCategory. Used by the C++-driven VRAM report.
|
||||
std::vector<std::tuple<std::string, int, size_t, size_t>>
|
||||
getBreakdownCategorized() const {
|
||||
std::lock_guard<std::mutex> lock(_mu);
|
||||
std::vector<std::tuple<std::string, int, size_t, size_t>> result;
|
||||
for (uint32_t i = 0; i < _slots.size(); ++i) {
|
||||
const Slot& s = _slots[i];
|
||||
if (!s.ptr && s.cap_bytes == 0) continue;
|
||||
PoolSlot slot = pool_key_slot(i);
|
||||
std::string name = std::string(slot_name(slot))
|
||||
+ sub_suffix(pool_key_sub(i));
|
||||
result.emplace_back(std::move(name), (int)slot_category(slot),
|
||||
s.used_bytes, s.cap_bytes);
|
||||
}
|
||||
for (auto& kv : _dyn)
|
||||
result.emplace_back(kv.first, (int)kv.second.cat,
|
||||
kv.second.slot.used_bytes, kv.second.slot.cap_bytes);
|
||||
for (auto& r : breakdown())
|
||||
result.emplace_back(r.name, (int)r.cat, r.used_bytes, r.cap_bytes);
|
||||
return result;
|
||||
}
|
||||
|
||||
@@ -449,6 +602,14 @@ public:
|
||||
};
|
||||
|
||||
|
||||
// Open a pool phase; the arena's slices from the previous one die here.
|
||||
// Called by the launcher that owns a phase, not by the engine layer -- see
|
||||
// POOL_ALIAS_TABLE in core/PoolSlots.h.
|
||||
inline void pool_begin_phase(PoolPhase p) {
|
||||
DevicePool::global().begin_phase(p);
|
||||
}
|
||||
|
||||
|
||||
// Free all device memory managed by the pool and scratch buffer.
|
||||
// Safe to call at VRAM pressure or before process exit.
|
||||
inline void freeAllDeviceMemory() {
|
||||
@@ -504,9 +665,10 @@ public:
|
||||
// Dynamic-keyed resize for Never-class scratch whose name is built at
|
||||
// runtime (prefix + index / suffix). Routes to the pool's side map with an
|
||||
// explicit category. See DevicePool::acquire_dynamic.
|
||||
void resize_dynamic(VramCategory cat, const std::string& name, int64_t numel) {
|
||||
void resize_dynamic(VramCategory cat, const std::string& name, int64_t numel,
|
||||
PoolPhase phase = PoolPhase::None) {
|
||||
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
|
||||
cat, name, (size_t)numel * sizeof(T));
|
||||
cat, name, (size_t)numel * sizeof(T), phase);
|
||||
_numel = numel;
|
||||
}
|
||||
|
||||
@@ -591,9 +753,10 @@ public:
|
||||
|
||||
// Dynamic-keyed resize (Never-class scratch, runtime-built name).
|
||||
void resize_dynamic(VramCategory cat, const std::string& name,
|
||||
int64_t numel_0, int64_t numel_1) {
|
||||
int64_t numel_0, int64_t numel_1,
|
||||
PoolPhase phase = PoolPhase::None) {
|
||||
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
|
||||
cat, name, (size_t)(numel_0 * numel_1) * sizeof(T));
|
||||
cat, name, (size_t)(numel_0 * numel_1) * sizeof(T), phase);
|
||||
_numel_0 = numel_0;
|
||||
_numel_1 = numel_1;
|
||||
}
|
||||
@@ -694,9 +857,10 @@ public:
|
||||
|
||||
// Dynamic-keyed resize (Never-class scratch, runtime-built name).
|
||||
void resize_dynamic(VramCategory cat, const std::string& name,
|
||||
int64_t numel_0, int64_t numel_1, int64_t numel_2) {
|
||||
int64_t numel_0, int64_t numel_1, int64_t numel_2,
|
||||
PoolPhase phase = PoolPhase::None) {
|
||||
_data_ptr = (T*)DevicePool::global().acquire_dynamic(
|
||||
cat, name, (size_t)(numel_0 * numel_1 * numel_2) * sizeof(T));
|
||||
cat, name, (size_t)(numel_0 * numel_1 * numel_2) * sizeof(T), phase);
|
||||
_numel_0 = numel_0;
|
||||
_numel_1 = numel_1;
|
||||
_numel_2 = numel_2;
|
||||
|
||||
@@ -33,6 +33,11 @@ void forward_3dgs(
|
||||
engine().sh_degree = sh_degree;
|
||||
engine().packed = packed;
|
||||
|
||||
// The stashed screen gradients live in the arena the intersect below
|
||||
// reuses, so this view is already dead. Dropping it turns a stale
|
||||
// fused-optim read into that step's "v_splats_s not stashed" error.
|
||||
engine().fwd.v_splats_s.clear();
|
||||
|
||||
// Build splats as DeviceTensorFloatND from typed device buffers.
|
||||
// For DeviceVector<float> (opacities), build a [N, 1] shape FND manually.
|
||||
DeviceTensorFloatND fnd_opac;
|
||||
|
||||
+29
-15
@@ -260,17 +260,20 @@ size_t engine_get_scratch_bytes() {
|
||||
// ===========================================================================
|
||||
|
||||
std::string engine_vram_report() {
|
||||
auto rows = DevicePool::global().getBreakdownCategorized();
|
||||
// Arena-backed rows own nothing, so their cap is 0 and the arena's one
|
||||
// allocation is a line of its own; the totals still add up.
|
||||
auto rows = DevicePool::global().breakdown();
|
||||
auto arena = DevicePool::global().arenaStats();
|
||||
const int kNCat = (int)VramCategory::Count;
|
||||
size_t used[kNCat] = {}, cap[kNCat] = {}, count[kNCat] = {};
|
||||
size_t used_total = 0, cap_total = 0;
|
||||
for (const auto& r : rows) {
|
||||
int c = std::get<1>(r);
|
||||
int c = (int)r.cat;
|
||||
count[c]++;
|
||||
used[c] += std::get<2>(r);
|
||||
cap[c] += std::get<3>(r);
|
||||
used_total += std::get<2>(r);
|
||||
cap_total += std::get<3>(r);
|
||||
used[c] += r.used_bytes;
|
||||
cap[c] += r.cap_bytes;
|
||||
used_total += r.used_bytes;
|
||||
cap_total += r.cap_bytes;
|
||||
}
|
||||
|
||||
std::string out;
|
||||
@@ -293,10 +296,13 @@ std::string engine_vram_report() {
|
||||
}
|
||||
size_t scratch = engine_get_scratch_bytes();
|
||||
add("%-30s %11zu %11.2f %11.2f\n", " pool total", rows.size(),
|
||||
mib(used_total), mib(cap_total));
|
||||
mib(used_total), mib(cap_total + arena.cap_bytes));
|
||||
if (arena.cap_bytes || arena.slots)
|
||||
add("%-30s %11zu %11s %11.2f\n", " of which alias arena", arena.slots,
|
||||
"-", mib(arena.cap_bytes));
|
||||
add("%-30s %11s %11s %11.2f\n", " scratch buffer", "-", "-", mib(scratch));
|
||||
add("%-30s %11s %11s %11.2f\n", " pool + scratch", "", "",
|
||||
mib(cap_total + scratch));
|
||||
mib(cap_total + arena.cap_bytes + scratch));
|
||||
|
||||
// The driver's figures place the pool against everything else the process
|
||||
// holds: staging buffers, backend allocations, NN models, GUI surfaces.
|
||||
@@ -308,16 +314,24 @@ std::string engine_vram_report() {
|
||||
add("%-30s %11s %11s %11.2f / %.2f\n", " device in use / total", "",
|
||||
"", mib(mem.used_bytes), mib(mem.total_bytes));
|
||||
|
||||
std::sort(rows.begin(), rows.end(), [](const auto& a, const auto& b) {
|
||||
return std::get<3>(a) > std::get<3>(b);
|
||||
// Aliased rows own nothing, so rank on whichever of the two is bigger.
|
||||
auto weight = [](const auto& r) {
|
||||
return std::max(r.used_bytes, r.cap_bytes);
|
||||
};
|
||||
std::sort(rows.begin(), rows.end(), [&](const auto& a, const auto& b) {
|
||||
return weight(a) > weight(b);
|
||||
});
|
||||
add("%-30s %11s %11s %11s\n", "individual buffers", "category", "used_MiB",
|
||||
"cap_MiB");
|
||||
for (size_t i = 0; i < rows.size(); i++) {
|
||||
if (std::get<3>(rows[i]) == 0) break;
|
||||
add(" %-28s %11s %11.2f %11.2f\n", std::get<0>(rows[i]).c_str(),
|
||||
to_string((VramCategory)std::get<1>(rows[i])),
|
||||
mib(std::get<2>(rows[i])), mib(std::get<3>(rows[i])));
|
||||
for (const auto& r : rows) {
|
||||
if (weight(r) == 0) break;
|
||||
const char* cat = to_string(r.cat);
|
||||
if (r.arena)
|
||||
add(" %-28s %11s %11.2f %11s\n", r.name.c_str(), cat,
|
||||
mib(r.used_bytes), "arena");
|
||||
else
|
||||
add(" %-28s %11s %11.2f %11.2f\n", r.name.c_str(), cat,
|
||||
mib(r.used_bytes), mib(r.cap_bytes));
|
||||
}
|
||||
add("[spirula-profile] --------------------------\n");
|
||||
return out;
|
||||
|
||||
@@ -140,6 +140,9 @@ inline std::tuple<
|
||||
std::optional<std::vector<DeviceTensorFloatND>> v_splats_w,
|
||||
std::optional<std::vector<DeviceTensorFloatND>> v_splats_s
|
||||
) {
|
||||
// Opens the raster-backward phase: the screen-gradient buffer below shares
|
||||
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::RasterBwd);
|
||||
if (!v_splats_w.has_value())
|
||||
v_splats_w = SplatPrimitive::WorldBuffer::zeros_pool(splats_w, PoolSlot::RasterBwdVWorld);
|
||||
if (!v_splats_s.has_value())
|
||||
|
||||
@@ -188,6 +188,9 @@ inline std::tuple<
|
||||
std::optional<std::vector<DeviceTensorFloatND>> v_splats_s,
|
||||
bool need_viewmat_grad
|
||||
) {
|
||||
// Opens the raster-backward phase: the screen-gradient buffer below shares
|
||||
// the tile intersector's arena (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::RasterBwd);
|
||||
if (!v_splats_w.has_value())
|
||||
v_splats_w = SplatPrimitive::WorldBuffer::zeros_pool(splats_w, PoolSlot::RasterBwdVWorld);
|
||||
if (!v_splats_s.has_value())
|
||||
|
||||
@@ -395,6 +395,11 @@ std::tuple<
|
||||
const uint32_t total_count = (uint32_t)depths.numel();
|
||||
const uint32_t N = packed ? total_count : total_count / I;
|
||||
|
||||
// Opens the tile-intersect phase: the key and count buffers below are
|
||||
// scratch that dies here, so they share the arena with the raster
|
||||
// backward (POOL_ALIAS_TABLE, core/PoolSlots.h).
|
||||
pool_begin_phase(PoolPhase::TileIsect);
|
||||
|
||||
intersect_check_image_size(image_width, image_height);
|
||||
uint32_t tile_width = _CEIL_DIV(image_width, TILE_SIZE_IX);
|
||||
uint32_t tile_height = _CEIL_DIV(image_height, TILE_SIZE_IY);
|
||||
|
||||
@@ -428,7 +428,7 @@ public:
|
||||
DeviceTensor2D<float> b;
|
||||
b.resize_dynamic(slot_category(key_prefix),
|
||||
std::string(slot_name(key_prefix)) + ".0",
|
||||
size, (int64_t)STRIDE);
|
||||
size, (int64_t)STRIDE, slot_phase(key_prefix));
|
||||
return { DeviceTensorFloatND(b) };
|
||||
}
|
||||
|
||||
@@ -537,17 +537,18 @@ public:
|
||||
int64_t size, std::array<int32_t, N> strides, PoolSlot key_prefix
|
||||
) {
|
||||
const VramCategory cat = slot_category(key_prefix);
|
||||
const PoolPhase ph = slot_phase(key_prefix);
|
||||
const std::string base = slot_name(key_prefix);
|
||||
std::vector<DeviceTensorFloatND> res;
|
||||
for (int i = 0; i < N; ++i) {
|
||||
if (strides[i] <= 0) { res.push_back(DeviceTensorFloatND()); continue; }
|
||||
std::string key = base + "." + std::to_string(i);
|
||||
switch (strides[i]) {
|
||||
case 1: { DeviceVector<float> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 2: { DeviceVector<float2> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 3: { DeviceVector<float3> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 4: { DeviceVector<float4> b; b.resize_dynamic(cat, key, size); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
default: { DeviceTensor2D<float> b; b.resize_dynamic(cat, key, size, (int64_t)strides[i]); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 1: { DeviceVector<float> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 2: { DeviceVector<float2> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 3: { DeviceVector<float3> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
case 4: { DeviceVector<float4> b; b.resize_dynamic(cat, key, size, ph); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
default: { DeviceTensor2D<float> b; b.resize_dynamic(cat, key, size, (int64_t)strides[i], ph); res.push_back(DeviceTensorFloatND(b)); break; }
|
||||
}
|
||||
}
|
||||
return res;
|
||||
|
||||
Reference in New Issue
Block a user