Compare commits

...
53 changed files with 4164 additions and 313 deletions
+1
View File
@@ -48,6 +48,7 @@ API and command-line option may change frequently.***
- [PiD](./docs/pid.md)
- [LongCat Image](./docs/longcat_image.md)
- [Z-Image](./docs/z_image.md)
- [MiniT2I](./docs/minit2i.md)
- [Ovis-Image](./docs/ovis_image.md)
- [Anima](./docs/anima.md)
- [ERNIE-Image](./docs/ernie_image.md)
+91
View File
@@ -51,6 +51,97 @@ Module names are case-insensitive. Hyphens and underscores in module names are i
sd-cli -m model.safetensors -p "a cat" --backend all=cuda0,te=cpu
```
## Multiple devices per module (layer split)
A `--backend` module assignment can list several devices separated by `&`:
```shell
sd-cli -m model.safetensors -p "a cat" --backend "diffusion=cuda0&cuda1"
```
The module's transformer blocks are then distributed across the listed devices
in contiguous ranges sized proportionally to each device's free memory (minus a
compute-buffer headroom of about 2 GiB per device), and the
module's graphs are executed with a `ggml_backend_sched` that runs each block
on the device holding its weights, copying the residual stream at the range
boundaries. The first device in the list is the module's main device: it also
holds the non-block tensors (embeddings, final norms, small sub-runners such as
CLIP models or projectors) and the graph inputs/outputs.
Layer split is supported for the `diffusion` and `te` modules. For `te` it
applies to the dominant text encoder (`t5xxl` or the LLM); other modules accept
only a single device. If the module has no recognizable transformer blocks, the
assignment falls back to the first listed device.
`--params-backend` accepts no device lists. If the module has no explicit
params assignment, each block range's parameters are loaded directly to (and,
with `--params-backend diffusion=disk`, released directly from) its own device;
an explicit assignment such as `te=cpu` keeps the parameters on that backend
and stages each range to its device on demand.
Layer split cannot be combined with `--max-vram` graph-cut segmentation or
`--stream-layers` for the split module; those are single-device mechanisms and
are disabled for it.
Use `--list-devices` to see the device names available on the system.
### Row split (`--split-mode row`)
`--split-mode` selects how a multi-device module distributes its weights:
`layer` (the default, described above) or `row`. It accepts a single mode or
per-module assignments:
```shell
sd-cli -m model.safetensors -p "a cat" --backend "diffusion=cuda0&cuda1" --split-mode row
sd-cli -m model.safetensors -p "a cat" --backend "diffusion=cuda0&cuda1,te=cuda0&cuda1" --split-mode diffusion=row,te=layer
```
In row mode the module keeps executing on its main (first listed) device, but
its transformer-block matmul weights are allocated in the backend's row-split
buffer type, which slices each weight's rows across the listed devices in
proportion to free memory and runs those matmuls on all devices in parallel.
Compared to a layer split this uses all GPUs within every layer (instead of
sequentially device by device) at the cost of a cross-device reduction per
matmul - usually the faster option when the devices have fast interconnect.
Row split requires backend support for split buffers and is currently
available on CUDA only; on other backends (or when the listed devices belong
to different backend registries) the module falls back to a layer split.
Embeddings, normalization weights, biases and other non-block tensors stay in
regular buffers on the main device.
Direct ("immediately") LoRA application cannot patch row-split tensors; with
`--split-mode row` the automatic LoRA mode selects runtime application, and an
explicit `--lora-apply-mode immediately` skips the split tensors with a
warning.
## Automatic placement (`--auto-fit`)
`--auto-fit` derives the `diffusion` / `te` / `vae` placements from the model
metadata and the per-device memory budgets, then feeds them into the same
backend assignment mechanism described above (the chosen specs are printed).
`--backend` and `--params-backend` are ignored while auto-fit is enabled.
```shell
sd-cli -m model.safetensors -p "a cat" --auto-fit
sd-cli -m model.safetensors -p "a cat" --auto-fit --max-vram cuda0=8,cuda1=14
sd-cli -m model.safetensors -p "a cat" --auto-fit --split-mode row
```
Budgets reuse `--max-vram`: a positive per-device value caps what auto-fit
plans with on that device, a negative value means "free memory minus that many
GiB", and with no budget set each device's free memory minus a 512 MiB margin
is used. (The same values still drive graph-cut segmented execution for
modules that end up on a single device.)
When everything fits resident, components are simply spread across the
available GPUs. When it does not, auto-fit switches to time-share mode: the
heavy components get `disk` params residency (loaded for their phase, freed
after), and a component too large for any single device is split across all
GPUs with the layer/row split mechanism (`--split-mode` selects which, layer
by default). Components that fit nowhere fall back to the CPU. If a VAE decode
still runs out of memory, tiling is enabled and the decode retried once.
## Modules
| Module | Purpose | Accepted names |
+59
View File
@@ -0,0 +1,59 @@
# Importance Matrix (imatrix) Quantization
## What is an Importance Matrix?
Quantization reduces the precision of a model's weights, decreasing its size and computational requirements. However, this can lead to a loss of quality. An importance matrix helps mitigate this by identifying which weights are *most* important for the model's performance. During quantization, these important weights are preserved with higher precision, while less important weights are quantized more aggressively. This allows for better overall quality at a given quantization level.
This originates from work done with language models in [llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/imatrix/README.md).
## Usage
The imatrix feature involves two main steps: *training* the matrix and *using* it during quantization.
### Training the Importance Matrix
To generate an imatrix, run stable-diffusion.cpp with the `--imat-out` flag, specifying the output filename. This process runs alongside normal image generation.
```bash
sd.exe [same exact parameters as normal generation] --imat-out imatrix.dat
```
* **`[same exact parameters as normal generation]`**: Use the same command-line arguments you would normally use for image generation (e.g., prompt, dimensions, sampling method, etc.).
* **`--imat-out imatrix.dat`**: Specifies the output file for the generated imatrix.
You can generate multiple images at once using the `-b` flag to speed up the training process.
### Continuing Training an Existing Matrix
If you want to refine an existing imatrix, use the `--imat-in` flag *in addition* to `--imat-out`. This will load the existing matrix and continue training it.
```bash
sd.exe [same exact parameters as normal generation] --imat-out imatrix.dat --imat-in imatrix.dat
```
With that, you can train and refine the imatrix while generating images like you'd normally do.
### Using Multiple Matrices
You can load and merge multiple imatrices together:
```bash
sd.exe [same exact parameters as normal generation] --imat-out imatrix.dat --imat-in imatrix.dat --imat-in imatrix2.dat
```
### Quantizing with an Importance Matrix
To quantize a model using a trained imatrix, use the `-M convert` option (or equivalent quantization command) and the `--imat-in` flag, specifying the imatrix file.
```bash
sd.exe -M convert [same exact parameters as normal quantization] --imat-in imatrix.dat
```
* **`[same exact parameters as normal quantization]`**: Use the same command-line arguments you would normally use for quantization (e.g., target quantization method, input/output filenames).
* **`--imat-in imatrix.dat`**: Specifies the imatrix file to use during quantization. You can specify multiple `--imat-in` flags to combine multiple matrices.
## Important Considerations
* The quality of the imatrix depends on the prompts and settings used during training. Use prompts and settings representative of the types of images you intend to generate for the best results.
* Experiment with different training parameters (e.g., number of images, prompt variations) to optimize the imatrix for your specific use case.
* The performance impact of training an imatrix during image generation or using an imatrix for quantization is negligible.
* Using already quantized models to train the imatrix seems to be working fine.
+48
View File
@@ -0,0 +1,48 @@
# How to Use
MiniT2I uses a MiniT2I diffusion transformer and `google/flan-t5-large` as the text encoder.
## Download weights
- Download MiniT2I diffusion model
- safetensors: https://huggingface.co/MiniT2I/MiniT2I/tree/main/minit2i-b-16/transformer (`diffusion_pytorch_model.safetensors`)
- Download flan-t5-large text encoder
- safetensors: https://huggingface.co/google/flan-t5-large/tree/main (`model.safetensors`)
## Examples
### Mac Metal
```
./bin/sd-cli \
--backend metal \
--diffusion-model ../models/minit2i/diffusion_pytorch_model.safetensors \
--t5xxl ../models/flan-t5-large/model.safetensors \
--prompt "a cat" \
--steps 100 \
--cfg-scale 6 \
--width 512 \
--height 512 \
--seed 42 \
--sampling-method euler \
--rng cpu \
--output minit2i_metal.png \
--threads 8
```
### CUDA with diffusion flash attention
```
./bin/sd-cli \
--diffusion-model ../models/minit2i/diffusion_pytorch_model.safetensors \
--t5xxl ../models/flan-t5-large/model.safetensors \
--prompt "a cat" \
--steps 100 \
--cfg-scale 6 \
--width 512 \
--height 512 \
--seed 42 \
--sampling-method euler \
--diffusion-fa \
--output minit2i_cuda.png
```
+73 -12
View File
@@ -1,6 +1,7 @@
#include <stdio.h>
#include <string.h>
#include <time.h>
#include <algorithm>
#include <cctype>
#include <filesystem>
#include <functional>
@@ -53,6 +54,9 @@ struct SDCliParams {
bool metadata_brief = false;
bool metadata_all = false;
std::string imatrix_out;
std::vector<std::string> imatrix_in;
bool normal_exit = false;
ArgOptions get_options() {
@@ -79,6 +83,11 @@ struct SDCliParams {
"path to write preview image to (default: ./preview.png). Multi-frame previews support .avi, .webm, and animated .webp",
0,
&preview_path},
{"",
"--imat-out",
"compute the imatrix for this run and save it to the provided path",
0,
&imatrix_out},
};
options.int_options = {
@@ -179,6 +188,14 @@ struct SDCliParams {
return -1;
};
auto on_imatrix_in_arg = [&](int argc, const char** argv, int index) {
if (++index >= argc) {
return -1;
}
imatrix_in.push_back(argv[index]);
return 1;
};
options.manual_options = {
{"-M",
"--mode",
@@ -192,6 +209,10 @@ struct SDCliParams {
"--help",
"show this help message and exit",
on_help_arg},
{"",
"--imat-in",
"load an imatrix file for quantization or continued collection; can be specified multiple times",
on_imatrix_in_arg},
};
return options;
@@ -253,6 +274,7 @@ struct SDCliParams {
<< " preview_fps: " << preview_fps << ",\n"
<< " taesd_preview: " << (taesd_preview ? "true" : "false") << ",\n"
<< " preview_noisy: " << (preview_noisy ? "true" : "false") << ",\n"
<< " imatrix_out: \"" << imatrix_out << "\",\n"
<< " metadata_raw: " << (metadata_raw ? "true" : "false") << ",\n"
<< " metadata_brief: " << (metadata_brief ? "true" : "false") << ",\n"
<< " metadata_all: " << (metadata_all ? "true" : "false") << "\n"
@@ -459,7 +481,8 @@ bool save_results(const SDCliParams& cli_params,
if (!img.data)
return false;
const int64_t metadata_seed = cli_params.mode == VID_GEN ? gen_params.seed : gen_params.seed + idx;
int images_per_batch = gen_params.batch_count > 0 ? std::max(1, num_results / gen_params.batch_count) : 1;
const int64_t metadata_seed = cli_params.mode == VID_GEN ? gen_params.seed : gen_params.seed + idx / images_per_batch;
std::string params = gen_params.embed_image_metadata
? get_image_params(ctx_params, gen_params, metadata_seed, cli_params.mode)
: "";
@@ -605,13 +628,32 @@ int main(int argc, const char* argv[]) {
LOG_DEBUG("%s", ctx_params.to_string().c_str());
LOG_DEBUG("%s", gen_params.to_string().c_str());
if (!cli_params.imatrix_out.empty()) {
if (fs::exists(cli_params.imatrix_out) &&
std::find(cli_params.imatrix_in.begin(), cli_params.imatrix_in.end(), cli_params.imatrix_out) == cli_params.imatrix_in.end()) {
LOG_WARN("imatrix file '%s' already exists and will be overwritten", cli_params.imatrix_out.c_str());
}
enable_imatrix_collection();
}
for (const auto& in_file : cli_params.imatrix_in) {
LOG_INFO("loading imatrix from '%s'", in_file.c_str());
if (!load_imatrix(in_file.c_str())) {
LOG_WARN("failed to load imatrix from '%s'", in_file.c_str());
}
}
if (cli_params.mode == CONVERT) {
bool success = convert(ctx_params.model_path.c_str(),
ctx_params.vae_path.c_str(),
cli_params.output_path.c_str(),
ctx_params.wtype,
ctx_params.tensor_type_rules.c_str(),
cli_params.convert_name);
bool success = convert_with_components(ctx_params.model_path.c_str(),
ctx_params.clip_l_path.c_str(),
ctx_params.clip_g_path.c_str(),
ctx_params.t5xxl_path.c_str(),
ctx_params.diffusion_model_path.c_str(),
ctx_params.vae_path.c_str(),
cli_params.output_path.c_str(),
ctx_params.wtype,
ctx_params.tensor_type_rules.c_str(),
cli_params.convert_name);
if (!success) {
LOG_ERROR("convert '%s'/'%s' to '%s' failed",
ctx_params.model_path.c_str(),
@@ -766,8 +808,12 @@ int main(int argc, const char* argv[]) {
if (cli_params.mode == IMG_GEN) {
sd_img_gen_params_t img_gen_params = gen_params.to_sd_img_gen_params_t();
num_results = gen_params.batch_count;
results.adopt(generate_image(sd_ctx.get(), &img_gen_params), num_results);
sd_image_t* generated_images = nullptr;
if (!generate_image(sd_ctx.get(), &img_gen_params, &generated_images, &num_results)) {
generated_images = nullptr;
num_results = 0;
}
results.adopt(generated_images, num_results);
} else if (cli_params.mode == VID_GEN) {
sd_vid_gen_params_t vid_gen_params = gen_params.to_sd_vid_gen_params_t();
sd_image_t* generated_video = nullptr;
@@ -802,12 +848,22 @@ int main(int argc, const char* argv[]) {
SDImageOwner current_image(results[i]);
results[i] = {0, 0, 0, nullptr};
for (int u = 0; u < gen_params.upscale_repeats; ++u) {
SDImageOwner upscaled_image(upscale(upscaler_ctx.get(), current_image.get(), upscale_factor));
if (upscaled_image.get().data == nullptr) {
sd_image_t* upscaled_images = nullptr;
int upscaled_count = 0;
bool upscale_ok = upscale(upscaler_ctx.get(),
current_image.get(),
upscale_factor,
&upscaled_images,
&upscaled_count);
if (!upscale_ok || upscaled_count <= 0 || upscaled_images[0].data == nullptr) {
free_sd_images(upscaled_images, upscaled_count);
LOG_ERROR("upscale failed");
break;
}
current_image = std::move(upscaled_image);
sd_image_t upscaled_image = upscaled_images[0];
upscaled_images[0] = {0, 0, 0, nullptr};
free_sd_images(upscaled_images, upscaled_count);
current_image.reset(upscaled_image);
}
results[i] = current_image.release(); // Set the final upscaled image as the result
}
@@ -819,6 +875,11 @@ int main(int argc, const char* argv[]) {
return 1;
}
if (!cli_params.imatrix_out.empty()) {
LOG_INFO("saving imatrix to '%s'", cli_params.imatrix_out.c_str());
save_imatrix(cli_params.imatrix_out.c_str());
}
free_sd_audio(generated_audio);
return 0;
+54 -2
View File
@@ -468,6 +468,13 @@ ArgOptions SDContextParams::get_options() {
"parameter backend assignment, e.g. disk, cpu, or diffusion=disk,clip=cpu",
(int)',',
&params_backend},
{"",
"--split-mode",
"weight distribution for modules assigned multiple devices (--backend \"diffusion=cuda0&cuda1\"): "
"layer (whole transformer blocks per device, default) or row (matmul rows split across devices, CUDA only). "
"Accepts a single mode or per-module assignments, e.g. row or diffusion=row,te=layer",
(int)',',
&split_mode},
{"",
"--rpc-servers",
"comma-separated list of RPC servers to connect to for offloading, in the format host:port, e.g. localhost:50052,192.168.1.3:50052",
@@ -501,6 +508,12 @@ ArgOptions SDContextParams::get_options() {
"--eager-load",
"load all params into the params backend at model-load time instead of lazily on first use (defaults to false)",
true, &eager_load},
{"",
"--auto-fit",
"pick the diffusion/te/vae device placements automatically from the model size and the per-device "
"memory budgets (--max-vram; defaults to free memory minus a small margin). Overrides --backend and "
"--params-backend; may split modules across GPUs (--split-mode still selects layer or row)",
true, &auto_fit},
{"",
"--force-sdxl-vae-conv-scale",
"force use of conv scale on sdxl vae",
@@ -663,6 +676,18 @@ ArgOptions SDContextParams::get_options() {
"but it usually offers faster inference speed and, in some cases, lower memory usage. "
"The at_runtime mode, on the other hand, is exactly the opposite.",
on_lora_apply_mode_arg},
{"",
"--list-devices",
"list available ggml backend devices (one 'name<TAB>description' per line) and exit; "
"the names are the device names accepted by --backend and --params-backend",
[](int /*argc*/, const char** /*argv*/, int /*index*/) {
size_t device_list_size = sd_list_devices(nullptr, 0);
std::vector<char> devices(device_list_size + 1);
sd_list_devices(devices.data(), devices.size());
fputs(devices.data(), stdout);
std::exit(0);
return 0;
}},
};
return options;
@@ -710,7 +735,18 @@ bool SDContextParams::resolve(SDMode mode) {
}
bool SDContextParams::validate(SDMode mode) {
if (mode != UPSCALE && mode != METADATA && model_path.length() == 0 && diffusion_model_path.length() == 0) {
if (mode == CONVERT) {
const bool has_convert_input = model_path.length() != 0 ||
clip_l_path.length() != 0 ||
clip_g_path.length() != 0 ||
t5xxl_path.length() != 0 ||
diffusion_model_path.length() != 0 ||
vae_path.length() != 0;
if (!has_convert_input) {
LOG_ERROR("error: convert mode needs at least one model input path\n");
return false;
}
} else if (mode != UPSCALE && mode != METADATA && model_path.length() == 0 && diffusion_model_path.length() == 0) {
LOG_ERROR("error: the following arguments are required: model_path/diffusion_model\n");
return false;
}
@@ -807,6 +843,8 @@ std::string SDContextParams::to_string() const {
<< " eager_load: " << (eager_load ? "true" : "false") << ",\n"
<< " backend: \"" << backend << "\",\n"
<< " params_backend: \"" << params_backend << "\",\n"
<< " split_mode: \"" << split_mode << "\",\n"
<< " auto_fit: " << (auto_fit ? "true" : "false") << ",\n"
<< " enable_mmap: " << (enable_mmap ? "true" : "false") << ",\n"
<< " control_net_cpu: " << (control_net_cpu ? "true" : "false") << ",\n"
<< " clip_on_cpu: " << (clip_on_cpu ? "true" : "false") << ",\n"
@@ -887,6 +925,8 @@ sd_ctx_params_t SDContextParams::to_sd_ctx_params_t(bool taesd_preview) {
sd_ctx_params.eager_load = eager_load;
sd_ctx_params.backend = effective_backend.c_str();
sd_ctx_params.params_backend = effective_params_backend.c_str();
sd_ctx_params.split_mode = split_mode.c_str();
sd_ctx_params.auto_fit = auto_fit;
sd_ctx_params.rpc_servers = rpc_servers.c_str();
return sd_ctx_params;
}
@@ -996,6 +1036,10 @@ ArgOptions SDGenerationParams::get_options() {
"--batch-count",
"batch count",
&batch_count},
{"",
"--qwen-image-layers",
"number of Qwen Image Layered layers; latent/output count is layers + 1 (default: 3)",
&qwen_image_layers},
{"",
"--video-frames",
"video frames (default: 1)",
@@ -1475,7 +1519,7 @@ ArgOptions SDGenerationParams::get_options() {
on_high_noise_sample_method_arg},
{"",
"--scheduler",
"denoiser sigma scheduler, one of [discrete, karras, exponential, ays, gits, smoothstep, sgm_uniform, simple, kl_optimal, lcm, bong_tangent, ltx2, logit_normal, flux2, flux], alias: normal=discrete, default: model-specific",
"denoiser sigma scheduler, one of [discrete, karras, exponential, ays, gits, smoothstep, sgm_uniform, simple, kl_optimal, lcm, bong_tangent, ltx2, logit_normal, flux2, flux, beta], alias: normal=discrete, default: model-specific",
on_scheduler_arg},
{"",
"--sigmas",
@@ -1816,6 +1860,7 @@ bool SDGenerationParams::from_json_str(
load_if_exists("width", width);
load_if_exists("height", height);
load_if_exists("batch_count", batch_count);
load_if_exists("qwen_image_layers", qwen_image_layers);
load_if_exists("video_frames", video_frames);
load_if_exists("fps", fps);
load_if_exists("upscale_repeats", upscale_repeats);
@@ -2240,6 +2285,11 @@ bool SDGenerationParams::validate(SDMode mode) {
return false;
}
if (qwen_image_layers < 0) {
LOG_ERROR("error: qwen_image_layers must be non-negative");
return false;
}
if (sample_params.sample_steps <= 0) {
LOG_ERROR("error: the sample_steps must be greater than 0\n");
return false;
@@ -2406,6 +2456,7 @@ sd_img_gen_params_t SDGenerationParams::to_sd_img_gen_params_t() {
params.strength = strength;
params.seed = seed;
params.batch_count = batch_count;
params.qwen_image_layers = qwen_image_layers;
params.control_image = control_image.get();
params.control_strength = control_strength;
params.pm_params = pm_params;
@@ -2531,6 +2582,7 @@ std::string SDGenerationParams::to_string() const {
<< " width: " << width << ",\n"
<< " height: " << height << ",\n"
<< " batch_count: " << batch_count << ",\n"
<< " qwen_image_layers: " << qwen_image_layers << ",\n"
<< " init_image_path: \"" << init_image_path << "\",\n"
<< " end_image_path: \"" << end_image_path << "\",\n"
<< " mask_image_path: \"" << mask_image_path << "\",\n"
+3
View File
@@ -151,6 +151,8 @@ struct SDContextParams {
bool eager_load = false;
std::string backend;
std::string params_backend;
std::string split_mode;
bool auto_fit = false;
std::string rpc_servers;
std::string effective_backend;
std::string effective_params_backend;
@@ -197,6 +199,7 @@ struct SDGenerationParams {
int width = -1;
int height = -1;
int batch_count = 1;
int qwen_image_layers = 3;
int64_t seed = 42;
float strength = 0.75f;
float control_strength = 0.9f;
+11 -3
View File
@@ -2,6 +2,7 @@
#include "async_jobs.h"
#include <algorithm>
#include <iomanip>
#include <sstream>
@@ -173,8 +174,13 @@ bool execute_img_gen_job(ServerRuntime& runtime,
{
std::lock_guard<std::mutex> lock(*runtime.sd_ctx_mutex);
sd_image_t* raw_results = generate_image(runtime.sd_ctx, &params);
results.adopt(raw_results, params.batch_count);
sd_image_t* raw_results = nullptr;
int num_results = 0;
if (!generate_image(runtime.sd_ctx, &params, &raw_results, &num_results)) {
raw_results = nullptr;
num_results = 0;
}
results.adopt(raw_results, num_results);
}
const int num_results = results.count();
@@ -190,6 +196,8 @@ bool execute_img_gen_job(ServerRuntime& runtime,
encoded_format = EncodedImageFormat::WEBP;
}
int batch_count = job.img_gen.gen_params.batch_count;
int images_per_batch = batch_count > 0 ? std::max(1, num_results / batch_count) : 1;
for (int i = 0; i < num_results; ++i) {
if (results[i].data == nullptr) {
continue;
@@ -198,7 +206,7 @@ bool execute_img_gen_job(ServerRuntime& runtime,
const std::string metadata = job.img_gen.gen_params.embed_image_metadata
? get_image_params(*runtime.ctx_params,
job.img_gen.gen_params,
job.img_gen.gen_params.seed + i)
job.img_gen.gen_params.seed + i / images_per_batch)
: "";
auto image_bytes = encode_image_to_vector(encoded_format,
results[i].data,
+13 -6
View File
@@ -229,8 +229,11 @@ static bool execute_sync_img_gen_request(ServerRuntime& runtime,
{
std::lock_guard<std::mutex> lock(*runtime.sd_ctx_mutex);
sd_image_t* raw_results = generate_image(runtime.sd_ctx, &img_gen_params);
num_results = request.gen_params.batch_count;
sd_image_t* raw_results = nullptr;
if (!generate_image(runtime.sd_ctx, &img_gen_params, &raw_results, &num_results)) {
raw_results = nullptr;
num_results = 0;
}
results.adopt(raw_results, num_results);
}
@@ -281,14 +284,16 @@ void register_openai_api_endpoints(httplib::Server& svr, ServerRuntime& rt) {
out["data"] = json::array();
out["output_format"] = request.output_format;
for (int i = 0; i < request.gen_params.batch_count; ++i) {
int result_count = results.count();
int images_per_batch = request.gen_params.batch_count > 0 ? std::max(1, result_count / request.gen_params.batch_count) : 1;
for (int i = 0; i < result_count; ++i) {
if (results[i].data == nullptr) {
continue;
}
std::string params = request.gen_params.embed_image_metadata
? get_image_params(*runtime->ctx_params,
request.gen_params,
request.gen_params.seed + i)
request.gen_params.seed + i / images_per_batch)
: "";
auto image_bytes = encode_image_to_vector(request.output_format == "jpeg"
? EncodedImageFormat::JPEG
@@ -353,14 +358,16 @@ void register_openai_api_endpoints(httplib::Server& svr, ServerRuntime& rt) {
out["data"] = json::array();
out["output_format"] = request.output_format;
for (int i = 0; i < request.gen_params.batch_count; ++i) {
int result_count = results.count();
int images_per_batch = request.gen_params.batch_count > 0 ? std::max(1, result_count / request.gen_params.batch_count) : 1;
for (int i = 0; i < result_count; ++i) {
if (results[i].data == nullptr) {
continue;
}
std::string params = request.gen_params.embed_image_metadata
? get_image_params(*runtime->ctx_params,
request.gen_params,
request.gen_params.seed + i)
request.gen_params.seed + i / images_per_batch)
: "";
auto image_bytes = encode_image_to_vector(request.output_format == "jpeg" ? EncodedImageFormat::JPEG : EncodedImageFormat::PNG,
results[i].data,
+7 -3
View File
@@ -292,8 +292,11 @@ void register_sdapi_endpoints(httplib::Server& svr, ServerRuntime& rt) {
{
std::lock_guard<std::mutex> lock(*runtime->sd_ctx_mutex);
sd_image_t* raw_results = generate_image(runtime->sd_ctx, &img_gen_params);
num_results = request.gen_params.batch_count;
sd_image_t* raw_results = nullptr;
if (!generate_image(runtime->sd_ctx, &img_gen_params, &raw_results, &num_results)) {
raw_results = nullptr;
num_results = 0;
}
results.adopt(raw_results, num_results);
}
@@ -308,6 +311,7 @@ void register_sdapi_endpoints(httplib::Server& svr, ServerRuntime& rt) {
out["parameters"] = j;
out["info"] = "";
int images_per_batch = request.gen_params.batch_count > 0 ? std::max(1, num_results / request.gen_params.batch_count) : 1;
for (int i = 0; i < num_results; ++i) {
if (results[i].data == nullptr) {
continue;
@@ -316,7 +320,7 @@ void register_sdapi_endpoints(httplib::Server& svr, ServerRuntime& rt) {
std::string params = request.gen_params.embed_image_metadata
? get_image_params(*runtime->ctx_params,
request.gen_params,
request.gen_params.seed + i)
request.gen_params.seed + i / images_per_batch)
: "";
auto image_bytes = encode_image_to_vector(EncodedImageFormat::PNG,
results[i].data,
+1
View File
@@ -126,6 +126,7 @@ static json make_img_gen_defaults_json(const SDGenerationParams& defaults, const
{"strength", defaults.strength},
{"seed", defaults.seed},
{"batch_count", defaults.batch_count},
{"qwen_image_layers", defaults.qwen_image_layers},
{"auto_resize_ref_image", defaults.auto_resize_ref_image},
{"increase_ref_index", defaults.increase_ref_index},
{"control_strength", defaults.control_strength},
+39 -4
View File
@@ -73,6 +73,7 @@ enum scheduler_t {
LOGIT_NORMAL_SCHEDULER,
FLUX2_SCHEDULER,
FLUX_SCHEDULER,
BETA_SCHEDULER,
SCHEDULER_COUNT
};
@@ -83,6 +84,7 @@ enum prediction_t {
FLOW_PRED,
FLUX_FLOW_PRED,
SEFI_FLOW_PRED,
MINIT2I_FLOW_PRED,
PREDICTION_COUNT
};
@@ -225,6 +227,8 @@ typedef struct {
bool eager_load; // Load all params into the params backend at model-load time instead of lazily on first use
const char* backend;
const char* params_backend;
const char* split_mode; // weight distribution for multi-device modules: layer (default) or row, or per-module assignments e.g. "diffusion=row"
bool auto_fit;
const char* rpc_servers;
} sd_ctx_params_t;
@@ -378,6 +382,7 @@ typedef struct {
sd_tiling_params_t vae_tiling_params;
sd_cache_params_t cache;
sd_hires_params_t hires;
int qwen_image_layers;
} sd_img_gen_params_t;
typedef struct {
@@ -406,14 +411,17 @@ typedef struct {
} sd_vid_gen_params_t;
typedef struct sd_ctx_t sd_ctx_t;
struct ggml_tensor;
typedef void (*sd_log_cb_t)(enum sd_log_level_t level, const char* text, void* data);
typedef void (*sd_progress_cb_t)(int step, int steps, float time, void* data);
typedef void (*sd_preview_cb_t)(int step, int frame_count, sd_image_t* frames, bool is_noisy, void* data);
typedef bool (*sd_graph_eval_callback_t)(struct ggml_tensor* t, bool ask, void* user_data);
SD_API void sd_set_log_callback(sd_log_cb_t sd_log_cb, void* data);
SD_API void sd_set_progress_callback(sd_progress_cb_t cb, void* data);
SD_API void sd_set_preview_callback(sd_preview_cb_t cb, enum preview_t mode, int interval, bool denoised, bool noisy, void* data);
SD_API void sd_set_backend_eval_callback(sd_graph_eval_callback_t cb, void* data);
SD_API int32_t sd_get_num_physical_cores();
SD_API const char* sd_get_system_info();
SD_API bool sd_ctx_supports_image_generation(const sd_ctx_t* sd_ctx);
@@ -454,7 +462,10 @@ SD_API enum scheduler_t sd_get_default_scheduler(const sd_ctx_t* sd_ctx, enum sa
SD_API void sd_img_gen_params_init(sd_img_gen_params_t* sd_img_gen_params);
SD_API char* sd_img_gen_params_to_str(const sd_img_gen_params_t* sd_img_gen_params);
SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* sd_img_gen_params);
SD_API bool generate_image(sd_ctx_t* sd_ctx,
const sd_img_gen_params_t* sd_img_gen_params,
sd_image_t** images_out,
int* num_images_out);
enum sd_cancel_mode_t {
// Stop the current generation as soon as possible.
@@ -484,9 +495,11 @@ SD_API upscaler_ctx_t* new_upscaler_ctx(const char* esrgan_path,
const char* params_backend);
SD_API void free_upscaler_ctx(upscaler_ctx_t* upscaler_ctx);
SD_API sd_image_t upscale(upscaler_ctx_t* upscaler_ctx,
sd_image_t input_image,
uint32_t upscale_factor);
SD_API bool upscale(upscaler_ctx_t* upscaler_ctx,
sd_image_t input_image,
uint32_t upscale_factor,
sd_image_t** images_out,
int* num_images_out);
SD_API int get_upscale_factor(upscaler_ctx_t* upscaler_ctx);
@@ -497,6 +510,17 @@ SD_API bool convert(const char* input_path,
const char* tensor_type_rules,
bool convert_name);
SD_API bool convert_with_components(const char* model_path,
const char* clip_l_path,
const char* clip_g_path,
const char* t5xxl_path,
const char* diffusion_model_path,
const char* vae_path,
const char* output_path,
enum sd_type_t output_type,
const char* tensor_type_rules,
bool convert_name);
SD_API bool preprocess_canny(sd_image_t image,
float high_threshold,
float low_threshold,
@@ -504,9 +528,20 @@ SD_API bool preprocess_canny(sd_image_t image,
float strong,
bool inverse);
SD_API bool load_imatrix(const char* imatrix_path);
SD_API void save_imatrix(const char* imatrix_path);
SD_API void enable_imatrix_collection(void);
SD_API void disable_imatrix_collection(void);
SD_API const char* sd_commit(void);
SD_API const char* sd_version(void);
// List available ggml backend devices, one `name<TAB>description` per line.
// The names are the device names accepted by the --backend / --params-backend
// assignment specs. Returns the number of bytes required, excluding the null
// terminator. Passing nullptr or buffer_size 0 only queries the required size.
SD_API size_t sd_list_devices(char* buffer, size_t buffer_size);
// for C API, caller needs to call free_sd_images to free the memory after use
// This helps avoid CRT problems on Windows when memory is allocated in the library but freed in the caller, which may use a different CRT.
SD_API void free_sd_images(sd_image_t* result_images, int num_images);
+234
View File
@@ -0,0 +1,234 @@
#!/usr/bin/env python3
"""Remove UTF-8 BOMs from files under a directory.
By default this scans the current working directory recursively and skips
repository areas that should not be touched by ordinary maintenance scripts.
Only files whose first three bytes are the UTF-8 BOM are rewritten.
"""
import argparse
import os
import shutil
import sys
import tempfile
from pathlib import Path
UTF8_BOM = b"\xef\xbb\xbf"
DEFAULT_EXCLUDED_DIR_NAMES = {
".git",
".hg",
".svn",
".mypy_cache",
".pytest_cache",
"__pycache__",
"test",
}
DEFAULT_EXCLUDED_DIR_PREFIXES = {
"build",
}
DEFAULT_EXCLUDED_REL_DIRS = {
"examples/server/frontend",
"ggml",
"models",
"src/vocab",
"thirdparty",
}
def rel_posix(path: Path, root: Path) -> str:
try:
return path.relative_to(root).as_posix()
except ValueError:
return path.as_posix()
def should_skip_dir(
path: Path,
root: Path,
excluded_rel_dirs: set[str],
excluded_names: set[str],
excluded_prefixes: set[str],
) -> bool:
rel = rel_posix(path, root)
return (
path.name in excluded_names
or rel in excluded_rel_dirs
or any(path.name.startswith(prefix) for prefix in excluded_prefixes)
)
def iter_files(
root: Path,
recursive: bool,
excluded_rel_dirs: set[str],
excluded_names: set[str],
excluded_prefixes: set[str],
follow_symlinks: bool,
):
if recursive:
for dirpath, dirnames, filenames in os.walk(root, followlinks=follow_symlinks):
current_dir = Path(dirpath)
dirnames[:] = [
name
for name in dirnames
if not should_skip_dir(
current_dir / name,
root,
excluded_rel_dirs,
excluded_names,
excluded_prefixes,
)
]
for filename in filenames:
path = current_dir / filename
if path.is_symlink() and not follow_symlinks:
continue
yield path
else:
for path in root.iterdir():
if path.is_file() and (follow_symlinks or not path.is_symlink()):
yield path
def has_utf8_bom(path: Path) -> bool:
with path.open("rb") as f:
return f.read(len(UTF8_BOM)) == UTF8_BOM
def strip_utf8_bom(path: Path) -> None:
tmp_path = None
try:
with path.open("rb") as src:
if src.read(len(UTF8_BOM)) != UTF8_BOM:
return
fd, tmp_name = tempfile.mkstemp(
prefix=f".{path.name}.",
suffix=".tmp",
dir=str(path.parent),
)
tmp_path = Path(tmp_name)
with os.fdopen(fd, "wb") as dst:
shutil.copyfileobj(src, dst, length=1024 * 1024)
shutil.copystat(path, tmp_path, follow_symlinks=False)
os.replace(tmp_path, path)
tmp_path = None
finally:
if tmp_path is not None:
try:
tmp_path.unlink()
except FileNotFoundError:
pass
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Scan files and convert UTF-8 BOM files to UTF-8 without BOM.",
)
parser.add_argument(
"root",
nargs="?",
default=".",
help="Directory to scan. Defaults to the current directory.",
)
parser.add_argument(
"-n",
"--dry-run",
action="store_true",
help="Only list files that would be converted.",
)
parser.add_argument(
"--no-recursive",
action="store_true",
help="Only scan files directly under root.",
)
parser.add_argument(
"--include-repo-excluded",
action="store_true",
help="Do not skip default repository excluded directories.",
)
parser.add_argument(
"--exclude-dir",
action="append",
default=[],
metavar="DIR",
help="Additional directory name or root-relative path to skip. Can be used multiple times.",
)
parser.add_argument(
"--follow-symlinks",
action="store_true",
help="Follow symlinked directories and files.",
)
parser.add_argument(
"-q",
"--quiet",
action="store_true",
help="Only print the final summary.",
)
return parser.parse_args()
def main() -> int:
args = parse_args()
root = Path(args.root).resolve()
if not root.is_dir():
print(f"error: not a directory: {root}", file=sys.stderr)
return 2
excluded_names = set()
excluded_rel_dirs = set()
excluded_prefixes = set()
if not args.include_repo_excluded:
excluded_names.update(DEFAULT_EXCLUDED_DIR_NAMES)
excluded_rel_dirs.update(DEFAULT_EXCLUDED_REL_DIRS)
excluded_prefixes.update(DEFAULT_EXCLUDED_DIR_PREFIXES)
for item in args.exclude_dir:
normalized = Path(item).as_posix().strip("/")
if "/" in normalized:
excluded_rel_dirs.add(normalized)
else:
excluded_names.add(normalized)
scanned = 0
converted = 0
errors = 0
for path in iter_files(
root=root,
recursive=not args.no_recursive,
excluded_rel_dirs=excluded_rel_dirs,
excluded_names=excluded_names,
excluded_prefixes=excluded_prefixes,
follow_symlinks=args.follow_symlinks,
):
scanned += 1
try:
if not has_utf8_bom(path):
continue
converted += 1
rel = rel_posix(path, root)
if args.dry_run:
if not args.quiet:
print(f"would convert: {rel}")
else:
strip_utf8_bom(path)
if not args.quiet:
print(f"converted: {rel}")
except OSError as exc:
errors += 1
print(f"error: {rel_posix(path, root)}: {exc}", file=sys.stderr)
action = "would convert" if args.dry_run else "converted"
print(f"scanned {scanned} file(s), {action} {converted}, errors {errors}")
return 1 if errors else 0
if __name__ == "__main__":
raise SystemExit(main())
+170 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_CONDITIONING_CONDITIONER_HPP__
#ifndef __SD_CONDITIONING_CONDITIONER_HPP__
#define __SD_CONDITIONING_CONDITIONER_HPP__
#include <cmath>
@@ -116,6 +116,8 @@ public:
virtual void get_param_tensors(std::map<std::string, ggml_tensor*>& tensors) = 0;
virtual void set_max_graph_vram_bytes(size_t max_vram_bytes) {}
virtual void set_stream_layers_enabled(bool enabled) {}
virtual void set_runtime_backends(const std::vector<ggml_backend_t>& backends) {}
virtual void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) {}
virtual void set_flash_attention_enabled(bool enabled) = 0;
virtual void set_weight_adapter(const std::shared_ptr<WeightAdapter>& adapter) {}
virtual void runner_done() {}
@@ -635,6 +637,18 @@ struct SD3CLIPEmbedder : public Conditioner {
}
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
if (t5) {
t5->set_runtime_backends(backends);
}
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
if (t5) {
t5->get_param_tensors(tensors, "text_encoders.t5xxl.transformer");
}
}
void set_flash_attention_enabled(bool enabled) override {
if (clip_l) {
clip_l->set_flash_attention_enabled(enabled);
@@ -994,6 +1008,18 @@ struct FluxCLIPEmbedder : public Conditioner {
}
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
if (t5) {
t5->set_runtime_backends(backends);
}
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
if (t5) {
t5->get_param_tensors(tensors, "text_encoders.t5xxl.transformer");
}
}
void set_flash_attention_enabled(bool enabled) override {
if (clip_l) {
clip_l->set_flash_attention_enabled(enabled);
@@ -1226,6 +1252,18 @@ struct T5CLIPEmbedder : public Conditioner {
}
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
if (t5) {
t5->set_runtime_backends(backends);
}
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
if (t5) {
t5->get_param_tensors(tensors, "text_encoders.t5xxl.transformer");
}
}
void set_flash_attention_enabled(bool enabled) override {
if (t5) {
t5->set_flash_attention_enabled(enabled);
@@ -1378,6 +1416,113 @@ struct T5CLIPEmbedder : public Conditioner {
}
};
struct MiniT2IConditioner : public Conditioner {
T5UniGramTokenizer tokenizer;
std::shared_ptr<T5Runner> t5;
size_t prompt_length = 256;
MiniT2IConditioner(ggml_backend_t backend,
const String2TensorStorage& tensor_storage_map = {},
std::shared_ptr<RunnerWeightManager> weight_manager = nullptr) {
bool use_t5 = false;
for (const auto& pair : tensor_storage_map) {
if (pair.first.find("text_encoders.t5xxl") != std::string::npos) {
use_t5 = true;
break;
}
}
if (!use_t5) {
LOG_WARN("IMPORTANT NOTICE: No MiniT2I T5 text encoder provided, cannot process prompts!");
return;
}
t5 = std::make_shared<T5Runner>(backend, tensor_storage_map, "text_encoders.t5xxl.transformer", false, weight_manager);
}
void get_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
if (t5) {
t5->get_param_tensors(tensors, "text_encoders.t5xxl.transformer");
}
}
void set_max_graph_vram_bytes(size_t max_vram_bytes) override {
if (t5) {
t5->set_max_graph_vram_bytes(max_vram_bytes);
}
}
void set_stream_layers_enabled(bool enabled) override {
if (t5) {
t5->set_stream_layers_enabled(enabled);
}
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
if (t5) {
t5->set_runtime_backends(backends);
}
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
if (t5) {
t5->get_param_tensors(tensors, "text_encoders.t5xxl.transformer");
}
}
void set_flash_attention_enabled(bool enabled) override {
if (t5) {
t5->set_flash_attention_enabled(enabled);
}
}
void set_weight_adapter(const std::shared_ptr<WeightAdapter>& adapter) override {
if (t5) {
t5->set_weight_adapter(adapter);
}
}
void runner_done() override {
if (t5) {
t5->runner_done();
}
}
SDCondition get_learned_condition(int n_threads,
const ConditionerParams& conditioner_params) override {
SDCondition result;
if (!t5) {
result.c_crossattn = sd::Tensor<float>::zeros({1024, static_cast<int64_t>(prompt_length)});
result.c_vector = sd::Tensor<float>::zeros({static_cast<int64_t>(prompt_length)});
return result;
}
std::vector<int> tokens = tokenizer.encode(conditioner_params.text);
if (tokens.size() > prompt_length) {
tokens.resize(prompt_length);
}
std::vector<float> mask(tokens.size(), 1.0f);
while (tokens.size() < prompt_length) {
tokens.push_back(tokenizer.PAD_TOKEN_ID);
mask.push_back(0.0f);
}
sd::Tensor<int32_t> input_ids({static_cast<int64_t>(tokens.size())}, tokens);
std::vector<float> t5_mask(mask.size(), 0.0f);
for (size_t i = 0; i < mask.size(); ++i) {
t5_mask[i] = mask[i] > 0.0f ? 0.0f : -HUGE_VALF;
}
sd::Tensor<float> hidden_states = t5->compute(n_threads,
input_ids,
sd::Tensor<float>::from_vector(t5_mask),
false,
true,
true);
GGML_ASSERT(!hidden_states.empty());
result.c_crossattn = std::move(hidden_states);
result.c_vector = sd::Tensor<float>::from_vector(mask);
return result;
}
};
struct AnimaConditioner : public Conditioner {
std::shared_ptr<BPETokenizer> qwen_tokenizer;
T5UniGramTokenizer t5_tokenizer;
@@ -1407,6 +1552,14 @@ struct AnimaConditioner : public Conditioner {
llm->set_stream_layers_enabled(enabled);
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
llm->set_runtime_backends(backends);
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
llm->get_param_tensors(tensors, "text_encoders.llm");
}
void set_flash_attention_enabled(bool enabled) override {
llm->set_flash_attention_enabled(enabled);
}
@@ -1552,6 +1705,14 @@ struct LLMEmbedder : public Conditioner {
llm->set_stream_layers_enabled(enabled);
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
llm->set_runtime_backends(backends);
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
llm->get_param_tensors(tensors, "text_encoders.llm");
}
void set_flash_attention_enabled(bool enabled) override {
llm->set_flash_attention_enabled(enabled);
}
@@ -2221,6 +2382,14 @@ struct LTXAVEmbedder : public Conditioner {
projector->set_max_graph_vram_bytes(max_vram_bytes);
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) override {
llm->set_runtime_backends(backends);
}
void get_layer_split_param_tensors(std::map<std::string, ggml_tensor*>& tensors) override {
llm->get_param_tensors(tensors, "text_encoders.llm");
}
void set_weight_adapter(const std::shared_ptr<WeightAdapter>& adapter) override {
llm->set_weight_adapter(adapter);
projector->set_weight_adapter(adapter);
+65 -20
View File
@@ -76,29 +76,22 @@ static bool load_tensors_for_export(ModelLoader& model_loader,
return success;
}
bool convert(const char* input_path,
const char* vae_path,
const char* output_path,
sd_type_t output_type,
const char* tensor_type_rules,
bool convert_name) {
ModelLoader model_loader;
if (!model_loader.init_from_file(input_path)) {
LOG_ERROR("init model loader from file failed: '%s'", input_path);
static bool init_convert_path(ModelLoader& model_loader, const char* path, const char* prefix, bool& loaded_any) {
if (path == nullptr || strlen(path) == 0) {
return true;
}
if (!model_loader.init_from_file(path, prefix)) {
LOG_ERROR("init model loader from file failed: '%s'", path);
return false;
}
loaded_any = true;
return true;
}
if (vae_path != nullptr && strlen(vae_path) > 0) {
if (!model_loader.init_from_file(vae_path, "vae.")) {
LOG_ERROR("init model loader from file failed: '%s'", vae_path);
return false;
}
}
if (convert_name) {
model_loader.convert_tensors_name();
}
static bool export_loaded_model(ModelLoader& model_loader,
const char* output_path,
sd_type_t output_type,
const char* tensor_type_rules) {
ggml_type type = sd_type_to_ggml_type(output_type);
bool output_is_safetensors = ends_with(output_path, ".safetensors");
TensorTypeRules type_rules = parse_tensor_type_rules(tensor_type_rules);
@@ -136,3 +129,55 @@ bool convert(const char* input_path,
ggml_free(ggml_ctx);
return success;
}
bool convert_with_components(const char* model_path,
const char* clip_l_path,
const char* clip_g_path,
const char* t5xxl_path,
const char* diffusion_model_path,
const char* vae_path,
const char* output_path,
sd_type_t output_type,
const char* tensor_type_rules,
bool convert_name) {
ModelLoader model_loader;
bool loaded_any = false;
if (!init_convert_path(model_loader, model_path, "", loaded_any) ||
!init_convert_path(model_loader, clip_l_path, "text_encoders.clip_l.transformer.", loaded_any) ||
!init_convert_path(model_loader, clip_g_path, "text_encoders.clip_g.transformer.", loaded_any) ||
!init_convert_path(model_loader, t5xxl_path, "text_encoders.t5xxl.transformer.", loaded_any) ||
!init_convert_path(model_loader, diffusion_model_path, "model.diffusion_model.", loaded_any) ||
!init_convert_path(model_loader, vae_path, "vae.", loaded_any)) {
return false;
}
if (!loaded_any) {
LOG_ERROR("no input model path provided for convert");
return false;
}
if (convert_name) {
model_loader.convert_tensors_name();
}
return export_loaded_model(model_loader, output_path, output_type, tensor_type_rules);
}
bool convert(const char* input_path,
const char* vae_path,
const char* output_path,
sd_type_t output_type,
const char* tensor_type_rules,
bool convert_name) {
return convert_with_components(input_path,
nullptr,
nullptr,
nullptr,
nullptr,
vae_path,
output_path,
output_type,
tensor_type_rules,
convert_name);
}
+390
View File
@@ -0,0 +1,390 @@
#include "backend_fit.h"
#include <algorithm>
#include <cctype>
#include <cstdint>
#include <utility>
#include <vector>
#include "core/ggml_extend_backend.h"
#include "core/util.h"
#include "ggml-backend.h"
namespace sd::backend_fit {
namespace {
constexpr int64_t MiB = 1024ll * 1024;
enum class ComponentKind {
DIT = 0,
VAE = 1,
CONDITIONER = 2,
};
struct Component {
ComponentKind kind;
const char* name;
int64_t params_bytes = 0;
int64_t reserve_bytes = 0;
bool splittable = false;
};
struct Device {
ggml_backend_dev_t dev = nullptr;
std::string name;
std::string description;
int64_t free_bytes = 0;
int64_t total_bytes = 0;
int64_t budget_bytes = 0;
};
struct Decision {
ComponentKind kind;
bool on_cpu = false;
std::vector<size_t> device_idxs;
};
struct Plan {
bool valid = false;
bool time_share = false;
std::vector<Decision> decisions;
};
bool classify_tensor(const std::string& name, ComponentKind& out) {
auto contains = [&](const char* s) { return name.find(s) != std::string::npos; };
if (contains("model.diffusion_model.") || contains("unet.")) {
out = ComponentKind::DIT;
return true;
}
if (contains("first_stage_model.") ||
name.rfind("vae.", 0) == 0 ||
name.rfind("tae.", 0) == 0) {
out = ComponentKind::VAE;
return true;
}
if (contains("text_encoders") ||
contains("cond_stage_model") ||
contains("te.text_model.") ||
contains("conditioner") ||
name.rfind("text_encoder.", 0) == 0 ||
name.rfind("text_embedding_projection.", 0) == 0 ||
contains(".aggregate_embed.")) {
out = ComponentKind::CONDITIONER;
return true;
}
return false;
}
std::vector<Component> estimate_components(ModelLoader& loader, ggml_type override_wtype) {
const auto& storage = loader.get_tensor_storage_map();
int64_t bytes[3] = {0, 0, 0};
for (const auto& [name, ts_const] : storage) {
TensorStorage ts = ts_const;
if (is_unused_tensor(ts.name)) {
continue;
}
ComponentKind kind;
if (!classify_tensor(ts.name, kind)) {
continue;
}
if (override_wtype != GGML_TYPE_COUNT &&
loader.tensor_should_be_converted(ts, override_wtype)) {
ts.type = override_wtype;
} else if (ts.expected_type != GGML_TYPE_COUNT && ts.expected_type != ts.type) {
ts.type = ts.expected_type;
}
bytes[int(kind)] += (int64_t)ts.nbytes() + 64;
}
std::vector<Component> out;
out.push_back({ComponentKind::DIT, "DiT", bytes[int(ComponentKind::DIT)], 2048 * MiB, true});
out.push_back({ComponentKind::VAE, "VAE", bytes[int(ComponentKind::VAE)], 1024 * MiB, false});
out.push_back({ComponentKind::CONDITIONER, "Conditioner", bytes[int(ComponentKind::CONDITIONER)], 2048 * MiB, true});
return out;
}
std::vector<Device> enumerate_gpu_devices(const sd::ggml_graph_cut::MaxVramAssignment& budgets) {
std::vector<Device> out;
for (size_t i = 0; i < ggml_backend_dev_count(); i++) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
if (ggml_backend_dev_type(dev) != GGML_BACKEND_DEVICE_TYPE_GPU) {
continue;
}
Device d;
d.dev = dev;
d.name = ggml_backend_dev_name(dev);
d.description = ggml_backend_dev_description(dev);
size_t free_bytes = 0, total_bytes = 0;
ggml_backend_dev_memory(dev, &free_bytes, &total_bytes);
d.free_bytes = (int64_t)free_bytes;
d.total_bytes = (int64_t)total_bytes;
std::string budget_key = d.name;
std::transform(budget_key.begin(), budget_key.end(), budget_key.begin(),
[](unsigned char c) { return (char)std::tolower(c); });
float gib = budgets.default_gib;
auto it = budgets.backend_gib.find(budget_key);
if (it != budgets.backend_gib.end()) {
gib = it->second;
}
if (gib > 0.f) {
d.budget_bytes = std::min<int64_t>((int64_t)(gib * 1024.0 * 1024.0 * 1024.0), d.free_bytes);
} else if (gib < 0.f) {
d.budget_bytes = d.free_bytes + (int64_t)(gib * 1024.0 * 1024.0 * 1024.0);
} else {
d.budget_bytes = d.free_bytes - 512 * MiB;
}
d.budget_bytes = std::max<int64_t>(d.budget_bytes, 0);
out.push_back(d);
}
return out;
}
Plan compute_plan(const std::vector<Component>& components, const std::vector<Device>& devices) {
Plan plan;
if (devices.empty()) {
return plan;
}
std::vector<size_t> order(components.size());
for (size_t i = 0; i < order.size(); i++) {
order[i] = i;
}
std::sort(order.begin(), order.end(), [&](size_t a, size_t b) {
return components[a].params_bytes > components[b].params_bytes;
});
{
std::vector<int64_t> params_sum(devices.size(), 0);
std::vector<int64_t> max_reserve(devices.size(), 0);
std::vector<Decision> decisions(components.size());
bool ok = true;
for (size_t ci : order) {
const Component& comp = components[ci];
decisions[ci].kind = comp.kind;
if (comp.params_bytes == 0) {
continue;
}
int best = -1;
for (size_t di = 0; di < devices.size(); di++) {
int64_t need = params_sum[di] + comp.params_bytes + std::max(max_reserve[di], comp.reserve_bytes);
if (need <= devices[di].budget_bytes &&
(best < 0 || devices[di].budget_bytes - params_sum[di] > devices[best].budget_bytes - params_sum[best])) {
best = (int)di;
}
}
if (best < 0) {
ok = false;
break;
}
params_sum[best] += comp.params_bytes;
max_reserve[best] = std::max(max_reserve[best], comp.reserve_bytes);
decisions[ci].device_idxs.push_back((size_t)best);
}
if (ok) {
plan.valid = true;
plan.time_share = false;
plan.decisions = std::move(decisions);
return plan;
}
}
plan.decisions.assign(components.size(), {});
for (size_t ci : order) {
const Component& comp = components[ci];
Decision& decision = plan.decisions[ci];
decision.kind = comp.kind;
if (comp.params_bytes == 0) {
continue;
}
int best = -1;
for (size_t di = 0; di < devices.size(); di++) {
if (comp.params_bytes + comp.reserve_bytes <= devices[di].budget_bytes &&
(best < 0 || devices[di].budget_bytes > devices[best].budget_bytes)) {
best = (int)di;
}
}
if (best >= 0) {
decision.device_idxs.push_back((size_t)best);
continue;
}
if (comp.splittable && devices.size() > 1) {
int64_t capacity = 0;
for (const Device& d : devices) {
capacity += std::max<int64_t>(d.budget_bytes - comp.reserve_bytes, 0);
}
if (comp.params_bytes <= capacity) {
std::vector<size_t> idxs(devices.size());
for (size_t i = 0; i < idxs.size(); i++) {
idxs[i] = i;
}
std::sort(idxs.begin(), idxs.end(), [&](size_t a, size_t b) {
return devices[a].budget_bytes > devices[b].budget_bytes;
});
decision.device_idxs = std::move(idxs);
continue;
}
}
decision.on_cpu = true;
}
plan.valid = true;
plan.time_share = true;
return plan;
}
void print_plan(const Plan& plan,
const std::vector<Component>& components,
const std::vector<Device>& devices) {
LOG_INFO("auto-fit plan%s:", plan.time_share ? " (time-share: params load per phase and free after)" : "");
LOG_INFO(" devices:");
for (const Device& d : devices) {
LOG_INFO(" %-12s %-32s free %6lld MiB, budget %6lld MiB",
d.name.c_str(), d.description.c_str(),
(long long)(d.free_bytes / MiB), (long long)(d.budget_bytes / MiB));
}
LOG_INFO(" components:");
for (size_t ci = 0; ci < components.size(); ci++) {
const Component& comp = components[ci];
const Decision& decision = plan.decisions[ci];
std::string target;
if (comp.params_bytes == 0) {
target = "(not present)";
} else if (decision.on_cpu) {
target = "CPU";
} else {
for (size_t k = 0; k < decision.device_idxs.size(); k++) {
if (k > 0) {
target += " & ";
}
target += devices[decision.device_idxs[k]].name;
}
if (decision.device_idxs.size() > 1) {
target += " (split)";
}
}
LOG_INFO(" %-12s params %6lld MiB, compute reserve %5lld MiB -> %s",
comp.name,
(long long)(comp.params_bytes / MiB),
(long long)(comp.reserve_bytes / MiB),
target.c_str());
}
}
void append_assignment(std::string& spec, const char* key, const std::string& value) {
if (!spec.empty()) {
spec += ",";
}
spec += key;
spec += "=";
spec += value;
}
void append_component_decision(const std::vector<Component>& components,
const std::vector<Device>& devices,
const Plan& plan,
ComponentKind kind,
const char* module_key,
std::string& runtime_spec,
std::string& params_spec) {
for (size_t ci = 0; ci < components.size(); ci++) {
if (components[ci].kind != kind || components[ci].params_bytes == 0) {
continue;
}
const Decision& decision = plan.decisions[ci];
if (decision.on_cpu) {
append_assignment(runtime_spec, module_key, "cpu");
return;
}
if (decision.device_idxs.empty()) {
return;
}
std::string device_list;
for (size_t k = 0; k < decision.device_idxs.size(); k++) {
if (k > 0) {
device_list += "&";
}
device_list += devices[decision.device_idxs[k]].name;
}
append_assignment(runtime_spec, module_key, device_list);
if (plan.time_share) {
append_assignment(params_spec, module_key, "disk");
}
return;
}
}
} // namespace
bool derive_backend_specs(ModelLoader& loader,
ggml_type override_wtype,
sd::ggml_graph_cut::MaxVramAssignment& budgets,
std::string& runtime_spec,
std::string& params_spec) {
if (!runtime_spec.empty() || !params_spec.empty()) {
LOG_WARN("--auto-fit is enabled; ignoring --backend / --params-backend");
}
{
std::string error;
if (!budgets.canonicalize_backend_keys(&error)) {
LOG_ERROR("%s", error.c_str());
return false;
}
}
auto components = estimate_components(loader, override_wtype);
auto devices = enumerate_gpu_devices(budgets);
auto plan = compute_plan(components, devices);
if (!plan.valid) {
LOG_WARN("auto-fit: no usable GPU devices; using the default backend");
runtime_spec.clear();
params_spec.clear();
return true;
}
print_plan(plan, components, devices);
std::string derived_runtime_spec;
std::string derived_params_spec;
append_component_decision(components, devices, plan, ComponentKind::DIT, "diffusion", derived_runtime_spec, derived_params_spec);
append_component_decision(components, devices, plan, ComponentKind::CONDITIONER, "te", derived_runtime_spec, derived_params_spec);
append_component_decision(components, devices, plan, ComponentKind::VAE, "vae", derived_runtime_spec, derived_params_spec);
runtime_spec = std::move(derived_runtime_spec);
params_spec = std::move(derived_params_spec);
LOG_INFO("auto-fit: --backend \"%s\"%s%s%s",
runtime_spec.empty() ? "(default)" : runtime_spec.c_str(),
params_spec.empty() ? "" : " --params-backend \"",
params_spec.c_str(),
params_spec.empty() ? "" : "\"");
return true;
}
bool prepare_vae_decode_retry_tiling(sd_tiling_params_t& tiling_params, bool prefer_temporal_tiling) {
if (prefer_temporal_tiling) {
if (tiling_params.temporal_tiling) {
return false;
}
tiling_params.temporal_tiling = true;
} else {
if (tiling_params.enabled) {
return false;
}
tiling_params.enabled = true;
if (tiling_params.tile_size_x <= 0) {
tiling_params.tile_size_x = 256;
}
if (tiling_params.tile_size_y <= 0) {
tiling_params.tile_size_y = 256;
}
}
LOG_WARN("auto-fit: VAE decode failed (likely out of memory); retrying with %s tiling",
tiling_params.temporal_tiling ? "temporal" : "spatial");
return true;
}
} // namespace sd::backend_fit
+23
View File
@@ -0,0 +1,23 @@
#ifndef __SD_BACKEND_FIT_H__
#define __SD_BACKEND_FIT_H__
#include <string>
#include "core/ggml_graph_cut.h"
#include "model_loader.h"
#include "stable-diffusion.h"
namespace sd::backend_fit {
bool derive_backend_specs(ModelLoader& loader,
ggml_type override_wtype,
sd::ggml_graph_cut::MaxVramAssignment& budgets,
std::string& runtime_spec,
std::string& params_spec);
bool prepare_vae_decode_retry_tiling(sd_tiling_params_t& tiling_params,
bool prefer_temporal_tiling);
} // namespace sd::backend_fit
#endif // __SD_BACKEND_FIT_H__
+224 -13
View File
@@ -391,7 +391,7 @@ __STATIC_INLINE__ uint8_t* ggml_tensor_to_sd_image(ggml_tensor* input, uint8_t*
int64_t width = input->ne[0];
int64_t height = input->ne[1];
int64_t channels = input->ne[2];
GGML_ASSERT(channels == 3 && input->type == GGML_TYPE_F32);
GGML_ASSERT(input->type == GGML_TYPE_F32);
if (image_data == nullptr) {
image_data = (uint8_t*)malloc(width * height * channels);
}
@@ -1038,6 +1038,7 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_linear(ggml_context* ctx,
}
__STATIC_INLINE__ ggml_tensor* ggml_ext_pad_ext(ggml_context* ctx,
ggml_backend_t backend,
ggml_tensor* x,
int lp0,
int rp0,
@@ -1063,7 +1064,17 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_pad_ext(ggml_context* ctx,
}
if (lp0 != 0 || rp0 != 0 || lp1 != 0 || rp1 != 0 || lp2 != 0 || rp2 != 0 || lp3 != 0 || rp3 != 0) {
x = ggml_pad_ext(ctx, x, lp0, rp0, lp1, rp1, lp2, rp2, lp3, rp3);
ggml_tensor* padded = ggml_pad_ext(ctx, x, lp0, rp0, lp1, rp1, lp2, rp2, lp3, rp3);
if (backend == nullptr || ggml_backend_supports_op(backend, padded)) {
x = padded;
} else {
// Some backends (e.g. Metal) only implement right-padding for
// GGML_OP_PAD (see #850): pad right by lp+rp instead, then roll
// the padding around to the left. shift < ne always holds because
// ne grew by lp+rp.
x = ggml_pad_ext(ctx, x, 0, lp0 + rp0, 0, lp1 + rp1, 0, lp2 + rp2, 0, lp3 + rp3);
x = ggml_roll(ctx, x, lp0, lp1, lp2, lp3);
}
}
return x;
}
@@ -1076,7 +1087,7 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_pad(ggml_context* ctx,
int p3 = 0,
bool circular_x = false,
bool circular_y = false) {
return ggml_ext_pad_ext(ctx, x, 0, p0, 0, p1, 0, p2, 0, p3, circular_x, circular_y);
return ggml_ext_pad_ext(ctx, nullptr, x, 0, p0, 0, p1, 0, p2, 0, p3, circular_x, circular_y);
}
// w: [OC,IC, KH, KW]
@@ -1105,7 +1116,7 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_conv_2d(ggml_context* ctx,
}
if ((p0 != 0 || p1 != 0) && (circular_x || circular_y)) {
x = ggml_ext_pad_ext(ctx, x, p0, p0, p1, p1, 0, 0, 0, 0, circular_x, circular_y);
x = ggml_ext_pad_ext(ctx, nullptr, x, p0, p0, p1, p1, 0, 0, 0, 0, circular_x, circular_y);
p0 = 0;
p1 = 0;
}
@@ -1130,6 +1141,7 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_conv_2d(ggml_context* ctx,
// b: [OC,]
// result: [N*OC, OD, OH, OW]
__STATIC_INLINE__ ggml_tensor* ggml_ext_conv_3d(ggml_context* ctx,
ggml_backend_t backend,
ggml_tensor* x,
ggml_tensor* w,
ggml_tensor* b,
@@ -1159,7 +1171,21 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_conv_3d(ggml_context* ctx,
x = ggml_cont(ctx, ggml_permute(ctx, x, 0, 1, 3, 2));
x = ggml_reshape_4d(ctx, x, im2col->ne[1], im2col->ne[2], OD, OC * N);
} else {
x = ggml_conv_3d(ctx, w, x, IC, s0, s1, s2, p0, p1, p2, d0, d1, d2);
// ggml_conv_3d decomposes into GGML_OP_IM2COL_3D, which some backends
// (e.g. Metal, see #850) do not implement. Fall back to
// GGML_OP_CONV_3D on those backends.
bool im2col_3d_supported = true;
if (backend != nullptr) {
ggml_tensor* im2col = ggml_im2col_3d(ctx, w, x, IC, s0, s1, s2, p0, p1, p2, d0, d1, d2, w->type);
im2col_3d_supported = ggml_backend_supports_op(backend, im2col);
}
if (im2col_3d_supported) {
x = ggml_conv_3d(ctx, w, x, IC, s0, s1, s2, p0, p1, p2, d0, d1, d2);
} else {
int64_t OC = w->ne[3] / IC;
int64_t N = x->ne[3] / IC;
x = ggml_conv_3d_direct(ctx, w, x, s0, s1, s2, p0, p1, p2, d0, d1, d2, (int)IC, (int)N, (int)OC);
}
}
if (b != nullptr) {
@@ -1362,6 +1388,9 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_attention_ext(ggml_context* ctx,
}
auto out = ggml_flash_attn_ext(ctx, q_in, k_in, v_in, mask_in, scale / kv_scale, 0, 0);
if (!ggml_backend_supports_op(backend, out)) {
return nullptr;
}
ggml_flash_attn_ext_set_prec(out, GGML_PREC_F32);
if (kv_scale != 1.0f) {
out = ggml_ext_scale(ctx, out, 1.0f / kv_scale);
@@ -1379,9 +1408,7 @@ __STATIC_INLINE__ ggml_tensor* ggml_ext_attention_ext(ggml_context* ctx,
if (can_use_flash_attn) {
kqv = build_kqv(q, k, v, mask);
if (!ggml_backend_supports_op(backend, kqv)) {
kqv = nullptr;
} else {
if (kqv != nullptr) {
kqv = ggml_view_4d(ctx,
kqv,
d_head,
@@ -1719,6 +1746,11 @@ protected:
bool stream_layers_enabled = false;
size_t observed_max_effective_budget_ = 0;
std::vector<ggml_backend_t> extra_runtime_backends; // borrowed (SDBackendManager-owned)
ggml_backend_sched_t sched = nullptr; // owned, multi-device only
ggml_backend_t cpu_fallback_backend = nullptr; // owned, sched requires a trailing CPU backend
bool multi_device_eval_callback_warned = false;
std::shared_ptr<WeightAdapter> weight_adapter = nullptr;
std::weak_ptr<RunnerWeightManager> weight_manager;
std::unordered_set<const ggml_tensor*> kept_compute_param_tensor_set;
@@ -1986,7 +2018,121 @@ protected:
return true;
}
// Pass explicit buffer types: synthesized defaults can make CUDA devices
// report supporting each other's buffers and skip a required copy.
bool ensure_sched(ggml_cgraph* gf) {
if (sched != nullptr) {
return true;
}
std::vector<ggml_backend_t> backends;
backends.reserve(extra_runtime_backends.size() + 2);
backends.push_back(runtime_backend);
for (ggml_backend_t backend : extra_runtime_backends) {
backends.push_back(backend);
}
if (cpu_fallback_backend == nullptr && !sd_backend_is_cpu(runtime_backend)) {
cpu_fallback_backend = sd_backend_cpu_init();
}
if (cpu_fallback_backend != nullptr) {
backends.push_back(cpu_fallback_backend);
}
std::vector<ggml_backend_buffer_type_t> bufts;
bufts.reserve(backends.size());
ggml_backend_dev_t main_dev = ggml_backend_get_device(runtime_backend);
for (ggml_backend_t backend : backends) {
ggml_backend_buffer_type_t buft = nullptr;
if (backend == cpu_fallback_backend && main_dev != nullptr) {
buft = ggml_backend_dev_host_buffer_type(main_dev);
}
if (buft == nullptr) {
buft = ggml_backend_get_default_buffer_type(backend);
}
bufts.push_back(buft);
}
size_t graph_size = MAX_GRAPH_SIZE;
if (gf != nullptr) {
graph_size = std::max<size_t>(graph_size, (size_t)ggml_graph_n_nodes(gf));
}
sched = ggml_backend_sched_new(backends.data(),
bufts.data(),
(int)backends.size(),
graph_size,
/*parallel=*/false,
/*op_offload=*/false);
if (sched == nullptr) {
LOG_ERROR("%s: failed to create backend sched", get_desc().c_str());
return false;
}
return true;
}
ggml_backend_t backend_for_weight(const ggml_tensor* tensor) const {
if (tensor == nullptr || tensor->buffer == nullptr) {
return nullptr;
}
if (ggml_backend_buffer_get_usage(tensor->buffer) != GGML_BACKEND_BUFFER_USAGE_WEIGHTS ||
ggml_backend_buffer_is_host(tensor->buffer)) {
return nullptr;
}
ggml_backend_dev_t dev = ggml_backend_buft_get_device(ggml_backend_buffer_get_type(tensor->buffer));
if (dev == nullptr) {
return nullptr;
}
if (ggml_backend_get_device(runtime_backend) == dev) {
return runtime_backend;
}
for (ggml_backend_t backend : extra_runtime_backends) {
if (ggml_backend_get_device(backend) == dev) {
return backend;
}
}
return nullptr;
}
// Weightless ops have no scheduler anchor, so pin them to the most recent
// weight device. Views must stay unpinned or cross-device copies can be
// skipped for their consumers.
void pin_multi_device_nodes(ggml_cgraph* gf) {
if (sched == nullptr || gf == nullptr) {
return;
}
ggml_backend_t current = runtime_backend;
const int n_nodes = ggml_graph_n_nodes(gf);
for (int i = 0; i < n_nodes; i++) {
ggml_tensor* node = ggml_graph_node(gf, i);
for (int s = 0; s < GGML_MAX_SRC; s++) {
ggml_backend_t weight_backend = backend_for_weight(node->src[s]);
if (weight_backend != nullptr) {
current = weight_backend;
}
}
if (node->op == GGML_OP_NONE || node->op == GGML_OP_VIEW || node->op == GGML_OP_RESHAPE ||
node->op == GGML_OP_PERMUTE || node->op == GGML_OP_TRANSPOSE) {
continue;
}
if (ggml_backend_supports_op(current, node)) {
ggml_backend_sched_set_tensor_backend(sched, node, current);
}
}
}
bool is_multi_device() const {
return !extra_runtime_backends.empty();
}
bool alloc_compute_buffer(ggml_cgraph* gf) {
if (is_multi_device()) {
// The sched replaces the gallocr. Do NOT ggml_backend_sched_reserve
// the graph here: reserve runs split_graph, which rewires the
// graph's src pointers to sched-internal copy tensors, and the
// later ggml_backend_sched_alloc_graph would split the already
// rewired graph, silently corrupting every cross-backend input. A
// graph must be split at most once; the alloc in execute_graph
// performs the real allocation.
return ensure_sched(gf);
}
if (compute_allocr != nullptr) {
return true;
}
@@ -2202,12 +2348,14 @@ protected:
plan.valid &&
max_graph_vram_bytes > 0 &&
plan.segments.size() > 1 &&
!sd_backend_is_cpu(runtime_backend);
!sd_backend_is_cpu(runtime_backend) &&
!is_multi_device();
}
bool can_attempt_graph_cut_segmented_compute() const {
return max_graph_vram_bytes > 0 &&
!sd_backend_is_cpu(runtime_backend);
!sd_backend_is_cpu(runtime_backend) &&
!is_multi_device();
}
bool resolve_graph_cut_plan(ggml_cgraph* gf,
@@ -2463,7 +2611,14 @@ protected:
};
ComputeBufferGuard compute_buffer_guard(this, free_compute_buffer);
if (!ggml_gallocr_alloc_graph(compute_allocr, gf)) {
if (is_multi_device()) {
ggml_backend_sched_reset(sched);
pin_multi_device_nodes(gf); // reset clears the pins; re-apply before alloc
if (!ggml_backend_sched_alloc_graph(sched, gf)) {
LOG_ERROR("%s sched alloc compute graph failed", get_desc().c_str());
return std::nullopt;
}
} else if (!ggml_gallocr_alloc_graph(compute_allocr, gf)) {
LOG_ERROR("%s alloc compute graph failed", get_desc().c_str());
return std::nullopt;
}
@@ -2472,8 +2627,27 @@ protected:
if (sd_backend_is_cpu(runtime_backend)) {
sd_backend_cpu_set_n_threads(runtime_backend, n_threads);
}
if (cpu_fallback_backend != nullptr) {
sd_backend_cpu_set_n_threads(cpu_fallback_backend, n_threads);
}
ggml_status status = ggml_backend_graph_compute(runtime_backend, gf);
ggml_status status;
if (is_multi_device()) {
if (sd_get_backend_eval_callback() != nullptr && !multi_device_eval_callback_warned) {
LOG_WARN("%s: eval callback is not supported with multiple runtime backends; ignoring",
get_desc().c_str());
multi_device_eval_callback_warned = true;
}
status = ggml_backend_sched_graph_compute(sched, gf);
if (status == GGML_STATUS_SUCCESS) {
ggml_backend_sched_synchronize(sched);
}
} else {
status = sd_backend_graph_compute_with_eval_callback(runtime_backend,
gf,
sd_get_backend_eval_callback(),
sd_get_backend_eval_callback_data());
}
if (status != GGML_STATUS_SUCCESS) {
LOG_ERROR("%s compute failed: %s", get_desc().c_str(), ggml_status_to_string(status));
return std::nullopt;
@@ -2650,6 +2824,10 @@ public:
free_params_ctx();
free_compute_ctx();
free_cache_ctx_and_buffer();
if (cpu_fallback_backend != nullptr) {
ggml_backend_free(cpu_fallback_backend);
cpu_fallback_backend = nullptr;
}
}
virtual GGMLRunnerContext get_context() {
@@ -2690,10 +2868,20 @@ public:
ggml_gallocr_free(compute_allocr);
compute_allocr = nullptr;
}
if (sched != nullptr) {
ggml_backend_sched_free(sched);
sched = nullptr;
}
}
// do copy after alloc graph
void set_backend_tensor_data(ggml_tensor* tensor, const void* data) {
if (is_multi_device()) {
// The sched only assigns a backend (and thus a buffer) to tensors
// that participate in the graph; flag standalone data tensors as
// inputs so they get one.
ggml_set_input(tensor);
}
backend_tensor_data_map[tensor] = data;
}
@@ -2829,8 +3017,31 @@ public:
}
void set_stream_layers_enabled(bool enabled) {
if (enabled && is_multi_device()) {
LOG_WARN("%s: --stream-layers is not supported with multiple runtime backends; ignoring",
get_desc().c_str());
return;
}
stream_layers_enabled = enabled;
}
void set_runtime_backends(const std::vector<ggml_backend_t>& backends) {
extra_runtime_backends.clear();
for (ggml_backend_t backend : backends) {
if (backend == nullptr || backend == runtime_backend) {
continue;
}
if (std::find(extra_runtime_backends.begin(), extra_runtime_backends.end(), backend) ==
extra_runtime_backends.end()) {
extra_runtime_backends.push_back(backend);
}
}
if (is_multi_device() && stream_layers_enabled) {
LOG_WARN("%s: --stream-layers is not supported with multiple runtime backends; ignoring",
get_desc().c_str());
stream_layers_enabled = false;
}
}
};
class GGMLBlock {
@@ -3365,7 +3576,7 @@ public:
b = ctx->weight_adapter->patch_weight(ctx->ggml_ctx, ctx->backend, b, prefix + "bias");
}
}
return ggml_ext_conv_3d(ctx->ggml_ctx, x, w, b, in_channels,
return ggml_ext_conv_3d(ctx->ggml_ctx, ctx->backend, x, w, b, in_channels,
std::get<2>(stride), std::get<1>(stride), std::get<0>(stride),
std::get<2>(padding), std::get<1>(padding), std::get<0>(padding),
std::get<2>(dilation), std::get<1>(dilation), std::get<0>(dilation),
+280 -8
View File
@@ -9,6 +9,7 @@
#include <vector>
#include "core/util.h"
#include "ggml/src/ggml-impl.h"
#include "stable-diffusion.h"
static std::string trim_copy(const std::string& value) {
@@ -110,7 +111,67 @@ static std::string resolve_first_device_by_type(enum ggml_backend_dev_type type)
if (dev == nullptr) {
return "";
}
return ggml_backend_dev_name(dev);
const char* dev_name = ggml_backend_dev_name(dev);
if (dev_name != nullptr && dev_name[0] != '\0') {
return dev_name;
}
ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev);
const char* reg_name = reg != nullptr ? ggml_backend_reg_name(reg) : nullptr;
return reg_name != nullptr ? reg_name : "";
}
static ggml_backend_dev_t resolve_first_device_by_registry_name(const std::string& name) {
std::string lower = lower_copy(trim_copy(name));
if (lower == "metal") {
lower = "mtl";
}
if (lower.empty()) {
return nullptr;
}
const size_t device_count = ggml_backend_dev_count();
for (size_t i = 0; i < device_count; ++i) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev);
if (reg == nullptr) {
continue;
}
const char* reg_name = ggml_backend_reg_name(reg);
if (reg_name != nullptr && lower_copy(reg_name) == lower) {
return dev;
}
}
return nullptr;
}
static ggml_backend_dev_t resolve_device_by_name(const std::string& name) {
const std::string lower = lower_copy(trim_copy(name));
if (lower.empty()) {
return nullptr;
}
const size_t device_count = ggml_backend_dev_count();
for (size_t i = 0; i < device_count; ++i) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
const char* dev_name = ggml_backend_dev_name(dev);
if (dev_name != nullptr && lower_copy(dev_name) == lower) {
return dev;
}
}
return nullptr;
}
static std::string backend_device_name(ggml_backend_dev_t dev) {
if (dev == nullptr) {
return "";
}
const char* name = ggml_backend_dev_name(dev);
if (name != nullptr && name[0] != '\0') {
return name;
}
ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev);
const char* reg_name = reg != nullptr ? ggml_backend_reg_name(reg) : nullptr;
return reg_name != nullptr ? reg_name : "";
}
static ggml_backend_buffer_t ggml_backend_tensor_buffer(const struct ggml_tensor* tensor) {
@@ -296,6 +357,10 @@ std::string sd_backend_resolve_name(const std::string& name) {
return resolve_first_device_by_type(GGML_BACKEND_DEVICE_TYPE_IGPU);
}
if (ggml_backend_dev_t dev = resolve_first_device_by_registry_name(requested)) {
return backend_device_name(dev);
}
const size_t device_count = ggml_backend_dev_count();
for (size_t i = 0; i < device_count; ++i) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
@@ -328,7 +393,20 @@ static ggml_backend_t init_named_backend(const std::string& name) {
return ggml_backend_init_best();
}
if (ggml_backend_dev_t dev = resolve_device_by_name(name)) {
return ggml_backend_dev_init(dev, nullptr);
}
if (ggml_backend_dev_t dev = resolve_first_device_by_registry_name(name)) {
return ggml_backend_dev_init(dev, nullptr);
}
std::string resolved = sd_backend_resolve_name(name);
if (ggml_backend_dev_t dev = resolve_device_by_name(resolved)) {
return ggml_backend_dev_init(dev, nullptr);
}
if (ggml_backend_dev_t dev = resolve_first_device_by_registry_name(resolved)) {
return ggml_backend_dev_init(dev, nullptr);
}
if (resolved.empty()) {
return nullptr;
}
@@ -364,6 +442,68 @@ bool sd_backend_cpu_set_n_threads(ggml_backend_t backend, int n_threads) {
return false;
}
static ggml_cgraph sd_ggml_graph_view(ggml_cgraph* cgraph0, int i0, int i1) {
ggml_cgraph cgraph = {
/*.size =*/0,
/*.n_nodes =*/i1 - i0,
/*.n_leafs =*/0,
/*.nodes =*/cgraph0->nodes + i0,
/*.grads =*/nullptr,
/*.grad_accs =*/nullptr,
/*.leafs =*/nullptr,
/*.use_counts =*/cgraph0->use_counts,
/*.visited_hash_set =*/cgraph0->visited_hash_set,
/*.order =*/cgraph0->order,
/*.uid =*/0,
};
return cgraph;
}
ggml_status sd_backend_graph_compute_with_eval_callback(ggml_backend_t backend,
ggml_cgraph* gf,
sd_graph_eval_callback_t callback_eval,
void* callback_eval_user_data) {
if (callback_eval == nullptr) {
return ggml_backend_graph_compute(backend, gf);
}
ggml_status status = GGML_STATUS_SUCCESS;
const int n_nodes = ggml_graph_n_nodes(gf);
bool stopped = false;
for (int j0 = 0; j0 < n_nodes; ++j0) {
ggml_tensor* t = ggml_graph_node(gf, j0);
bool need = callback_eval(t, true, callback_eval_user_data);
int j1 = j0;
while (!need && j1 < n_nodes - 1) {
t = ggml_graph_node(gf, ++j1);
need = callback_eval(t, true, callback_eval_user_data);
}
ggml_cgraph gv = sd_ggml_graph_view(gf, j0, j1 + 1);
status = ggml_backend_graph_compute_async(backend, &gv);
if (status != GGML_STATUS_SUCCESS) {
break;
}
ggml_backend_synchronize(backend);
if (need && !callback_eval(t, false, callback_eval_user_data)) {
stopped = true;
break;
}
j0 = j1;
}
ggml_backend_synchronize(backend);
if (stopped && status == GGML_STATUS_SUCCESS) {
status = GGML_STATUS_ABORTED;
}
return status;
}
const char* sd_get_system_info() {
static std::string cache_info = []() -> std::string {
ggml_backend_load_all_once();
@@ -525,12 +665,52 @@ SDBackendManager::~SDBackendManager() {
void SDBackendManager::reset() {
backends_.clear();
runtime_assignment_ = {};
params_assignment_ = {};
runtime_assignment_ = {};
params_assignment_ = {};
split_mode_assignment_ = {};
}
static std::vector<std::string> split_device_list(const std::string& value) {
std::vector<std::string> names;
for (const std::string& raw : split_copy(value, '&')) {
const std::string name = trim_copy(raw);
if (!name.empty()) {
names.push_back(name);
}
}
return names;
}
static std::string primary_device_name(const std::string& value) {
std::vector<std::string> names = split_device_list(value);
return names.empty() ? std::string() : names.front();
}
ggml_backend_t SDBackendManager::runtime_backend(SDBackendModule module) {
return init_cached_backend(runtime_assignment_.get(module));
return init_cached_backend(primary_device_name(runtime_assignment_.get(module)));
}
std::vector<ggml_backend_t> SDBackendManager::runtime_backends(SDBackendModule module) {
std::vector<ggml_backend_t> backends;
for (const std::string& name : split_device_list(runtime_assignment_.get(module))) {
ggml_backend_t backend = init_cached_backend(name);
if (backend == nullptr) {
LOG_ERROR("failed to initialize backend '%s' for module %s",
name.c_str(),
sd_backend_module_name(module));
continue;
}
if (std::find(backends.begin(), backends.end(), backend) == backends.end()) {
backends.push_back(backend);
}
}
if (backends.empty()) {
ggml_backend_t backend = runtime_backend(module);
if (backend != nullptr) {
backends.push_back(backend);
}
}
return backends;
}
ggml_backend_t SDBackendManager::params_backend(SDBackendModule module) {
@@ -556,6 +736,10 @@ bool SDBackendManager::params_backend_is_disk(SDBackendModule module) const {
return is_disk_backend_token(params_assignment_.get(module));
}
bool SDBackendManager::params_backend_follows_runtime(SDBackendModule module) const {
return params_assignment_.get(module).empty();
}
bool SDBackendManager::runtime_backend_supports_host_buffer(SDBackendModule module) {
ggml_backend_t backend = runtime_backend(module);
if (backend == nullptr) {
@@ -575,6 +759,7 @@ bool SDBackendManager::runtime_backend_supports_host_buffer(SDBackendModule modu
bool SDBackendManager::init(const char* backend_spec,
const char* params_backend_spec,
const char* split_mode_spec,
std::string* error) {
reset();
@@ -584,12 +769,53 @@ bool SDBackendManager::init(const char* backend_spec,
if (!sd_parse_backend_assignment(SAFE_STR(params_backend_spec), &params_assignment_, error)) {
return false;
}
if (!sd_parse_backend_assignment(SAFE_STR(split_mode_spec), &split_mode_assignment_, error)) {
return false;
}
return validate(error);
}
SDSplitMode SDBackendManager::split_mode(SDBackendModule module) const {
return lower_copy(trim_copy(split_mode_assignment_.get(module))) == "row" ? SDSplitMode::ROW
: SDSplitMode::LAYER;
}
ggml_backend_buffer_type_t SDBackendManager::split_buffer_type(ggml_backend_t backend,
const std::vector<float>& tensor_split) {
if (backend == nullptr) {
return nullptr;
}
ggml_backend_dev_t dev = ggml_backend_get_device(backend);
if (dev == nullptr) {
return nullptr;
}
ggml_backend_reg_t reg = ggml_backend_dev_backend_reg(dev);
if (reg == nullptr) {
return nullptr;
}
auto fn = (ggml_backend_split_buffer_type_t)ggml_backend_reg_get_proc_address(reg, "ggml_backend_split_buffer_type");
if (fn == nullptr) {
return nullptr;
}
int main_device = -1;
const size_t dev_count = ggml_backend_reg_dev_count(reg);
for (size_t i = 0; i < dev_count; ++i) {
if (ggml_backend_reg_dev_get(reg, i) == dev) {
main_device = (int)i;
break;
}
}
if (main_device < 0) {
return nullptr;
}
std::vector<float> padded_split(std::max<size_t>(tensor_split.size(), 64), 0.0f);
std::copy(tensor_split.begin(), tensor_split.end(), padded_split.begin());
return fn(main_device, padded_split.data());
}
bool SDBackendManager::validate(std::string* error) const {
auto validate_runtime_name = [&](const std::string& name) -> bool {
auto validate_single_runtime_name = [&](const std::string& name) -> bool {
if (is_default_backend_token(name)) {
return true;
}
@@ -599,7 +825,7 @@ bool SDBackendManager::validate(std::string* error) const {
}
return false;
}
if (!sd_backend_resolve_name(name).empty()) {
if (!sd_backend_resolve_name(name).empty() || resolve_first_device_by_registry_name(name) != nullptr) {
return true;
}
if (error != nullptr) {
@@ -607,15 +833,56 @@ bool SDBackendManager::validate(std::string* error) const {
}
return false;
};
auto validate_runtime_name = [&](const std::string& name) -> bool {
if (name.find('&') == std::string::npos) {
return validate_single_runtime_name(name);
}
std::vector<std::string> names = split_device_list(name);
if (names.empty()) {
if (error != nullptr) {
*error = "invalid backend device list '" + name + "'";
}
return false;
}
for (const std::string& entry : names) {
if (is_default_backend_token(entry)) {
if (error != nullptr) {
*error = "default backend token is not allowed in a device list '" + name + "'";
}
return false;
}
if (!validate_single_runtime_name(entry)) {
return false;
}
}
return true;
};
auto validate_params_name = [&](const std::string& name) -> bool {
if (is_disk_backend_token(name)) {
return true;
}
return validate_runtime_name(name);
if (name.find('&') != std::string::npos) {
if (error != nullptr) {
*error = "params_backend does not accept device lists ('" + name + "')";
}
return false;
}
return validate_single_runtime_name(name);
};
auto validate_split_mode_name = [&](const std::string& name) -> bool {
const std::string lower = lower_copy(trim_copy(name));
if (lower.empty() || lower == "layer" || lower == "row") {
return true;
}
if (error != nullptr) {
*error = "invalid split mode '" + name + "' (expected layer or row)";
}
return false;
};
if (!validate_runtime_name(runtime_assignment_.default_name) ||
!validate_params_name(params_assignment_.default_name)) {
!validate_params_name(params_assignment_.default_name) ||
!validate_split_mode_name(split_mode_assignment_.default_name)) {
return false;
}
for (const auto& kv : runtime_assignment_.module_names) {
@@ -628,6 +895,11 @@ bool SDBackendManager::validate(std::string* error) const {
return false;
}
}
for (const auto& kv : split_mode_assignment_.module_names) {
if (!validate_split_mode_name(kv.second)) {
return false;
}
}
return true;
}
+20
View File
@@ -6,9 +6,11 @@
#include <memory>
#include <string>
#include <unordered_map>
#include <vector>
#include "ggml-backend.h"
#include "ggml.h"
#include "stable-diffusion.h"
enum class SDBackendModule {
DIFFUSION,
@@ -36,10 +38,16 @@ struct SDBackendHandleDeleter {
using SDBackendHandle = std::unique_ptr<struct ggml_backend, SDBackendHandleDeleter>;
enum class SDSplitMode {
LAYER,
ROW,
};
class SDBackendManager {
private:
SDBackendAssignment runtime_assignment_;
SDBackendAssignment params_assignment_;
SDBackendAssignment split_mode_assignment_;
std::unordered_map<std::string, SDBackendHandle> backends_;
public:
@@ -51,15 +59,23 @@ public:
bool init(const char* backend_spec,
const char* params_backend_spec,
const char* split_mode_spec,
std::string* error);
void reset();
ggml_backend_t runtime_backend(SDBackendModule module);
ggml_backend_t params_backend(SDBackendModule module);
std::vector<ggml_backend_t> runtime_backends(SDBackendModule module);
SDSplitMode split_mode(SDBackendModule module) const;
ggml_backend_buffer_type_t split_buffer_type(ggml_backend_t backend,
const std::vector<float>& tensor_split);
bool runtime_backend_is_cpu(SDBackendModule module);
bool params_backend_is_cpu(SDBackendModule module);
bool params_backend_is_disk(SDBackendModule module) const;
bool params_backend_follows_runtime(SDBackendModule module) const;
bool runtime_backend_supports_host_buffer(SDBackendModule module);
private:
@@ -71,6 +87,10 @@ bool sd_backend_is(ggml_backend_t backend, const std::string& name);
bool sd_backend_is_cpu(ggml_backend_t backend);
ggml_backend_t sd_backend_cpu_init();
bool sd_backend_cpu_set_n_threads(ggml_backend_t backend_cpu, int n_threads);
ggml_status sd_backend_graph_compute_with_eval_callback(ggml_backend_t backend,
ggml_cgraph* gf,
sd_graph_eval_callback_t callback_eval,
void* callback_eval_user_data);
std::string sd_backend_resolve_name(const std::string& name);
const char* sd_backend_module_name(SDBackendModule module);
void ggml_ext_im_set_f32_1d(const struct ggml_tensor* tensor, int i, float value);
+221
View File
@@ -0,0 +1,221 @@
#include "core/layer_split_partition.h"
#include <algorithm>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include "core/util.h"
namespace sd {
static bool layer_split_path_segment_starts_at(const std::string& name, size_t pos) {
return pos == 0 || name[pos - 1] == '.';
}
static bool layer_split_has_path_segment(const std::string& name, const char* segment) {
size_t pos = name.find(segment);
while (pos != std::string::npos) {
if (layer_split_path_segment_starts_at(name, pos)) {
return true;
}
pos = name.find(segment, pos + 1);
}
return false;
}
int layer_split_tensor_block_index(const std::string& name) {
static const char* unet_block_segments[] = {"input_blocks.", "output_blocks.", "middle_block.",
"down_blocks.", "up_blocks.", "mid_block."};
for (const char* segment : unet_block_segments) {
if (layer_split_has_path_segment(name, segment)) {
return -1;
}
}
static const char* block_keywords[] = {"transformer_blocks.", "joint_blocks.", "double_blocks.",
"single_blocks.", "blocks.", "block.", "layers."};
for (const char* keyword : block_keywords) {
size_t pos = name.find(keyword);
while (pos != std::string::npos) {
if (!layer_split_path_segment_starts_at(name, pos)) {
pos = name.find(keyword, pos + 1);
continue;
}
pos += std::strlen(keyword);
size_t end = pos;
while (end < name.size() && name[end] >= '0' && name[end] <= '9') {
end++;
}
if (end > pos && (end == name.size() || name[end] == '.')) {
return std::atoi(name.substr(pos, end - pos).c_str());
}
break;
}
}
return -1;
}
std::string layer_split_backend_device_display_name(ggml_backend_t backend) {
ggml_backend_dev_t dev = ggml_backend_get_device(backend);
const char* name = dev != nullptr ? ggml_backend_dev_name(dev) : ggml_backend_name(backend);
return name != nullptr ? name : "unknown";
}
static bool layer_split_backend_supports_tensor(ggml_backend_t backend, const ggml_tensor* tensor) {
return backend != nullptr && tensor != nullptr && ggml_backend_supports_op(backend, tensor);
}
static size_t layer_split_supported_target(const std::string& desc,
const std::string& tensor_name,
const ggml_tensor* tensor,
const std::vector<ggml_backend_t>& backends,
size_t preferred) {
if (tensor == nullptr || backends.empty()) {
return preferred;
}
size_t preferred_safe = std::min(preferred, backends.size() - 1);
if (layer_split_backend_supports_tensor(backends[preferred_safe], tensor)) {
return preferred_safe;
}
for (size_t i = 0; i < backends.size(); i++) {
if (layer_split_backend_supports_tensor(backends[i], tensor)) {
LOG_WARN("%s layer split: moving tensor '%s' from %s to %s because the preferred backend cannot run op=%s type=%s nbytes=%.2f MB",
desc.c_str(),
tensor_name.c_str(),
layer_split_backend_device_display_name(backends[preferred_safe]).c_str(),
layer_split_backend_device_display_name(backends[i]).c_str(),
ggml_op_name(tensor->op),
ggml_type_name(tensor->type),
ggml_nbytes(tensor) / (1024.0 * 1024.0));
return i;
}
}
LOG_WARN("%s layer split: tensor '%s' is not supported by any split backend: op=%s type=%s nbytes=%.2f MB",
desc.c_str(),
tensor_name.c_str(),
ggml_op_name(tensor->op),
ggml_type_name(tensor->type),
ggml_nbytes(tensor) / (1024.0 * 1024.0));
return preferred_safe;
}
std::vector<std::map<std::string, ggml_tensor*>> partition_layer_split_tensors(
const std::string& desc,
const std::map<std::string, ggml_tensor*>& tensors,
const std::map<std::string, ggml_tensor*>& split_tensors,
const std::vector<ggml_backend_t>& backends) {
std::vector<std::map<std::string, ggml_tensor*>> partitions(backends.size());
if (backends.empty()) {
LOG_WARN("%s: no backend available for a layer split", desc.c_str());
return partitions;
}
std::map<int, int64_t> block_bytes;
std::map<std::string, size_t> non_block_targets;
std::vector<int64_t> other_bytes_by_backend(backends.size(), 0);
int64_t total_block_bytes = 0;
int64_t total_other_bytes = 0;
int n_blocks = 0;
for (const auto& kv : tensors) {
int64_t bytes = (int64_t)ggml_nbytes(kv.second);
int idx = split_tensors.count(kv.first) != 0 ? layer_split_tensor_block_index(kv.first) : -1;
if (idx >= 0) {
block_bytes[idx] += bytes;
total_block_bytes += bytes;
n_blocks = std::max(n_blocks, idx + 1);
} else {
size_t target = layer_split_supported_target(desc, kv.first, kv.second, backends, 0);
non_block_targets[kv.first] = target;
other_bytes_by_backend[target] += bytes;
total_other_bytes += bytes;
}
}
if (n_blocks == 0) {
LOG_WARN("%s: no transformer blocks found for a layer split; keeping tensors on compatible backends starting from %s",
desc.c_str(),
layer_split_backend_device_display_name(backends[0]).c_str());
for (const auto& kv : tensors) {
size_t target = 0;
auto target_it = non_block_targets.find(kv.first);
if (target_it != non_block_targets.end()) {
target = target_it->second;
}
partitions[target][kv.first] = kv.second;
}
return partitions;
}
// Reserve compute headroom and subtract each device's actual non-block
// bytes from its block budget.
constexpr int64_t compute_headroom_bytes = 2ll * 1024 * 1024 * 1024;
std::vector<double> device_weights(backends.size(), 1.0);
double weight_sum = 0.0;
for (size_t i = 0; i < backends.size(); i++) {
ggml_backend_dev_t dev = ggml_backend_get_device(backends[i]);
size_t free_bytes = 0, total_bytes = 0;
if (dev != nullptr) {
ggml_backend_dev_memory(dev, &free_bytes, &total_bytes);
}
// Keep a small share even for tight devices instead of dropping them.
int64_t usable_bytes = std::max<int64_t>((int64_t)free_bytes - compute_headroom_bytes,
(int64_t)free_bytes / 8);
device_weights[i] = usable_bytes > 0 ? (double)usable_bytes : 1.0;
weight_sum += device_weights[i];
}
std::vector<int64_t> block_budgets(backends.size(), 0);
const int64_t total_bytes = total_block_bytes + total_other_bytes;
for (size_t i = 0; i < backends.size(); i++) {
int64_t budget = (int64_t)((double)total_bytes * device_weights[i] / weight_sum);
budget = std::max<int64_t>(budget - other_bytes_by_backend[i], 0);
block_budgets[i] = budget;
}
std::vector<int> boundaries(backends.size(), n_blocks);
size_t current = 0;
int64_t used = 0;
for (int b = 0; b < n_blocks; b++) {
int64_t bytes = block_bytes.count(b) != 0 ? block_bytes[b] : 0;
if (current + 1 < backends.size() && used > 0 && used + bytes > block_budgets[current]) {
boundaries[current] = b;
current++;
used = 0;
}
used += bytes;
}
for (const auto& kv : tensors) {
size_t target = 0;
int idx = split_tensors.count(kv.first) != 0 ? layer_split_tensor_block_index(kv.first) : -1;
if (idx >= 0) {
while (target < boundaries.size() && idx >= boundaries[target]) {
target++;
}
target = std::min(target, backends.size() - 1);
target = layer_split_supported_target(desc, kv.first, kv.second, backends, target);
} else {
auto target_it = non_block_targets.find(kv.first);
if (target_it != non_block_targets.end()) {
target = target_it->second;
}
}
partitions[target][kv.first] = kv.second;
}
int range_start = 0;
for (size_t i = 0; i < backends.size(); i++) {
int range_end = boundaries[i];
const char* non_block_suffix = other_bytes_by_backend[i] > 0 ? " + non-block tensors" : "";
LOG_INFO("%s layer split: %s <- blocks [%d, %d)%s",
desc.c_str(),
layer_split_backend_device_display_name(backends[i]).c_str(),
range_start,
range_end,
non_block_suffix);
range_start = range_end;
}
return partitions;
}
} // namespace sd
+24
View File
@@ -0,0 +1,24 @@
#ifndef __SD_CORE_LAYER_SPLIT_PARTITION_H__
#define __SD_CORE_LAYER_SPLIT_PARTITION_H__
#include <map>
#include <string>
#include <vector>
#include "ggml-backend.h"
#include "ggml.h"
namespace sd {
std::string layer_split_backend_device_display_name(ggml_backend_t backend);
int layer_split_tensor_block_index(const std::string& name);
std::vector<std::map<std::string, ggml_tensor*>> partition_layer_split_tensors(
const std::string& desc,
const std::map<std::string, ggml_tensor*>& tensors,
const std::map<std::string, ggml_tensor*>& split_tensors,
const std::vector<ggml_backend_t>& backends);
} // namespace sd
#endif // __SD_CORE_LAYER_SPLIT_PARTITION_H__
+42
View File
@@ -4,6 +4,8 @@
#include <cmath>
#include <codecvt>
#include <cstdarg>
#include <cstdlib>
#include <cstring>
#include <exception>
#include <fstream>
#include <locale>
@@ -25,6 +27,7 @@
#include <unistd.h>
#endif
#include "ggml-backend.h"
#include "ggml.h"
#include "stable-diffusion.h"
@@ -346,6 +349,9 @@ int sd_preview_interval = 1;
bool sd_preview_denoised = true;
bool sd_preview_noisy = false;
static sd_graph_eval_callback_t sd_backend_eval_cb = nullptr;
static void* sd_backend_eval_cb_data = nullptr;
std::u32string utf8_to_utf32(const std::string& utf8_str) {
std::wstring_convert<std::codecvt_utf8<char32_t>, char32_t> converter;
return converter.from_bytes(utf8_str);
@@ -629,6 +635,11 @@ void sd_set_preview_callback(sd_preview_cb_t cb, preview_t mode, int interval, b
sd_preview_noisy = noisy;
}
void sd_set_backend_eval_callback(sd_graph_eval_callback_t cb, void* data) {
sd_backend_eval_cb = cb;
sd_backend_eval_cb_data = data;
}
sd_preview_cb_t sd_get_preview_callback() {
return sd_preview_cb;
}
@@ -649,6 +660,14 @@ bool sd_should_preview_noisy() {
return sd_preview_noisy;
}
sd_graph_eval_callback_t sd_get_backend_eval_callback() {
return sd_backend_eval_cb;
}
void* sd_get_backend_eval_callback_data() {
return sd_backend_eval_cb_data;
}
sd_progress_cb_t sd_get_progress_callback() {
return sd_progress_cb;
}
@@ -981,3 +1000,26 @@ std::vector<std::pair<std::string, float>> split_quotation_attention(
}
return result;
}
size_t sd_list_devices(char* buffer, size_t buffer_size) {
if (ggml_backend_dev_count() == 0) {
// dynamic-backend builds discover their backend modules at runtime
ggml_backend_load_all();
}
std::ostringstream oss;
for (size_t i = 0; i < ggml_backend_dev_count(); i++) {
ggml_backend_dev_t dev = ggml_backend_dev_get(i);
const char* name = ggml_backend_dev_name(dev);
const char* desc = ggml_backend_dev_description(dev);
oss << (name ? name : "") << '\t' << (desc ? desc : "") << '\n';
}
std::string devices = oss.str();
if (buffer != nullptr && buffer_size > 0) {
size_t copy_size = std::min(devices.size(), buffer_size - 1);
memcpy(buffer, devices.data(), copy_size);
buffer[copy_size] = '\0';
}
return devices.size();
}
+3
View File
@@ -98,6 +98,9 @@ int sd_get_preview_interval();
bool sd_should_preview_denoised();
bool sd_should_preview_noisy();
sd_graph_eval_callback_t sd_get_backend_eval_callback();
void* sd_get_backend_eval_callback_data();
// test if the backend is a specific one, e.g. "CUDA", "ROCm", "Vulkan" etc.
bool sd_backend_is(ggml_backend_t backend, const std::string& name);
+11 -1
View File
@@ -36,6 +36,7 @@ enum SDVersion {
VERSION_WAN2_2_I2V,
VERSION_WAN2_2_TI2V,
VERSION_QWEN_IMAGE,
VERSION_QWEN_IMAGE_LAYERED,
VERSION_ANIMA,
VERSION_FLUX2,
VERSION_FLUX2_KLEIN,
@@ -46,6 +47,7 @@ enum SDVersion {
VERSION_OVIS_IMAGE,
VERSION_ERNIE_IMAGE,
VERSION_LENS,
VERSION_MINIT2I,
VERSION_LONGCAT,
VERSION_PID,
VERSION_IDEOGRAM4,
@@ -126,7 +128,7 @@ static inline bool sd_version_is_wan(SDVersion version) {
}
static inline bool sd_version_is_qwen_image(SDVersion version) {
if (version == VERSION_QWEN_IMAGE) {
if (version == VERSION_QWEN_IMAGE || version == VERSION_QWEN_IMAGE_LAYERED) {
return true;
}
return false;
@@ -174,6 +176,13 @@ static inline bool sd_version_is_lens(SDVersion version) {
return false;
}
static inline bool sd_version_is_minit2i(SDVersion version) {
if (version == VERSION_MINIT2I) {
return true;
}
return false;
}
static inline bool sd_version_is_pid(SDVersion version) {
if (version == VERSION_PID) {
return true;
@@ -247,6 +256,7 @@ static inline bool sd_version_is_dit(SDVersion version) {
sd_version_is_boogu_image(version) ||
sd_version_is_ernie_image(version) ||
sd_version_is_lens(version) ||
sd_version_is_minit2i(version) ||
sd_version_is_longcat(version) ||
sd_version_is_pid(version) ||
sd_version_is_ideogram4(version) ||
+83 -61
View File
@@ -12,6 +12,16 @@ namespace Rope {
ErnieImage,
};
enum class RefIndexMode {
FIXED,
INCREASE,
DECREASE,
};
__STATIC_INLINE__ RefIndexMode ref_index_mode_from_bool(bool increase_ref_index) {
return increase_ref_index ? RefIndexMode::INCREASE : RefIndexMode::FIXED;
}
template <class T>
__STATIC_INLINE__ std::vector<T> linspace(T start, T end, int num) {
std::vector<T> result(num);
@@ -346,7 +356,7 @@ namespace Rope {
int axes_dim_num,
int start_index,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index,
RefIndexMode ref_index_mode,
float ref_index_scale,
bool scale_rope,
int base_offset = 0) {
@@ -357,13 +367,15 @@ namespace Rope {
for (ggml_tensor* ref : ref_latents) {
int h_offset = 0;
int w_offset = 0;
if (!increase_ref_index) {
if (ref_index_mode == RefIndexMode::FIXED) {
if (ref->ne[1] + curr_h_offset > ref->ne[0] + curr_w_offset) {
w_offset = curr_w_offset;
} else {
h_offset = curr_h_offset;
}
scale_rope = false;
} else if (ref_index_mode == RefIndexMode::DECREASE) {
index--;
}
auto ref_ids = gen_flux_img_ids(static_cast<int>(ref->ne[1]),
@@ -377,7 +389,7 @@ namespace Rope {
scale_rope);
ids = concat_ids(ids, ref_ids, bs);
if (increase_ref_index) {
if (ref_index_mode == RefIndexMode::INCREASE) {
index++;
}
@@ -395,7 +407,7 @@ namespace Rope {
int context_len,
std::set<int> txt_arange_dims,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index,
RefIndexMode ref_index_mode,
float ref_index_scale,
bool is_longcat) {
int x_index = is_longcat ? 1 : 0;
@@ -406,7 +418,7 @@ namespace Rope {
auto ids = concat_ids(txt_ids, img_ids, bs);
if (ref_latents.size() > 0) {
auto refs_ids = gen_refs_ids(patch_size, bs, axes_dim_num, x_index + 1, ref_latents, increase_ref_index, ref_index_scale, false, offset);
auto refs_ids = gen_refs_ids(patch_size, bs, axes_dim_num, x_index + 1, ref_latents, ref_index_mode, ref_index_scale, false, offset);
ids = concat_ids(ids, refs_ids, bs);
}
return ids;
@@ -420,7 +432,7 @@ namespace Rope {
int context_len,
std::set<int> txt_arange_dims,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index,
RefIndexMode ref_index_mode,
float ref_index_scale,
int theta,
bool circular_h,
@@ -435,7 +447,7 @@ namespace Rope {
context_len,
txt_arange_dims,
ref_latents,
increase_ref_index,
ref_index_mode,
ref_index_scale,
is_longcat);
std::vector<std::vector<int>> wrap_dims;
@@ -481,17 +493,64 @@ namespace Rope {
return embed_nd(ids, bs, static_cast<float>(theta), axes_dim, wrap_dims);
}
__STATIC_INLINE__ std::vector<std::vector<float>> gen_qwen_image_ids(int h,
__STATIC_INLINE__ std::vector<std::vector<float>> gen_vid_ids(int t,
int h,
int w,
int pt,
int ph,
int pw,
int bs,
int t_offset = 0,
int h_offset = 0,
int w_offset = 0,
bool scale_rope = false) {
int t_len = (t + (pt / 2)) / pt;
int h_len = (h + (ph / 2)) / ph;
int w_len = (w + (pw / 2)) / pw;
std::vector<std::vector<float>> vid_ids(t_len * h_len * w_len, std::vector<float>(3, 0.0));
if (scale_rope) {
h_offset -= h_len / 2;
w_offset -= w_len / 2;
}
std::vector<float> t_ids = linspace<float>(1.f * t_offset, 1.f * t_len - 1 + t_offset, t_len);
std::vector<float> h_ids = linspace<float>(1.f * h_offset, 1.f * h_len - 1 + h_offset, h_len);
std::vector<float> w_ids = linspace<float>(1.f * w_offset, 1.f * w_len - 1 + w_offset, w_len);
for (int i = 0; i < t_len; ++i) {
for (int j = 0; j < h_len; ++j) {
for (int k = 0; k < w_len; ++k) {
int idx = i * h_len * w_len + j * w_len + k;
vid_ids[idx][0] = t_ids[i];
vid_ids[idx][1] = h_ids[j];
vid_ids[idx][2] = w_ids[k];
}
}
}
std::vector<std::vector<float>> vid_ids_repeated(bs * vid_ids.size(), std::vector<float>(3));
for (int i = 0; i < bs; ++i) {
for (int j = 0; j < vid_ids.size(); ++j) {
vid_ids_repeated[i * vid_ids.size() + j] = vid_ids[j];
}
}
return vid_ids_repeated;
}
__STATIC_INLINE__ std::vector<std::vector<float>> gen_qwen_image_ids(int t,
int h,
int w,
int patch_size,
int bs,
int context_len,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index) {
RefIndexMode ref_index_mode) {
int h_len = (h + (patch_size / 2)) / patch_size;
int w_len = (w + (patch_size / 2)) / patch_size;
int txt_id_start = std::max(h_len, w_len);
auto txt_ids = linspace<float>(1.f * txt_id_start, 1.f * context_len + txt_id_start, context_len);
int txt_id_start = std::max(h_len, w_len) / 2;
auto txt_ids = linspace<float>(1.f * txt_id_start, 1.f * txt_id_start + context_len - 1, context_len);
std::vector<std::vector<float>> txt_ids_repeated(bs * context_len, std::vector<float>(3));
for (int i = 0; i < bs; ++i) {
for (int j = 0; j < txt_ids.size(); ++j) {
@@ -499,28 +558,30 @@ namespace Rope {
}
}
int axes_dim_num = 3;
auto img_ids = gen_flux_img_ids(h, w, patch_size, bs, axes_dim_num, 0, 0, 0, true);
auto img_ids = gen_vid_ids(t, h, w, 1, patch_size, patch_size, bs, 0, 0, 0, true);
auto ids = concat_ids(txt_ids_repeated, img_ids, bs);
if (ref_latents.size() > 0) {
auto refs_ids = gen_refs_ids(patch_size, bs, axes_dim_num, 1, ref_latents, increase_ref_index, 1.f, true);
ids = concat_ids(ids, refs_ids, bs);
int ref_start_index = ref_index_mode == RefIndexMode::DECREASE ? 0 : 1;
auto refs_ids = gen_refs_ids(patch_size, bs, axes_dim_num, ref_start_index, ref_latents, ref_index_mode, 1.f, true);
ids = concat_ids(ids, refs_ids, bs);
}
return ids;
}
// Generate qwen_image positional embeddings
__STATIC_INLINE__ std::vector<float> gen_qwen_image_pe(int h,
__STATIC_INLINE__ std::vector<float> gen_qwen_image_pe(int t,
int h,
int w,
int patch_size,
int bs,
int context_len,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index,
RefIndexMode ref_index_mode,
int theta,
bool circular_h,
bool circular_w,
const std::vector<int>& axes_dim) {
std::vector<std::vector<float>> ids = gen_qwen_image_ids(h, w, patch_size, bs, context_len, ref_latents, increase_ref_index);
std::vector<std::vector<float>> ids = gen_qwen_image_ids(t, h, w, patch_size, bs, context_len, ref_latents, ref_index_mode);
std::vector<std::vector<int>> wrap_dims;
// This logic simply stores the (pad and patch_adjusted) sizes of images so we can make sure rope correctly tiles
if ((circular_h || circular_w) && bs > 0 && axes_dim.size() >= 3) {
@@ -533,7 +594,7 @@ namespace Rope {
// Track per-token wrap lengths for the row/column axes so only spatial tokens become periodic.
wrap_dims.assign(axes_dim.size(), std::vector<int>(total_tokens / bs, 0));
size_t cursor = context_len; // ignore text tokens
const size_t img_tokens = static_cast<size_t>(h_len) * static_cast<size_t>(w_len);
const size_t img_tokens = static_cast<size_t>(t) * static_cast<size_t>(h_len) * static_cast<size_t>(w_len);
for (size_t token_i = 0; token_i < img_tokens; ++token_i) {
if (circular_h) {
wrap_dims[1][cursor + token_i] = h_len;
@@ -684,46 +745,6 @@ namespace Rope {
return embed_nd(ids, bs, static_cast<float>(theta), axes_dim, wrap_dims, EmbedNDLayout::ErnieImage);
}
__STATIC_INLINE__ std::vector<std::vector<float>> gen_vid_ids(int t,
int h,
int w,
int pt,
int ph,
int pw,
int bs,
int t_offset = 0,
int h_offset = 0,
int w_offset = 0) {
int t_len = (t + (pt / 2)) / pt;
int h_len = (h + (ph / 2)) / ph;
int w_len = (w + (pw / 2)) / pw;
std::vector<std::vector<float>> vid_ids(t_len * h_len * w_len, std::vector<float>(3, 0.0));
std::vector<float> t_ids = linspace<float>(1.f * t_offset, 1.f * t_len - 1 + t_offset, t_len);
std::vector<float> h_ids = linspace<float>(1.f * h_offset, 1.f * h_len - 1 + h_offset, h_len);
std::vector<float> w_ids = linspace<float>(1.f * w_offset, 1.f * w_len - 1 + w_offset, w_len);
for (int i = 0; i < t_len; ++i) {
for (int j = 0; j < h_len; ++j) {
for (int k = 0; k < w_len; ++k) {
int idx = i * h_len * w_len + j * w_len + k;
vid_ids[idx][0] = t_ids[i];
vid_ids[idx][1] = h_ids[j];
vid_ids[idx][2] = w_ids[k];
}
}
}
std::vector<std::vector<float>> vid_ids_repeated(bs * vid_ids.size(), std::vector<float>(3));
for (int i = 0; i < bs; ++i) {
for (int j = 0; j < vid_ids.size(); ++j) {
vid_ids_repeated[i * vid_ids.size() + j] = vid_ids[j];
}
}
return vid_ids_repeated;
}
// Generate wan positional embeddings
__STATIC_INLINE__ std::vector<float> gen_wan_pe(int t,
int h,
@@ -785,7 +806,8 @@ namespace Rope {
int context_len,
int seq_multi_of,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index) {
RefIndexMode ref_index_mode) {
SD_UNUSED(ref_index_mode);
int padded_context_len = context_len + bound_mod(context_len, seq_multi_of);
auto txt_ids = std::vector<std::vector<float>>(bs * padded_context_len, std::vector<float>(3, 0.0f));
for (int i = 0; i < bs * padded_context_len; i++) {
@@ -816,12 +838,12 @@ namespace Rope {
int context_len,
int seq_multi_of,
const std::vector<ggml_tensor*>& ref_latents,
bool increase_ref_index,
RefIndexMode ref_index_mode,
int theta,
bool circular_h,
bool circular_w,
const std::vector<int>& axes_dim) {
std::vector<std::vector<float>> ids = gen_z_image_ids(h, w, patch_size, bs, context_len, seq_multi_of, ref_latents, increase_ref_index);
std::vector<std::vector<float>> ids = gen_z_image_ids(h, w, patch_size, bs, context_len, seq_multi_of, ref_latents, ref_index_mode);
std::vector<std::vector<int>> wrap_dims;
if ((circular_h || circular_w) && bs > 0 && axes_dim.size() >= 3) {
int pad_h = (patch_size - (h % patch_size)) % patch_size;
+1 -1
View File
@@ -612,7 +612,7 @@ namespace Anima {
0,
{},
empty_ref_latents,
false,
Rope::RefIndexMode::FIXED,
1.0f,
false);
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_DIFFUSION_CONTROL_HPP__
#ifndef __SD_MODEL_DIFFUSION_CONTROL_HPP__
#define __SD_MODEL_DIFFUSION_CONTROL_HPP__
#include "model/common/block.hpp"
+12 -12
View File
@@ -135,23 +135,23 @@ namespace DiT {
return x;
}
inline ggml_tensor* unpatchify(ggml_context* ctx,
ggml_tensor* x,
int64_t t_len,
int64_t h_len,
int64_t w_len,
int pt,
int ph,
int pw) {
// x: [N, t_len*h_len*w_len, pt*ph*pw*C]
inline ggml_tensor* unpatchify_3d(ggml_context* ctx,
ggml_tensor* x,
int64_t t_len,
int64_t h_len,
int64_t w_len,
int pt,
int ph,
int pw) {
// x: [N, t_len*h_len*w_len, C*pt*ph*pw]
// return: [N*C, t_len*pt, h_len*ph, w_len*pw]
int64_t N = x->ne[3];
int64_t N = x->ne[2];
int64_t C = x->ne[0] / pt / ph / pw;
GGML_ASSERT(C * pt * ph * pw == x->ne[0]);
x = ggml_reshape_4d(ctx, x, C, pw * ph * pt, w_len * h_len * t_len, N); // [N, t_len*h_len*w_len, pt*ph*pw, C]
x = ggml_ext_cont(ctx, ggml_ext_torch_permute(ctx, x, 1, 2, 0, 3)); // [N, C, t_len*h_len*w_len, pt*ph*pw]
x = ggml_reshape_4d(ctx, x, pw * ph * pt, C, w_len * h_len * t_len, N); // [N, t_len*h_len*w_len, C, pt*ph*pw]
x = ggml_ext_cont(ctx, ggml_ext_torch_permute(ctx, x, 0, 2, 1, 3)); // [N, C, t_len*h_len*w_len, pt*ph*pw]
x = ggml_reshape_4d(ctx, x, pw, ph * pt, w_len, h_len * t_len * C * N); // [N*C*t_len*h_len, w_len, pt*ph, pw]
x = ggml_ext_cont(ctx, ggml_ext_torch_permute(ctx, x, 0, 2, 1, 3)); // [N*C*t_len*h_len, pt*ph, w_len, pw]
x = ggml_reshape_4d(ctx, x, pw * w_len, ph, pt, h_len * t_len * C * N); // [N*C*t_len*h_len, pt, ph, w_len*pw]
+8 -8
View File
@@ -1485,7 +1485,7 @@ namespace Flux {
const sd::Tensor<float>& y_tensor = {},
const sd::Tensor<float>& guidance_tensor = {},
const std::vector<sd::Tensor<float>>& ref_latents_tensor = {},
bool increase_ref_index = false,
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::FIXED,
std::vector<int> skip_layers = {},
const sd::Tensor<float>& pulid_id_tensor = {},
float pulid_id_weight = 1.0f) {
@@ -1527,8 +1527,8 @@ namespace Flux {
}
std::set<int> txt_arange_dims;
if (sd_version_is_flux2(version) || sd_version_is_sefi_image(version)) {
txt_arange_dims = {3};
increase_ref_index = true;
txt_arange_dims = {3};
ref_index_mode = Rope::RefIndexMode::INCREASE;
} else if (version == VERSION_OVIS_IMAGE) {
txt_arange_dims = {1, 2};
}
@@ -1539,7 +1539,7 @@ namespace Flux {
static_cast<int>(context->ne[1]),
txt_arange_dims,
ref_latents,
increase_ref_index,
ref_index_mode,
config.ref_index_scale,
config.theta,
circular_y_enabled,
@@ -1599,7 +1599,7 @@ namespace Flux {
const sd::Tensor<float>& y = {},
const sd::Tensor<float>& guidance = {},
const std::vector<sd::Tensor<float>>& ref_latents = {},
bool increase_ref_index = false,
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::FIXED,
std::vector<int> skip_layers = std::vector<int>(),
const sd::Tensor<float>& pulid_id = {},
float pulid_id_weight = 1.0f) {
@@ -1610,7 +1610,7 @@ namespace Flux {
// guidance: [N, ]
// pulid_id: empty (no injection) or [N, num_id_tokens=32, kv_dim=2048]
auto get_graph = [&]() -> ggml_cgraph* {
return build_graph(x, timesteps, context, c_concat, y, guidance, ref_latents, increase_ref_index, skip_layers, pulid_id, pulid_id_weight);
return build_graph(x, timesteps, context, c_concat, y, guidance, ref_latents, ref_index_mode, skip_layers, pulid_id, pulid_id_weight);
};
auto result = restore_trailing_singleton_dims(GGMLRunner::compute<float>(get_graph, n_threads, false, false, false), x.dim());
@@ -1632,7 +1632,7 @@ namespace Flux {
tensor_or_empty(diffusion_params.y),
tensor_or_empty(extra->guidance),
diffusion_params.ref_latents ? *diffusion_params.ref_latents : empty_ref_latents,
diffusion_params.increase_ref_index,
diffusion_params.ref_index_mode,
extra->skip_layers ? *extra->skip_layers : empty_skip_layers,
tensor_or_empty(extra->pulid_id),
extra->pulid_id_weight);
@@ -1683,7 +1683,7 @@ namespace Flux {
{},
guidance,
{},
false);
Rope::RefIndexMode::FIXED);
int64_t t1 = ggml_time_ms();
GGML_ASSERT(!out_opt.empty());
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_DIFFUSION_HIDREAM_O1_HPP__
#ifndef __SD_MODEL_DIFFUSION_HIDREAM_O1_HPP__
#define __SD_MODEL_DIFFUSION_HIDREAM_O1_HPP__
#include <algorithm>
+611
View File
@@ -0,0 +1,611 @@
#ifndef __SD_MODEL_DIFFUSION_MINIT2I_HPP__
#define __SD_MODEL_DIFFUSION_MINIT2I_HPP__
#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <memory>
#include <string>
#include <vector>
#include "core/ggml_extend.hpp"
#include "model/common/rope.hpp"
#include "model/diffusion/dit.hpp"
#include "model/diffusion/model.hpp"
#include "model_loader.h"
namespace MiniT2I {
constexpr int MINIT2I_GRAPH_SIZE = 196608;
struct MiniT2IConfig {
int64_t image_size = 512;
int64_t patch_size = 16;
int64_t in_channels = 3;
int64_t txt_input_size = 1024;
int64_t hidden_size = 768;
int64_t txt_hidden_size = 768;
int64_t cond_vec_size = 768;
int64_t depth_double = 17;
int64_t txt_preamble_depth = 2;
int64_t num_heads = 12;
int64_t head_dim = 64;
float mlp_ratio = 2.6667f;
int64_t pca_channels = 128;
int64_t prompt_length = 256;
int64_t n_T = 100;
float cfg_interval_start = 0.0f;
float cfg_interval_end = 1.0f;
static MiniT2IConfig detect_from_weights(const String2TensorStorage& tensor_storage_map, const std::string& prefix) {
MiniT2IConfig config;
config.depth_double = 0;
config.txt_preamble_depth = 0;
for (const auto& [name, tensor_storage] : tensor_storage_map) {
if (!starts_with(name, prefix)) {
continue;
}
if (ends_with(name, "img_embedder.proj1.weight") && tensor_storage.n_dims == 4) {
config.patch_size = tensor_storage.ne[0];
config.in_channels = tensor_storage.ne[2];
config.pca_channels = tensor_storage.ne[3];
} else if (ends_with(name, "img_embedder.proj2.weight") && tensor_storage.n_dims == 4) {
config.pca_channels = tensor_storage.ne[2];
config.hidden_size = tensor_storage.ne[3];
} else if (ends_with(name, "txt_embedder.weight") && tensor_storage.n_dims == 2) {
config.txt_input_size = tensor_storage.ne[0];
config.txt_hidden_size = tensor_storage.ne[1];
} else if (ends_with(name, "pooled_embedder.weight") && tensor_storage.n_dims == 2) {
config.cond_vec_size = tensor_storage.ne[1];
} else if (ends_with(name, "double_blocks.0.img_qkv.weight") && tensor_storage.n_dims == 2) {
int64_t inner3 = tensor_storage.ne[1];
int64_t inner = inner3 / 3;
config.hidden_size = tensor_storage.ne[0];
if (config.hidden_size == 768) {
config.num_heads = 12;
config.head_dim = 64;
} else if (config.hidden_size == 1248) {
config.num_heads = 24;
config.head_dim = 52;
} else if (inner > 0) {
config.head_dim = 64;
config.num_heads = std::max<int64_t>(1, inner / config.head_dim);
}
} else if (ends_with(name, "final_layer.linear.weight") && tensor_storage.n_dims == 2) {
int64_t patch_area = config.patch_size * config.patch_size;
config.hidden_size = tensor_storage.ne[0];
config.in_channels = patch_area > 0 ? tensor_storage.ne[1] / patch_area : config.in_channels;
} else if (ends_with(name, "mask_token") && tensor_storage.n_dims >= 2) {
config.prompt_length = tensor_storage.ne[1];
}
size_t pos = name.find("double_blocks.");
if (pos != std::string::npos) {
auto items = split_string(name.substr(pos), '.');
if (items.size() > 1) {
int64_t idx = atoi(items[1].c_str());
config.depth_double = std::max<int64_t>(config.depth_double, idx + 1);
}
}
pos = name.find("txt_preamble_blocks.");
if (pos != std::string::npos) {
auto items = split_string(name.substr(pos), '.');
if (items.size() > 1) {
int64_t idx = atoi(items[1].c_str());
config.txt_preamble_depth = std::max<int64_t>(config.txt_preamble_depth, idx + 1);
}
}
}
if (config.depth_double <= 0) {
config.depth_double = config.hidden_size == 1248 ? 23 : 17;
}
if (config.txt_preamble_depth <= 0) {
config.txt_preamble_depth = 2;
}
if (config.head_dim <= 0 || config.num_heads <= 0) {
config.head_dim = config.hidden_size == 1248 ? 52 : 64;
config.num_heads = config.hidden_size / config.head_dim;
}
LOG_DEBUG("minit2i: hidden_size=%" PRId64 ", txt_hidden_size=%" PRId64 ", heads=%" PRId64 ", head_dim=%" PRId64 ", double_blocks=%" PRId64 ", txt_blocks=%" PRId64 ", patch=%" PRId64 ", in_channels=%" PRId64,
config.hidden_size,
config.txt_hidden_size,
config.num_heads,
config.head_dim,
config.depth_double,
config.txt_preamble_depth,
config.patch_size,
config.in_channels);
return config;
}
};
inline std::vector<float> make_2d_sincos_pos_embed(int grid_size, int dim) {
GGML_ASSERT(dim % 4 == 0);
int half_dim = dim / 2;
int quarter = half_dim / 2;
std::vector<float> out(static_cast<size_t>(grid_size) * grid_size * dim);
std::vector<float> omega(quarter);
for (int i = 0; i < quarter; ++i) {
omega[i] = 1.0f / std::pow(10000.0f, static_cast<float>(i) / static_cast<float>(quarter));
}
for (int y = 0; y < grid_size; ++y) {
for (int x = 0; x < grid_size; ++x) {
size_t base = static_cast<size_t>(y * grid_size + x) * dim;
for (int i = 0; i < quarter; ++i) {
float ay = y * omega[i];
float ax = x * omega[i];
out[base + i] = std::sin(ax);
out[base + quarter + i] = std::cos(ax);
out[base + half_dim + i] = std::sin(ay);
out[base + half_dim + quarter + i] = std::cos(ay);
}
}
}
return out;
}
inline std::vector<float> make_text_rope(int length, int head_dim) {
return Rope::flatten(Rope::rope(Rope::linspace(0.f, static_cast<float>(length - 1), length), head_dim, 10000.f));
}
inline std::vector<float> make_vision_rope(int side, int head_dim) {
GGML_ASSERT(head_dim % 4 == 0);
int dim = head_dim / 2;
int quarter = dim / 2;
int length = side * side;
std::vector<float> out(static_cast<size_t>(length) * (head_dim / 2) * 4);
std::vector<float> freqs(quarter);
for (int i = 0; i < quarter; ++i) {
freqs[i] = 1.0f / std::pow(10000.0f, static_cast<float>(2 * i) / static_cast<float>(dim));
}
for (int y = 0; y < side; ++y) {
for (int x = 0; x < side; ++x) {
int pos = y * side + x;
size_t base = static_cast<size_t>(pos) * (head_dim / 2) * 4;
for (int i = 0; i < quarter; ++i) {
float ay = y * freqs[i];
float ax = x * freqs[i];
float angles[2] = {ay, ax};
for (int axis = 0; axis < 2; ++axis) {
int j = axis * quarter + i;
out[base + 4 * j] = std::cos(angles[axis]);
out[base + 4 * j + 1] = -std::sin(angles[axis]);
out[base + 4 * j + 2] = std::sin(angles[axis]);
out[base + 4 * j + 3] = std::cos(angles[axis]);
}
}
}
}
return out;
}
struct SwiGLUMlp : public GGMLBlock {
SwiGLUMlp(int64_t in_features, int64_t hidden_features) {
int64_t hidden_dim = ((hidden_features + 7) / 8) * 8;
blocks["w1"] = std::make_shared<Linear>(in_features, hidden_dim, false);
blocks["w3"] = std::make_shared<Linear>(in_features, hidden_dim, false);
blocks["w2"] = std::make_shared<Linear>(hidden_dim, in_features, false);
}
ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* x) {
auto w1 = std::dynamic_pointer_cast<Linear>(blocks["w1"]);
auto w3 = std::dynamic_pointer_cast<Linear>(blocks["w3"]);
auto w2 = std::dynamic_pointer_cast<Linear>(blocks["w2"]);
auto gate = ggml_silu(ctx->ggml_ctx, w1->forward(ctx, x));
auto up = w3->forward(ctx, x);
return w2->forward(ctx, ggml_mul(ctx->ggml_ctx, gate, up));
}
};
struct BottleneckPatchEmbed : public GGMLBlock {
int64_t patch_size;
BottleneckPatchEmbed(int64_t patch_size, int64_t in_channels, int64_t pca_channels, int64_t hidden_size)
: patch_size(patch_size) {
blocks["proj1"] = std::make_shared<Conv2d>(in_channels,
pca_channels,
std::pair<int, int>{static_cast<int>(patch_size), static_cast<int>(patch_size)},
std::pair<int, int>{static_cast<int>(patch_size), static_cast<int>(patch_size)},
std::pair<int, int>{0, 0},
std::pair<int, int>{1, 1},
false);
blocks["proj2"] = std::make_shared<Conv2d>(pca_channels,
hidden_size,
std::pair<int, int>{1, 1},
std::pair<int, int>{1, 1},
std::pair<int, int>{0, 0},
std::pair<int, int>{1, 1},
true);
}
ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* x) {
auto proj1 = std::dynamic_pointer_cast<Conv2d>(blocks["proj1"]);
auto proj2 = std::dynamic_pointer_cast<Conv2d>(blocks["proj2"]);
x = proj1->forward(ctx, x);
x = proj2->forward(ctx, x);
x = ggml_reshape_3d(ctx->ggml_ctx, x, x->ne[0] * x->ne[1], x->ne[2], x->ne[3]);
x = ggml_cont(ctx->ggml_ctx, ggml_permute(ctx->ggml_ctx, x, 1, 0, 2, 3));
return x;
}
};
struct TimestepEmbedder : public GGMLBlock {
int frequency_embedding_size;
TimestepEmbedder(int64_t hidden_size, int frequency_embedding_size = 256)
: frequency_embedding_size(frequency_embedding_size) {
blocks["mlp.0"] = std::make_shared<Linear>(frequency_embedding_size, hidden_size, true, true);
blocks["mlp.2"] = std::make_shared<Linear>(hidden_size, hidden_size, true, true);
}
ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* t) {
auto mlp_0 = std::dynamic_pointer_cast<Linear>(blocks["mlp.0"]);
auto mlp_2 = std::dynamic_pointer_cast<Linear>(blocks["mlp.2"]);
auto t_emb = ggml_ext_timestep_embedding(ctx->ggml_ctx, t, frequency_embedding_size, 10000, 1.0f);
t_emb = mlp_0->forward(ctx, t_emb);
t_emb = ggml_silu_inplace(ctx->ggml_ctx, t_emb);
return mlp_2->forward(ctx, t_emb);
}
};
inline std::vector<ggml_tensor*> split_qkv(ggml_context* ctx, ggml_tensor* qkv, int64_t num_heads, int64_t head_dim) {
int64_t N = qkv->ne[2];
int64_t L = qkv->ne[1];
auto q = ggml_view_4d(ctx, qkv, head_dim, num_heads, L, N,
qkv->nb[0] * head_dim, qkv->nb[1], qkv->nb[2], 0);
auto k = ggml_view_4d(ctx, qkv, head_dim, num_heads, L, N,
qkv->nb[0] * head_dim, qkv->nb[1], qkv->nb[2], qkv->nb[0] * head_dim * num_heads);
auto v = ggml_view_4d(ctx, qkv, head_dim, num_heads, L, N,
qkv->nb[0] * head_dim, qkv->nb[1], qkv->nb[2], qkv->nb[0] * head_dim * num_heads * 2);
return {q, k, v};
}
struct PlainTextTransformerBlock : public GGMLBlock {
int64_t num_heads;
int64_t head_dim;
PlainTextTransformerBlock(int64_t hidden_size, int64_t num_heads, int64_t head_dim, float mlp_ratio)
: num_heads(num_heads), head_dim(head_dim) {
int64_t inner_dim = num_heads * head_dim;
blocks["norm1"] = std::make_shared<RMSNorm>(hidden_size, 1e-6f);
blocks["norm2"] = std::make_shared<RMSNorm>(hidden_size, 1e-6f);
blocks["qkv"] = std::make_shared<Linear>(hidden_size, inner_dim * 3, true);
blocks["attn_proj"] = std::make_shared<Linear>(inner_dim, hidden_size, true);
blocks["mlp"] = std::make_shared<SwiGLUMlp>(hidden_size, static_cast<int64_t>(hidden_size * mlp_ratio));
blocks["q_norm"] = std::make_shared<RMSNorm>(head_dim, 1e-6f);
blocks["k_norm"] = std::make_shared<RMSNorm>(head_dim, 1e-6f);
}
ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* txt, ggml_tensor* pe) {
auto norm1 = std::dynamic_pointer_cast<RMSNorm>(blocks["norm1"]);
auto norm2 = std::dynamic_pointer_cast<RMSNorm>(blocks["norm2"]);
auto qkv_proj = std::dynamic_pointer_cast<Linear>(blocks["qkv"]);
auto attn_proj = std::dynamic_pointer_cast<Linear>(blocks["attn_proj"]);
auto mlp = std::dynamic_pointer_cast<SwiGLUMlp>(blocks["mlp"]);
auto q_norm = std::dynamic_pointer_cast<RMSNorm>(blocks["q_norm"]);
auto k_norm = std::dynamic_pointer_cast<RMSNorm>(blocks["k_norm"]);
auto qkv = split_qkv(ctx->ggml_ctx, qkv_proj->forward(ctx, norm1->forward(ctx, txt)), num_heads, head_dim);
auto q = q_norm->forward(ctx, qkv[0]);
auto k = k_norm->forward(ctx, qkv[1]);
auto v = qkv[2];
auto out = Rope::attention(ctx, q, k, v, pe, nullptr, 1.0f, false);
txt = ggml_add(ctx->ggml_ctx, txt, attn_proj->forward(ctx, out));
txt = ggml_add(ctx->ggml_ctx, txt, mlp->forward(ctx, norm2->forward(ctx, txt)));
return txt;
}
};
struct DoubleStreamDiTBlock : public GGMLBlock {
int64_t num_heads;
int64_t head_dim;
DoubleStreamDiTBlock(int64_t hidden_size, int64_t txt_hidden_size, int64_t num_heads, int64_t head_dim, float mlp_ratio)
: num_heads(num_heads), head_dim(head_dim) {
int64_t inner_dim = num_heads * head_dim;
blocks["img_norm1"] = std::make_shared<RMSNorm>(hidden_size, 1e-6f);
blocks["img_norm2"] = std::make_shared<RMSNorm>(hidden_size, 1e-6f);
blocks["txt_norm1"] = std::make_shared<RMSNorm>(txt_hidden_size, 1e-6f);
blocks["txt_norm2"] = std::make_shared<RMSNorm>(txt_hidden_size, 1e-6f);
blocks["img_qkv"] = std::make_shared<Linear>(hidden_size, inner_dim * 3, true);
blocks["txt_qkv"] = std::make_shared<Linear>(txt_hidden_size, inner_dim * 3, true);
blocks["q_norm"] = std::make_shared<RMSNorm>(head_dim, 1e-6f);
blocks["k_norm"] = std::make_shared<RMSNorm>(head_dim, 1e-6f);
blocks["img_attn_proj"] = std::make_shared<Linear>(inner_dim, hidden_size, true);
blocks["txt_attn_proj"] = std::make_shared<Linear>(inner_dim, txt_hidden_size, true);
blocks["img_mlp"] = std::make_shared<SwiGLUMlp>(hidden_size, static_cast<int64_t>(hidden_size * mlp_ratio));
blocks["txt_mlp"] = std::make_shared<SwiGLUMlp>(txt_hidden_size, static_cast<int64_t>(txt_hidden_size * mlp_ratio));
}
std::pair<ggml_tensor*, ggml_tensor*> forward(GGMLRunnerContext* ctx,
ggml_tensor* img,
ggml_tensor* txt,
ggml_tensor* pe) {
auto img_norm1 = std::dynamic_pointer_cast<RMSNorm>(blocks["img_norm1"]);
auto img_norm2 = std::dynamic_pointer_cast<RMSNorm>(blocks["img_norm2"]);
auto txt_norm1 = std::dynamic_pointer_cast<RMSNorm>(blocks["txt_norm1"]);
auto txt_norm2 = std::dynamic_pointer_cast<RMSNorm>(blocks["txt_norm2"]);
auto img_qkv_p = std::dynamic_pointer_cast<Linear>(blocks["img_qkv"]);
auto txt_qkv_p = std::dynamic_pointer_cast<Linear>(blocks["txt_qkv"]);
auto q_norm = std::dynamic_pointer_cast<RMSNorm>(blocks["q_norm"]);
auto k_norm = std::dynamic_pointer_cast<RMSNorm>(blocks["k_norm"]);
auto img_proj = std::dynamic_pointer_cast<Linear>(blocks["img_attn_proj"]);
auto txt_proj = std::dynamic_pointer_cast<Linear>(blocks["txt_attn_proj"]);
auto img_mlp = std::dynamic_pointer_cast<SwiGLUMlp>(blocks["img_mlp"]);
auto txt_mlp = std::dynamic_pointer_cast<SwiGLUMlp>(blocks["txt_mlp"]);
int64_t li = img->ne[1];
int64_t lt = txt->ne[1];
auto img_qkv = split_qkv(ctx->ggml_ctx, img_qkv_p->forward(ctx, img_norm1->forward(ctx, img)), num_heads, head_dim);
auto txt_qkv = split_qkv(ctx->ggml_ctx, txt_qkv_p->forward(ctx, txt_norm1->forward(ctx, txt)), num_heads, head_dim);
auto q = ggml_concat(ctx->ggml_ctx, q_norm->forward(ctx, txt_qkv[0]), q_norm->forward(ctx, img_qkv[0]), 2);
auto k = ggml_concat(ctx->ggml_ctx, k_norm->forward(ctx, txt_qkv[1]), k_norm->forward(ctx, img_qkv[1]), 2);
auto v = ggml_concat(ctx->ggml_ctx, txt_qkv[2], img_qkv[2], 2);
auto out = Rope::attention(ctx, q, k, v, pe, nullptr, 1.0f, false);
auto out_txt = ggml_ext_slice(ctx->ggml_ctx, out, 1, 0, lt);
auto out_img = ggml_ext_slice(ctx->ggml_ctx, out, 1, lt, lt + li);
img = ggml_add(ctx->ggml_ctx, img, img_proj->forward(ctx, out_img));
txt = ggml_add(ctx->ggml_ctx, txt, txt_proj->forward(ctx, out_txt));
img = ggml_add(ctx->ggml_ctx, img, img_mlp->forward(ctx, img_norm2->forward(ctx, img)));
txt = ggml_add(ctx->ggml_ctx, txt, txt_mlp->forward(ctx, txt_norm2->forward(ctx, txt)));
return {img, txt};
}
};
struct FinalLayer : public GGMLBlock {
FinalLayer(int64_t hidden_size, int64_t patch_size, int64_t out_channels) {
blocks["norm_final"] = std::make_shared<RMSNorm>(hidden_size, 1e-6f);
blocks["linear"] = std::make_shared<Linear>(hidden_size, patch_size * patch_size * out_channels, true);
}
ggml_tensor* forward(GGMLRunnerContext* ctx, ggml_tensor* x) {
auto norm_final = std::dynamic_pointer_cast<RMSNorm>(blocks["norm_final"]);
auto linear = std::dynamic_pointer_cast<Linear>(blocks["linear"]);
return linear->forward(ctx, norm_final->forward(ctx, x));
}
};
struct MMJiT : public GGMLBlock {
MiniT2IConfig config;
MMJiT(const MiniT2IConfig& config)
: config(config) {
blocks["img_embedder"] = std::make_shared<BottleneckPatchEmbed>(config.patch_size, config.in_channels, config.pca_channels, config.hidden_size);
blocks["txt_embedder"] = std::make_shared<Linear>(config.txt_input_size, config.txt_hidden_size, false);
blocks["t_embedder"] = std::make_shared<TimestepEmbedder>(config.cond_vec_size);
blocks["pooled_embedder"] = std::make_shared<Linear>(config.txt_input_size, config.cond_vec_size, false);
for (int64_t i = 0; i < config.txt_preamble_depth; ++i) {
blocks["txt_preamble_blocks." + std::to_string(i)] = std::make_shared<PlainTextTransformerBlock>(config.txt_hidden_size, config.num_heads, config.head_dim, config.mlp_ratio);
}
for (int64_t i = 0; i < config.depth_double; ++i) {
blocks["double_blocks." + std::to_string(i)] = std::make_shared<DoubleStreamDiTBlock>(config.hidden_size, config.txt_hidden_size, config.num_heads, config.head_dim, config.mlp_ratio);
}
blocks["final_layer"] = std::make_shared<FinalLayer>(config.hidden_size, config.patch_size, config.in_channels);
}
void init_params(ggml_context* ctx, const String2TensorStorage& tensor_storage_map = {}, const std::string prefix = "") override {
GGMLBlock::init_params(ctx, tensor_storage_map, prefix);
enum ggml_type wtype = get_type(prefix + "mask_token", tensor_storage_map, GGML_TYPE_F32);
params["mask_token"] = ggml_new_tensor_3d(ctx, wtype, config.txt_input_size, 1, 1);
}
ggml_tensor* apply_text_mask(GGMLRunnerContext* ctx, ggml_tensor* context, ggml_tensor* mask) {
if (mask == nullptr) {
return context;
}
mask = ggml_reshape_3d(ctx->ggml_ctx, mask, 1, mask->ne[0], mask->ne[1]);
mask = ggml_repeat(ctx->ggml_ctx, mask, context);
auto keep = ggml_mul(ctx->ggml_ctx, context, mask);
auto inv = ggml_sub(ctx->ggml_ctx, ggml_ext_ones_like(ctx->ggml_ctx, mask), mask);
auto mask_token = ggml_repeat(ctx->ggml_ctx, params["mask_token"], context);
return ggml_add(ctx->ggml_ctx, keep, ggml_mul(ctx->ggml_ctx, mask_token, inv));
}
ggml_tensor* pool_context(GGMLRunnerContext* ctx, ggml_tensor* context) {
int64_t dim = context->ne[0];
int64_t len = context->ne[1];
int64_t N = context->ne[2];
auto x = ggml_cont(ctx->ggml_ctx, ggml_permute(ctx->ggml_ctx, context, 1, 0, 2, 3));
x = ggml_reshape_3d(ctx->ggml_ctx, x, len, dim, N);
x = ggml_mean(ctx->ggml_ctx, x);
x = ggml_reshape_2d(ctx->ggml_ctx, x, dim, N);
return x;
}
ggml_tensor* forward(GGMLRunnerContext* ctx,
ggml_tensor* img,
ggml_tensor* context,
ggml_tensor* mask,
ggml_tensor* pos_embed,
ggml_tensor* txt_pe,
ggml_tensor* joint_pe) {
auto img_embedder = std::dynamic_pointer_cast<BottleneckPatchEmbed>(blocks["img_embedder"]);
auto txt_embedder = std::dynamic_pointer_cast<Linear>(blocks["txt_embedder"]);
auto final_layer = std::dynamic_pointer_cast<FinalLayer>(blocks["final_layer"]);
int64_t W = img->ne[0];
int64_t H = img->ne[1];
int64_t hp = H / config.patch_size;
int64_t wp = W / config.patch_size;
context = apply_text_mask(ctx, context, mask);
auto x = img_embedder->forward(ctx, img);
x = ggml_add(ctx->ggml_ctx, x, pos_embed);
auto txt = txt_embedder->forward(ctx, context);
for (int64_t i = 0; i < config.txt_preamble_depth; ++i) {
auto block = std::dynamic_pointer_cast<PlainTextTransformerBlock>(blocks["txt_preamble_blocks." + std::to_string(i)]);
txt = block->forward(ctx, txt, txt_pe);
sd::ggml_graph_cut::mark_graph_cut(txt, "minit2i.txt_preamble_blocks." + std::to_string(i), "txt");
}
for (int64_t i = 0; i < config.depth_double; ++i) {
auto block = std::dynamic_pointer_cast<DoubleStreamDiTBlock>(blocks["double_blocks." + std::to_string(i)]);
auto out = block->forward(ctx, x, txt, joint_pe);
x = out.first;
txt = out.second;
sd::ggml_graph_cut::mark_graph_cut(x, "minit2i.double_blocks." + std::to_string(i), "x");
sd::ggml_graph_cut::mark_graph_cut(txt, "minit2i.double_blocks." + std::to_string(i), "txt");
}
auto combined = ggml_concat(ctx->ggml_ctx, txt, x, 1);
auto out = final_layer->forward(ctx, combined);
auto img_out = ggml_ext_slice(ctx->ggml_ctx, out, 1, txt->ne[1], txt->ne[1] + x->ne[1]);
return DiT::unpatchify(ctx->ggml_ctx, img_out, hp, wp, static_cast<int>(config.patch_size), static_cast<int>(config.patch_size), false);
}
};
struct MiniT2IRunner : public DiffusionModelRunner {
MiniT2IConfig config;
MMJiT model;
ggml_context* position_cache_ctx = nullptr;
ggml_backend_buffer_t position_cache_buffer = nullptr;
ggml_tensor* cached_pos_embed = nullptr;
ggml_tensor* cached_txt_pe = nullptr;
ggml_tensor* cached_joint_pe = nullptr;
int64_t cached_img_side = -1;
int64_t cached_txt_len = -1;
int64_t cached_hidden_size = -1;
int64_t cached_head_dim = -1;
MiniT2IRunner(ggml_backend_t backend,
const String2TensorStorage& tensor_storage_map = {},
const std::string prefix = "",
std::shared_ptr<RunnerWeightManager> weight_manager = nullptr)
: DiffusionModelRunner(backend, prefix, weight_manager),
config(MiniT2IConfig::detect_from_weights(tensor_storage_map, this->prefix)),
model(config) {
model.init(params_ctx, tensor_storage_map, this->prefix);
}
~MiniT2IRunner() override {
free_position_cache();
}
std::string get_desc() override {
return "MiniT2I";
}
void get_param_tensors(std::map<std::string, ggml_tensor*>& tensors, const std::string& prefix) override {
model.get_param_tensors(tensors, prefix);
}
void free_position_cache() {
if (position_cache_buffer != nullptr) {
ggml_backend_buffer_free(position_cache_buffer);
position_cache_buffer = nullptr;
}
if (position_cache_ctx != nullptr) {
ggml_free(position_cache_ctx);
position_cache_ctx = nullptr;
}
cached_pos_embed = nullptr;
cached_txt_pe = nullptr;
cached_joint_pe = nullptr;
cached_img_side = -1;
cached_txt_len = -1;
cached_hidden_size = -1;
cached_head_dim = -1;
}
void ensure_position_cache(int64_t img_side, int64_t txt_len) {
if (cached_img_side == img_side &&
cached_txt_len == txt_len &&
cached_hidden_size == config.hidden_size &&
cached_head_dim == config.head_dim &&
cached_pos_embed != nullptr &&
cached_txt_pe != nullptr &&
cached_joint_pe != nullptr) {
return;
}
free_position_cache();
auto pos_embed_vec = make_2d_sincos_pos_embed(static_cast<int>(img_side), static_cast<int>(config.hidden_size));
auto txt_pe_vec = make_text_rope(static_cast<int>(txt_len), static_cast<int>(config.head_dim));
auto img_pe_vec = make_vision_rope(static_cast<int>(img_side), static_cast<int>(config.head_dim));
auto joint_pe_vec = txt_pe_vec;
joint_pe_vec.insert(joint_pe_vec.end(), img_pe_vec.begin(), img_pe_vec.end());
ggml_init_params params;
params.mem_size = static_cast<size_t>(3 * ggml_tensor_overhead());
params.mem_buffer = nullptr;
params.no_alloc = true;
position_cache_ctx = ggml_init(params);
GGML_ASSERT(position_cache_ctx != nullptr);
cached_pos_embed = ggml_new_tensor_3d(position_cache_ctx, GGML_TYPE_F32, config.hidden_size, img_side * img_side, 1);
ggml_set_name(cached_pos_embed, "minit2i.pos_embed");
cached_txt_pe = ggml_new_tensor_4d(position_cache_ctx, GGML_TYPE_F32, 2, 2, config.head_dim / 2, txt_len);
ggml_set_name(cached_txt_pe, "minit2i.txt_pe");
cached_joint_pe = ggml_new_tensor_4d(position_cache_ctx, GGML_TYPE_F32, 2, 2, config.head_dim / 2, txt_len + img_side * img_side);
ggml_set_name(cached_joint_pe, "minit2i.joint_pe");
position_cache_buffer = ggml_backend_alloc_ctx_tensors(position_cache_ctx, runtime_backend);
GGML_ASSERT(position_cache_buffer != nullptr);
ggml_backend_buffer_set_usage(position_cache_buffer, GGML_BACKEND_BUFFER_USAGE_WEIGHTS);
ggml_backend_tensor_set(cached_pos_embed, pos_embed_vec.data(), 0, ggml_nbytes(cached_pos_embed));
ggml_backend_tensor_set(cached_txt_pe, txt_pe_vec.data(), 0, ggml_nbytes(cached_txt_pe));
ggml_backend_tensor_set(cached_joint_pe, joint_pe_vec.data(), 0, ggml_nbytes(cached_joint_pe));
ggml_backend_synchronize(runtime_backend);
cached_img_side = img_side;
cached_txt_len = txt_len;
cached_hidden_size = config.hidden_size;
cached_head_dim = config.head_dim;
}
ggml_cgraph* build_graph(const sd::Tensor<float>& x_tensor,
const sd::Tensor<float>& timesteps_tensor,
const sd::Tensor<float>& context_tensor,
const sd::Tensor<float>& mask_tensor) {
ggml_cgraph* gf = new_graph_custom(MINIT2I_GRAPH_SIZE);
ggml_tensor* x = make_input(x_tensor);
ggml_tensor* context = make_input(context_tensor);
ggml_tensor* mask = make_input(mask_tensor);
SD_UNUSED(timesteps_tensor);
int64_t W = x->ne[0];
int64_t H = x->ne[1];
int64_t img_side = H / config.patch_size;
int64_t txt_len = context->ne[1];
ensure_position_cache(img_side, txt_len);
auto runner_ctx = get_context();
auto out = model.forward(&runner_ctx, x, context, mask, cached_pos_embed, cached_txt_pe, cached_joint_pe);
ggml_build_forward_expand(gf, out);
return gf;
}
sd::Tensor<float> compute(int n_threads,
const sd::Tensor<float>& x,
const sd::Tensor<float>& timesteps,
const sd::Tensor<float>& context,
const sd::Tensor<float>& mask) {
auto get_graph = [&]() -> ggml_cgraph* {
return build_graph(x, timesteps, context, mask);
};
return restore_trailing_singleton_dims(GGMLRunner::compute<float>(get_graph, n_threads, false, false, false), x.dim());
}
sd::Tensor<float> compute(int n_threads,
const DiffusionParams& diffusion_params) override {
GGML_ASSERT(diffusion_params.x != nullptr);
GGML_ASSERT(diffusion_params.timesteps != nullptr);
GGML_ASSERT(diffusion_params.context != nullptr);
const auto* extra = diffusion_extra_as<MiniT2IDiffusionExtra>(diffusion_params);
GGML_ASSERT(extra->mask != nullptr);
return compute(n_threads,
*diffusion_params.x,
*diffusion_params.timesteps,
*diffusion_params.context,
*extra->mask);
}
};
} // namespace MiniT2I
#endif // __SD_MODEL_DIFFUSION_MINIT2I_HPP__
+9 -3
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_DIFFUSION_MODEL_HPP__
#ifndef __SD_MODEL_DIFFUSION_MODEL_HPP__
#define __SD_MODEL_DIFFUSION_MODEL_HPP__
#include <string>
@@ -7,6 +7,7 @@
#include "core/ggml_extend.hpp"
#include "core/tensor_ggml.hpp"
#include "model/common/rope.hpp"
#include "model_manager.h"
struct UNetDiffusionExtra {
@@ -52,6 +53,10 @@ struct LTXAVDiffusionExtra {
const sd::Tensor<float>* video_positions = nullptr;
};
struct MiniT2IDiffusionExtra {
const sd::Tensor<float>* mask = nullptr;
};
using DiffusionExtraParams = std::variant<std::monostate,
UNetDiffusionExtra,
SkipLayerDiffusionExtra,
@@ -59,7 +64,8 @@ using DiffusionExtraParams = std::variant<std::monostate,
AnimaDiffusionExtra,
WanDiffusionExtra,
HiDreamO1DiffusionExtra,
LTXAVDiffusionExtra>;
LTXAVDiffusionExtra,
MiniT2IDiffusionExtra>;
struct DiffusionParams {
const sd::Tensor<float>* x = nullptr;
@@ -68,7 +74,7 @@ struct DiffusionParams {
const sd::Tensor<float>* c_concat = nullptr;
const sd::Tensor<float>* y = nullptr;
const std::vector<sd::Tensor<float>>* ref_latents = nullptr;
bool increase_ref_index = false;
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::FIXED;
DiffusionExtraParams extra = std::monostate{};
};
+93 -27
View File
@@ -4,6 +4,7 @@
#include <memory>
#include "model/common/block.hpp"
#include "model/diffusion/dit.hpp"
#include "model/diffusion/flux.hpp"
#include "model/diffusion/model.hpp"
#include "model_loader.h"
@@ -23,6 +24,7 @@ namespace Qwen {
std::vector<int> axes_dim = {16, 56, 56};
int axes_dim_sum = 128;
bool zero_cond_t = false;
bool use_additional_t_cond = false;
static QwenImageConfig detect_from_weights(const String2TensorStorage& tensor_storage_map, const std::string& prefix) {
QwenImageConfig config;
@@ -88,19 +90,33 @@ namespace Qwen {
};
struct QwenTimestepProjEmbeddings : public GGMLBlock {
protected:
bool use_additional_t_cond = false;
public:
QwenTimestepProjEmbeddings(int64_t embedding_dim) {
QwenTimestepProjEmbeddings(int64_t embedding_dim, bool use_additional_t_cond = false)
: use_additional_t_cond(use_additional_t_cond) {
blocks["timestep_embedder"] = std::shared_ptr<GGMLBlock>(new TimestepEmbedding(256, embedding_dim));
if (use_additional_t_cond) {
blocks["addition_t_embedding"] = std::shared_ptr<GGMLBlock>(new Embedding(2, embedding_dim));
}
}
ggml_tensor* forward(GGMLRunnerContext* ctx,
ggml_tensor* timesteps) {
ggml_tensor* timesteps,
ggml_tensor* addition_t_cond = nullptr) {
// timesteps: [N,]
// return: [N, embedding_dim]
auto timestep_embedder = std::dynamic_pointer_cast<TimestepEmbedding>(blocks["timestep_embedder"]);
auto timesteps_proj = ggml_ext_timestep_embedding(ctx->ggml_ctx, timesteps, 256, 10000, 1.f);
auto timesteps_emb = timestep_embedder->forward(ctx, timesteps_proj);
if (use_additional_t_cond) {
GGML_ASSERT(addition_t_cond != nullptr);
auto addition_t_embedding = std::dynamic_pointer_cast<Embedding>(blocks["addition_t_embedding"]);
auto addition_t_emb = addition_t_embedding->forward(ctx, addition_t_cond);
timesteps_emb = ggml_add(ctx->ggml_ctx, timesteps_emb, addition_t_emb);
}
return timesteps_emb;
}
};
@@ -402,7 +418,7 @@ namespace Qwen {
QwenImageModel(QwenImageConfig config)
: config(config) {
int64_t inner_dim = config.num_attention_heads * config.attention_head_dim;
blocks["time_text_embed"] = std::shared_ptr<GGMLBlock>(new QwenTimestepProjEmbeddings(inner_dim));
blocks["time_text_embed"] = std::shared_ptr<GGMLBlock>(new QwenTimestepProjEmbeddings(inner_dim, config.use_additional_t_cond));
blocks["txt_norm"] = std::shared_ptr<GGMLBlock>(new RMSNorm(config.joint_attention_dim, 1e-6f));
blocks["img_in"] = std::shared_ptr<GGMLBlock>(new Linear(config.in_channels, inner_dim));
blocks["txt_in"] = std::shared_ptr<GGMLBlock>(new Linear(config.joint_attention_dim, inner_dim));
@@ -424,6 +440,7 @@ namespace Qwen {
ggml_tensor* forward_orig(GGMLRunnerContext* ctx,
ggml_tensor* x,
ggml_tensor* timestep,
ggml_tensor* addition_t_cond,
ggml_tensor* context,
ggml_tensor* pe,
ggml_tensor* modulate_index = nullptr) {
@@ -434,9 +451,9 @@ namespace Qwen {
auto norm_out = std::dynamic_pointer_cast<AdaLayerNormContinuous>(blocks["norm_out"]);
auto proj_out = std::dynamic_pointer_cast<Linear>(blocks["proj_out"]);
auto t_emb = time_text_embed->forward(ctx, timestep);
auto t_emb = time_text_embed->forward(ctx, timestep, addition_t_cond);
if (config.zero_cond_t) {
auto t_emb_0 = time_text_embed->forward(ctx, ggml_ext_zeros_like(ctx->ggml_ctx, timestep));
auto t_emb_0 = time_text_embed->forward(ctx, ggml_ext_zeros_like(ctx->ggml_ctx, timestep), addition_t_cond);
t_emb = ggml_concat(ctx->ggml_ctx, t_emb, t_emb_0, 1);
}
auto img = img_in->forward(ctx, x);
@@ -469,33 +486,50 @@ namespace Qwen {
ggml_tensor* forward(GGMLRunnerContext* ctx,
ggml_tensor* x,
ggml_tensor* timestep,
ggml_tensor* addition_t_cond,
ggml_tensor* context,
ggml_tensor* pe,
std::vector<ggml_tensor*> ref_latents = {},
ggml_tensor* modulate_index = nullptr) {
// Forward pass of DiT.
// x: [N, C, H, W]
// x: [N, C, H, W] or [N*C, T, H, W]
// timestep: [N,]
// context: [N, L, D]
// pe: [L, d_head/2, 2, 2]
// return: [N, C, H, W]
// return: [N, C, H, W] or [N*C, T, H, W]
int64_t W = x->ne[0];
int64_t H = x->ne[1];
int64_t C = x->ne[2];
int64_t N = x->ne[3];
int64_t W = x->ne[0];
int64_t H = x->ne[1];
int64_t T = 1;
int64_t N = addition_t_cond != nullptr ? addition_t_cond->ne[0] : x->ne[3];
bool has_time_axis = false;
if (x->ne[3] != 1) {
T = x->ne[2];
has_time_axis = true;
}
auto img = DiT::pad_and_patchify(ctx, x, config.patch_size, config.patch_size);
auto patchify_input = [&](ggml_tensor* input) -> ggml_tensor* {
input = DiT::pad_to_patch_size(ctx, input, config.patch_size, config.patch_size);
if (!has_time_axis) {
return DiT::patchify(ctx->ggml_ctx, input, config.patch_size, config.patch_size);
}
if (input->ne[3] == 1) {
input = ggml_reshape_4d(ctx->ggml_ctx, input, input->ne[0], input->ne[1], 1, input->ne[2]);
}
return DiT::patchify(ctx->ggml_ctx, input, 1, config.patch_size, config.patch_size, N);
};
auto img = patchify_input(x);
int64_t img_tokens = img->ne[1];
if (ref_latents.size() > 0) {
for (ggml_tensor* ref : ref_latents) {
ref = DiT::pad_and_patchify(ctx, ref, config.patch_size, config.patch_size);
ref = patchify_input(ref);
img = ggml_concat(ctx->ggml_ctx, img, ref, 1);
}
}
auto out = forward_orig(ctx, img, timestep, context, pe, modulate_index); // [N, h_len*w_len, ph*pw*C]
auto out = forward_orig(ctx, img, timestep, addition_t_cond, context, pe, modulate_index); // [N, h_len*w_len, ph*pw*C]
if (out->ne[1] > img_tokens) {
out = ggml_cont(ctx->ggml_ctx, ggml_permute(ctx->ggml_ctx, out, 0, 2, 1, 3)); // [num_tokens, N, C * patch_size * patch_size]
@@ -503,7 +537,17 @@ namespace Qwen {
out = ggml_cont(ctx->ggml_ctx, ggml_permute(ctx->ggml_ctx, out, 0, 2, 1, 3)); // [N, h*w, C * patch_size * patch_size]
}
out = DiT::unpatchify_and_crop(ctx->ggml_ctx, out, H, W, config.patch_size, config.patch_size); // [N, C, H, W]
if (has_time_axis) {
int pad_h = (config.patch_size - H % config.patch_size) % config.patch_size;
int pad_w = (config.patch_size - W % config.patch_size) % config.patch_size;
int h_len = static_cast<int>((H + pad_h) / config.patch_size);
int w_len = static_cast<int>((W + pad_w) / config.patch_size);
out = DiT::unpatchify_3d(ctx->ggml_ctx, out, T, h_len, w_len, 1, config.patch_size, config.patch_size);
out = ggml_ext_slice(ctx->ggml_ctx, out, 1, 0, H); // [N*C, T, H, W + pad_w]
out = ggml_ext_slice(ctx->ggml_ctx, out, 0, 0, W); // [N*C, T, H, W]
} else {
out = DiT::unpatchify_and_crop(ctx->ggml_ctx, out, H, W, config.patch_size, config.patch_size); // [N, C, H, W]
}
return out;
}
@@ -515,6 +559,7 @@ namespace Qwen {
QwenImageModel qwen_image;
std::vector<float> pe_vec;
std::vector<float> modulate_index_vec;
std::vector<int32_t> additional_t_cond_vec;
SDVersion version;
QwenImageRunner(ggml_backend_t backend,
@@ -524,9 +569,13 @@ namespace Qwen {
bool zero_cond_t = false,
std::shared_ptr<RunnerWeightManager> weight_manager = nullptr)
: DiffusionModelRunner(backend, prefix, weight_manager),
config(QwenImageConfig::detect_from_weights(tensor_storage_map, prefix)) {
config(QwenImageConfig::detect_from_weights(tensor_storage_map, prefix)),
version(version) {
config.zero_cond_t = config.zero_cond_t || zero_cond_t;
qwen_image = QwenImageModel(config);
if (version == VERSION_QWEN_IMAGE_LAYERED) {
config.use_additional_t_cond = true;
}
qwen_image = QwenImageModel(config);
qwen_image.init(params_ctx, tensor_storage_map, prefix);
}
@@ -542,11 +591,11 @@ namespace Qwen {
const sd::Tensor<float>& timesteps_tensor,
const sd::Tensor<float>& context_tensor,
const std::vector<sd::Tensor<float>>& ref_latents_tensor = {},
bool increase_ref_index = false) {
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::INCREASE) {
ggml_cgraph* gf = new_graph_custom(QWEN_IMAGE_GRAPH_SIZE);
ggml_tensor* x = make_input(x_tensor);
ggml_tensor* timesteps = make_input(timesteps_tensor);
GGML_ASSERT(x->ne[3] == 1);
GGML_ASSERT(x->ne[3] == 1 || x_tensor.dim() == 5);
GGML_ASSERT(!context_tensor.empty());
ggml_tensor* context = make_input(context_tensor);
std::vector<ggml_tensor*> ref_latents;
@@ -555,13 +604,29 @@ namespace Qwen {
ref_latents.push_back(make_input(ref_latent_tensor));
}
pe_vec = Rope::gen_qwen_image_pe(static_cast<int>(x->ne[1]),
int batch_size = static_cast<int>(x->ne[3]);
int time_len = 1;
if (x_tensor.dim() == 5) {
time_len = static_cast<int>(x_tensor.shape()[2]);
batch_size = static_cast<int>(x_tensor.shape()[4]);
}
ggml_tensor* addition_t_cond = nullptr;
if (version == VERSION_QWEN_IMAGE_LAYERED) {
additional_t_cond_vec.assign(static_cast<size_t>(batch_size), 0);
addition_t_cond = ggml_new_tensor_1d(compute_ctx, GGML_TYPE_I32, batch_size);
set_backend_tensor_data(addition_t_cond, additional_t_cond_vec.data());
ref_index_mode = Rope::RefIndexMode::DECREASE;
}
pe_vec = Rope::gen_qwen_image_pe(time_len,
static_cast<int>(x->ne[1]),
static_cast<int>(x->ne[0]),
config.patch_size,
static_cast<int>(x->ne[3]),
batch_size,
static_cast<int>(context->ne[1]),
ref_latents,
increase_ref_index,
ref_index_mode,
config.theta,
circular_y_enabled,
circular_x_enabled,
@@ -604,6 +669,7 @@ namespace Qwen {
ggml_tensor* out = qwen_image.forward(&runner_ctx,
x,
timesteps,
addition_t_cond,
context,
pe,
ref_latents,
@@ -619,12 +685,12 @@ namespace Qwen {
const sd::Tensor<float>& timesteps,
const sd::Tensor<float>& context,
const std::vector<sd::Tensor<float>>& ref_latents = {},
bool increase_ref_index = false) {
// x: [N, in_channels, h, w]
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::INCREASE) {
// x: [N, C, H, W] or [N*C, T, H, W]
// timesteps: [N, ]
// context: [N, max_position, hidden_size]
auto get_graph = [&]() -> ggml_cgraph* {
return build_graph(x, timesteps, context, ref_latents, increase_ref_index);
return build_graph(x, timesteps, context, ref_latents, ref_index_mode);
};
return restore_trailing_singleton_dims(GGMLRunner::compute<float>(get_graph, n_threads, false, false, false), x.dim());
@@ -640,7 +706,7 @@ namespace Qwen {
*diffusion_params.timesteps,
tensor_or_empty(diffusion_params.context),
diffusion_params.ref_latents ? *diffusion_params.ref_latents : empty_ref_latents,
diffusion_params.increase_ref_index);
diffusion_params.ref_index_mode);
}
void test() {
@@ -674,7 +740,7 @@ namespace Qwen {
timesteps,
context,
{},
false);
Rope::RefIndexMode::FIXED);
int64_t t1 = ggml_time_ms();
GGML_ASSERT(!out_opt.empty());
+6 -6
View File
@@ -575,7 +575,7 @@ namespace ZImage {
const sd::Tensor<float>& timesteps_tensor,
const sd::Tensor<float>& context_tensor,
const std::vector<sd::Tensor<float>>& ref_latents_tensor = {},
bool increase_ref_index = false) {
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::FIXED) {
ggml_cgraph* gf = new_graph_custom(Z_IMAGE_GRAPH_SIZE);
ggml_tensor* x = make_input(x_tensor);
ggml_tensor* timesteps = make_input(timesteps_tensor);
@@ -595,7 +595,7 @@ namespace ZImage {
static_cast<int>(context->ne[1]),
SEQ_MULTI_OF,
ref_latents,
increase_ref_index,
ref_index_mode,
config.theta,
circular_y_enabled,
circular_x_enabled,
@@ -626,12 +626,12 @@ namespace ZImage {
const sd::Tensor<float>& timesteps,
const sd::Tensor<float>& context,
const std::vector<sd::Tensor<float>>& ref_latents = {},
bool increase_ref_index = false) {
Rope::RefIndexMode ref_index_mode = Rope::RefIndexMode::FIXED) {
// x: [N, in_channels, h, w]
// timesteps: [N, ]
// context: [N, max_position, hidden_size]
auto get_graph = [&]() -> ggml_cgraph* {
return build_graph(x, timesteps, context, ref_latents, increase_ref_index);
return build_graph(x, timesteps, context, ref_latents, ref_index_mode);
};
return restore_trailing_singleton_dims(GGMLRunner::compute<float>(get_graph, n_threads, false, false, false), x.dim());
@@ -647,7 +647,7 @@ namespace ZImage {
*diffusion_params.timesteps,
tensor_or_empty(diffusion_params.context),
diffusion_params.ref_latents ? *diffusion_params.ref_latents : empty_ref_latents,
diffusion_params.increase_ref_index);
diffusion_params.ref_index_mode);
}
void test() {
@@ -681,7 +681,7 @@ namespace ZImage {
timesteps,
context,
{},
false);
Rope::RefIndexMode::FIXED);
int64_t t1 = ggml_time_ms();
GGML_ASSERT(!out_opt.empty());
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_TE_CLIP_HPP__
#ifndef __SD_MODEL_TE_CLIP_HPP__
#define __SD_MODEL_TE_CLIP_HPP__
#include "core/ggml_extend.hpp"
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_TE_LLM_HPP__
#ifndef __SD_MODEL_TE_LLM_HPP__
#define __SD_MODEL_TE_LLM_HPP__
#include <algorithm>
+56 -3
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_TE_T5_HPP__
#ifndef __SD_MODEL_TE_T5_HPP__
#define __SD_MODEL_TE_T5_HPP__
#include <cfloat>
@@ -26,13 +26,66 @@ struct T5Config {
static T5Config detect_from_weights(const String2TensorStorage& tensor_storage_map,
const std::string& prefix,
bool is_umt5 = false) {
(void)tensor_storage_map;
(void)prefix;
T5Config config;
if (is_umt5) {
config.vocab_size = 256384;
config.relative_attention = false;
}
auto find_tensor = [&](const std::string& suffix) -> const TensorStorage* {
auto it = tensor_storage_map.find(prefix + "." + suffix);
if (it != tensor_storage_map.end()) {
return &it->second;
}
it = tensor_storage_map.find(prefix + suffix);
if (it != tensor_storage_map.end()) {
return &it->second;
}
return nullptr;
};
if (const TensorStorage* shared = find_tensor("shared.weight")) {
if (shared->n_dims == 2) {
config.vocab_size = shared->ne[1];
config.model_dim = shared->ne[0];
}
}
if (const TensorStorage* q = find_tensor("encoder.block.0.layer.0.SelfAttention.q.weight")) {
if (q->n_dims == 2) {
config.model_dim = q->ne[0];
int64_t inner_dim = q->ne[1];
// Flan-T5/T5 uses d_kv=64 for common sizes.
if (inner_dim % 64 == 0) {
config.num_heads = inner_dim / 64;
}
}
}
if (const TensorStorage* wi = find_tensor("encoder.block.0.layer.1.DenseReluDense.wi_0.weight")) {
if (wi->n_dims == 2) {
config.model_dim = wi->ne[0];
config.ff_dim = wi->ne[1];
}
}
int64_t detected_layers = 0;
for (const auto& [name, _] : tensor_storage_map) {
std::string base = prefix;
if (!base.empty() && base.back() != '.') {
base += ".";
}
std::string layer_prefix = base + "encoder.block.";
if (!starts_with(name, layer_prefix)) {
continue;
}
size_t pos = layer_prefix.size();
size_t dot = name.find('.', pos);
if (dot == std::string::npos) {
continue;
}
int64_t layer = atoi(name.substr(pos, dot - pos).c_str());
detected_layers = std::max(detected_layers, layer + 1);
}
if (detected_layers > 0) {
config.num_layers = detected_layers;
}
return config;
}
};
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_UPSCALER_ESRGAN_HPP__
#ifndef __SD_MODEL_UPSCALER_ESRGAN_HPP__
#define __SD_MODEL_UPSCALER_ESRGAN_HPP__
#include <algorithm>
+1 -1
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_UPSCALER_LTX_LATENT_UPSCALER_HPP__
#ifndef __SD_MODEL_UPSCALER_LTX_LATENT_UPSCALER_HPP__
#define __SD_MODEL_UPSCALER_LTX_LATENT_UPSCALER_HPP__
#include <algorithm>
+3 -2
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_VAE_LTX_AUDIO_VAE_HPP__
#ifndef __SD_MODEL_VAE_LTX_AUDIO_VAE_HPP__
#define __SD_MODEL_VAE_LTX_AUDIO_VAE_HPP__
#include <cmath>
@@ -214,7 +214,7 @@ namespace LTXV {
auto x = ggml_reshape_3d(ctx, waveform, time, 1, channels * batch);
if (left_pad > 0) {
x = ggml_pad_ext(ctx, x, static_cast<int>(left_pad), 0, 0, 0, 0, 0, 0, 0);
x = ggml_ext_pad_ext(ctx, runner_ctx->backend, x, static_cast<int>(left_pad), 0, 0, 0, 0, 0, 0, 0);
}
auto frames = ggml_conv_1d(ctx, forward_basis, x, hop_length, 0, 1);
@@ -451,6 +451,7 @@ namespace LTXV {
int pad_h = kernel_size.first - 1;
int pad_w = kernel_size.second - 1;
x = ggml_ext_pad_ext(ctx->ggml_ctx,
ctx->backend,
x,
pad_w / 2,
pad_w - pad_w / 2,
+2 -2
View File
@@ -408,7 +408,7 @@ public:
h = conv->forward(ctx, h);
for (int j = 0; j < num_blocks; j++) {
auto block = std::dynamic_pointer_cast<MemBlock>(blocks[std::to_string(index++)]);
auto mem = ggml_pad_ext(ctx->ggml_ctx, h, 0, 0, 0, 0, 0, 0, 1, 0);
auto mem = ggml_ext_pad_ext(ctx->ggml_ctx, ctx->backend, h, 0, 0, 0, 0, 0, 0, 1, 0);
mem = ggml_view_4d(ctx->ggml_ctx, mem, h->ne[0], h->ne[1], h->ne[2], h->ne[3], h->nb[1], h->nb[2], h->nb[3], 0);
h = block->forward(ctx, h, mem);
}
@@ -479,7 +479,7 @@ public:
int index = 3;
for (int i = 0; i < num_layers; i++) {
for (int j = 0; j < num_blocks; j++) {
auto mem = ggml_pad_ext(ctx->ggml_ctx, h, 0, 0, 0, 0, 0, 0, 1, 0);
auto mem = ggml_ext_pad_ext(ctx->ggml_ctx, ctx->backend, h, 0, 0, 0, 0, 0, 0, 1, 0);
mem = ggml_view_4d(ctx->ggml_ctx, mem, h->ne[0], h->ne[1], h->ne[2], h->ne[3], h->nb[1], h->nb[2], h->nb[3], 0);
if (is_wide) {
auto block = std::dynamic_pointer_cast<WideMemBlock>(blocks[std::to_string(index++)]);
+2 -2
View File
@@ -1,4 +1,4 @@
#ifndef __SD_MODEL_VAE_VAE_HPP__
#ifndef __SD_MODEL_VAE_VAE_HPP__
#define __SD_MODEL_VAE_VAE_HPP__
#include "core/tensor_ggml.hpp"
@@ -78,7 +78,7 @@ public:
scale_factor = 16;
} else if (sd_version_uses_flux2_vae(version)) {
scale_factor = 16;
} else if (version == VERSION_CHROMA_RADIANCE || version == VERSION_HIDREAM_O1) {
} else if (version == VERSION_CHROMA_RADIANCE || version == VERSION_HIDREAM_O1 || sd_version_is_minit2i(version)) {
scale_factor = 1;
}
return scale_factor;
+34 -27
View File
@@ -72,8 +72,8 @@ namespace WAN {
lp2 -= (int)cache_x->ne[2];
}
x = ggml_ext_pad_ext(ctx->ggml_ctx, x, lp0, rp0, lp1, rp1, lp2, rp2, 0, 0, ctx->circular_x_enabled, ctx->circular_y_enabled);
return ggml_ext_conv_3d(ctx->ggml_ctx, x, w, b, in_channels,
x = ggml_ext_pad_ext(ctx->ggml_ctx, ctx->backend, x, lp0, rp0, lp1, rp1, lp2, rp2, 0, 0, ctx->circular_x_enabled, ctx->circular_y_enabled);
return ggml_ext_conv_3d(ctx->ggml_ctx, ctx->backend, x, w, b, in_channels,
std::get<2>(stride), std::get<1>(stride), std::get<0>(stride),
0, 0, 0,
std::get<2>(dilation), std::get<1>(dilation), std::get<0>(dilation));
@@ -195,7 +195,7 @@ namespace WAN {
2);
}
if (chunk_idx == 1 && cache_x->ne[2] < 2) { // Rep
cache_x = ggml_pad_ext(ctx->ggml_ctx, cache_x, 0, 0, 0, 0, (int)cache_x->ne[2], 0, 0, 0);
cache_x = ggml_ext_pad_ext(ctx->ggml_ctx, ctx->backend, cache_x, 0, 0, 0, 0, (int)cache_x->ne[2], 0, 0, 0);
// aka cache_x = torch.cat([torch.zeros_like(cache_x).to(cache_x.device),cache_x],dim=2)
}
if (chunk_idx == 1) {
@@ -283,7 +283,7 @@ namespace WAN {
int pad_t = (factor_t - T % factor_t) % factor_t;
x = ggml_pad_ext(ctx->ggml_ctx, x, 0, 0, 0, 0, pad_t, 0, 0, 0);
x = ggml_ext_pad_ext(ctx->ggml_ctx, ctx->backend, x, 0, 0, 0, 0, pad_t, 0, 0, 0);
T = x->ne[2];
x = ggml_reshape_4d(ctx->ggml_ctx, x, W * H, factor_t, T / factor_t, C); // [C, T/factor_t, factor_t, H*W]
@@ -631,6 +631,7 @@ namespace WAN {
class Encoder3d : public GGMLBlock {
protected:
bool wan2_2;
int64_t in_channels;
int64_t dim;
int64_t z_dim;
std::vector<int> dim_mult;
@@ -641,12 +642,14 @@ namespace WAN {
public:
Encoder3d(int64_t dim = 128,
int64_t z_dim = 4,
int64_t in_channels = 3,
std::vector<int> dim_mult = {1, 2, 4, 4},
int num_res_blocks = 2,
std::vector<bool> temperal_downsample = {false, true, true},
bool wan2_2 = false,
bool is_2D = false)
: dim(dim),
: in_channels(in_channels),
dim(dim),
z_dim(z_dim),
dim_mult(dim_mult),
num_res_blocks(num_res_blocks),
@@ -659,11 +662,10 @@ namespace WAN {
dims.push_back(dim * u);
}
int64_t input_dim = wan2_2 ? 12 : 3;
if (is_2D) {
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new Conv2dBut3d(input_dim, dims[0], {3, 3}, {1, 1}, {1, 1}));
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new Conv2dBut3d(in_channels, dims[0], {3, 3}, {1, 1}, {1, 1}));
} else {
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new CausalConv3d(input_dim, dims[0], {3, 3, 3}, {1, 1, 1}, {1, 1, 1}));
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new CausalConv3d(in_channels, dims[0], {3, 3, 3}, {1, 1, 1}, {1, 1, 1}));
}
int index = 0;
@@ -812,6 +814,7 @@ namespace WAN {
class Decoder3d : public GGMLBlock {
protected:
bool wan2_2;
int64_t out_channels;
int64_t dim;
int64_t z_dim;
std::vector<int> dim_mult;
@@ -822,12 +825,14 @@ namespace WAN {
public:
Decoder3d(int64_t dim = 128,
int64_t z_dim = 4,
int64_t out_channels = 3,
std::vector<int> dim_mult = {1, 2, 4, 4},
int num_res_blocks = 2,
std::vector<bool> temperal_upsample = {true, true, false},
bool wan2_2 = false,
bool is_2D = false)
: dim(dim),
: out_channels(out_channels),
dim(dim),
z_dim(z_dim),
dim_mult(dim_mult),
num_res_blocks(num_res_blocks),
@@ -889,7 +894,7 @@ namespace WAN {
// output blocks
blocks["head.0"] = std::shared_ptr<GGMLBlock>(new RMS_norm(out_dim));
int64_t final_dim = wan2_2 ? 12 : 3;
int64_t final_dim = out_channels;
// head.1 is nn.SiLU()
if (is_2D) {
blocks["head.2"] = std::shared_ptr<GGMLBlock>(new Conv2dBut3d(out_dim, final_dim, {3, 3}, {1, 1}, {1, 1}));
@@ -1002,6 +1007,8 @@ namespace WAN {
public:
bool wan2_2 = false;
bool decode_only = true;
int64_t input_channels = 3;
int patch_size = 1;
int64_t dim = 96;
int64_t dec_dim = 96;
int64_t z_dim = 16;
@@ -1026,16 +1033,22 @@ namespace WAN {
}
public:
WanVAE(bool decode_only = true, bool wan2_2 = false, bool is_2D = false)
: decode_only(decode_only), wan2_2(wan2_2), is_2D(is_2D) {
WanVAE(bool decode_only = true, SDVersion version = VERSION_WAN2, bool is_2D = false)
: decode_only(decode_only),
wan2_2(version == VERSION_WAN2_2_TI2V),
is_2D(is_2D) {
// attn_scales is always []
if (wan2_2) {
dim = 160;
dec_dim = 256;
z_dim = 48;
dim = 160;
dec_dim = 256;
z_dim = 48;
input_channels = 12;
patch_size = 2;
_conv_num = 34;
_enc_conv_num = 26;
} else if (version == VERSION_QWEN_IMAGE_LAYERED) {
input_channels = 4;
}
if (is_2D) {
@@ -1044,14 +1057,14 @@ namespace WAN {
}
if (!decode_only) {
blocks["encoder"] = std::shared_ptr<GGMLBlock>(new Encoder3d(dim, z_dim * 2, dim_mult, num_res_blocks, temperal_downsample, wan2_2, is_2D));
blocks["encoder"] = std::shared_ptr<GGMLBlock>(new Encoder3d(dim, z_dim * 2, input_channels, dim_mult, num_res_blocks, temperal_downsample, wan2_2, is_2D));
if (is_2D) {
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new Conv2dBut3d(z_dim * 2, z_dim * 2, {1, 1}));
} else {
blocks["conv1"] = std::shared_ptr<GGMLBlock>(new CausalConv3d(z_dim * 2, z_dim * 2, {1, 1, 1}));
}
}
blocks["decoder"] = std::shared_ptr<GGMLBlock>(new Decoder3d(dec_dim, z_dim, dim_mult, num_res_blocks, temperal_upsample, wan2_2, is_2D));
blocks["decoder"] = std::shared_ptr<GGMLBlock>(new Decoder3d(dec_dim, z_dim, input_channels, dim_mult, num_res_blocks, temperal_upsample, wan2_2, is_2D));
if (is_2D) {
blocks["conv2"] = std::shared_ptr<GGMLBlock>(new Conv2dBut3d(z_dim, z_dim, {1, 1}));
} else {
@@ -1125,9 +1138,7 @@ namespace WAN {
clear_cache();
if (wan2_2) {
x = patchify(ctx->ggml_ctx, x, 2, b);
}
x = patchify(ctx->ggml_ctx, x, patch_size, b);
// sd::ggml_graph_cut::mark_graph_cut(x, "wan_vae.encode.prelude", "x");
auto encoder = std::dynamic_pointer_cast<Encoder3d>(blocks["encoder"]);
@@ -1202,9 +1213,7 @@ namespace WAN {
}
}
}
if (wan2_2) {
out = unpatchify(ctx->ggml_ctx, out, 2, b);
}
out = unpatchify(ctx->ggml_ctx, out, patch_size, b);
// sd::ggml_graph_cut::mark_graph_cut(out, "wan_vae.decode.final", "out");
clear_cache();
return out;
@@ -1225,9 +1234,7 @@ namespace WAN {
auto in = ggml_ext_slice(ctx->ggml_ctx, x, 2, i, i + 1); // [b*c, 1, h, w]
_conv_idx = 0;
auto out = decoder->forward(ctx, in, b, _feat_map, _conv_idx, i);
if (wan2_2) {
out = unpatchify(ctx->ggml_ctx, out, 2, b);
}
out = unpatchify(ctx->ggml_ctx, out, patch_size, b);
// sd::ggml_graph_cut::mark_graph_cut(out, "wan_vae.decode_partial.final", "out");
return out;
}
@@ -1257,7 +1264,7 @@ namespace WAN {
if (is_2D) {
LOG_DEBUG("USING 2D VAE");
}
ae = WanVAE(decode_only, version == VERSION_WAN2_2_TI2V, is_2D);
ae = WanVAE(decode_only, version, is_2D);
ae.init(params_ctx, tensor_storage_map, prefix);
}
+16 -4
View File
@@ -20,6 +20,7 @@
#include "model_io/torch_legacy_io.h"
#include "model_io/torch_zip_io.h"
#include "model_loader.h"
#include "runtime/imatrix.h"
#include "stable-diffusion.h"
#include "core/ggml_extend_backend.h"
@@ -156,7 +157,8 @@ void convert_tensor(void* src,
void* dst,
ggml_type dst_type,
int nrows,
int n_per_row) {
int n_per_row,
std::vector<float> imatrix = {}) {
int n = nrows * n_per_row;
if (src_type == dst_type) {
size_t nbytes = n * ggml_type_size(src_type) / ggml_blck_size(src_type);
@@ -165,7 +167,7 @@ void convert_tensor(void* src,
if (dst_type == GGML_TYPE_F16) {
ggml_fp32_to_fp16_row((float*)src, (ggml_fp16_t*)dst, n);
} else {
std::vector<float> imatrix(n_per_row, 1.0f); // dummy importance matrix
imatrix.resize(n_per_row, 1.0f);
const float* im = imatrix.data();
ggml_quantize_chunk(dst_type, (float*)src, dst, 0, nrows, n_per_row, im);
}
@@ -195,7 +197,7 @@ void convert_tensor(void* src,
if (dst_type == GGML_TYPE_F16) {
ggml_fp32_to_fp16_row((float*)src_data_f32, (ggml_fp16_t*)dst, n);
} else {
std::vector<float> imatrix(n_per_row, 1.0f); // dummy importance matrix
imatrix.resize(n_per_row, 1.0f);
const float* im = imatrix.data();
ggml_quantize_chunk(dst_type, (float*)src_data_f32, dst, 0, nrows, n_per_row, im);
}
@@ -470,7 +472,13 @@ SDVersion ModelLoader::get_sd_version() {
tensor_storage_map.find("model.diffusion_model.transformer_blocks.0.img_mlp.w1.weight") != tensor_storage_map.end()) {
return VERSION_LENS;
}
if (tensor_storage.name.find("net.img_embedder.proj1.weight") != std::string::npos) {
return VERSION_MINIT2I;
}
if (tensor_storage.name.find("model.diffusion_model.transformer_blocks.0.img_mod.1.weight") != std::string::npos) {
if (tensor_storage_map.find("model.diffusion_model.time_text_embed.addition_t_embedding.weight") != tensor_storage_map.end()) {
return VERSION_QWEN_IMAGE_LAYERED;
}
return VERSION_QWEN_IMAGE;
}
if (tensor_storage.name.find("llm_adapter.blocks.0.cross_attn.q_proj.weight") != std::string::npos) {
@@ -970,6 +978,7 @@ bool ModelLoader::load_tensors(on_new_tensor_cb_t on_new_tensor_cb,
size_t total_tensors_processed = 0;
const int64_t t_start = start_time;
int last_n_threads = 1;
SDVersion imatrix_version = (version_ == VERSION_COUNT) ? get_sd_version() : version_;
for (size_t file_index = 0; file_index < file_data.size(); ++file_index) {
auto& fdata = file_data[file_index];
@@ -1154,12 +1163,15 @@ bool ModelLoader::load_tensors(on_new_tensor_cb_t on_new_tensor_cb,
failed = true;
return;
}
std::string processed_name = convert_tensor_name(tensor_storage.name, imatrix_version);
std::vector<float> imatrix = get_imatrix_collector().get_values(processed_name);
convert_tensor((void*)target_buf,
tensor_storage.type,
convert_buf,
dst_tensor->type,
(int)tensor_storage.nelements() / (int)tensor_storage.ne[0],
(int)tensor_storage.ne[0]);
(int)tensor_storage.ne[0],
std::move(imatrix));
} else {
convert_buf = read_buf;
}
+2
View File
@@ -27,6 +27,8 @@ struct MmapTensorStore {
std::shared_ptr<struct ggml_backend_buffer> mmbuffer;
};
bool is_unused_tensor(const std::string& name);
class ModelLoader {
protected:
SDVersion version_ = VERSION_COUNT;
+61 -13
View File
@@ -100,12 +100,41 @@ size_t estimate_tensors_size(const std::map<std::string, ggml_tensor*>& tensors)
return size;
}
void ModelManager::set_split_buffer_type(ggml_backend_t compute_backend, ggml_backend_buffer_type_t split_buft) {
if (compute_backend == nullptr) {
return;
}
if (split_buft == nullptr) {
split_buffer_types_.erase(compute_backend);
return;
}
split_buffer_types_[compute_backend] = split_buft;
}
bool ModelManager::tensor_shape_supports_split_buffer(const ggml_tensor* tensor) {
return tensor != nullptr &&
tensor->view_src == nullptr &&
ggml_is_contiguous(tensor) &&
ggml_n_dims(tensor) == 2 &&
tensor->ne[0] >= 256 &&
tensor->ne[1] >= 256;
}
ggml_backend_buffer_type_t ModelManager::split_buffer_type_for(const TensorState& state) const {
if (!state.allow_split_buffer || !tensor_shape_supports_split_buffer(state.tensor)) {
return nullptr;
}
auto it = split_buffer_types_.find(state.compute_backend);
return it != split_buffer_types_.end() ? it->second : nullptr;
}
bool ModelManager::register_param_tensors(const std::string& desc,
std::map<std::string, ggml_tensor*> tensors,
ResidencyMode residency_mode,
ggml_backend_t compute_backend,
ggml_backend_t params_backend,
size_t* registered_tensor_size) {
size_t* registered_tensor_size,
bool allow_split_buffer) {
if (desc.empty()) {
LOG_ERROR("model manager tensor desc is empty");
return false;
@@ -129,13 +158,14 @@ bool ModelManager::register_param_tensors(const std::string& desc,
}
ggml_set_name(tensor, name.c_str());
auto state = std::make_unique<TensorState>();
state->name = name;
state->tensor = tensor;
state->desc = desc;
state->residency_mode = residency_mode;
state->compute_backend = compute_backend;
state->params_backend = params_backend;
auto state = std::make_unique<TensorState>();
state->name = name;
state->tensor = tensor;
state->desc = desc;
state->residency_mode = residency_mode;
state->compute_backend = compute_backend;
state->params_backend = params_backend;
state->allow_split_buffer = allow_split_buffer;
new_states.push_back(std::move(state));
}
@@ -237,7 +267,7 @@ bool ModelManager::load_tensors_to_params_backend(const std::vector<TensorState*
}
bool ModelManager::stage_tensors_to_compute_backend(const std::vector<TensorState*>& states) {
std::map<ggml_backend_t, std::vector<TensorState*>> states_by_compute_backend;
std::map<std::pair<ggml_backend_t, ggml_backend_buffer_type_t>, std::vector<TensorState*>> states_by_staging_target;
for (TensorState* state : states) {
if (state == nullptr || should_ignore(*state) || is_optional_missing_tensor(state->name)) {
continue;
@@ -257,11 +287,16 @@ bool ModelManager::stage_tensors_to_compute_backend(const std::vector<TensorStat
LOG_ERROR("model manager tensor '%s' is not loaded to params backend", state->name.c_str());
return false;
}
states_by_compute_backend[state->compute_backend].push_back(state);
ggml_backend_buffer_type_t staging_buft = split_buffer_type_for(*state);
if (staging_buft == nullptr) {
staging_buft = ggml_backend_get_default_buffer_type(state->compute_backend);
}
states_by_staging_target[{state->compute_backend, staging_buft}].push_back(state);
}
for (const auto& pair : states_by_compute_backend) {
ggml_backend_t compute_backend = pair.first;
for (const auto& pair : states_by_staging_target) {
ggml_backend_t compute_backend = pair.first.first;
ggml_backend_buffer_type_t staging_buft = pair.first.second;
const std::vector<TensorState*>& states = pair.second;
if (states.empty()) {
continue;
@@ -285,7 +320,7 @@ bool ModelManager::stage_tensors_to_compute_backend(const std::vector<TensorStat
staged_tensors.push_back({state, staging_tensor});
}
ggml_backend_buffer_t compute_buffer = ggml_backend_alloc_ctx_tensors(staging_ctx, compute_backend);
ggml_backend_buffer_t compute_buffer = ggml_backend_alloc_ctx_tensors_from_buft(staging_ctx, staging_buft);
if (compute_buffer == nullptr) {
LOG_ERROR("model manager alloc compute params backend buffer failed, num_tensors = %zu",
staged_tensors.size());
@@ -350,6 +385,17 @@ bool ModelManager::apply_loras_to_params(const std::vector<TensorState*>& states
LOG_ERROR("model manager compute backend is null for lora target tensor '%s'", state->name.c_str());
return false;
}
if (state->tensor->buffer != nullptr &&
ggml_backend_buffer_get_type(state->tensor->buffer) == split_buffer_type_for(*state)) {
if (!warned_split_lora_skip_) {
LOG_WARN(
"model manager skipping direct lora application to row-split tensors "
"(use --lora-apply-mode at_runtime with row split)");
warned_split_lora_skip_ = true;
}
state->applied_lora_epoch = current_lora_epoch_;
continue;
}
if (state->tensor->data == nullptr) {
LOG_ERROR("model manager lora target tensor '%s' is not prepared", state->name.c_str());
return false;
@@ -694,6 +740,8 @@ ggml_backend_buffer_type_t ModelManager::params_buffer_type_for(const TensorStat
if (compute_dev != nullptr) {
params_buft = ggml_backend_dev_host_buffer_type(compute_dev);
}
} else if (state.params_backend == state.compute_backend) {
params_buft = split_buffer_type_for(state);
}
if (params_buft == nullptr) {
params_buft = ggml_backend_get_default_buffer_type(state.params_backend);
+9 -1
View File
@@ -36,6 +36,7 @@ private:
ResidencyMode residency_mode = ResidencyMode::ParamBackend;
ggml_backend_t compute_backend = nullptr;
ggml_backend_t params_backend = nullptr;
bool allow_split_buffer = false;
bool metadata_validated = false;
int active_prepare_count = 0;
@@ -63,6 +64,8 @@ private:
std::map<std::string, TensorState*> tensor_states_by_name_;
std::vector<std::unique_ptr<ParamsStorageBlock>> params_storage_blocks_;
std::vector<std::unique_ptr<ComputeStagingBlock>> compute_staging_blocks_;
std::map<ggml_backend_t, ggml_backend_buffer_type_t> split_buffer_types_;
bool warned_split_lora_skip_ = false;
std::set<std::string> common_ignore_tensors_;
std::vector<LoraSpec> loras_;
SDVersion lora_version_ = VERSION_COUNT;
@@ -91,6 +94,7 @@ private:
bool stage_tensors_to_compute_backend(const std::vector<TensorState*>& states);
ggml_backend_buffer_type_t params_buffer_type_for(const TensorState& state) const;
ggml_backend_buffer_type_t split_buffer_type_for(const TensorState& state) const;
void release_compute_staging_blocks(bool force = false,
const std::unordered_set<TensorState*>* target_states = nullptr);
void release_params_storage_blocks(bool force = false,
@@ -114,6 +118,9 @@ public:
void set_writable_mmap(bool writable_mmap) { writable_mmap_ = writable_mmap; }
void set_common_ignore_tensors(std::set<std::string> ignore_tensors);
void set_loras(std::vector<LoraSpec> loras, SDVersion version);
void set_split_buffer_type(ggml_backend_t compute_backend, ggml_backend_buffer_type_t split_buft);
static bool tensor_shape_supports_split_buffer(const ggml_tensor* tensor);
std::set<std::string> tensor_names() const;
@@ -122,7 +129,8 @@ public:
ResidencyMode residency_mode,
ggml_backend_t compute_backend,
ggml_backend_t params_backend,
size_t* registered_tensor_size = nullptr);
size_t* registered_tensor_size = nullptr,
bool allow_split_buffer = false);
template <typename Runner>
bool register_runner_params(const std::string& desc,
+197
View File
@@ -302,6 +302,137 @@ struct KarrasScheduler : SigmaScheduler {
}
};
struct BetaScheduler : SigmaScheduler {
static constexpr double alpha = 0.6;
static constexpr double beta = 0.6;
static double log_beta(double a, double b) {
return std::lgamma(a) + std::lgamma(b) - std::lgamma(a + b);
}
static double incbeta(double x, double a, double b) {
if (x <= 0.0) {
return 0.0;
}
if (x >= 1.0) {
return 1.0;
}
// Continued fraction approximation using Lentz's method.
const int max_iter = 200;
const double epsilon = 3.0e-7;
const double tiny = 1e-30;
const double qab = a + b;
const double qap = a + 1.0;
const double qam = a - 1.0;
double c = 1.0;
double d = 1.0 - qab * x / qap;
if (std::abs(d) < tiny) {
d = tiny;
}
d = 1.0 / d;
double h = d;
for (int m = 1; m <= max_iter; m++) {
const int m2 = 2 * m;
double aa = m * (b - m) * x / ((qam + m2) * (a + m2));
d = 1.0 + aa * d;
if (std::abs(d) < tiny) {
d = tiny;
}
c = 1.0 + aa / c;
if (std::abs(c) < tiny) {
c = tiny;
}
d = 1.0 / d;
h *= d * c;
aa = -(a + m) * (qab + m) * x / ((a + m2) * (qap + m2));
d = 1.0 + aa * d;
if (std::abs(d) < tiny) {
d = tiny;
}
c = 1.0 + aa / c;
if (std::abs(c) < tiny) {
c = tiny;
}
d = 1.0 / d;
const double del = d * c;
h *= del;
if (std::abs(del - 1.0) < epsilon) {
break;
}
}
return std::exp(a * std::log(x) + b * std::log(1.0 - x) - log_beta(a, b)) / a * h;
}
static double beta_cdf(double x, double a, double b) {
if (x == 0.0) {
return 0.0;
}
if (x == 1.0) {
return 1.0;
}
if (x < (a + 1.0) / (a + b + 2.0)) {
return incbeta(x, a, b);
}
return 1.0 - incbeta(1.0 - x, b, a);
}
static double beta_ppf(double u, double a, double b, int max_iter = 30) {
double x = 0.5;
for (int i = 0; i < max_iter; i++) {
const double f = beta_cdf(x, a, b) - u;
if (std::abs(f) < 1e-10) {
break;
}
const double df = std::exp((a - 1.0) * std::log(x) + (b - 1.0) * std::log(1.0 - x) - log_beta(a, b));
x -= f / df;
if (x <= 0.0) {
x = 1e-10;
}
if (x >= 1.0) {
x = 1.0 - 1e-10;
}
}
return x;
}
std::vector<float> get_sigmas(uint32_t n, float /*sigma_min*/, float /*sigma_max*/, t_to_sigma_t t_to_sigma) override {
std::vector<float> result;
result.reserve(n + 1);
const int t_max = TIMESTEPS - 1;
if (n == 0) {
return result;
} else if (n == 1) {
result.push_back(t_to_sigma(static_cast<float>(t_max)));
result.push_back(0.f);
return result;
}
int last_t = -1;
for (uint32_t i = 0; i < n; i++) {
const double u = 1.0 - static_cast<double>(i) / static_cast<double>(n);
const double t_cont = beta_ppf(u, alpha, beta) * t_max;
const int t = static_cast<int>(std::lround(t_cont));
if (t != last_t) {
result.push_back(t_to_sigma(static_cast<float>(t)));
last_t = t;
}
}
result.push_back(0.f);
return result;
}
};
struct SimpleScheduler : SigmaScheduler {
std::vector<float> get_sigmas(uint32_t n, float sigma_min, float sigma_max, t_to_sigma_t t_to_sigma) override {
std::vector<float> result_sigmas;
@@ -895,6 +1026,10 @@ struct Denoiser {
LOG_INFO("get_sigmas with Karras scheduler");
scheduler = std::make_shared<KarrasScheduler>();
break;
case BETA_SCHEDULER:
LOG_INFO("get_sigmas with Beta scheduler");
scheduler = std::make_shared<BetaScheduler>();
break;
case EXPONENTIAL_SCHEDULER:
LOG_INFO("get_sigmas exponential scheduler");
scheduler = std::make_shared<ExponentialScheduler>();
@@ -1203,6 +1338,68 @@ struct SefiFlowDenoiser : public FluxFlowDenoiser {
}
};
// MiniT2I predicts x0 directly and integrates a linear flow ODE:
// x_{t+dt} = x_t + (x0 - x_t)/(1 - t) * dt, t in [0, 1), x0 = start = noise * 2.
// Mapping sigma = 1 - t makes the generic Euler update
// x += (x - denoised)/sigma * (sigma_next - sigma)
// exactly reproduce that step when denoised == x0. To make the generic
// `denoised = pred * c_out + x * c_skip` yield x0 from the model's raw x0
// prediction we use c_skip = 0, c_out = 1, c_in = 1. Sigmas run linearly 1 -> 0.
struct MiniT2IFlowDenoiser : public Denoiser {
float sigma_min() override {
return 0.0f;
}
float sigma_max() override {
return 1.0f;
}
float sigma_to_t(float sigma) override {
return 1.0f - sigma;
}
float t_to_sigma(float t) override {
return 1.0f - t;
}
std::vector<float> get_scalings(float sigma) override {
SD_UNUSED(sigma);
float c_skip = 0.0f;
float c_out = 1.0f;
float c_in = 1.0f;
return {c_skip, c_out, c_in};
}
sd::Tensor<float> noise_scaling(float sigma,
const sd::Tensor<float>& noise,
const sd::Tensor<float>& latent) override {
SD_UNUSED(sigma);
SD_UNUSED(latent);
// Sampling starts from x0_init = noise * 2 (see MiniT2I reference).
return noise * 2.0f;
}
sd::Tensor<float> inverse_noise_scaling(float sigma, const sd::Tensor<float>& latent) override {
SD_UNUSED(sigma);
return latent;
}
std::vector<float> get_sigmas(uint32_t n, int image_seq_len, scheduler_t scheduler_type, SDVersion version, const char* extra_sample_args = nullptr) override {
SD_UNUSED(image_seq_len);
SD_UNUSED(scheduler_type);
SD_UNUSED(version);
SD_UNUSED(extra_sample_args);
// Uniform t schedule 0 -> 1 => sigma 1 -> 0, matching the reference loop.
std::vector<float> sigmas;
sigmas.reserve(n + 1);
for (uint32_t i = 0; i < n; ++i) {
sigmas.push_back(1.0f - static_cast<float>(i) / static_cast<float>(n));
}
sigmas.push_back(0.0f);
return sigmas;
}
};
typedef std::function<sd::guidance::GuiderOutput(const sd::Tensor<float>&, float, int)> denoise_cb_t;
static std::pair<float, float> get_ancestral_step(float sigma_from,
+308
View File
@@ -0,0 +1,308 @@
#include "runtime/imatrix.h"
/* Adapted from llama.cpp (credits: Kawrakow). */
#include "core/util.h"
#include "ggml-backend.h"
#include "ggml.h"
#include "stable-diffusion.h"
#include <cmath>
#include <cstdlib>
#include <cstring>
static IMatrixCollector imatrix_collector;
IMatrixCollector& get_imatrix_collector() {
return imatrix_collector;
}
// remove any prefix and suffixes from the name
// CUDA0#blk.0.attn_k.weight#0 => blk.0.attn_k.weight
static std::string filter_tensor_name(const char* name) {
std::string wname;
const char* p = strchr(name, '#');
if (p != NULL) {
p = p + 1;
const char* q = strchr(p, '#');
if (q != NULL) {
wname = std::string(p, q - p);
} else {
wname = p;
}
} else {
wname = name;
}
return wname;
}
bool IMatrixCollector::collect_imatrix(struct ggml_tensor* t, bool ask, void* user_data) {
GGML_UNUSED(user_data);
if (t == nullptr) {
return false;
}
if (t->op != GGML_OP_MUL_MAT && t->op != GGML_OP_MUL_MAT_ID) {
return false;
}
const struct ggml_tensor* src0 = t->src[0];
const struct ggml_tensor* src1 = t->src[1];
if (src0 == nullptr || src1 == nullptr) {
return false;
}
std::string wname = filter_tensor_name(src0->name);
// when ask is true, the scheduler wants to know if we are interested in data from this tensor
// if we return true, a follow-up call will be made with ask=false in which we can do the actual collection
if (ask) {
if (t->op == GGML_OP_MUL_MAT_ID) {
return true; // collect all indirect matrix multiplications
}
// why are small batches ignored (<16 tokens)?
// if (src1->ne[1] < 16 || src1->type != GGML_TYPE_F32) return false;
if (!(wname.substr(0, 6) == "model." || wname.substr(0, 17) == "cond_stage_model." || wname.substr(0, 14) == "text_encoders.")) {
return false;
}
return true;
}
std::lock_guard<std::mutex> lock(mutex_);
// copy the data from the GPU memory if needed
const bool is_host = src1->buffer == NULL || ggml_backend_buffer_is_host(src1->buffer);
if (!is_host) {
src1_data_.resize(ggml_nelements(src1));
ggml_backend_tensor_get(src1, src1_data_.data(), 0, ggml_nbytes(src1));
}
const float* data = is_host ? (const float*)src1->data : src1_data_.data();
// this has been adapted to the new format of storing merged experts in a single 3d tensor
// ref: https://github.com/ggml-org/llama.cpp/pull/6387
if (t->op == GGML_OP_MUL_MAT_ID) {
// ids -> [n_experts_used, n_tokens]
// src1 -> [cols, n_expert_used, n_tokens]
const ggml_tensor* ids = t->src[2];
const int n_as = static_cast<int>(src0->ne[2]);
const int n_ids = static_cast<int>(ids->ne[0]);
// the top-k selected expert ids are stored in the ids tensor
// for simplicity, always copy ids to host, because it is small
// take into account that ids is not contiguous!
GGML_ASSERT(ids->ne[1] == src1->ne[2]);
ids_.resize(ggml_nbytes(ids));
ggml_backend_tensor_get(ids, ids_.data(), 0, ggml_nbytes(ids));
auto& e = stats_[wname];
++e.ncall;
if (e.values.empty()) {
e.values.resize(src1->ne[0] * n_as, 0);
e.counts.resize(src1->ne[0] * n_as, 0);
} else if (e.values.size() != (size_t)src1->ne[0] * n_as) {
LOG_ERROR("inconsistent size for %s (%d vs %d)\n", wname.c_str(), (int)e.values.size(), (int)src1->ne[0] * n_as);
exit(1); // GGML_ABORT("fatal error");
}
// loop over all possible experts, regardless if they are used or not in the batch
for (int ex = 0; ex < n_as; ++ex) {
size_t e_start = ex * src1->ne[0];
for (int idx = 0; idx < n_ids; ++idx) {
for (int row = 0; row < (int)src1->ne[2]; ++row) {
const int excur = *(const int32_t*)(ids_.data() + row * ids->nb[1] + idx * ids->nb[0]);
GGML_ASSERT(excur >= 0 && excur < n_as); // sanity check
if (excur != ex)
continue;
const int64_t i11 = idx % src1->ne[1];
const int64_t i12 = row;
const float* x = (const float*)((const char*)data + i11 * src1->nb[1] + i12 * src1->nb[2]);
for (int j = 0; j < (int)src1->ne[0]; ++j) {
e.values[e_start + j] += x[j] * x[j];
e.counts[e_start + j]++;
if (!std::isfinite(e.values[e_start + j])) {
LOG_ERROR("%f detected in %s\n", e.values[e_start + j], wname.c_str());
exit(1);
}
}
}
}
}
} else {
auto& e = stats_[wname];
if (e.values.empty()) {
e.values.resize(src1->ne[0], 0);
e.counts.resize(src1->ne[0], 0);
} else if (e.values.size() != (size_t)src1->ne[0]) {
LOG_WARN("inconsistent size for %s (%d vs %d)\n", wname.c_str(), (int)e.values.size(), (int)src1->ne[0]);
exit(1); // GGML_ABORT("fatal error");
}
++e.ncall;
for (int row = 0; row < (int)src1->ne[1]; ++row) {
const float* x = data + row * src1->ne[0];
for (int j = 0; j < (int)src1->ne[0]; ++j) {
if (std::isfinite(x[j])) {
e.values[j] += x[j] * x[j];
e.counts[j]++;
if (!std::isfinite(e.values[j])) {
LOG_WARN("%f detected in %s\n", e.values[j], wname.c_str());
exit(1);
}
} else {
// Likely something from an attention mask?
}
}
}
}
return true;
}
bool load_imatrix(const char* imatrix_path) {
return imatrix_collector.load_imatrix(imatrix_path);
}
void save_imatrix(const char* imatrix_path) {
imatrix_collector.save_imatrix(imatrix_path);
}
static bool collect_imatrix(struct ggml_tensor* t, bool ask, void* user_data) {
return imatrix_collector.collect_imatrix(t, ask, user_data);
}
void enable_imatrix_collection() {
sd_set_backend_eval_callback(collect_imatrix, nullptr);
}
void disable_imatrix_collection() {
sd_set_backend_eval_callback(nullptr, nullptr);
}
void IMatrixCollector::save_imatrix(std::string fname, int ncall) const {
if (ncall > 0) {
fname += ".at_";
fname += std::to_string(ncall);
}
// avoid writing imatrix entries that do not have full data
// this can happen with MoE models where some of the experts end up not being exercised by the provided training data
int n_entries = 0;
std::vector<std::string> to_store;
for (const auto& kv : stats_) {
const int n_all = static_cast<int>(kv.second.counts.size());
if (n_all == 0) {
continue;
}
int n_zeros = 0;
for (const int c : kv.second.counts) {
if (c == 0) {
n_zeros++;
}
}
if (n_zeros == n_all) {
LOG_WARN("entry '%40s' has no data - skipping\n", kv.first.c_str());
continue;
}
if (n_zeros > 0) {
LOG_WARN("entry '%40s' has partial data (%.2f%%) - skipping\n", kv.first.c_str(), 100.0f * (n_all - n_zeros) / n_all);
continue;
}
n_entries++;
to_store.push_back(kv.first);
}
if (to_store.size() < stats_.size()) {
LOG_WARN("storing only %zu out of %zu entries\n", to_store.size(), stats_.size());
}
std::ofstream out(fname, std::ios::binary);
out.write((const char*)&n_entries, sizeof(n_entries));
for (const auto& name : to_store) {
const auto& stat = stats_.at(name);
int len = static_cast<int>(name.size());
out.write((const char*)&len, sizeof(len));
out.write(name.c_str(), len);
out.write((const char*)&stat.ncall, sizeof(stat.ncall));
int nval = static_cast<int>(stat.values.size());
out.write((const char*)&nval, sizeof(nval));
if (nval > 0) {
std::vector<float> tmp(nval);
for (int i = 0; i < nval; i++) {
tmp[i] = (stat.values[i] / static_cast<float>(stat.counts[i])) * static_cast<float>(stat.ncall);
}
out.write((const char*)tmp.data(), nval * sizeof(float));
}
}
// Write the number of call the matrix was computed with
out.write((const char*)&last_call_, sizeof(last_call_));
}
bool IMatrixCollector::load_imatrix(const char* fname) {
std::ifstream in(fname, std::ios::binary);
if (!in) {
LOG_ERROR("failed to open %s\n", fname);
return false;
}
int n_entries;
in.read((char*)&n_entries, sizeof(n_entries));
if (in.fail() || n_entries < 1) {
LOG_ERROR("no data in file %s\n", fname);
return false;
}
for (int i = 0; i < n_entries; ++i) {
int len;
in.read((char*)&len, sizeof(len));
std::vector<char> name_as_vec(len + 1);
in.read((char*)name_as_vec.data(), len);
if (in.fail()) {
LOG_ERROR("failed reading name for entry %d from %s\n", i + 1, fname);
return false;
}
name_as_vec[len] = 0;
std::string name{name_as_vec.data()};
auto& e = stats_[std::move(name)];
int ncall;
in.read((char*)&ncall, sizeof(ncall));
int nval;
in.read((char*)&nval, sizeof(nval));
if (in.fail() || nval < 1) {
LOG_ERROR("failed reading number of values for entry %d\n", i);
stats_ = {};
return false;
}
if (e.values.empty()) {
e.values.resize(nval, 0);
e.counts.resize(nval, 0);
}
std::vector<float> tmp(nval);
in.read((char*)tmp.data(), nval * sizeof(float));
if (in.fail()) {
LOG_ERROR("failed reading data for entry %d\n", i);
stats_ = {};
return false;
}
// Recreate the state as expected by save_imatrix(), and correct for weighted sum.
for (int i = 0; i < nval; i++) {
e.values[i] += tmp[i];
e.counts[i] += ncall;
}
e.ncall += ncall;
}
return true;
}
+45
View File
@@ -0,0 +1,45 @@
#ifndef __SD_RUNTIME_IMATRIX_H__
#define __SD_RUNTIME_IMATRIX_H__
#include <fstream>
#include <mutex>
#include <string>
#include <unordered_map>
#include <vector>
/* Adapted from llama.cpp (credits: Kawrakow). */
struct ggml_tensor;
struct IMatrixStats {
std::vector<float> values{};
std::vector<int> counts{};
int ncall = 0;
};
class IMatrixCollector {
private:
std::unordered_map<std::string, IMatrixStats> stats_ = {};
std::mutex mutex_;
int last_call_ = 0;
std::vector<float> src1_data_;
std::vector<char> ids_; // the expert ids from ggml_mul_mat_id
public:
IMatrixCollector() = default;
bool collect_imatrix(struct ggml_tensor* t, bool ask, void* user_data);
void save_imatrix(std::string fname, int ncall = -1) const;
bool load_imatrix(const char* fname);
std::vector<float> get_values(const std::string& key) const {
auto it = stats_.find(key);
if (it != stats_.end()) {
return it->second.values;
} else {
return {};
}
}
};
IMatrixCollector& get_imatrix_collector();
#endif // __SD_RUNTIME_IMATRIX_H__
+456 -60
View File
@@ -2,11 +2,14 @@
#include <cmath>
#include <cstdlib>
#include <set>
#include <type_traits>
#include <unordered_set>
#include <utility>
#include <vector>
#include "core/ggml_extend.hpp"
#include "core/ggml_graph_cut.h"
#include "core/layer_split_partition.h"
#include "core/rng.hpp"
#include "core/rng_mt19937.hpp"
@@ -17,6 +20,7 @@
#include "stable-diffusion.h"
#include "conditioning/conditioner.hpp"
#include "core/backend_fit.h"
#include "extensions/generation_extension.h"
#include "model/adapter/lora.hpp"
#include "model/diffusion/anima.hpp"
@@ -29,6 +33,7 @@
#include "model/diffusion/krea2.hpp"
#include "model/diffusion/lens.hpp"
#include "model/diffusion/ltxv.hpp"
#include "model/diffusion/minit2i.hpp"
#include "model/diffusion/mmdit.hpp"
#include "model/diffusion/model.hpp"
#include "model/diffusion/pid.hpp"
@@ -83,6 +88,7 @@ const char* model_version_to_str[] = {
"Wan 2.2 I2V",
"Wan 2.2 TI2V",
"Qwen Image",
"Qwen Image Layered",
"Anima",
"Flux.2",
"Flux.2 klein",
@@ -93,6 +99,7 @@ const char* model_version_to_str[] = {
"Ovis Image",
"Ernie Image",
"Lens",
"MiniT2I",
"Longcat-Image",
"PiD",
"Ideogram 4",
@@ -167,6 +174,13 @@ static float get_cache_reuse_threshold(const sd_cache_params_t& params) {
/*=============================================== StableDiffusionGGML ================================================*/
template <typename T, typename = void>
struct has_set_runtime_backends : std::false_type {};
template <typename T>
struct has_set_runtime_backends<T,
std::void_t<decltype(std::declval<T&>().set_runtime_backends(
std::declval<const std::vector<ggml_backend_t>&>()))>> : std::true_type {};
static_assert(std::atomic<sd_cancel_mode_t>::is_always_lock_free,
"sd_cancel_mode_t must be lock-free");
@@ -205,6 +219,8 @@ public:
bool eager_load = false;
std::string backend_spec;
std::string params_backend_spec;
std::string split_mode_spec;
bool auto_fit_enabled = false;
bool is_using_v_parameterization = false;
bool is_using_edm_v_parameterization = false;
@@ -272,18 +288,230 @@ public:
if (model_manager == nullptr) {
return true;
}
ModelManager::ResidencyMode residency_mode =
backend_manager.params_backend_is_disk(module) ? ModelManager::ResidencyMode::Disk : ModelManager::ResidencyMode::ParamBackend;
std::vector<ggml_backend_t> module_backends = backend_manager.runtime_backends(module);
if (module_backends.size() > 1) {
if constexpr (has_set_runtime_backends<T>::value) {
if (module == SDBackendModule::DIFFUSION || module == SDBackendModule::TE) {
if (backend_manager.split_mode(module) == SDSplitMode::ROW) {
return register_row_split_runner_params(desc,
model,
module,
module_backends,
std::move(group_tensors),
residency_mode,
params_mem_size);
}
return register_layer_split_runner_params(desc,
model,
module,
module_backends,
std::move(group_tensors),
residency_mode,
params_mem_size);
}
}
LOG_WARN("%s module does not support multiple runtime backends; using %s",
sd_backend_module_name(module),
sd::layer_split_backend_device_display_name(module_backends[0]).c_str());
}
return model_manager->register_param_tensors(desc,
std::move(group_tensors),
backend_manager.params_backend_is_disk(module) ? ModelManager::ResidencyMode::Disk : ModelManager::ResidencyMode::ParamBackend,
residency_mode,
backend_for(module),
params_backend_for(module),
params_mem_size);
}
template <typename T>
bool register_row_split_runner_params(const std::string& desc,
const std::shared_ptr<T>& model,
SDBackendModule module,
const std::vector<ggml_backend_t>& module_backends,
std::map<std::string, ggml_tensor*> group_tensors,
ModelManager::ResidencyMode residency_mode,
size_t* params_mem_size) {
ggml_backend_t main_backend = module_backends[0];
auto fall_back_to_layer_split = [&](const char* reason) {
LOG_WARN("%s: row split unavailable (%s); falling back to layer split", desc.c_str(), reason);
return register_layer_split_runner_params(desc,
model,
module,
module_backends,
std::move(group_tensors),
residency_mode,
params_mem_size);
};
ggml_backend_dev_t main_dev = ggml_backend_get_device(main_backend);
ggml_backend_reg_t reg = main_dev != nullptr ? ggml_backend_dev_backend_reg(main_dev) : nullptr;
if (reg == nullptr) {
return fall_back_to_layer_split("no backend registry");
}
const size_t reg_dev_count = ggml_backend_reg_dev_count(reg);
std::vector<float> tensor_split(reg_dev_count, 0.0f);
constexpr int64_t compute_headroom_bytes = 2ll * 1024 * 1024 * 1024;
for (ggml_backend_t backend : module_backends) {
ggml_backend_dev_t dev = ggml_backend_get_device(backend);
int reg_index = -1;
for (size_t i = 0; i < reg_dev_count; i++) {
if (ggml_backend_reg_dev_get(reg, i) == dev) {
reg_index = (int)i;
break;
}
}
if (reg_index < 0) {
return fall_back_to_layer_split("devices span different backend registries");
}
size_t free_bytes = 0, total_bytes = 0;
ggml_backend_dev_memory(dev, &free_bytes, &total_bytes);
int64_t usable_bytes = std::max<int64_t>((int64_t)free_bytes - compute_headroom_bytes,
(int64_t)free_bytes / 8);
tensor_split[reg_index] = usable_bytes > 0 ? (float)((double)usable_bytes / (1024.0 * 1024.0)) : 1.0f;
}
ggml_backend_buffer_type_t split_buft = backend_manager.split_buffer_type(main_backend, tensor_split);
if (split_buft == nullptr) {
return fall_back_to_layer_split("backend has no split buffer type");
}
model_manager->set_split_buffer_type(main_backend, split_buft);
std::map<std::string, ggml_tensor*> split_tensors;
if constexpr (std::is_base_of_v<Conditioner, T>) {
model->get_layer_split_param_tensors(split_tensors);
} else {
split_tensors = group_tensors;
}
std::map<std::string, ggml_tensor*> row_split_map;
std::map<std::string, ggml_tensor*> regular_map;
size_t row_split_bytes = 0;
for (const auto& kv : group_tensors) {
if (split_tensors.count(kv.first) != 0 &&
sd::layer_split_tensor_block_index(kv.first) >= 0 &&
ModelManager::tensor_shape_supports_split_buffer(kv.second)) {
row_split_map[kv.first] = kv.second;
row_split_bytes += ggml_nbytes(kv.second);
} else {
regular_map[kv.first] = kv.second;
}
}
if (row_split_map.empty()) {
return fall_back_to_layer_split("no row-splittable transformer block weights found");
}
LOG_INFO("%s row split: %zu tensors (%.1f MB) split across %zu devices (main %s)",
desc.c_str(),
row_split_map.size(),
row_split_bytes / (1024.f * 1024.f),
module_backends.size(),
sd::layer_split_backend_device_display_name(main_backend).c_str());
if (!model_manager->register_param_tensors(desc,
std::move(row_split_map),
residency_mode,
main_backend,
params_backend_for(module),
params_mem_size,
/*allow_split_buffer=*/true)) {
return false;
}
return model_manager->register_param_tensors(desc,
std::move(regular_map),
residency_mode,
main_backend,
params_backend_for(module),
params_mem_size);
}
// Register each layer-split partition with its compute backend; the
// ModelManager handles allocation, staging, and LoRA by backend.
template <typename T>
bool register_layer_split_runner_params(const std::string& desc,
const std::shared_ptr<T>& model,
SDBackendModule module,
const std::vector<ggml_backend_t>& module_backends,
std::map<std::string, ggml_tensor*> group_tensors,
ModelManager::ResidencyMode residency_mode,
size_t* params_mem_size) {
bool has_cpu_device = false;
for (ggml_backend_t backend : module_backends) {
has_cpu_device = has_cpu_device || sd_backend_is_cpu(backend);
}
if (has_cpu_device) {
// The scheduler reserves the CPU slot for its fallback backend, and
// CPU weight participation is what --params-backend <module>=cpu is
// for; a CPU device in a split list is almost certainly a mistake.
LOG_WARN(
"%s: layer split across a CPU device is not supported; using %s "
"(use --params-backend %s=cpu to keep weights in RAM)",
desc.c_str(),
sd::layer_split_backend_device_display_name(module_backends[0]).c_str(),
sd_backend_module_name(module));
return model_manager->register_param_tensors(desc,
std::move(group_tensors),
residency_mode,
module_backends[0],
params_backend_for(module),
params_mem_size);
}
std::map<std::string, ggml_tensor*> split_tensors;
if constexpr (std::is_base_of_v<Conditioner, T>) {
model->get_layer_split_param_tensors(split_tensors);
} else {
split_tensors = group_tensors;
}
auto partitions = sd::partition_layer_split_tensors(desc, group_tensors, split_tensors, module_backends);
bool is_split = false;
for (size_t i = 1; i < partitions.size(); i++) {
if (!partitions[i].empty()) {
is_split = true;
break;
}
}
if (!is_split) {
return model_manager->register_param_tensors(desc,
std::move(group_tensors),
residency_mode,
module_backends[0],
params_backend_for(module),
params_mem_size);
}
model->set_runtime_backends(module_backends);
const bool params_follow_runtime = backend_manager.params_backend_follows_runtime(module) ||
backend_manager.params_backend_is_disk(module);
for (size_t i = 0; i < module_backends.size(); i++) {
if (partitions[i].empty()) {
continue;
}
ggml_backend_t partition_params_backend =
params_follow_runtime ? module_backends[i] : params_backend_for(module);
if (partition_params_backend == nullptr) {
return false;
}
if (!model_manager->register_param_tensors(desc,
std::move(partitions[i]),
residency_mode,
module_backends[i],
partition_params_backend,
params_mem_size)) {
return false;
}
}
return true;
}
bool init_backend() {
std::string error;
if (!backend_manager.init(backend_spec.c_str(),
params_backend_spec.c_str(),
split_mode_spec.c_str(),
&error)) {
LOG_ERROR("backend config failed: %s", error.c_str());
return false;
@@ -291,6 +519,16 @@ public:
return ensure_backend_pair(SDBackendModule::DIFFUSION);
}
bool row_split_active() {
for (SDBackendModule module : {SDBackendModule::DIFFUSION, SDBackendModule::TE}) {
if (backend_manager.split_mode(module) == SDSplitMode::ROW &&
backend_manager.runtime_backends(module).size() > 1) {
return true;
}
}
return false;
}
std::shared_ptr<RNG> get_rng(rng_type_t rng_type) {
if (rng_type == STD_DEFAULT_RNG) {
return std::make_shared<STDDefaultRNG>();
@@ -349,6 +587,8 @@ public:
eager_load = sd_ctx_params->eager_load;
backend_spec = SAFE_STR(sd_ctx_params->backend);
params_backend_spec = SAFE_STR(sd_ctx_params->params_backend);
split_mode_spec = SAFE_STR(sd_ctx_params->split_mode);
auto_fit_enabled = sd_ctx_params->auto_fit;
max_vram_assignment.reset(0.f);
{
std::string error;
@@ -374,21 +614,6 @@ public:
ggml_log_set(ggml_log_callback_default, nullptr);
if (!init_backend()) {
return false;
}
{
std::string error;
if (!max_vram_assignment.canonicalize_backend_keys(&error)) {
LOG_ERROR("%s", error.c_str());
return false;
}
}
if (stream_layers && !backend_manager.params_backend_is_cpu(SDBackendModule::DIFFUSION)) {
LOG_WARN("--stream-layers has no effect unless diffusion params backend is cpu; ignoring");
stream_layers = false;
}
model_manager = std::make_shared<ModelManager>();
model_manager->set_n_threads(n_threads);
model_manager->set_enable_mmap(enable_mmap);
@@ -536,6 +761,31 @@ public:
model_loader.set_wtype_override(wtype, tensor_type_rules);
}
if (auto_fit_enabled) {
if (!sd::backend_fit::derive_backend_specs(model_loader,
wtype,
max_vram_assignment,
backend_spec,
params_backend_spec)) {
return false;
}
}
if (!init_backend()) {
return false;
}
{
std::string error;
if (!max_vram_assignment.canonicalize_backend_keys(&error)) {
LOG_ERROR("%s", error.c_str());
return false;
}
}
if (stream_layers && !backend_manager.params_backend_is_cpu(SDBackendModule::DIFFUSION)) {
LOG_WARN("--stream-layers has no effect unless diffusion params backend is cpu; ignoring");
stream_layers = false;
}
std::map<ggml_type, uint32_t> wtype_stat = model_loader.get_wtype_stat();
std::map<ggml_type, uint32_t> conditioner_wtype_stat = model_loader.get_conditioner_wtype_stat();
std::map<ggml_type, uint32_t> diffusion_model_wtype_stat = model_loader.get_diffusion_model_wtype_stat();
@@ -577,12 +827,17 @@ public:
// Avoid full-model LoRA merge buffers on constrained setups.
const bool params_offloaded = params_backend_for(SDBackendModule::DIFFUSION) != backend_for(SDBackendModule::DIFFUSION);
const bool streaming_constrained = stream_layers || params_offloaded;
if (have_quantized_weight || streaming_constrained) {
if (have_quantized_weight || streaming_constrained || row_split_active()) {
apply_lora_immediately = false;
} else {
apply_lora_immediately = true;
}
} else if (sd_ctx_params->lora_apply_mode == LORA_APPLY_IMMEDIATELY) {
if (row_split_active()) {
LOG_WARN(
"row-split tensors do not support the immediately LoRA apply mode; "
"LoRAs will not be applied to them (use --lora-apply-mode at_runtime)");
}
apply_lora_immediately = true;
} else {
apply_lora_immediately = false;
@@ -752,13 +1007,14 @@ public:
}
}
} else if (sd_version_is_qwen_image(version)) {
cond_stage_model = std::make_shared<LLMEmbedder>(backend_for(SDBackendModule::TE),
bool enable_vision = version != VERSION_QWEN_IMAGE_LAYERED;
cond_stage_model = std::make_shared<LLMEmbedder>(backend_for(SDBackendModule::TE),
tensor_storage_map,
version,
"",
true,
enable_vision,
model_manager);
diffusion_model = std::make_shared<Qwen::QwenImageRunner>(backend_for(SDBackendModule::DIFFUSION),
diffusion_model = std::make_shared<Qwen::QwenImageRunner>(backend_for(SDBackendModule::DIFFUSION),
tensor_storage_map,
"model.diffusion_model",
version,
@@ -785,6 +1041,14 @@ public:
tensor_storage_map,
"model",
model_manager);
} else if (sd_version_is_minit2i(version)) {
cond_stage_model = std::make_shared<MiniT2IConditioner>(backend_for(SDBackendModule::TE),
tensor_storage_map,
model_manager);
diffusion_model = std::make_shared<MiniT2I::MiniT2IRunner>(backend_for(SDBackendModule::DIFFUSION),
tensor_storage_map,
"model.diffusion_model.model.net",
model_manager);
} else if (sd_version_is_anima(version)) {
cond_stage_model = std::make_shared<AnimaConditioner>(backend_for(SDBackendModule::TE),
tensor_storage_map,
@@ -958,7 +1222,7 @@ public:
}
};
if (version == VERSION_CHROMA_RADIANCE || version == VERSION_HIDREAM_O1) {
if (version == VERSION_CHROMA_RADIANCE || version == VERSION_HIDREAM_O1 || sd_version_is_minit2i(version)) {
LOG_INFO("using FakeVAE");
first_stage_model = std::make_shared<FakeVAE>(version,
backend_for(SDBackendModule::VAE),
@@ -1299,6 +1563,8 @@ public:
}
} else if (sd_version_is_sefi_image(version)) {
pred_type = SEFI_FLOW_PRED;
} else if (sd_version_is_minit2i(version)) {
pred_type = MINIT2I_FLOW_PRED;
} else {
pred_type = EPS_PRED;
}
@@ -1336,6 +1602,11 @@ public:
denoiser = std::make_shared<SefiFlowDenoiser>();
break;
}
case MINIT2I_FLOW_PRED: {
LOG_INFO("running in MiniT2I FLOW mode");
denoiser = std::make_shared<MiniT2IFlowDenoiser>();
break;
}
default: {
LOG_ERROR("Unknown predition type %i", pred_type);
return false;
@@ -2032,12 +2303,13 @@ public:
}
int64_t last_progress_us = ggml_time_us();
sd::Tensor<float> x_t = !noise.empty()
? denoiser->noise_scaling(sigmas[0], noise, init_latent)
: init_latent;
sd::Tensor<float> denoised = x_t;
SamplePreviewContext preview = prepare_sample_preview_context();
sd::Tensor<float> x_t = !noise.empty()
? denoiser->noise_scaling(sigmas[0], noise, init_latent)
: init_latent;
sd::Tensor<float> denoised = x_t;
auto denoise = [&](const sd::Tensor<float>& x, float sigma, int step) -> sd::guidance::GuiderOutput {
if (get_cancel_flag() == SD_CANCEL_ALL) {
LOG_DEBUG("cancelling generation");
@@ -2100,9 +2372,9 @@ public:
sd_sample::SampleStepCacheDispatcher step_cache(cache_runtime, step, sigma);
std::vector<sd::Tensor<float>> controls;
DiffusionParams diffusion_params;
diffusion_params.x = &noised_input;
diffusion_params.timesteps = &timesteps_tensor;
diffusion_params.increase_ref_index = increase_ref_index;
diffusion_params.x = &noised_input;
diffusion_params.timesteps = &timesteps_tensor;
diffusion_params.ref_index_mode = Rope::ref_index_mode_from_bool(increase_ref_index);
sd::guidance::GuidanceInput step_guidance_input;
step_guidance_input.step = step;
step_guidance_input.schedule_size = sigmas.size();
@@ -2155,6 +2427,9 @@ public:
audio_length,
frame_rate,
video_positions.empty() ? nullptr : &video_positions};
} else if (sd_version_is_minit2i(version)) {
diffusion_params.extra = MiniT2IDiffusionExtra{
condition.c_vector.empty() ? nullptr : &condition.c_vector};
} else {
diffusion_params.extra = std::monostate{};
}
@@ -2335,6 +2610,8 @@ public:
latent_channel = 3;
} else if (version == VERSION_CHROMA_RADIANCE) {
latent_channel = 3;
} else if (sd_version_is_minit2i(version)) {
latent_channel = 3;
} else if (sd_version_is_pid(version)) {
latent_channel = 3;
} else if (sd_version_is_sefi_image(version)) {
@@ -2348,6 +2625,10 @@ public:
return latent_channel;
}
int get_image_channels() const {
return version == VERSION_QWEN_IMAGE_LAYERED ? 4 : 3;
}
int get_image_seq_len(int h, int w) {
int vae_scale_factor = get_vae_scale_factor();
return (h / vae_scale_factor) * (w / vae_scale_factor);
@@ -2416,12 +2697,21 @@ public:
}
sd::Tensor<float> decode_first_stage(const sd::Tensor<float>& x, bool decode_video = false) {
if (sd_version_is_pid(version)) {
if (sd_version_is_pid(version) || sd_version_is_minit2i(version)) {
return sd::ops::clamp((x + 1.f) * 0.5f, 0.0f, 1.0f);
}
auto latents = first_stage_model->diffusion_to_vae_latents(x);
first_stage_model->set_temporal_tiling_enabled(vae_tiling_params.temporal_tiling);
return first_stage_model->decode(n_threads, latents, vae_tiling_params, decode_video, circular_x, circular_y);
auto decoded = first_stage_model->decode(n_threads, latents, vae_tiling_params, decode_video, circular_x, circular_y);
if (decoded.empty() && auto_fit_enabled) {
bool prefer_temporal_tiling = decode_video && std::dynamic_pointer_cast<LTXVideoVAE>(first_stage_model) != nullptr;
if (sd::backend_fit::prepare_vae_decode_retry_tiling(vae_tiling_params, prefer_temporal_tiling)) {
first_stage_model->free_compute_buffer();
first_stage_model->set_temporal_tiling_enabled(vae_tiling_params.temporal_tiling);
decoded = first_stage_model->decode(n_threads, latents, vae_tiling_params, decode_video, circular_x, circular_y);
}
}
return decoded;
}
sd::Tensor<float> normalize_ltx_video_latents(const sd::Tensor<float>& x) {
@@ -2562,6 +2852,7 @@ const char* scheduler_to_str[] = {
"logit_normal",
"flux2",
"flux",
"beta",
};
const char* sd_scheduler_name(enum scheduler_t scheduler) {
@@ -2590,6 +2881,7 @@ const char* prediction_to_str[] = {
"sd3_flow",
"flux_flow",
"sefi_flow",
"minit2i_flow",
};
const char* sd_prediction_name(enum prediction_t prediction) {
@@ -2775,6 +3067,8 @@ void sd_ctx_params_init(sd_ctx_params_t* sd_ctx_params) {
sd_ctx_params->vae_format = SD_VAE_FORMAT_AUTO;
sd_ctx_params->backend = nullptr;
sd_ctx_params->params_backend = nullptr;
sd_ctx_params->split_mode = nullptr;
sd_ctx_params->auto_fit = false;
sd_ctx_params->rpc_servers = nullptr;
sd_ctx_params->pulid_weights_path = nullptr;
}
@@ -2814,6 +3108,8 @@ char* sd_ctx_params_to_str(const sd_ctx_params_t* sd_ctx_params) {
"eager_load: %s\n"
"backend: %s\n"
"params_backend: %s\n"
"split_mode: %s\n"
"auto_fit: %s\n"
"flash_attn: %s\n"
"diffusion_flash_attn: %s\n"
"circular_x: %s\n"
@@ -2850,6 +3146,8 @@ char* sd_ctx_params_to_str(const sd_ctx_params_t* sd_ctx_params) {
BOOL_STR(sd_ctx_params->eager_load),
SAFE_STR(sd_ctx_params->backend),
SAFE_STR(sd_ctx_params->params_backend),
SAFE_STR(sd_ctx_params->split_mode),
BOOL_STR(sd_ctx_params->auto_fit),
BOOL_STR(sd_ctx_params->flash_attn),
BOOL_STR(sd_ctx_params->diffusion_flash_attn),
BOOL_STR(sd_ctx_params->circular_x),
@@ -2933,6 +3231,7 @@ void sd_img_gen_params_init(sd_img_gen_params_t* sd_img_gen_params) {
sd_img_gen_params->seed = -1;
sd_img_gen_params->batch_count = 1;
sd_img_gen_params->control_strength = 0.9f;
sd_img_gen_params->qwen_image_layers = 3;
sd_img_gen_params->pm_params = {nullptr, 0, nullptr, 20.f};
sd_img_gen_params->pulid_params = {nullptr, 1.0f};
sd_img_gen_params->vae_tiling_params = {false, false, 0, 0, 0.5f, 0.0f, 0.0f, nullptr};
@@ -2959,6 +3258,7 @@ char* sd_img_gen_params_to_str(const sd_img_gen_params_t* sd_img_gen_params) {
"seed: %" PRId64
"\n"
"batch_count: %d\n"
"qwen_image_layers: %d\n"
"ref_images_count: %d\n"
"auto_resize_ref_image: %s\n"
"increase_ref_index: %s\n"
@@ -2975,6 +3275,7 @@ char* sd_img_gen_params_to_str(const sd_img_gen_params_t* sd_img_gen_params) {
sd_img_gen_params->strength,
sd_img_gen_params->seed,
sd_img_gen_params->batch_count,
sd_img_gen_params->qwen_image_layers,
sd_img_gen_params->ref_images_count,
BOOL_STR(sd_img_gen_params->auto_resize_ref_image),
BOOL_STR(sd_img_gen_params->increase_ref_index),
@@ -3242,6 +3543,7 @@ struct GenerationRequest {
bool has_ref_images = false;
const sd_cache_params_t* cache_params = nullptr;
int batch_count = 1;
int qwen_image_layers = 3;
int shifted_timestep = 0;
float strength = 1.f;
float control_strength = 0.f;
@@ -3267,6 +3569,7 @@ struct GenerationRequest {
diffusion_model_down_factor = sd_ctx->sd->get_diffusion_model_down_factor();
seed = sd_img_gen_params->seed;
batch_count = sd_img_gen_params->batch_count;
qwen_image_layers = std::max(0, sd_img_gen_params->qwen_image_layers);
clip_skip = sd_img_gen_params->clip_skip;
shifted_timestep = sd_img_gen_params->sample_params.shifted_timestep;
strength = sd_img_gen_params->strength;
@@ -3971,6 +4274,34 @@ public:
ImageVaeAxesGuard& operator=(const ImageVaeAxesGuard&) = delete;
};
static sd::Tensor<float> ensure_image_tensor_channels(sd::Tensor<float> image, int channels) {
if (image.empty()) {
return image;
}
GGML_ASSERT(image.dim() == 4);
int64_t current_channels = image.shape()[2];
if (current_channels == channels) {
return image;
}
if (channels == 4) {
sd::Tensor<float> alpha = sd::full<float>({image.shape()[0], image.shape()[1], 1, image.shape()[3]}, 1.f);
if (current_channels == 3) {
return sd::ops::concat(image, alpha, 2);
}
if (current_channels == 1) {
sd::Tensor<float> rgb = sd::ops::concat(image, image, 2);
rgb = sd::ops::concat(rgb, image, 2);
return sd::ops::concat(rgb, alpha, 2);
}
}
if (channels == 3 && current_channels >= 3) {
return sd::ops::slice(image, 2, 0, 3);
}
GGML_ABORT("cannot convert image tensor from %lld to %d channels",
(long long)current_channels,
channels);
}
static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd_ctx_t* sd_ctx,
const sd_img_gen_params_t* sd_img_gen_params,
GenerationRequest* request,
@@ -3980,6 +4311,7 @@ static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd
sd::Tensor<float> init_image_tensor;
sd::Tensor<float> control_image_tensor;
sd::Tensor<float> mask_image_tensor;
int image_channels = sd_ctx->sd->get_image_channels();
if (sd_img_gen_params->init_image.data != nullptr) {
LOG_INFO("IMG2IMG");
@@ -3996,7 +4328,8 @@ static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd
plan->sample_steps = static_cast<int>(plan->sigmas.size() - 1);
}
init_image_tensor = sd_image_to_tensor(sd_img_gen_params->init_image, request->width, request->height);
init_image_tensor = ensure_image_tensor_channels(sd_image_to_tensor(sd_img_gen_params->init_image, request->width, request->height),
image_channels);
}
if (sd_img_gen_params->mask_image.data != nullptr) {
@@ -4028,7 +4361,11 @@ static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd
sd::Tensor<float> init_latent;
sd::Tensor<float> control_latent;
if (init_image_tensor.empty()) {
init_latent = sd_ctx->sd->generate_init_latent(request->width, request->height);
if (sd_ctx->sd->version == VERSION_QWEN_IMAGE_LAYERED) {
init_latent = sd_ctx->sd->generate_init_latent(request->width, request->height, request->qwen_image_layers + 1, true);
} else {
init_latent = sd_ctx->sd->generate_init_latent(request->width, request->height);
}
} else {
init_latent = sd_ctx->sd->encode_first_stage(init_image_tensor);
if (init_latent.empty()) {
@@ -4047,12 +4384,13 @@ static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd
std::vector<sd::Tensor<float>> ref_images;
for (int i = 0; i < sd_img_gen_params->ref_images_count; i++) {
ref_images.push_back(sd_image_to_tensor(sd_img_gen_params->ref_images[i]));
ref_images.push_back(ensure_image_tensor_channels(sd_image_to_tensor(sd_img_gen_params->ref_images[i]),
image_channels));
}
if (ref_images.empty() && sd_version_is_unet_edit(sd_ctx->sd->version)) {
LOG_WARN("This model needs at least one reference image; using an empty reference");
ref_images.push_back(sd::zeros<float>({request->width, request->height, 3, 1}));
ref_images.push_back(sd::zeros<float>({request->width, request->height, image_channels, 1}));
request->guidance.img_cfg = request->guidance.txt_cfg;
request->use_img_uncond = false;
}
@@ -4078,7 +4416,10 @@ static std::optional<ImageGenerationLatents> prepare_image_generation_latents(sd
vae_width = round(vae_width / factor) * factor;
auto resized_ref_img = sd::ops::interpolate(ref_images[i],
{static_cast<int>(vae_width), static_cast<int>(vae_height), 3, 1});
{static_cast<int>(vae_width),
static_cast<int>(vae_height),
ref_images[i].shape()[2],
ref_images[i].shape()[3]});
LOG_DEBUG("resize vae ref image %d from %" PRId64 "x%" PRId64 " to %" PRId64 "x%" PRId64,
static_cast<int>(i),
@@ -4223,6 +4564,11 @@ static std::optional<ImageGenerationEmbeds> prepare_image_generation_embeds(sd_c
if (request->use_uncond || request->use_high_noise_uncond) {
if (sd_version_is_ideogram4(sd_ctx->sd->version)) {
uncond.c_vector = sd::Tensor<float>::from_vector({1.0f});
} else if (sd_version_is_minit2i(sd_ctx->sd->version)) {
// MiniT2I derives the unconditional signal from the same T5 hidden
// states with a zeroed prompt mask, so no extra text encode is needed.
uncond.c_crossattn = cond.c_crossattn;
uncond.c_vector = sd::Tensor<float>::zeros_like(cond.c_vector);
} else {
bool zero_out_masked = false;
if (sd_version_is_sdxl(sd_ctx->sd->version) &&
@@ -4278,7 +4624,8 @@ static std::optional<ImageGenerationEmbeds> prepare_image_generation_embeds(sd_c
static sd_image_t* decode_image_outputs(sd_ctx_t* sd_ctx,
const GenerationRequest& request,
const std::vector<sd::Tensor<float>>& final_latents) {
const std::vector<sd::Tensor<float>>& final_latents,
int* num_images_out) {
if (final_latents.empty()) {
LOG_ERROR("no latent images to decode");
return nullptr;
@@ -4302,13 +4649,41 @@ static sd_image_t* decode_image_outputs(sd_ctx_t* sd_ctx,
cancelled = true;
break;
}
int64_t t1 = ggml_time_ms();
sd::Tensor<float> image = sd_ctx->sd->decode_first_stage(final_latents[i]);
if (image.empty()) {
LOG_ERROR("decode_first_stage failed for latent %" PRId64, i + 1);
return nullptr;
int64_t t1 = ggml_time_ms();
if (sd_ctx->sd->version == VERSION_QWEN_IMAGE_LAYERED) {
int qwen_image_latent_layers = request.qwen_image_layers + 1;
if (final_latents[i].dim() < 5 || final_latents[i].shape()[2] < qwen_image_latent_layers) {
LOG_ERROR("qwen image layered expected at least %d latent layers, got shape dim=%d",
qwen_image_latent_layers,
final_latents[i].dim());
return nullptr;
}
for (int layer_index = 0; layer_index < qwen_image_latent_layers; layer_index++) {
if (sd_ctx->sd->get_cancel_flag() == SD_CANCEL_ALL) {
LOG_ERROR("cancelling latent decodings");
cancelled = true;
break;
}
sd::Tensor<float> layer_latent = sd::ops::slice(final_latents[i], 2, layer_index, layer_index + 1);
layer_latent.squeeze_(2);
sd::Tensor<float> image = sd_ctx->sd->decode_first_stage(layer_latent);
if (image.empty()) {
LOG_ERROR("decode_first_stage failed for latent %zu layer %d", i + 1, layer_index + 1);
return nullptr;
}
decoded_images.push_back(std::move(image));
}
if (cancelled) {
break;
}
} else {
sd::Tensor<float> image = sd_ctx->sd->decode_first_stage(final_latents[i]);
if (image.empty()) {
LOG_ERROR("decode_first_stage failed for latent %" PRId64, i + 1);
return nullptr;
}
decoded_images.push_back(std::move(image));
}
decoded_images.push_back(std::move(image));
int64_t t2 = ggml_time_ms();
LOG_INFO("latent %zu decoded, taking %.2fs", i + 1, (t2 - t1) * 1.0f / 1000);
}
@@ -4320,11 +4695,14 @@ static sd_image_t* decode_image_outputs(sd_ctx_t* sd_ctx,
return nullptr;
}
sd_image_t* result_images = (sd_image_t*)calloc(request.batch_count, sizeof(sd_image_t));
int image_count = static_cast<int>(decoded_images.size());
sd_image_t* result_images = (sd_image_t*)calloc(image_count, sizeof(sd_image_t));
if (result_images == nullptr) {
return nullptr;
}
memset(result_images, 0, request.batch_count * sizeof(sd_image_t));
if (num_images_out != nullptr) {
*num_images_out = image_count;
}
for (size_t i = 0; i < decoded_images.size(); i++) {
result_images[i] = tensor_to_sd_image(decoded_images[i]);
@@ -4517,9 +4895,18 @@ static std::vector<float> make_hires_sigma_schedule(sd_ctx_t* sd_ctx,
sigmas.end());
}
SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* sd_img_gen_params) {
SD_API bool generate_image(sd_ctx_t* sd_ctx,
const sd_img_gen_params_t* sd_img_gen_params,
sd_image_t** images_out,
int* num_images_out) {
if (images_out != nullptr) {
*images_out = nullptr;
}
if (num_images_out != nullptr) {
*num_images_out = 0;
}
if (sd_ctx == nullptr || sd_img_gen_params == nullptr) {
return nullptr;
return false;
}
sd_ctx->sd->reset_cancel_flag();
@@ -4542,7 +4929,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
&request,
&plan);
if (!latents_opt.has_value()) {
return nullptr;
return false;
}
ImageGenerationLatents latents = std::move(*latents_opt);
@@ -4552,7 +4939,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
&plan,
&latents);
if (!embeds_opt.has_value()) {
return nullptr;
return false;
}
ImageGenerationEmbeds embeds = std::move(*embeds_opt);
@@ -4562,7 +4949,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
sd_cancel_mode_t cancel = sd_ctx->sd->get_cancel_flag();
if (cancel == SD_CANCEL_ALL) {
LOG_ERROR("cancelling generation");
return nullptr;
return false;
}
if (cancel == SD_CANCEL_NEW_LATENTS) {
LOG_INFO("cancelling new latent generation, returning %zu/%d completed latents",
@@ -4614,7 +5001,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
b + 1,
request.batch_count,
(sampling_end - sampling_start) * 1.0f / 1000);
return nullptr;
return false;
}
int64_t denoise_end = ggml_time_ms();
LOG_INFO("generating %zu latent images completed, taking %.2fs",
@@ -4622,13 +5009,13 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
(denoise_end - denoise_start) * 1.0f / 1000);
if (final_latents.empty()) {
LOG_ERROR("no latent images generated");
return nullptr;
return false;
}
if (request.hires.enabled && request.hires.target_width > 0) {
if (sd_ctx->sd->get_cancel_flag() == SD_CANCEL_ALL) {
LOG_ERROR("cancelling generation before hires fix");
return nullptr;
return false;
}
LOG_INFO("hires fix: upscaling to %dx%d", request.hires.target_width, request.hires.target_height);
@@ -4636,7 +5023,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
if (request.hires.upscaler == SD_HIRES_UPSCALER_MODEL) {
if (sd_ctx->sd->get_cancel_flag() == SD_CANCEL_ALL) {
LOG_ERROR("cancelling generation before hires model load");
return nullptr;
return false;
}
LOG_INFO("hires fix: loading model upscaler from '%s'", request.hires.model_path);
hires_upscaler = std::make_unique<UpscalerGGML>(sd_ctx->sd->n_threads,
@@ -4649,7 +5036,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
if (!hires_upscaler->load_from_file(request.hires.model_path,
sd_ctx->sd->n_threads)) {
LOG_ERROR("load hires model upscaler failed");
return nullptr;
return false;
}
}
@@ -4673,7 +5060,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
for (int b = 0; b < (int)final_latents.size(); b++) {
if (sd_ctx->sd->get_cancel_flag() == SD_CANCEL_ALL) {
LOG_ERROR("cancelling generation during hires fix");
return nullptr;
return false;
}
int64_t cur_seed = request.seed + b;
sd_ctx->sd->rng->manual_seed(cur_seed);
@@ -4684,7 +5071,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
request,
hires_upscaler.get());
if (upscaled.empty()) {
return nullptr;
return false;
}
sd::Tensor<float> noise = sd::randn_like<float>(upscaled, sd_ctx->sd->rng);
@@ -4738,7 +5125,7 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
b + 1,
(int)final_latents.size(),
(hires_sample_end - hires_sample_start) * 1.0f / 1000);
return nullptr;
return false;
}
int64_t hires_denoise_end = ggml_time_ms();
LOG_INFO("hires fix completed, taking %.2fs", (hires_denoise_end - hires_denoise_start) * 1.0f / 1000);
@@ -4746,16 +5133,25 @@ SD_API sd_image_t* generate_image(sd_ctx_t* sd_ctx, const sd_img_gen_params_t* s
final_latents = std::move(hires_final_latents);
}
auto result = decode_image_outputs(sd_ctx, request, final_latents);
int num_images = 0;
auto result = decode_image_outputs(sd_ctx, request, final_latents, &num_images);
if (result == nullptr) {
return nullptr;
return false;
}
sd_ctx->sd->lora_stat();
int64_t t1 = ggml_time_ms();
LOG_INFO("generate_image completed in %.2fs", (t1 - t0) * 1.0f / 1000);
return result;
if (num_images_out != nullptr) {
*num_images_out = num_images;
}
if (images_out != nullptr) {
*images_out = result;
} else {
free_sd_images(result, num_images);
}
return true;
}
static std::optional<ImageGenerationLatents> prepare_video_generation_latents(sd_ctx_t* sd_ctx,
+37 -2
View File
@@ -4,6 +4,7 @@
#include "model_loader.h"
#include "stable-diffusion.h"
#include <cstdlib>
#include <utility>
UpscalerGGML::UpscalerGGML(int n_threads,
@@ -45,6 +46,7 @@ bool UpscalerGGML::load_from_file(const std::string& esrgan_path,
std::string error;
if (!backend_manager.init(backend_spec.c_str(),
params_backend_spec.c_str(),
/*split_mode_spec=*/nullptr,
&error)) {
LOG_ERROR("upscaler backend config failed: %s", error.c_str());
return false;
@@ -198,8 +200,41 @@ upscaler_ctx_t* new_upscaler_ctx(const char* esrgan_path_c_str,
return upscaler_ctx;
}
sd_image_t upscale(upscaler_ctx_t* upscaler_ctx, sd_image_t input_image, uint32_t upscale_factor) {
return upscaler_ctx->upscaler->upscale(input_image, upscale_factor);
bool upscale(upscaler_ctx_t* upscaler_ctx,
sd_image_t input_image,
uint32_t upscale_factor,
sd_image_t** images_out,
int* num_images_out) {
if (images_out != nullptr) {
*images_out = nullptr;
}
if (num_images_out != nullptr) {
*num_images_out = 0;
}
if (upscaler_ctx == nullptr || upscaler_ctx->upscaler == nullptr) {
return false;
}
sd_image_t* result_images = (sd_image_t*)calloc(1, sizeof(sd_image_t));
if (result_images == nullptr) {
return false;
}
result_images[0] = upscaler_ctx->upscaler->upscale(input_image, upscale_factor);
if (result_images[0].data == nullptr) {
free(result_images);
return false;
}
if (num_images_out != nullptr) {
*num_images_out = 1;
}
if (images_out != nullptr) {
*images_out = result_images;
} else {
free_sd_images(result_images, 1);
}
return true;
}
int get_upscale_factor(upscaler_ctx_t* upscaler_ctx) {