Multi-LoRA Tinker Gateway
Read the docs: Multi-LoRA training.
serve_qwen3_30b_a3b_tinker.py: prepare Qwen3-30B-A3B and launch the gateway.run_multi_tenant_example.py: check marker memorization for one client or adapter isolation across concurrent tenants.
Layout
One 8-GPU node, disaggregated (multi-LoRA forbids --colocate):
- 4 training GPUs: TP2 for the dense layers, EP4 for the 128 routed experts.
- 4 sampling GPUs: two SGLang engines of 2 GPUs each, serving adapter versions by name.
- 4 adapter slots (
--multi-lora-n-adapters), rank up to 32, covering attention (linear_qkv,linear_proj), the per-expert MoE projections (linear_fc1,linear_fc2), and the output layer (output_layer) so the cookbook's defaulttrain_unembed=Trueis servable.
The example enables attention, MLP, and output-head training. Client SDK flags
must match the server's selected groups. See LoRA target selection
for --target-modules attn,mlp,unembed; use attn,mlp to disable output-head
training. Tinker accepts only these group names and rejects --exclude-modules.
Run
The gateway implements the tinker==0.26.2 wire schema (newer SDKs renamed protobuf fields); install that exact version on the serving node and the client:
pip install "tinker==0.26.2"
Start the gateway:
python examples/multi_lora/serve_qwen3_30b_a3b_tinker.py prepare # once per node
python examples/multi_lora/serve_qwen3_30b_a3b_tinker.py serve # Tinker API on :10613
Checkpoints default to <output_dir>/checkpoints/<run_id>; use --save-dir to choose another root.
Install tinker on the client, then run the marker checks:
# one client: train, save for sampler, sample back the marker
python examples/multi_lora/run_multi_tenant_example.py --base-model /root/models/Qwen3-30B-A3B --mode single
# four tenants training concurrently on the same prompt with different markers;
# passing means the adapters stayed isolated end to end
python examples/multi_lora/run_multi_tenant_example.py --base-model /root/models/Qwen3-30B-A3B --mode multi --clients 4
Supported inputs
Training accepts text with 1-D loss inputs. 2-D soft targets, including SDFT,
are not supported. Sampling requires a /sampler_weights/ path returned by
save_weights_for_sampler(); /weights/ training checkpoints cannot be sampled directly.
Failure handling
A terminal failure of forward_backward, optim_step, or load_state ends
training for that model, including commands already queued behind it.
This includes content validation failures with a valid model and sequence.
Create a new model and restore a saved checkpoint to continue; completed
futures and published checkpoints keep their results.
Known request-local failures of forward or sampling leave model training
available. Checkpoint load/save execution failures, including filesystem errors,
invalidate the shared trainer cell and stop the server.
Saving sampler weights commits an immutable directory;
it does not call the inference engines. Sampling loads that snapshot from disk
on demand, including after cache eviction. An engine load failure fails the
sampling request; it leaves the snapshot and training state intact. Unknown
trainer execution failures invalidate the shared trainer cell and stop the server.
This gateway provides failure isolation, not automatic training recovery. Checkpoints persist; futures, deduplication, and unsaved accumulation do not survive a server restart.
Sampler snapshots
Training and inference must use the same base checkpoint. Tinker engines load
that frozen base at startup and serve without trainer weight updates; dummy
loading and update_weights: true are rejected. Ordinary full-model and
single-LoRA training continue to use the existing weight updater.
--tinker-checkpoint-root must be on storage shared by the trainers, gateway,
and every inference engine. All trainer ranks participate in adapter gathering;
rank 0 writes the tensors and config, then META.json after all ranks finish.
Existing sampler versions cannot be overwritten. Saving between forward_backward
and optim_step neither applies nor discards pending gradients.
Training checkpoint saves and loads are serialized within the gateway. Overwriting deletes the previous checkpoint before writing the new one; a failed overwrite does not preserve the previous checkpoint. Checkpoints are ordinary directories; overwriting does not retain hidden versions.