fix(ci): install test dependencies before training pins (#3580)

This commit is contained in:
Yueming Yuan
2026-09-21 17:08:17 -07:00
committed by GitHub
parent 29a8aa5b1f
commit a66064f190
6 changed files with 7 additions and 12 deletions
+1
View File
@@ -166,6 +166,7 @@ jobs:
v=$(python -m pip show "$pkg" 2>/dev/null | awk '/^Version:/{print $2}' || true)
if [ -n "$v" ]; then pinned="$pinned $pkg==$v"; fi
done
python -m pip install -r "$GITHUB_WORKSPACE/examples/multi_lora/requirements.txt" --break-system-packages
python -m pip install -r "$GITHUB_WORKSPACE/requirements.txt" --break-system-packages
for spec in $pinned; do
pkg=${spec%%==*}
+1 -1
View File
@@ -72,7 +72,7 @@ A test normally runs at its home stage. A [dispatch label](/developer/ci/01-labe
Without a dispatch label nothing leaves its home stage, so scheduled and called runs are unaffected. Whatever stage actually executes a run is also the stage its performance baseline is keyed on, keeping the generations' numbers apart.
**Dependency boundary.** CUDA stages start from dependencies baked into `radixark/miles`, reconcile Miles runtime dependencies from `requirements.txt`, update the SGLang and Megatron-LM checkouts to the selected refs, and expose all three source trees through `PYTHONPATH`; they do not rebuild or install the Miles, SGLang, or Megatron-LM source trees after the container starts. The hosted CPU stages install dependencies from `requirements.txt` and the fully pinned `tests/ci/requirements-ci-cpu.txt`, then expose the Miles, SGLang, and Megatron-LM source trees through `PYTHONPATH` without editable installs or inline package lists. The ROCm stage instead uses the SGLang and Megatron-LM versions baked into `rocm/sgl-dev`, unless the run names a ref for either.
**Dependency boundary.** CUDA stages start from dependencies baked into `radixark/miles`, install CUDA test dependencies from `examples/multi_lora/requirements.txt`, override them with Miles runtime dependencies from `requirements.txt`, restore the image’s cuDNN pins, update the SGLang and Megatron-LM checkouts to the selected refs, and expose all three source trees through `PYTHONPATH`; they do not rebuild or install the Miles, SGLang, or Megatron-LM source trees after the container starts. The hosted CPU stages install dependencies from `requirements.txt` and the fully pinned `tests/ci/requirements-ci-cpu.txt`, then expose the Miles, SGLang, and Megatron-LM source trees through `PYTHONPATH` without editable installs or inline package lists. The ROCm stage instead uses the SGLang and Megatron-LM versions baked into `rocm/sgl-dev`, unless the run names a ref for either.
CUDA and CPU dependency refs resolve in this order: explicit dispatch input or PR-body directive, committed `release-lock.json`, then the moving `sglang-miles` / `miles-main` branch heads. A called release run therefore checks out its requested Miles `ref` and consumes the lockfile on that ref unless an explicit override exists. ROCm checks out the requested Miles ref but keeps the dependencies baked into its image unless the run names a ref for one.
+2
View File
@@ -0,0 +1,2 @@
tinker==0.26.2
tinker_cookbook[math-rl] @ git+https://github.com/thinking-machines-lab/tinker-cookbook@1f962eda3a2c
+2 -2
View File
@@ -3,9 +3,9 @@
The cookbook's sl_loop (SFT) and rl_loop (GRPO) are the executable definition
of the Tinker wire contract; passing them is the gateway's acceptance bar.
Requires, next to the pinned SDK (tinker==0.26.2):
Install the client dependencies:
pip install git+https://github.com/thinking-machines-lab/tinker-cookbook@1f962eda3a2c
pip install -r examples/multi_lora/requirements.txt
``--base-model`` must be both the name this gateway serves (--tinker-base-model)
and a HuggingFace name the cookbook can resolve a tokenizer and renderer for.
@@ -12,13 +12,6 @@ register_cuda_ci(
hardware=["hopper"],
)
COOKBOOK_PIN = "git+https://github.com/thinking-machines-lab/tinker-cookbook@1f962eda3a2c"
def prepare():
prepare_gateway()
U.exec_command_cpu(f"pip install tinker==0.26.2 {COOKBOOK_PIN}")
def execute():
with running_gateway() as base_url:
@@ -29,5 +22,5 @@ def execute():
if __name__ == "__main__":
prepare()
prepare_gateway()
execute()
-1
View File
@@ -18,7 +18,6 @@ SERVE_TIMEOUT_S = 1200
def prepare_gateway():
U.exec_command_cpu("mkdir -p /root/models")
U.exec_command_cpu(f"hf download {BASE_MODEL} --local-dir /root/models/{MODEL_NAME}")
U.exec_command_cpu("pip install tinker==0.26.2")
def _wait_for_gateway(server: subprocess.Popen) -> None: