Commit Graph
2569 Commits
Author SHA1 Message Date
Tom Chen 3c2f9e5046 Refuse an optimizer a state reset cannot check before resetting it
Resolves op27-30 #1
Resolves op27-17 #3

A take-over that finds no checkpoint zeroes the scheduler and resets the
optimizer, and only then checks the state it rebuilt against the moments adam
keeps. A Muon member rebuilds a `momentum_buffer`, so the check refused it
after the trainer had already been mutated, and the surviving trainer was left
half changed. Refuse the optimizer among the other preconditions, before
anything is touched.
backup/20260902-200358-drop354/op35-354
2026-09-03 20:35:32 +08:00
Tom Chen 38c12e959f Answer with the rollout an FSDP trainer really restored
Resolves op27-27 #1

The FSDP actor answered `init` with `args.start_rollout_id`, which it defaults
to zero before loading and which `finalize_load` overwrites only from a
checkpoint's `meta.json`. A checkpoint with a missing or unreadable `meta.json`
therefore left an explicitly requested `--start-rollout-id` in place, and
`create_training_model` compared that request against itself and logged
nothing where the two positions differ. Carry the position the checkpoint
restored and answer with it.
backup/20260902-200358-drop353/op35-353
2026-09-03 20:35:32 +08:00
Tom Chen def5e71ef9 Refuse a reset a chunked optimizer state offload outlives
Resolves op27-24 #1
Resolves op27-31 #1

The reset guarded only the deprecated `offload_optimizer_states`, which MCore
normalizes to false once `--chunked-optimizer-state-offload` is the active
mode. The chunked offloader keeps the canonical moments in its own cpu store,
which is not the mapping the reset clears, so a trainer hot restarted before
its first checkpoint passed the self-check and then prefetched the moments of
the rollouts it had just discarded. Refuse the mode instead.
2026-09-03 20:35:32 +08:00
Tom Chen e201fb3354 Say why a multi-LoRA data source restores nothing
Resolves op27-22 #1

`create_training_models` runs before the adapters named on the command line
are registered, and it loads the rollout executor, so `sources` is still empty
and the run was reported as serving no adapter at all. That is not what
happened: the adapters are configured, and `reconcile()` has not filled the
sources yet. Say that instead.

Unit tests cover the message a run whose adapters are not reconciled yet is
given, and the restore of every source once they are.
backup/20260902-200358-drop351/op35-351
2026-09-03 20:35:32 +08:00
Tom Chen c43dab3073 Let every reload converge before a failed one is reported
Resolves op27-21 #1

`load_state` reloads the ranks of a cell, and the cells of a controller, with
a fail-fast `asyncio.gather`, and it does not kill a worker whose reload
failed. The first failure therefore returned while the other reloads were
still running their distributed checkpoint collectives, so the controller went
idle and an orchestration retry could interleave with them. Gather the reloads
with `gather_and_raise_first`, which waits for all of them and then raises the
first failure.
2026-09-03 20:35:32 +08:00
Tom Chen 807ad3bbc3 Put a take-over back on the random stream the run had reached
Resolves op27-17 #2

A take-over that finds no checkpoint reseeded from the arguments, which is the
stream the process stood on before `build_model_and_optimizer` ran. Parameter
initialization advances both the host generators and Megatron's CUDA RNG
tracker, so the surviving trainer resumed from a point the discarded run had
already consumed, and a replay could diverge silently. Snapshot the generators
once the model and the optimizer are built, and restore that snapshot instead.
2026-09-03 20:35:32 +08:00
Tom Chen 770668e9d2 Let a bridge run be taken over before it has saved anything
Resolves op27-17 #1

The no-checkpoint reload path demanded `--finetune`, `--no-load-optim` and
`--no-load-rng`, and a `ckpt_step` taken from `--ref-ckpt-step`. Only the
non-bridge branch of `resolve_args_checkpoint_load` sets those; the bridge
branch rewrites `load` and `start_rollout_id` alone. A non-LoRA bridge run
taken over before its first save therefore hit the assertion on its ordinary
cold-start path. Require the flags only where the cold start actually sets
them.
2026-09-03 20:35:32 +08:00
Tom Chen 36b28a5253 Authenticate every request the sglang api client posts
Resolves op27-13 #1

`SGLangApiClient` kept the engine's `api_key` but posted without an
`Authorization` header, so with `--sglang-api-key` or a model-level key every
`_make_request` endpoint answered 401. `abort_all_requests` goes through that
path, so a hot restart could not abort the generations still in flight and the
take-over failed. Send the same headers the client already computes for its
other calls.
2026-09-03 20:35:32 +08:00
Tom Chen c67ca5ee7c Refuse a hot restart request that names no component
Resolves op27-11 #1

`parse_hot_restart` dropped every blank token, so a non-empty request such as
`--hot-restart ,` parsed to an empty list. Every caller reads that list's
truthiness, so the malformed request became an ordinary launch: the backend
check that only Kubernetes serves a hot restart was skipped, and on Ray the
launch went on to `_clean_up_previous_run`, which kills the host's SGLang,
Miles and Ray processes. Reject a non-empty value that names no component.

Unit tests cover the separators and whitespace that name no component, the empty
value that is still an ordinary launch, and a trailing separator that is not.
2026-09-03 20:35:32 +08:00
Tom Chen ad0dd9dc24 Refuse a reload that would lose a periodically updated reference
Resolves op27-9 #2

`--ref-update-interval` backs the actor up as the new reference in memory, but
`save_model` persists only the actor, its optimizer and its scheduler, so a
reload reloads the reference from the static `--ref-load` instead. The run
would then compute KL against the reference it started from while the actor
stands at a rollout that was trained against a newer one. Reject the
combination up front, alongside the other reload preconditions.
2026-09-03 20:35:32 +08:00
Tom Chen 7218a61e74 Derive a trainer's reload directory after its own overrides
Resolves op27-9 #1

`compute_trainer_args` only re-resolved the checkpoint load when a
`--megatron-config` was given, so a critic synthesized from `--use-critic`
kept the `requested_load` that `validate_args` had derived from the actor's
`--load`. A hot restart reads `requested_load`, so the critic would look for
its state under the actor's checkpoint directory and either fail to load or
restore the wrong model and optimizer. Re-resolve whenever a trainer's
overrides moved `load` away from the value the base arguments were resolved
from.
2026-09-03 20:35:31 +08:00
Tom Chen a6026eed44 Take a worker over only once it says it is not inited
Resolves op27-7 #1.

The wait that lets a hot restart adopt a rollout executor asked
is_initialized(), which is only true in the INITED state. A process
still running the init of the previous script, or one whose init failed
and left it half-built, therefore answered false and was let through as
the replacement; the script then called init() on it and the init-once
guard refused, failing the hot restart.

The rollout executor now reports its init state, and the wait goes on
until that state is NOT_INITED, which is the only state a process the
script may initialize itself is in.
2026-09-03 20:35:31 +08:00
Tom Chen 97768f4d1d Stop the external agent when a hot restart takes the fleet over
Resolves op27-6 #1.

Taking an inference controller over aborted the generations the fleet
was running, but said nothing to the agent backend behind
--custom-agent-function-path. That loop keeps issuing fresh completion
requests after the script that started it is gone, so it takes the fleet
and the agent-server slots the new rollout executor is about to ask for.

The take-over now calls the same abort hook the ordinary abort path
calls, on the same bounded budget as the rest of the take-over.

Unit tests cover the hook called after the fleet abort, the cold start that
calls none, and the take-over budget it runs on.
2026-09-03 20:35:31 +08:00
Tom Chen b99316551e Read the checkpoint of every trainer a config declares
Resolves op27-5 #1.

Deciding whether a take-over starts a run over asserted that the run has
no --megatron-config, and read the tracker under the base --load. A
configured single-actor run is supported everywhere else on this path,
so a hot restart of one aborted right there, before the event log was
ever reset, and the base --load holds no tracker to read anyway once the
per-trainer checkpoint directories are derived from it.

The question is now asked of each declared trainer through the arguments
compute_trainer_args resolves for it, and the run counts as resumable if
any of them has a checkpoint to resume from.
2026-09-03 20:35:31 +08:00
Tom Chen e5617e3ddb Guard the whole of a trainer's init, not its common prefix
Resolves op27-2 #1.

The init-once guard sat on _init_common, which is only the first step of
what a concrete trainer does in init: the megatron and fsdp actors go on
to build parallel state, the model, tracking and the tokenizer after it
returns. An actor that threw in any of that was left reporting
is_initialized() true, so a recovery would take a half-built process for
one it can resume and drive it as a trainer that has no model.

init is now a concrete method on the base that guards the whole thing
and hands off to the _init each backend implements, so the state the
guard reports covers everything init does.
2026-09-03 20:35:31 +08:00
Tom Chen 7f9e6058d2 Renew a reporter only on a snapshot the hub took in
Resolves op25-10 #1.

The last ingest time was stamped before the ownership check, so a
snapshot rejected for claiming another reporter's cells still renewed
the sender's lease. A reporter that keeps sending such snapshots is
never dropped, and the cells it registered before, which may long be
gone, stay in the run's view of the fleet.

The stamp moves after the replacement, so only a snapshot that was
actually taken in renews the reporter.

Unit tests cover the lease of a reporter whose snapshot was refused, the
cells dropped with it, and the snapshot that does renew its sender.
2026-09-03 20:35:31 +08:00
Tom Chen 2e71088fdc Check a registered cell announces exactly the workers it carries
Resolves op25-8 #1.

The snapshot check read the worker names off cell.workers and compared
the cells they belong to as a set, so it never looked at the
CellInfo.worker_names the run is served from. A cell that announced an
empty list, a name no worker of it has, or the same name twice passed,
and the rollout reconciliation that indexes worker_names[0] and looks
addresses up by those names then failed on a snapshot the hub had
already taken in.

The two lists now have to be one duplicate-free mapping of each other.

Unit tests cover a cell announcing nothing, announcing a worker it does not
carry and carrying one twice, and the membership a refused snapshot leaves
alone.
2026-09-03 20:35:31 +08:00
Tom Chen 4e479f4c10 Follow the registration hub across a restart
Resolves op25-2 #2.

The reporter waits for the hub once, before its loop, and the handle
pins the boot uuid it saw there. When the inference controller process
that hosts the hub restarts, every later ingest raises
ServerRestartedError, the generic handler logs it and the loop retries
against the same pin forever, so the new hub never learns about the
cells this deployment reports and the run waits for engines that are
running.

A restarted hub is now awaited again with allow_server_uuid_change, which
is the only way the handle drops the pin, and reporting resumes against
the process that answers now.

Unit tests cover the readiness awaited again with the pin dropped, the
reporting that resumes against the new process, and the ordinary failure
that leaves the pin alone.
2026-09-03 20:35:31 +08:00
Tom Chen 75cc753092 Restore the event log only from a checkpoint the run asked to resume
Resolves op21-5 #4.

A run whose --load holds no checkpoint has that argument rewritten to
--ref-load, so the event log restore read the tracker of the reference
weights and copied their debug_events into the live directory. A fresh
finetune then started with the checksum and witness events of another
run, under rollout ids it is about to reuse, which is exactly the mixed
history the fault tolerance analysis reads as an inconsistency.

The restore now keys off args.requested_load, the load the user actually
asked for, which parsing already records. Runs driven by --megatron-config
never resolved that field on the base namespace, so it is set there too.
2026-09-03 20:35:31 +08:00
Tom Chen be8441baa9 Keep a secondary cancellation failure on Python 3.10
Resolves op35-156 #1.

add_note landed in Python 3.11, and setup.py still claims 3.10 support, so on
3.10 the attempt to annotate the primary failure raised AttributeError from
the error path itself and replaced the very failure it was describing. The
annotation now follows the capability check the rest of the repo uses, and
falls back to logging the secondary traceback.

Unit tests drive both branches with an exception that answers no add_note, and
check the primary failure still reaches the caller unreplaced.
2026-09-03 20:35:31 +08:00
Tom Chen 155cd7f1b3 Own every driver's teardown in a disposer
Resolves op35-150 #1.

Every driver released its objects on its last lines, so a raise from training,
evaluation, or any of the dispose calls left the Ray CommandActor without a
shutdown request and the run without its process trees reaped. Each driver now
takes a disposer its entry point opens for it and registers each object's
teardown where the object is created, so an exit path nobody thought of still
releases everything in reverse creation order. `with_exit_stack` is what opens
it: `asyncio.run(with_exit_stack(train, args))` replaces both the bare
`asyncio.run` and the `finally: finish_tracking()` that only three of the four
drivers had.

The two teardowns every driver shares are no longer every driver's business.
`init_orchestration_script` is what starts the tracking and launches the worker
manager, so it takes the disposer too and registers `finish_tracking` and
`shutdown_worker_manager` itself; no driver names either any more.

The async driver's trailing `await eval_dispatcher.drain()` becomes a
registration made where the dispatcher is created. That is after the executor
and the models are registered, so the drain runs before them on the way out and
an in-flight eval still finds the engines and the model it is reading. The drain
any raise above it used to skip now happens on every exit path.

A teardown that fails no longer stops the ones behind it: the stack runs every
callback and chains what they raise onto the failure that started the unwind.

The AST test that guards this selected no script at all, because no driver has
named `launch_worker_manager` since `init_orchestration_script` took the launch
over. It selects on that name now, and asserts each driver hands its disposer to
the shared composition root and is entered through `with_exit_stack`.
2026-09-03 20:25:17 +08:00
Tom Chen 81653c5528 Treat an unparsable release as belonging to no run
Resolves op35-148 #2.

ReleaseName.parse builds a validated model, so a release whose name is legal
for Helm but not for Miles - a 34-character run id, say - raises out of the
candidate loop. That loop runs outside both try blocks of the failure path,
so a single such neighbour replaced the diagnosis and the original SystemExit
with a validation traceback.

Unit tests cover a release whose name is legal for Helm but not for Miles,
and the diagnosis of the failed run that survives such a neighbour.
2026-09-03 08:25:24 +08:00
Tom Chen 2940ba2daa Diagnose a failed run against the releases of that run only
Resolves op35-148 #1.

The diagnosis listed every release of the namespace and only then kept the
ones of this run. helm list truncates to its 256-release maximum before
returning, so a busy namespace can drop the sibling trainer or inference
release of the very run that failed. Listing now passes a name filter, which
helm applies before that maximum.

Unit tests cover the filter passed to helm, a listing without one, the run
prefix every release of a run starts with, and the filter the diagnosis
builds from it.
2026-09-03 08:23:39 +08:00
Tom Chen d5d5258064 Require the rollout id placeholder in debug dump templates
Resolves op35-131 #3.

Both dump entry points formatted the user template with str.format, which
silently drops a keyword the template never references. A literal file name
was therefore accepted, and every rollout - and, in a multi-policy run, every
policy - rewrote the same path. Formatting now goes through one helper that
asserts the placeholder is present.
2026-09-03 08:23:39 +08:00
Tom Chen 6e568274f8 Refuse a policy that spells the name eval
Resolves op35-131 #2.

A debug dump names an eval rollout eval_N and a policy's training
rollout <model id>_N, so a run declaring the model id eval wrote its
training dumps, trajectories and dashboard columns onto the paths of the
shared eval dumps and one silently replaced the other. Reject that model
id where the megatron config is validated.
2026-09-03 08:23:39 +08:00
Tom Chen b9bfab9550 Report the unfiltered reward from the verifiers rollout
Resolves op35-127 #1.

The verifiers rollout calls the dynamic filter without recording what the
groups scored first, so rollout/raw_reward_unfiltered was missing under
that backend while every other one reported it. Feed the gatherer the
group's samples before the filter sees them.
2026-09-03 08:23:39 +08:00
Tom Chen 8cb8f4cb20 Let a rollout fn tear itself down synchronously
Resolves op35-79 #1.

The base classes declare dispose as a coroutine, but an out-of-tree
rollout or checkpoint eval fn written against the older contract
overrides it with a plain method. Awaiting its None raised a TypeError
in the shutdown path, after the backend had already been torn down and
with nothing left to salvage the run. Await the hook's result only when it
is awaitable, through a small maybe_await helper.
2026-09-03 08:23:39 +08:00
Tom Chen dd3269ad88 Keep the namespace in the documented runs root
Resolves op35-55 #1.

The one infra.yaml an administrator is told to write hardcoded a runs
root with no namespace in it, and that file drives every chart in every
namespace, so following the document undid the ${NAMESPACE} default the
charts ship and let two namespaces running the same run ID write over
each other. Put the variable back into the example, and say where it may
be used and that it is the only one.
2026-09-03 08:23:39 +08:00
Tom Chen 5685881f35 Judge the runs root by the mount that actually provides it
Resolves op35-54 #4.

Every mount covering the runs root was collected into a writable list
and a read-only list, and one writable entry anywhere passed the check.
Kubernetes gives the path to the most specific mount, so a read-only
/data/runs nested under a writable /data was accepted and then failed
the run at its first write. Pick that most specific mount and judge its
source and its read-only flag alone.

The chart tests now nest a read-only mount and an emptyDir inside a
shared one, and render the reverse nesting that must still pass.
2026-09-03 08:23:39 +08:00
Tom Chen c25e57ab5c Compare the runs root against a cleaned mount path
Resolves op35-54 #3.

The containment test built its prefix straight off the configured mount
path, so a mount written as /cluster-storage/ produced the prefix
/cluster-storage// and the chart refused a runs root that plainly lives
under it. Clean both paths first, and trim the one trailing slash a
cleaned root still carries, so a mount at / keeps covering everything.

The chart tests now render a mount written both ways and pin the
refusals the cleaning must leave standing.
2026-09-03 08:23:39 +08:00
Tom Chen 94dd0bc1bb Refuse a runs root that lives in a single pod
Resolves op35-54 #1.

The runs root check only asked whether some mount covered the path and
whether that mount was writable, so an emptyDir volume passed it. Every
pod would then get its own empty copy, the orchestrator's state file
would be invisible to the launcher, and the run would wait for a verdict
nobody could ever read. Reject an emptyDir under the runs root, and say
so where the folder convention is documented.
2026-09-02 23:58:34 +08:00
Tom Chen 9d818b1334 Skip only the trace context ray itself injects
Resolves op35-45 #1.

Ray injects the trace context as an unannotated keyword-only parameter
defaulting to None. Matching on the name and the keyword-only kind alone
also swallowed a worker's own required keyword-only parameter of that
name, which then took no annotation check, joined no request model, and
left the method uncallable either way. Match the whole injected shape so
a parameter the worker declares itself stays on the wire.

Unit tests cover a worker's own keyword-only parameter of that name, required and
defaulted, beside the exact shape ray injects.
2026-09-02 23:58:34 +08:00
Tom Chen 67dc50fa29 Let a displaced command job pod fail the step
Resolves op35-25 #1.

The ignore rule was written to forgive a pod the cluster evicted before it
ran, so that a displacement would not spend the job's one attempt. It cannot
tell that pod apart from one evicted halfway through the command, though, and
forgiving the second one leaves the job short of a completion, so Kubernetes
starts a replacement and runs the command again. A command step is
at-most-once: doing its work twice is worse than reporting a displacement as a
failure the caller can retry. Drop the rule and let any pod failure end the
job. Name the replacement policy the rule used to imply as well, because
without a pod failure policy it defaults to TerminatingOrFailed, which starts
a replacement the moment the evicted pod begins terminating and before the
failure ends the job; Failed waits for the terminal pod, leaving one pod per
completion index.
2026-09-02 23:58:34 +08:00
Tom Chen 29f4f721d1 Resolve the mooncake store config in one place, environment included
Resolves op35-23 #1.

A kubernetes run that names no store configuration is given one, and those
invented defaults - a loopback master, tcp, 2 GiB of each capacity - outrank
MOONCAKE_MASTER, MOONCAKE_PROTOCOL and the capacity variables, because
explicit init kwargs win over the environment everywhere the store is built.
A pod configured by its platform through those variables would therefore
dial loopback with the wrong protocol and capacity. Move the resolver that
turns init kwargs into a store config out of object_store.py and into
object_store_config.py, where one table of init kwarg, environment variable
and built-in default now serves both it and the launcher's defaults, so the
order - explicit kwargs over the environment over the default - is written
once and a defaulted field the environment names is taken from there.

Unit tests cover the capacities the environment names, the defaults it
leaves alone, and a kubernetes run whose store configuration comes from its
platform.
2026-09-02 23:58:34 +08:00
Tom Chen 71273e5808 Let help be printed by the parse that still requires what is required
Resolves op35-19 #1.

The probe that finds the user-provided functions parses the command line
with every required argument relaxed, and argparse runs the help action
during that parse, so miles --help printed a usage listing
--rollout-batch-size as optional and then exited. Take the help action out
of the probe instead, leaving it to the real parse, whose usage says what
the run actually requires.

Unit tests cover the suppressed parse, the restoration of both help spellings
(including after a failure), and a parser that declares no help action.
2026-09-02 23:58:34 +08:00
Tom Chen 39ef3fcf75 Release every inference pod of a trainer even if one release fails
Resolves op34-4 #1.

The release patch tests the scheduling gate it removes, so the apiserver
answers 422 whenever the cached pod the reconcile is acting on has already
been released. That exception aborted the loop over one trainer's inference
pods, leaving the pods behind it gated until the watch happened to refresh
the cache, which breaks the at-least-once and idempotent contract the
reconcile loop is built on. Release every pod of the trainer independently
and log whatever any one of them raised, so a stale cache costs that pod a
retry on the next pass rather than costing its neighbours their release.
2026-09-02 23:58:34 +08:00
Tom Chen d79b6b014d Render a pool environment's gpu id sentinel as a variable too
Resolves op34-3 #1.

A sub-node pool is rendered with sentinel gpu ids, and the two probing
contexts a values entry builds its environment with carry the same
is_sub_node, so an environment value derived from the gpu ids passes the
equality check with the sentinel in it. Only the command was rewritten into
kubelet placeholders, so such a pod would read the literal 987654322 rather
than the base gpu id it was actually given. Substitute the environment
values the same way; the chart already defines MILES_BASE_GPU_ID ahead of a
pool's own environment, so the kubelet expands it.
2026-09-02 23:58:34 +08:00
Tom Chen 38f81c1e74 Keep resolve_hardware in the v1 command utilities
Resolves op33-2 #1.

The v1 compatibility module is what the launch script guide tells older ray
launchers to import, but it never re-exported resolve_hardware, which the v1
API carried and which every launcher that leaves --hardware on auto calls as
U.resolve_hardware(self). Such a script raises AttributeError while building
its config. Re-export the function the current module already defines.

Unit tests cover the re-export, the function working through this module,
and every name __all__ advertises being one the module has.
2026-09-02 23:58:33 +08:00
Tom Chen cc427b8169 Hot restart a run into the wandb run it was already reporting to
Resolves op32-10 #1.

The hot restart example stabilizes only the wandb group: the helper derives
it from the run id, but the launcher mints a fresh --wandb-run-id on every
invocation when the argv carries none. The relaunched orchestration script
therefore opens a second wandb run while the surviving trainers keep writing
to the first one, splitting one job's metrics in two. Name the wandb run
after the run id, as the hot restart e2e scenario already does; primary
initialization resumes such a run rather than starting it over.
2026-09-02 23:58:33 +08:00
Tom Chen e9a0a63057 Group the solver verifier metrics under the run that produced them
Resolves op32-7 #1.

The multi policy example asks for the default wandb arguments without
handing over the run id it was configured with, so with WANDB_API_KEY set
the helper mints a fresh random id for the wandb group while the output
directories keep using --run-id. Restarting the same run then scatters its
history over unrelated wandb groups. Pass the configured run id, the way
every other example does.
2026-09-02 23:58:33 +08:00
Tom Chen ce524a7e3e Name the run the split multi policy example installs
Resolves op32-3 #1.

Example 2 of the split deployment README installs its releases without
ever setting MILES_SCRIPT_RUN_ID or MILES_SCRIPT_RUN_UUID; only example 1
exports them. Run on its own, the multi policy script asserts on the empty
run uuid inside RunAddressBook.of_config() before the first release is
installed. Export a run id and a fresh uuid for this example too.
2026-09-02 23:58:33 +08:00
Tom Chen 7d3c9d8006 Run the split deployment examples as modules
Resolves op32-2 #1.

Both scripts import their shared address book as `examples.…`, and
running a file by path puts its own directory on `sys.path` instead of
the repository root, so the documented command died on
`ModuleNotFoundError: No module named 'examples'` before installing
anything.
2026-09-02 23:58:33 +08:00
Tom Chen bc937144af Identify a rendered object by its api group, not its version
Resolves op29-21 #1.

A custom resource answers on every version it serves, so the same object
written once as `example.com/v1` and once as `example.com/v1beta1` took
two identities, slipped past the duplicate check and made the upgrade
diff depend on which spelling Helm applied last. The group still keeps a
crd apart from a built-in of the same kind and name.

The scaling exemption and the stateful set lookup read an identity too, so
both now match on the api group as well.
2026-09-02 23:58:33 +08:00
Tom Chen 9562bded9a Say which fault tolerance scenarios delete their dumps
Resolves op29-13 #2.

Only the comparison pipeline removes the dump directory when it ends.
The soak scenarios clear a stale directory before they start and leave
everything a successful run produced, so the blanket claim sent readers
looking for dumps that are still there.
2026-09-02 23:46:29 +08:00
Tom Chen ec746ed759 Name the external engine objects after the run installing them
Resolves op29-10 #1.

The Service, the StatefulSet, their selector and the engine DNS
addresses were all the fixed name `external-sglang`, so two runs of the
test overlapping in one namespace, or one starting before the previous
release finished being deleted, fought over the same objects. The name
now carries the run id, exactly as the release does.
2026-09-02 23:46:29 +08:00
Tom Chen f9aef66a14 Refuse an override map that names one argument twice
Resolves op28-27 #1.

`coerce_dict_to_args` resolves every accepted alias to its target dest,
so `add_bias_linear` and `disable_bias_linear` both land on the same
key. The dict comprehension kept whichever came last, making the model
settings of a `--megatron-config` override depend on the order the YAML
happened to list them.
2026-09-02 23:46:29 +08:00
Tom Chen 94de896c4d Keep the policy identity in the CI history record
Resolves op28-25 #1.

The whitelist is matched by suffix, so a policy-prefixed metric was
found and then stored under the bare canonical key. Two policies of one
run therefore wrote one series holding step 0 twice, and the gate's
selector refused the duplicate step instead of checking either policy.
The record now keeps the key as logged.

Unit tests cover a policy-prefixed key kept as logged, two policies staying
two series, an unprefixed run unchanged, and the whitelist that still
refuses everything else.
2026-09-02 23:46:29 +08:00
Tom Chen f5fcd7d65d Reset the boolean model flags a policy script leaves out
Resolves op28-22 #1.

Per policy overrides carried only the options the model script spells
out, so a policy omitting a `store_true` flag that another policy of the
run declares kept the leader's value: a Qwen3 verifier next to a Qwen2.5
solver inherited `add_qkv_bias=True` and no longer matched its own
checkpoint. The override map now states the parser default for every
boolean flag some sibling declares and this one does not.
2026-09-02 23:46:29 +08:00
Tom Chen 18057ffe13 Forward the policy id into the composed buffer metrics call
Resolves op28-20 #2.

`DefaultMultiDataBuffer.get_metrics` received the policy the drain asked
about and then dropped it, unlike the neighbouring `get`. A custom
`DataBuffer` that selects by policy therefore saw `None` and could
attribute its metrics to the wrong policy.

Unit tests cover the policy reaching the composed buffer, each policy being
asked about itself only, and the metrics that come back.
2026-09-02 23:46:29 +08:00
Tom Chen eeefe3867c Accept the epsilon name the deprecated megatron parser produces
Resolves op28-14 #2

Model scripts spell the normalization epsilon --norm-epsilon, and its
dest is layernorm_epsilon on the current megatron parser but
norm_epsilon on the deprecated one this run still supports
(DEPRECATED_MEGATRON_COMPATIBLE=1). Only the first name was allowed, so
under the deprecated parser a policy config carrying its model's epsilon
was refused.

Allow both names; each is admitted only where the parser in use actually
declares it.
2026-09-02 23:46:29 +08:00