Resolves op27-30 #1
Resolves op27-17 #3
A take-over that finds no checkpoint zeroes the scheduler and resets the
optimizer, and only then checks the state it rebuilt against the moments adam
keeps. A Muon member rebuilds a `momentum_buffer`, so the check refused it
after the trainer had already been mutated, and the surviving trainer was left
half changed. Refuse the optimizer among the other preconditions, before
anything is touched.
Resolves op27-27 #1
The FSDP actor answered `init` with `args.start_rollout_id`, which it defaults
to zero before loading and which `finalize_load` overwrites only from a
checkpoint's `meta.json`. A checkpoint with a missing or unreadable `meta.json`
therefore left an explicitly requested `--start-rollout-id` in place, and
`create_training_model` compared that request against itself and logged
nothing where the two positions differ. Carry the position the checkpoint
restored and answer with it.
Resolves op27-24 #1
Resolves op27-31 #1
The reset guarded only the deprecated `offload_optimizer_states`, which MCore
normalizes to false once `--chunked-optimizer-state-offload` is the active
mode. The chunked offloader keeps the canonical moments in its own cpu store,
which is not the mapping the reset clears, so a trainer hot restarted before
its first checkpoint passed the self-check and then prefetched the moments of
the rollouts it had just discarded. Refuse the mode instead.
Resolves op27-22 #1
`create_training_models` runs before the adapters named on the command line
are registered, and it loads the rollout executor, so `sources` is still empty
and the run was reported as serving no adapter at all. That is not what
happened: the adapters are configured, and `reconcile()` has not filled the
sources yet. Say that instead.
Unit tests cover the message a run whose adapters are not reconciled yet is
given, and the restore of every source once they are.
Resolves op27-21 #1
`load_state` reloads the ranks of a cell, and the cells of a controller, with
a fail-fast `asyncio.gather`, and it does not kill a worker whose reload
failed. The first failure therefore returned while the other reloads were
still running their distributed checkpoint collectives, so the controller went
idle and an orchestration retry could interleave with them. Gather the reloads
with `gather_and_raise_first`, which waits for all of them and then raises the
first failure.
Resolves op27-17 #2
A take-over that finds no checkpoint reseeded from the arguments, which is the
stream the process stood on before `build_model_and_optimizer` ran. Parameter
initialization advances both the host generators and Megatron's CUDA RNG
tracker, so the surviving trainer resumed from a point the discarded run had
already consumed, and a replay could diverge silently. Snapshot the generators
once the model and the optimizer are built, and restore that snapshot instead.
Resolves op27-17 #1
The no-checkpoint reload path demanded `--finetune`, `--no-load-optim` and
`--no-load-rng`, and a `ckpt_step` taken from `--ref-ckpt-step`. Only the
non-bridge branch of `resolve_args_checkpoint_load` sets those; the bridge
branch rewrites `load` and `start_rollout_id` alone. A non-LoRA bridge run
taken over before its first save therefore hit the assertion on its ordinary
cold-start path. Require the flags only where the cold start actually sets
them.
Resolves op27-13 #1
`SGLangApiClient` kept the engine's `api_key` but posted without an
`Authorization` header, so with `--sglang-api-key` or a model-level key every
`_make_request` endpoint answered 401. `abort_all_requests` goes through that
path, so a hot restart could not abort the generations still in flight and the
take-over failed. Send the same headers the client already computes for its
other calls.
Resolves op27-11 #1
`parse_hot_restart` dropped every blank token, so a non-empty request such as
`--hot-restart ,` parsed to an empty list. Every caller reads that list's
truthiness, so the malformed request became an ordinary launch: the backend
check that only Kubernetes serves a hot restart was skipped, and on Ray the
launch went on to `_clean_up_previous_run`, which kills the host's SGLang,
Miles and Ray processes. Reject a non-empty value that names no component.
Unit tests cover the separators and whitespace that name no component, the empty
value that is still an ordinary launch, and a trailing separator that is not.
Resolves op27-9 #2
`--ref-update-interval` backs the actor up as the new reference in memory, but
`save_model` persists only the actor, its optimizer and its scheduler, so a
reload reloads the reference from the static `--ref-load` instead. The run
would then compute KL against the reference it started from while the actor
stands at a rollout that was trained against a newer one. Reject the
combination up front, alongside the other reload preconditions.
Resolves op27-9 #1
`compute_trainer_args` only re-resolved the checkpoint load when a
`--megatron-config` was given, so a critic synthesized from `--use-critic`
kept the `requested_load` that `validate_args` had derived from the actor's
`--load`. A hot restart reads `requested_load`, so the critic would look for
its state under the actor's checkpoint directory and either fail to load or
restore the wrong model and optimizer. Re-resolve whenever a trainer's
overrides moved `load` away from the value the base arguments were resolved
from.
Resolves op27-7 #1.
The wait that lets a hot restart adopt a rollout executor asked
is_initialized(), which is only true in the INITED state. A process
still running the init of the previous script, or one whose init failed
and left it half-built, therefore answered false and was let through as
the replacement; the script then called init() on it and the init-once
guard refused, failing the hot restart.
The rollout executor now reports its init state, and the wait goes on
until that state is NOT_INITED, which is the only state a process the
script may initialize itself is in.
Resolves op27-6 #1.
Taking an inference controller over aborted the generations the fleet
was running, but said nothing to the agent backend behind
--custom-agent-function-path. That loop keeps issuing fresh completion
requests after the script that started it is gone, so it takes the fleet
and the agent-server slots the new rollout executor is about to ask for.
The take-over now calls the same abort hook the ordinary abort path
calls, on the same bounded budget as the rest of the take-over.
Unit tests cover the hook called after the fleet abort, the cold start that
calls none, and the take-over budget it runs on.
Resolves op27-5 #1.
Deciding whether a take-over starts a run over asserted that the run has
no --megatron-config, and read the tracker under the base --load. A
configured single-actor run is supported everywhere else on this path,
so a hot restart of one aborted right there, before the event log was
ever reset, and the base --load holds no tracker to read anyway once the
per-trainer checkpoint directories are derived from it.
The question is now asked of each declared trainer through the arguments
compute_trainer_args resolves for it, and the run counts as resumable if
any of them has a checkpoint to resume from.
Resolves op27-2 #1.
The init-once guard sat on _init_common, which is only the first step of
what a concrete trainer does in init: the megatron and fsdp actors go on
to build parallel state, the model, tracking and the tokenizer after it
returns. An actor that threw in any of that was left reporting
is_initialized() true, so a recovery would take a half-built process for
one it can resume and drive it as a trainer that has no model.
init is now a concrete method on the base that guards the whole thing
and hands off to the _init each backend implements, so the state the
guard reports covers everything init does.
Resolves op25-10 #1.
The last ingest time was stamped before the ownership check, so a
snapshot rejected for claiming another reporter's cells still renewed
the sender's lease. A reporter that keeps sending such snapshots is
never dropped, and the cells it registered before, which may long be
gone, stay in the run's view of the fleet.
The stamp moves after the replacement, so only a snapshot that was
actually taken in renews the reporter.
Unit tests cover the lease of a reporter whose snapshot was refused, the
cells dropped with it, and the snapshot that does renew its sender.
Resolves op25-8 #1.
The snapshot check read the worker names off cell.workers and compared
the cells they belong to as a set, so it never looked at the
CellInfo.worker_names the run is served from. A cell that announced an
empty list, a name no worker of it has, or the same name twice passed,
and the rollout reconciliation that indexes worker_names[0] and looks
addresses up by those names then failed on a snapshot the hub had
already taken in.
The two lists now have to be one duplicate-free mapping of each other.
Unit tests cover a cell announcing nothing, announcing a worker it does not
carry and carrying one twice, and the membership a refused snapshot leaves
alone.
Resolves op25-2 #2.
The reporter waits for the hub once, before its loop, and the handle
pins the boot uuid it saw there. When the inference controller process
that hosts the hub restarts, every later ingest raises
ServerRestartedError, the generic handler logs it and the loop retries
against the same pin forever, so the new hub never learns about the
cells this deployment reports and the run waits for engines that are
running.
A restarted hub is now awaited again with allow_server_uuid_change, which
is the only way the handle drops the pin, and reporting resumes against
the process that answers now.
Unit tests cover the readiness awaited again with the pin dropped, the
reporting that resumes against the new process, and the ordinary failure
that leaves the pin alone.
Resolves op21-5 #4.
A run whose --load holds no checkpoint has that argument rewritten to
--ref-load, so the event log restore read the tracker of the reference
weights and copied their debug_events into the live directory. A fresh
finetune then started with the checksum and witness events of another
run, under rollout ids it is about to reuse, which is exactly the mixed
history the fault tolerance analysis reads as an inconsistency.
The restore now keys off args.requested_load, the load the user actually
asked for, which parsing already records. Runs driven by --megatron-config
never resolved that field on the base namespace, so it is set there too.
Resolves op35-156 #1.
add_note landed in Python 3.11, and setup.py still claims 3.10 support, so on
3.10 the attempt to annotate the primary failure raised AttributeError from
the error path itself and replaced the very failure it was describing. The
annotation now follows the capability check the rest of the repo uses, and
falls back to logging the secondary traceback.
Unit tests drive both branches with an exception that answers no add_note, and
check the primary failure still reaches the caller unreplaced.
Resolves op35-150 #1.
Every driver released its objects on its last lines, so a raise from training,
evaluation, or any of the dispose calls left the Ray CommandActor without a
shutdown request and the run without its process trees reaped. Each driver now
takes a disposer its entry point opens for it and registers each object's
teardown where the object is created, so an exit path nobody thought of still
releases everything in reverse creation order. `with_exit_stack` is what opens
it: `asyncio.run(with_exit_stack(train, args))` replaces both the bare
`asyncio.run` and the `finally: finish_tracking()` that only three of the four
drivers had.
The two teardowns every driver shares are no longer every driver's business.
`init_orchestration_script` is what starts the tracking and launches the worker
manager, so it takes the disposer too and registers `finish_tracking` and
`shutdown_worker_manager` itself; no driver names either any more.
The async driver's trailing `await eval_dispatcher.drain()` becomes a
registration made where the dispatcher is created. That is after the executor
and the models are registered, so the drain runs before them on the way out and
an in-flight eval still finds the engines and the model it is reading. The drain
any raise above it used to skip now happens on every exit path.
A teardown that fails no longer stops the ones behind it: the stack runs every
callback and chains what they raise onto the failure that started the unwind.
The AST test that guards this selected no script at all, because no driver has
named `launch_worker_manager` since `init_orchestration_script` took the launch
over. It selects on that name now, and asserts each driver hands its disposer to
the shared composition root and is entered through `with_exit_stack`.
Resolves op35-148 #2.
ReleaseName.parse builds a validated model, so a release whose name is legal
for Helm but not for Miles - a 34-character run id, say - raises out of the
candidate loop. That loop runs outside both try blocks of the failure path,
so a single such neighbour replaced the diagnosis and the original SystemExit
with a validation traceback.
Unit tests cover a release whose name is legal for Helm but not for Miles,
and the diagnosis of the failed run that survives such a neighbour.
Resolves op35-148 #1.
The diagnosis listed every release of the namespace and only then kept the
ones of this run. helm list truncates to its 256-release maximum before
returning, so a busy namespace can drop the sibling trainer or inference
release of the very run that failed. Listing now passes a name filter, which
helm applies before that maximum.
Unit tests cover the filter passed to helm, a listing without one, the run
prefix every release of a run starts with, and the filter the diagnosis
builds from it.
Resolves op35-131 #3.
Both dump entry points formatted the user template with str.format, which
silently drops a keyword the template never references. A literal file name
was therefore accepted, and every rollout - and, in a multi-policy run, every
policy - rewrote the same path. Formatting now goes through one helper that
asserts the placeholder is present.
Resolves op35-131 #2.
A debug dump names an eval rollout eval_N and a policy's training
rollout <model id>_N, so a run declaring the model id eval wrote its
training dumps, trajectories and dashboard columns onto the paths of the
shared eval dumps and one silently replaced the other. Reject that model
id where the megatron config is validated.
Resolves op35-127 #1.
The verifiers rollout calls the dynamic filter without recording what the
groups scored first, so rollout/raw_reward_unfiltered was missing under
that backend while every other one reported it. Feed the gatherer the
group's samples before the filter sees them.
Resolves op35-79 #1.
The base classes declare dispose as a coroutine, but an out-of-tree
rollout or checkpoint eval fn written against the older contract
overrides it with a plain method. Awaiting its None raised a TypeError
in the shutdown path, after the backend had already been torn down and
with nothing left to salvage the run. Await the hook's result only when it
is awaitable, through a small maybe_await helper.
Resolves op35-55 #1.
The one infra.yaml an administrator is told to write hardcoded a runs
root with no namespace in it, and that file drives every chart in every
namespace, so following the document undid the ${NAMESPACE} default the
charts ship and let two namespaces running the same run ID write over
each other. Put the variable back into the example, and say where it may
be used and that it is the only one.
Resolves op35-54 #4.
Every mount covering the runs root was collected into a writable list
and a read-only list, and one writable entry anywhere passed the check.
Kubernetes gives the path to the most specific mount, so a read-only
/data/runs nested under a writable /data was accepted and then failed
the run at its first write. Pick that most specific mount and judge its
source and its read-only flag alone.
The chart tests now nest a read-only mount and an emptyDir inside a
shared one, and render the reverse nesting that must still pass.
Resolves op35-54 #3.
The containment test built its prefix straight off the configured mount
path, so a mount written as /cluster-storage/ produced the prefix
/cluster-storage// and the chart refused a runs root that plainly lives
under it. Clean both paths first, and trim the one trailing slash a
cleaned root still carries, so a mount at / keeps covering everything.
The chart tests now render a mount written both ways and pin the
refusals the cleaning must leave standing.
Resolves op35-54 #1.
The runs root check only asked whether some mount covered the path and
whether that mount was writable, so an emptyDir volume passed it. Every
pod would then get its own empty copy, the orchestrator's state file
would be invisible to the launcher, and the run would wait for a verdict
nobody could ever read. Reject an emptyDir under the runs root, and say
so where the folder convention is documented.
Resolves op35-45 #1.
Ray injects the trace context as an unannotated keyword-only parameter
defaulting to None. Matching on the name and the keyword-only kind alone
also swallowed a worker's own required keyword-only parameter of that
name, which then took no annotation check, joined no request model, and
left the method uncallable either way. Match the whole injected shape so
a parameter the worker declares itself stays on the wire.
Unit tests cover a worker's own keyword-only parameter of that name, required and
defaulted, beside the exact shape ray injects.
Resolves op35-25 #1.
The ignore rule was written to forgive a pod the cluster evicted before it
ran, so that a displacement would not spend the job's one attempt. It cannot
tell that pod apart from one evicted halfway through the command, though, and
forgiving the second one leaves the job short of a completion, so Kubernetes
starts a replacement and runs the command again. A command step is
at-most-once: doing its work twice is worse than reporting a displacement as a
failure the caller can retry. Drop the rule and let any pod failure end the
job. Name the replacement policy the rule used to imply as well, because
without a pod failure policy it defaults to TerminatingOrFailed, which starts
a replacement the moment the evicted pod begins terminating and before the
failure ends the job; Failed waits for the terminal pod, leaving one pod per
completion index.
Resolves op35-23 #1.
A kubernetes run that names no store configuration is given one, and those
invented defaults - a loopback master, tcp, 2 GiB of each capacity - outrank
MOONCAKE_MASTER, MOONCAKE_PROTOCOL and the capacity variables, because
explicit init kwargs win over the environment everywhere the store is built.
A pod configured by its platform through those variables would therefore
dial loopback with the wrong protocol and capacity. Move the resolver that
turns init kwargs into a store config out of object_store.py and into
object_store_config.py, where one table of init kwarg, environment variable
and built-in default now serves both it and the launcher's defaults, so the
order - explicit kwargs over the environment over the default - is written
once and a defaulted field the environment names is taken from there.
Unit tests cover the capacities the environment names, the defaults it
leaves alone, and a kubernetes run whose store configuration comes from its
platform.
Resolves op35-19 #1.
The probe that finds the user-provided functions parses the command line
with every required argument relaxed, and argparse runs the help action
during that parse, so miles --help printed a usage listing
--rollout-batch-size as optional and then exited. Take the help action out
of the probe instead, leaving it to the real parse, whose usage says what
the run actually requires.
Unit tests cover the suppressed parse, the restoration of both help spellings
(including after a failure), and a parser that declares no help action.
Resolves op34-4 #1.
The release patch tests the scheduling gate it removes, so the apiserver
answers 422 whenever the cached pod the reconcile is acting on has already
been released. That exception aborted the loop over one trainer's inference
pods, leaving the pods behind it gated until the watch happened to refresh
the cache, which breaks the at-least-once and idempotent contract the
reconcile loop is built on. Release every pod of the trainer independently
and log whatever any one of them raised, so a stale cache costs that pod a
retry on the next pass rather than costing its neighbours their release.
Resolves op34-3 #1.
A sub-node pool is rendered with sentinel gpu ids, and the two probing
contexts a values entry builds its environment with carry the same
is_sub_node, so an environment value derived from the gpu ids passes the
equality check with the sentinel in it. Only the command was rewritten into
kubelet placeholders, so such a pod would read the literal 987654322 rather
than the base gpu id it was actually given. Substitute the environment
values the same way; the chart already defines MILES_BASE_GPU_ID ahead of a
pool's own environment, so the kubelet expands it.
Resolves op33-2 #1.
The v1 compatibility module is what the launch script guide tells older ray
launchers to import, but it never re-exported resolve_hardware, which the v1
API carried and which every launcher that leaves --hardware on auto calls as
U.resolve_hardware(self). Such a script raises AttributeError while building
its config. Re-export the function the current module already defines.
Unit tests cover the re-export, the function working through this module,
and every name __all__ advertises being one the module has.
Resolves op32-10 #1.
The hot restart example stabilizes only the wandb group: the helper derives
it from the run id, but the launcher mints a fresh --wandb-run-id on every
invocation when the argv carries none. The relaunched orchestration script
therefore opens a second wandb run while the surviving trainers keep writing
to the first one, splitting one job's metrics in two. Name the wandb run
after the run id, as the hot restart e2e scenario already does; primary
initialization resumes such a run rather than starting it over.
Resolves op32-7 #1.
The multi policy example asks for the default wandb arguments without
handing over the run id it was configured with, so with WANDB_API_KEY set
the helper mints a fresh random id for the wandb group while the output
directories keep using --run-id. Restarting the same run then scatters its
history over unrelated wandb groups. Pass the configured run id, the way
every other example does.
Resolves op32-3 #1.
Example 2 of the split deployment README installs its releases without
ever setting MILES_SCRIPT_RUN_ID or MILES_SCRIPT_RUN_UUID; only example 1
exports them. Run on its own, the multi policy script asserts on the empty
run uuid inside RunAddressBook.of_config() before the first release is
installed. Export a run id and a fresh uuid for this example too.
Resolves op32-2 #1.
Both scripts import their shared address book as `examples.…`, and
running a file by path puts its own directory on `sys.path` instead of
the repository root, so the documented command died on
`ModuleNotFoundError: No module named 'examples'` before installing
anything.
Resolves op29-21 #1.
A custom resource answers on every version it serves, so the same object
written once as `example.com/v1` and once as `example.com/v1beta1` took
two identities, slipped past the duplicate check and made the upgrade
diff depend on which spelling Helm applied last. The group still keeps a
crd apart from a built-in of the same kind and name.
The scaling exemption and the stateful set lookup read an identity too, so
both now match on the api group as well.
Resolves op29-13 #2.
Only the comparison pipeline removes the dump directory when it ends.
The soak scenarios clear a stale directory before they start and leave
everything a successful run produced, so the blanket claim sent readers
looking for dumps that are still there.
Resolves op29-10 #1.
The Service, the StatefulSet, their selector and the engine DNS
addresses were all the fixed name `external-sglang`, so two runs of the
test overlapping in one namespace, or one starting before the previous
release finished being deleted, fought over the same objects. The name
now carries the run id, exactly as the release does.
Resolves op28-27 #1.
`coerce_dict_to_args` resolves every accepted alias to its target dest,
so `add_bias_linear` and `disable_bias_linear` both land on the same
key. The dict comprehension kept whichever came last, making the model
settings of a `--megatron-config` override depend on the order the YAML
happened to list them.
Resolves op28-25 #1.
The whitelist is matched by suffix, so a policy-prefixed metric was
found and then stored under the bare canonical key. Two policies of one
run therefore wrote one series holding step 0 twice, and the gate's
selector refused the duplicate step instead of checking either policy.
The record now keeps the key as logged.
Unit tests cover a policy-prefixed key kept as logged, two policies staying
two series, an unprefixed run unchanged, and the whitelist that still
refuses everything else.
Resolves op28-22 #1.
Per policy overrides carried only the options the model script spells
out, so a policy omitting a `store_true` flag that another policy of the
run declares kept the leader's value: a Qwen3 verifier next to a Qwen2.5
solver inherited `add_qkv_bias=True` and no longer matched its own
checkpoint. The override map now states the parser default for every
boolean flag some sibling declares and this one does not.
Resolves op28-20 #2.
`DefaultMultiDataBuffer.get_metrics` received the policy the drain asked
about and then dropped it, unlike the neighbouring `get`. A custom
`DataBuffer` that selects by policy therefore saw `None` and could
attribute its metrics to the wrong policy.
Unit tests cover the policy reaching the composed buffer, each policy being
asked about itself only, and the metrics that come back.
Resolves op28-14 #2
Model scripts spell the normalization epsilon --norm-epsilon, and its
dest is layernorm_epsilon on the current megatron parser but
norm_epsilon on the deprecated one this run still supports
(DEPRECATED_MEGATRON_COMPATIBLE=1). Only the first name was allowed, so
under the deprecated parser a policy config carrying its model's epsilon
was refused.
Allow both names; each is admitted only where the parser in use actually
declares it.