100 Commits
Author SHA1 Message Date
Shi-Dong bc6191a6e4 dashboard: show per-family match counts while filtering metrics (#3696) 2026-09-24 16:43:48 -07:00
Shi-Dong 5510af6752 Fix trainer-rollout diagnostics with rollout logprobs enabled (#3655) 2026-09-24 10:12:47 -07:00
Shi-Dong 4eebfbab3f Add experimental Nemotron Workplace Assistant RL example (#3658) 2026-09-24 00:25:15 -07:00
Shi-Dong a20947e65c dashboard: compute chart extents without spreading into Math.min/max (#3654) 2026-09-23 21:53:10 -07:00
Shi-Dong 8ec0a119af fix(args): treat the fully-async rollout path as the mode that selects it (#3587)
--fully-async selects FullyAsyncRolloutFn, but naming that class through
--rollout-function-path selected the same producer while leaving the mode off, so
the run skipped every --fully-async check (colocate, partial rollout, legacy
rollout v1, pause mode, multi-LoRA) and train.py's async-driver guard.

Normalize the path spelling into the flag before validation runs. Only the exact
class is recognized; a subclass still passes --fully-async explicitly.
2026-09-21 22:29:04 -07:00
Shi-Dongandyueming-yuan e2a5a3e592 fix(rollout): isolate cancelled groups in fully async generation (#3319)
Co-authored-by: yueming-yuan <yym022502@gmail.com>
2026-09-18 17:22:57 -07:00
Shi-Dong 97122c1ed1 fix: guard connection allocations during offloaded weight sync (#3129)
Signed-off-by: Shi Dong <shi.dong@radixark.ai>
2026-09-14 17:42:13 -07:00
9a9c3eb386 fix: align MTP loss masks with input tokens before context-parallel slicing (#3219)
Co-authored-by: Jiajun Li <jiajun.li@radixark.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 13:20:12 -07:00
Shi-Dong e5c03bda7e Fix MTP gradient detachment in AutoBridge mode (#3211) 2026-09-11 13:26:15 -07:00
Shi-DongandZhichenzzz f01f9a3166 fix: use host weights when syncing an offloaded actor (#3128)
Signed-off-by: Shi Dong <shi.dong@radixark.ai>
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-09-09 15:29:59 -07:00
Shi-DongandZhichenzzz f1eac9bad4 Fix missing Nemotron-H Mamba convolution mappings (#3163)
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-09-09 14:43:08 -07:00
Shi-Dong 021d311e89 Fix Nemotron-H forward compatibility when MTP is disabled (#3164) 2026-09-09 14:42:53 -07:00
Shi-Dong d8dd0dec8d Scan full responses for repetition metrics (#3115) 2026-09-09 12:04:50 -07:00
Shi-Dong d2fc97ce58 Tolerate missing optional SGLang Anthropic helpers (#3114) 2026-09-04 14:48:36 -07:00
Shi-Dong 428d8882d0 dashboard: pin the anatomy lane to a duplicated index's first leaf (#3086) 2026-09-03 17:42:08 -07:00
Shi-Dong 36d1b8856d Dashboard: handle multiple TITO leaves per sample index (#2881) 2026-09-02 13:35:23 -07:00
Shi-Dong 2a7d68bc7f fix(lora): pass config to the bridge value head instead of sequence_parallel (#2869) 2026-09-01 14:32:21 -07:00
Shi-DongandZhichenzzz 542213b962 dashboard: give a step with no samples the summary schema (#2817)
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-08-31 13:32:19 -07:00
Shi-DongandZhichenzzz 9805fecc7f dashboard: land Rollouts on the newest step that has data (#2815)
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-08-31 13:32:10 -07:00
Shi-Dongandyueming-yuan 290feb9e1a fix(megatron): keep SFT logits in model precision (#2818)
Co-authored-by: yueming-yuan <yym022502@gmail.com>
2026-08-30 22:35:58 -07:00
Shi-Dong 721c020bc3 dashboard: scroll the token strip instead of paging through it (#2794) 2026-08-29 00:02:54 -07:00
Shi-Dong e9cf74d1c3 feat(metrics): log per-rollout response lengths (#2765) 2026-08-27 14:54:56 -07:00
Shi-Dong de78d4279b fix(metrics): scope episode rewards by adapter (#2778) 2026-08-27 14:05:11 -07:00
Shi-Dong bfe13117c3 Add Terminus 2 compaction training example (#2741) 2026-08-25 12:44:01 -07:00
Shi-Dong f2b7c79298 Log compaction-aware rollout metrics (#2710) 2026-08-24 19:38:29 -07:00
Shi-Dong b0f8a74737 docs: add serial commas throughout the README (#2593) 2026-08-24 13:31:18 -07:00
Shi-Dong 782c7ad1e1 fix: use wandb.sdk.lib.runid.generate_id for the wandb group suffix (#2566) 2026-08-15 13:12:24 -07:00
Shi-Dong 9fc8571042 docs: delete the FAQ page (#2523) 2026-08-14 22:34:14 -07:00
Shi-Dong d05814b21b docs: quick-start pre-flight checks, disk sizing, and honest timing (#2537) 2026-08-13 21:37:10 -07:00
Shi-Dong 9075e867f9 fix(swe-agent example): stop logging unmeasured agent metrics as zero (#2526) 2026-08-13 15:01:15 -07:00
Shi-Dong bc7c0fa80d dashboard: keep sample status/reward chips visible on the Tokens tab (#2490) 2026-08-12 15:57:19 -07:00
Shi-Dong 4cc536f3b6 Add a PPO example for the shared actor/critic setup (#2367) 2026-08-12 15:20:01 -07:00
Shi-Dong 07ed809957 [example] GLM-5.2 744B-A40B LoRA agentic launcher (TB2 on Daytona) (#2280) 2026-08-12 14:38:27 -07:00
Shi-Dong 79eeaec6de fix(dashboard): zero the trainer log-probs the loss masks out (#2476) 2026-08-12 10:53:10 -07:00
Shi-Dongandguapisolo 63d2190ae2 fix(rollout): normalize rewards per rollout (#2369)
Co-authored-by: guapisolo <guapisolo@gmail.com>
2026-08-12 09:35:42 -07:00
Shi-Dong 95ecc7120c fix(rollout): group session v2 leaf samples (#2368) 2026-08-11 20:39:43 -07:00
Shi-Dong 82916b7ce1 docs(args): correct the offload flags' help text (#2390) 2026-08-11 17:03:57 -07:00
Shi-Dong dc3bf8f58b scripts: download the DAPO dataset the NPU recipe trains on (#2402) 2026-08-11 17:00:25 -07:00
Shi-Dong 8183f2fa0b docs: fix the gsm8k download on the reproducibility page (#2399) 2026-08-11 16:55:46 -07:00
Shi-Dong fd73c3a3dc docs: download the DAPO jsonl mirror the launcher expects (#2398) 2026-08-11 16:41:41 -07:00
Shi-DongandZhichenzzz 0daa9ff926 fix(fsdp): stop store_true from shadowing bool defaults in FSDPArgs (#2384)
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-08-11 15:19:35 -07:00
Shi-Dong 93b77615ff fix: drop duplicated rematerialize validation call (#2382) 2026-08-11 11:42:52 -07:00
Shi-Dong 8058efb77b docs: split FAQ out of Resources and link the blog to LMSYS (#2374) 2026-08-11 11:11:34 -07:00
Shi-Dong 3a6634f850 docs: trim repo README to banner and nav links (#2225) 2026-08-10 15:27:04 -07:00
77ca8abb15 docs: add GLM-5.2 model page and update supported-models tables (#2216)
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-08-10 14:20:05 -07:00
Shi-Dong 7af821e2a0 docs: refresh the homepage supported-models table (#2271) 2026-08-10 13:30:17 -07:00
Shi-Dong 3f09ea7533 docs: polish the Quick Start page (#2298) 2026-08-10 13:05:53 -07:00
Shi-Dong 5e14ec7cf6 docs: update homepage Core features section per the v0.1 feature list (#2264) 2026-08-10 11:55:18 -07:00
Shi-Dong 355cac452d scripts: enable the Miles dashboard in the quick-start launcher (#2300) 2026-08-10 11:27:34 -07:00
Shi-Dong 0113aa5ab8 Fix padded -1 index handling in GLM-5 sparse-MLA tilelang kernels (#2079) 2026-08-07 23:37:20 -07:00
Shi-Dong 03491d0431 fix: preserve routing-replay state around MTP spec creation (#2226) 2026-08-06 15:29:19 -07:00
Shi-Dong 925f3471fc fix: derive --critic-save from --save so PPO critic checkpoints are not silently skipped (#2224) 2026-08-06 15:29:16 -07:00
Shi-Dong 7bca0ac5ff examples: rename swe-agent to swe-agent-harbor-docker (#2233)
The Daytona variant is named swe-agent-harbor-daytona, so the original example being called just swe-agent read as if it were the generic one rather than the Docker-sandbox sibling. Rename the directory for symmetry and sweep every filesystem path that referred to it.

Content-preserving: all six files are byte-identical apart from the README title and its own self-references. docs/docs.json has no swe-agent page, so no Mintlify route changes. mini-swe-agent and swe-agent-harbor-daytona are deliberately untouched.
2026-08-06 15:00:48 -07:00
Shi-Dong 7cfa84260b examples: make the agent-server trial timeout configurable (#2228)
AGENT_TRIAL_TIMEOUT was hard-coded in the SWE-agent rollout function, so it could not be raised in step with the agent server's own --agent-timeout. Expose it as an environment variable and document the ordering constraint: the client-side ceiling must sit above the server-side agent budget, otherwise the client aborts a trial the server would still have graded.
2026-08-06 14:56:52 -07:00
Shi-Dong 41b9ae23d7 examples: add swe-agent-harbor-daytona (Harbor sandboxes on Daytona) (#1919)
Adds examples/experimental/swe-agent-harbor-daytona: a Daytona-backed variant of the Harbor SWE-agent example, for teams that cannot run local Docker sandboxes on the training nodes.

Documentation and a launcher script only; it reuses the existing examples/swe-agent/run.py training entrypoint rather than duplicating it.
2026-08-06 14:53:54 -07:00
Shi-Dong 38c467fbbd Revert "Reject --disable-weights-backuper for LoRA + colocate + offload-train" (#2077) (#2203) 2026-08-04 17:50:25 -07:00
Shi-Dong 7546154083 Reject --disable-weights-backuper for LoRA + colocate + offload-train (#2077) 2026-08-04 15:37:48 -07:00
Shi-Dong 86954a70c6 fix: skip the --dump-details processor dump when it cannot serialise (#2134)
Bespoke processors (e.g. Inkling's) implement only what the rollout path needs
(`extract_media`, `__call__`); they are not `ProcessorMixin` and have no
`save_pretrained`. Any run on that model family with `--dump-details` therefore
died in `RolloutManager.__init__`, which also broke `--use-miles-dashboard`
(it asserts `--dump-details`).

Guard on the method instead of on truthiness. Nothing reads the processor dump
(the dashboard only consumes the tokenizer dump), so skipping it when it cannot
be serialised is safe and leaves the tokenizer dump intact. `hasattr` also
covers the `processor is None` case the previous truthiness check handled.
2026-08-03 23:35:05 -07:00
Shi-Dong 6c5e4cdc5a fix(docs): grad_norm is logged before clipping, not after (#2132) 2026-08-03 14:51:46 -07:00
Shi-Dong 7401900847 scripts, examples: stop forcing the deprecated Miles router in launchers (#2015) 2026-08-03 12:57:29 -07:00
Shi-Dong bda10138e9 session: apply the trained LoRA adapter to session-server rollouts (#2075) 2026-08-02 22:32:50 -07:00
Shi-Dong bc232eb88d fix: require explicit off-policy correction for async PPO training (#1829) 2026-07-28 18:26:36 -07:00
Shi-Dong 3d2e0e411c fix: run PPO GAE over trainable tokens only (#1827) 2026-07-28 18:25:48 -07:00
Shi-Dong 7e436d0dc6 Remove the experimental swe-agent examples (#1918) 2026-07-28 14:53:08 -07:00
Shi-Dong 32cf119f86 Add examples/swe-agent: GLM-4.7-Flash agentic training with Harbor (#1741) 2026-07-28 13:42:25 -07:00
Shi-DongandJiajun Li 587ec4e351 docs: add Claude general code style rule (#1831)
Co-authored-by: Jiajun Li <guapisolo@gmail.com>
2026-07-28 12:45:23 -07:00
Shi-Dong 1c12d6b492 fix: allow TITO session rollback to the empty checkpoint (#1826) 2026-07-28 00:24:00 -07:00
Shi-Dong deb3eef08c fix(session): drop upstream Server/Date so aiohttp clients can read replies (#1828) 2026-07-27 22:38:07 -07:00
Shi-Dong dfc66ff387 docs: fix inaccuracies and expand explanations in Quick Start (#1698) 2026-07-26 22:50:56 -07:00
Shi-Dong 4ccd9ff8cd docs: remove nonexistent --sglang-log-dir flag and /tmp/sglang log path (#1802) 2026-07-26 16:51:30 -07:00
Shi-Dong 31e7fbe2b0 fix(eval): guard eval reward aggregation against None rewards (#1706) 2026-07-18 23:07:20 +08:00
Shi-DongandShi Dong 913633d7cd openenv: TB2 agentic RL adapter + GLM-4.7-Flash launcher (#1487)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-07-17 15:02:18 +08:00
Shi-Dong a6d81cc709 docs: document the optional agent abort hook (#1695) 2026-07-17 09:50:39 +08:00
9d65de1f7f rollout: decouple oversampling-abort teardown into a pluggable agent hook (#1639)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
Co-authored-by: maocheng23 <35615230+maocheng23@users.noreply.github.com>
2026-07-16 13:54:08 +08:00
Shi-DongandShi Dong 6811653228 docs: fix incorrect GLM-5 conversion claim and FSDP2 backend path (#1681)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-07-15 10:52:18 -07:00
Shi-Dong 0f2445fdaf docs: correct cli-reference argument defaults (#1490) 2026-07-15 20:39:09 +08:00
Shi-Dong a472996e32 rollout: subtract agent-reported tool time from throughput accounting (#1663) 2026-07-15 20:30:41 +08:00
Shi-Dong bf5d45a04c fix(rollout): stop merge_samples at routing-replay gap from aborted turns (#1672) 2026-07-15 13:12:54 +08:00
Shi-Dong 254091fd9f Add Shi-Dong to megatron/sglang backends and rollout/session (#1580) 2026-07-06 14:47:18 +09:00
Shi-DongandShi Dong 0d59bda149 Fix: Remove trailing comma from help attribute causing tuple bug (#1503)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-29 09:13:22 -07:00
Shi-Dong e80610bfd0 Remove deprecated --use-miles-router flag from launch scripts (#1480) 2026-06-26 12:02:47 +09:00
Shi-Dong 1a488e437e Remove deprecated --use-miles-router from recipe docs (#1479) 2026-06-26 12:02:18 +09:00
Shi-Dong 138522565a Fix minor docs issues (#1478) 2026-06-25 23:25:57 +09:00
Shi-Dong fc9bf98c2e spawn router/session-server subprocesses instead of forking (#1367) 2026-06-22 18:32:39 +08:00
Shi-Dong 969c2a70b8 [fix] stop merging agentic turns at first non-COMPLETED turn (#1323) 2026-06-13 10:07:08 +08:00
Shi-DongandShi Dong c8b6697f2d Remove eval/terminal_bench example (#1258)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-03 21:54:02 -07:00
Shi-DongandShi Dong 30af8d8be3 ci: fix and re-enable test_run_megatron_worker_main (#1269)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-03 20:39:56 -07:00
Shi-DongandShi Dong 067ebff6e4 Re-enable test_qwen3_4B_ckpt.py (#1271)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-03 17:44:57 -07:00
Shi-DongandShi Dong 1993ac2aad Add @Shi-Dong as code owner for backends, rollout, utils (#1247)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-01 21:37:20 -07:00
Shi-DongandShi Dong be392027bc ci: re-enable test_run_megatron parallelism-equivalence test (#1270)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-06-01 11:19:51 -07:00
Shi-DongandShi Dong 1c529248f9 tau-bench: let user simulator use any litellm provider (e.g. DeepSeek) (#1266)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-05-31 23:30:38 -07:00
Shi-Dong ffdf753324 Fix geo3k_vlm Megatron crash on non-tensor multimodal field (#1265) 2026-05-30 11:11:27 +08:00
Shi-DongandYueming Yuan 98231dea67 Default geo3k_vlm_multi_turn to megatron backend (#1260)
Co-authored-by: Yueming Yuan <yym022502@gmail.com>
2026-05-30 11:11:02 +08:00
Shi-Dong 7deb4b744a Add 2-node TB2 training example targeting the harbor-private agent server (#1236) 2026-05-29 19:18:58 +08:00
Shi-Dong 1bb1b2bb79 Add Terminal-Bench eval driver targeting miles_agent_server (DeepSeek-V4-Pro example) (#1225) 2026-05-29 19:18:32 +08:00
Shi-Dong 7a6cf48e6f add Code of Conduct (#1145) 2026-05-19 08:58:40 +08:00
Shi-DongandShi Dong a8b51c5e02 Point README Documentation link at the Mintlify docs site (#1128)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
2026-05-14 14:42:20 -07:00
cafbd11ac6 docs: fix --rollout-sample-filter-path arg help signature (#1056)
Co-authored-by: Shi Dong <shi.dong@radixark.ai>
Co-authored-by: Zhichenzzz <zczeng@uw.edu>
2026-05-01 23:24:55 -07:00
Shi-Dong 9a6fc9e6f9 docs: fix --rollout-function-path arg help signature (#1055) 2026-04-29 15:38:14 -07:00
Shi-Dong a772c33026 Add zero_std all_zero_ratio and all_one_ratio metrics (#1034) 2026-04-22 22:46:39 -07:00