Shi-Dong
bc6191a6e4
dashboard: show per-family match counts while filtering metrics ( #3696 )
2026-09-24 16:43:48 -07:00
Shi-Dong
5510af6752
Fix trainer-rollout diagnostics with rollout logprobs enabled ( #3655 )
2026-09-24 10:12:47 -07:00
Shi-Dong
4eebfbab3f
Add experimental Nemotron Workplace Assistant RL example ( #3658 )
2026-09-24 00:25:15 -07:00
Shi-Dong
a20947e65c
dashboard: compute chart extents without spreading into Math.min/max ( #3654 )
2026-09-23 21:53:10 -07:00
Shi-Dong
8ec0a119af
fix(args): treat the fully-async rollout path as the mode that selects it ( #3587 )
...
--fully-async selects FullyAsyncRolloutFn, but naming that class through
--rollout-function-path selected the same producer while leaving the mode off, so
the run skipped every --fully-async check (colocate, partial rollout, legacy
rollout v1, pause mode, multi-LoRA) and train.py's async-driver guard.
Normalize the path spelling into the flag before validation runs. Only the exact
class is recognized; a subclass still passes --fully-async explicitly.
2026-09-21 22:29:04 -07:00
Shi-Dong and yueming-yuan
e2a5a3e592
fix(rollout): isolate cancelled groups in fully async generation ( #3319 )
...
Co-authored-by: yueming-yuan <yym022502@gmail.com >
2026-09-18 17:22:57 -07:00
Shi-Dong
97122c1ed1
fix: guard connection allocations during offloaded weight sync ( #3129 )
...
Signed-off-by: Shi Dong <shi.dong@radixark.ai >
2026-09-14 17:42:13 -07:00
9a9c3eb386
fix: align MTP loss masks with input tokens before context-parallel slicing ( #3219 )
...
Co-authored-by: Jiajun Li <jiajun.li@radixark.ai >
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
2026-09-14 13:20:12 -07:00
Shi-Dong
e5c03bda7e
Fix MTP gradient detachment in AutoBridge mode ( #3211 )
2026-09-11 13:26:15 -07:00
Shi-Dong and Zhichenzzz
f01f9a3166
fix: use host weights when syncing an offloaded actor ( #3128 )
...
Signed-off-by: Shi Dong <shi.dong@radixark.ai >
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-09-09 15:29:59 -07:00
Shi-Dong and Zhichenzzz
f1eac9bad4
Fix missing Nemotron-H Mamba convolution mappings ( #3163 )
...
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-09-09 14:43:08 -07:00
Shi-Dong
021d311e89
Fix Nemotron-H forward compatibility when MTP is disabled ( #3164 )
2026-09-09 14:42:53 -07:00
Shi-Dong
d8dd0dec8d
Scan full responses for repetition metrics ( #3115 )
2026-09-09 12:04:50 -07:00
Shi-Dong
d2fc97ce58
Tolerate missing optional SGLang Anthropic helpers ( #3114 )
2026-09-04 14:48:36 -07:00
Shi-Dong
428d8882d0
dashboard: pin the anatomy lane to a duplicated index's first leaf ( #3086 )
2026-09-03 17:42:08 -07:00
Shi-Dong
36d1b8856d
Dashboard: handle multiple TITO leaves per sample index ( #2881 )
2026-09-02 13:35:23 -07:00
Shi-Dong
2a7d68bc7f
fix(lora): pass config to the bridge value head instead of sequence_parallel ( #2869 )
2026-09-01 14:32:21 -07:00
Shi-Dong and Zhichenzzz
542213b962
dashboard: give a step with no samples the summary schema ( #2817 )
...
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-08-31 13:32:19 -07:00
Shi-Dong and Zhichenzzz
9805fecc7f
dashboard: land Rollouts on the newest step that has data ( #2815 )
...
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-08-31 13:32:10 -07:00
Shi-Dong and yueming-yuan
290feb9e1a
fix(megatron): keep SFT logits in model precision ( #2818 )
...
Co-authored-by: yueming-yuan <yym022502@gmail.com >
2026-08-30 22:35:58 -07:00
Shi-Dong
721c020bc3
dashboard: scroll the token strip instead of paging through it ( #2794 )
2026-08-29 00:02:54 -07:00
Shi-Dong
e9cf74d1c3
feat(metrics): log per-rollout response lengths ( #2765 )
2026-08-27 14:54:56 -07:00
Shi-Dong
de78d4279b
fix(metrics): scope episode rewards by adapter ( #2778 )
2026-08-27 14:05:11 -07:00
Shi-Dong
bfe13117c3
Add Terminus 2 compaction training example ( #2741 )
2026-08-25 12:44:01 -07:00
Shi-Dong
f2b7c79298
Log compaction-aware rollout metrics ( #2710 )
2026-08-24 19:38:29 -07:00
Shi-Dong
b0f8a74737
docs: add serial commas throughout the README ( #2593 )
2026-08-24 13:31:18 -07:00
Shi-Dong
782c7ad1e1
fix: use wandb.sdk.lib.runid.generate_id for the wandb group suffix ( #2566 )
2026-08-15 13:12:24 -07:00
Shi-Dong
9fc8571042
docs: delete the FAQ page ( #2523 )
2026-08-14 22:34:14 -07:00
Shi-Dong
d05814b21b
docs: quick-start pre-flight checks, disk sizing, and honest timing ( #2537 )
2026-08-13 21:37:10 -07:00
Shi-Dong
9075e867f9
fix(swe-agent example): stop logging unmeasured agent metrics as zero ( #2526 )
2026-08-13 15:01:15 -07:00
Shi-Dong
bc7c0fa80d
dashboard: keep sample status/reward chips visible on the Tokens tab ( #2490 )
2026-08-12 15:57:19 -07:00
Shi-Dong
4cc536f3b6
Add a PPO example for the shared actor/critic setup ( #2367 )
2026-08-12 15:20:01 -07:00
Shi-Dong
07ed809957
[example] GLM-5.2 744B-A40B LoRA agentic launcher (TB2 on Daytona) ( #2280 )
2026-08-12 14:38:27 -07:00
Shi-Dong
79eeaec6de
fix(dashboard): zero the trainer log-probs the loss masks out ( #2476 )
2026-08-12 10:53:10 -07:00
Shi-Dong and guapisolo
63d2190ae2
fix(rollout): normalize rewards per rollout ( #2369 )
...
Co-authored-by: guapisolo <guapisolo@gmail.com >
2026-08-12 09:35:42 -07:00
Shi-Dong
95ecc7120c
fix(rollout): group session v2 leaf samples ( #2368 )
2026-08-11 20:39:43 -07:00
Shi-Dong
82916b7ce1
docs(args): correct the offload flags' help text ( #2390 )
2026-08-11 17:03:57 -07:00
Shi-Dong
dc3bf8f58b
scripts: download the DAPO dataset the NPU recipe trains on ( #2402 )
2026-08-11 17:00:25 -07:00
Shi-Dong
8183f2fa0b
docs: fix the gsm8k download on the reproducibility page ( #2399 )
2026-08-11 16:55:46 -07:00
Shi-Dong
fd73c3a3dc
docs: download the DAPO jsonl mirror the launcher expects ( #2398 )
2026-08-11 16:41:41 -07:00
Shi-Dong and Zhichenzzz
0daa9ff926
fix(fsdp): stop store_true from shadowing bool defaults in FSDPArgs ( #2384 )
...
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-08-11 15:19:35 -07:00
Shi-Dong
93b77615ff
fix: drop duplicated rematerialize validation call ( #2382 )
2026-08-11 11:42:52 -07:00
Shi-Dong
8058efb77b
docs: split FAQ out of Resources and link the blog to LMSYS ( #2374 )
2026-08-11 11:11:34 -07:00
Shi-Dong
3a6634f850
docs: trim repo README to banner and nav links ( #2225 )
2026-08-10 15:27:04 -07:00
77ca8abb15
docs: add GLM-5.2 model page and update supported-models tables ( #2216 )
...
Co-authored-by: yueming-yuan <yym022502@gmail.com >
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-08-10 14:20:05 -07:00
Shi-Dong
7af821e2a0
docs: refresh the homepage supported-models table ( #2271 )
2026-08-10 13:30:17 -07:00
Shi-Dong
3f09ea7533
docs: polish the Quick Start page ( #2298 )
2026-08-10 13:05:53 -07:00
Shi-Dong
5e14ec7cf6
docs: update homepage Core features section per the v0.1 feature list ( #2264 )
2026-08-10 11:55:18 -07:00
Shi-Dong
355cac452d
scripts: enable the Miles dashboard in the quick-start launcher ( #2300 )
2026-08-10 11:27:34 -07:00
Shi-Dong
0113aa5ab8
Fix padded -1 index handling in GLM-5 sparse-MLA tilelang kernels ( #2079 )
2026-08-07 23:37:20 -07:00
Shi-Dong
03491d0431
fix: preserve routing-replay state around MTP spec creation ( #2226 )
2026-08-06 15:29:19 -07:00
Shi-Dong
925f3471fc
fix: derive --critic-save from --save so PPO critic checkpoints are not silently skipped ( #2224 )
2026-08-06 15:29:16 -07:00
Shi-Dong
7bca0ac5ff
examples: rename swe-agent to swe-agent-harbor-docker ( #2233 )
...
The Daytona variant is named swe-agent-harbor-daytona, so the original example being called just swe-agent read as if it were the generic one rather than the Docker-sandbox sibling. Rename the directory for symmetry and sweep every filesystem path that referred to it.
Content-preserving: all six files are byte-identical apart from the README title and its own self-references. docs/docs.json has no swe-agent page, so no Mintlify route changes. mini-swe-agent and swe-agent-harbor-daytona are deliberately untouched.
2026-08-06 15:00:48 -07:00
Shi-Dong
7cfa84260b
examples: make the agent-server trial timeout configurable ( #2228 )
...
AGENT_TRIAL_TIMEOUT was hard-coded in the SWE-agent rollout function, so it could not be raised in step with the agent server's own --agent-timeout. Expose it as an environment variable and document the ordering constraint: the client-side ceiling must sit above the server-side agent budget, otherwise the client aborts a trial the server would still have graded.
2026-08-06 14:56:52 -07:00
Shi-Dong
41b9ae23d7
examples: add swe-agent-harbor-daytona (Harbor sandboxes on Daytona) ( #1919 )
...
Adds examples/experimental/swe-agent-harbor-daytona: a Daytona-backed variant of the Harbor SWE-agent example, for teams that cannot run local Docker sandboxes on the training nodes.
Documentation and a launcher script only; it reuses the existing examples/swe-agent/run.py training entrypoint rather than duplicating it.
2026-08-06 14:53:54 -07:00
Shi-Dong
38c467fbbd
Revert "Reject --disable-weights-backuper for LoRA + colocate + offload-train" ( #2077 ) ( #2203 )
2026-08-04 17:50:25 -07:00
Shi-Dong
7546154083
Reject --disable-weights-backuper for LoRA + colocate + offload-train ( #2077 )
2026-08-04 15:37:48 -07:00
Shi-Dong
86954a70c6
fix: skip the --dump-details processor dump when it cannot serialise ( #2134 )
...
Bespoke processors (e.g. Inkling's) implement only what the rollout path needs
(`extract_media`, `__call__`); they are not `ProcessorMixin` and have no
`save_pretrained`. Any run on that model family with `--dump-details` therefore
died in `RolloutManager.__init__`, which also broke `--use-miles-dashboard`
(it asserts `--dump-details`).
Guard on the method instead of on truthiness. Nothing reads the processor dump
(the dashboard only consumes the tokenizer dump), so skipping it when it cannot
be serialised is safe and leaves the tokenizer dump intact. `hasattr` also
covers the `processor is None` case the previous truthiness check handled.
2026-08-03 23:35:05 -07:00
Shi-Dong
6c5e4cdc5a
fix(docs): grad_norm is logged before clipping, not after ( #2132 )
2026-08-03 14:51:46 -07:00
Shi-Dong
7401900847
scripts, examples: stop forcing the deprecated Miles router in launchers ( #2015 )
2026-08-03 12:57:29 -07:00
Shi-Dong
bda10138e9
session: apply the trained LoRA adapter to session-server rollouts ( #2075 )
2026-08-02 22:32:50 -07:00
Shi-Dong
bc232eb88d
fix: require explicit off-policy correction for async PPO training ( #1829 )
2026-07-28 18:26:36 -07:00
Shi-Dong
3d2e0e411c
fix: run PPO GAE over trainable tokens only ( #1827 )
2026-07-28 18:25:48 -07:00
Shi-Dong
7e436d0dc6
Remove the experimental swe-agent examples ( #1918 )
2026-07-28 14:53:08 -07:00
Shi-Dong
32cf119f86
Add examples/swe-agent: GLM-4.7-Flash agentic training with Harbor ( #1741 )
2026-07-28 13:42:25 -07:00
Shi-Dong and Jiajun Li
587ec4e351
docs: add Claude general code style rule ( #1831 )
...
Co-authored-by: Jiajun Li <guapisolo@gmail.com >
2026-07-28 12:45:23 -07:00
Shi-Dong
1c12d6b492
fix: allow TITO session rollback to the empty checkpoint ( #1826 )
2026-07-28 00:24:00 -07:00
Shi-Dong
deb3eef08c
fix(session): drop upstream Server/Date so aiohttp clients can read replies ( #1828 )
2026-07-27 22:38:07 -07:00
Shi-Dong
dfc66ff387
docs: fix inaccuracies and expand explanations in Quick Start ( #1698 )
2026-07-26 22:50:56 -07:00
Shi-Dong
4ccd9ff8cd
docs: remove nonexistent --sglang-log-dir flag and /tmp/sglang log path ( #1802 )
2026-07-26 16:51:30 -07:00
Shi-Dong
31e7fbe2b0
fix(eval): guard eval reward aggregation against None rewards ( #1706 )
2026-07-18 23:07:20 +08:00
Shi-Dong and Shi Dong
913633d7cd
openenv: TB2 agentic RL adapter + GLM-4.7-Flash launcher ( #1487 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-07-17 15:02:18 +08:00
Shi-Dong
a6d81cc709
docs: document the optional agent abort hook ( #1695 )
2026-07-17 09:50:39 +08:00
9d65de1f7f
rollout: decouple oversampling-abort teardown into a pluggable agent hook ( #1639 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
Co-authored-by: maocheng23 <35615230+maocheng23@users.noreply.github.com >
2026-07-16 13:54:08 +08:00
Shi-Dong and Shi Dong
6811653228
docs: fix incorrect GLM-5 conversion claim and FSDP2 backend path ( #1681 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-07-15 10:52:18 -07:00
Shi-Dong
0f2445fdaf
docs: correct cli-reference argument defaults ( #1490 )
2026-07-15 20:39:09 +08:00
Shi-Dong
a472996e32
rollout: subtract agent-reported tool time from throughput accounting ( #1663 )
2026-07-15 20:30:41 +08:00
Shi-Dong
bf5d45a04c
fix(rollout): stop merge_samples at routing-replay gap from aborted turns ( #1672 )
2026-07-15 13:12:54 +08:00
Shi-Dong
254091fd9f
Add Shi-Dong to megatron/sglang backends and rollout/session ( #1580 )
2026-07-06 14:47:18 +09:00
Shi-Dong and Shi Dong
0d59bda149
Fix: Remove trailing comma from help attribute causing tuple bug ( #1503 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-29 09:13:22 -07:00
Shi-Dong
e80610bfd0
Remove deprecated --use-miles-router flag from launch scripts ( #1480 )
2026-06-26 12:02:47 +09:00
Shi-Dong
1a488e437e
Remove deprecated --use-miles-router from recipe docs ( #1479 )
2026-06-26 12:02:18 +09:00
Shi-Dong
138522565a
Fix minor docs issues ( #1478 )
2026-06-25 23:25:57 +09:00
Shi-Dong
fc9bf98c2e
spawn router/session-server subprocesses instead of forking ( #1367 )
2026-06-22 18:32:39 +08:00
Shi-Dong
969c2a70b8
[fix] stop merging agentic turns at first non-COMPLETED turn ( #1323 )
2026-06-13 10:07:08 +08:00
Shi-Dong and Shi Dong
c8b6697f2d
Remove eval/terminal_bench example ( #1258 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-03 21:54:02 -07:00
Shi-Dong and Shi Dong
30af8d8be3
ci: fix and re-enable test_run_megatron_worker_main ( #1269 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-03 20:39:56 -07:00
Shi-Dong and Shi Dong
067ebff6e4
Re-enable test_qwen3_4B_ckpt.py ( #1271 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-03 17:44:57 -07:00
Shi-Dong and Shi Dong
1993ac2aad
Add @Shi-Dong as code owner for backends, rollout, utils ( #1247 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-01 21:37:20 -07:00
Shi-Dong and Shi Dong
be392027bc
ci: re-enable test_run_megatron parallelism-equivalence test ( #1270 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-06-01 11:19:51 -07:00
Shi-Dong and Shi Dong
1c529248f9
tau-bench: let user simulator use any litellm provider (e.g. DeepSeek) ( #1266 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-05-31 23:30:38 -07:00
Shi-Dong
ffdf753324
Fix geo3k_vlm Megatron crash on non-tensor multimodal field ( #1265 )
2026-05-30 11:11:27 +08:00
Shi-Dong and Yueming Yuan
98231dea67
Default geo3k_vlm_multi_turn to megatron backend ( #1260 )
...
Co-authored-by: Yueming Yuan <yym022502@gmail.com >
2026-05-30 11:11:02 +08:00
Shi-Dong
7deb4b744a
Add 2-node TB2 training example targeting the harbor-private agent server ( #1236 )
2026-05-29 19:18:58 +08:00
Shi-Dong
1bb1b2bb79
Add Terminal-Bench eval driver targeting miles_agent_server (DeepSeek-V4-Pro example) ( #1225 )
2026-05-29 19:18:32 +08:00
Shi-Dong
7a6cf48e6f
add Code of Conduct ( #1145 )
2026-05-19 08:58:40 +08:00
Shi-Dong and Shi Dong
a8b51c5e02
Point README Documentation link at the Mintlify docs site ( #1128 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
2026-05-14 14:42:20 -07:00
cafbd11ac6
docs: fix --rollout-sample-filter-path arg help signature ( #1056 )
...
Co-authored-by: Shi Dong <shi.dong@radixark.ai >
Co-authored-by: Zhichenzzz <zczeng@uw.edu >
2026-05-01 23:24:55 -07:00
Shi-Dong
9a6fc9e6f9
docs: fix --rollout-function-path arg help signature ( #1055 )
2026-04-29 15:38:14 -07:00
Shi-Dong
a772c33026
Add zero_std all_zero_ratio and all_one_ratio metrics ( #1034 )
2026-04-22 22:46:39 -07:00