mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
- Compute each full-vocab normalizer once and free the teacher's fp32 copy before materializing the student's; document the memory/communication cost. - Stop writing back into the removed skip_lm_loss field so configs survive dataclasses.replace()/asdict() round-trips; read kd_loss_alpha == 1.0 instead. - Tie the ghost-token clamp floor to the dtype's eps and document its effects. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>