From 9038b71f09468dd23e39781ab046221edd2fa036 Mon Sep 17 00:00:00 2001 From: Jenny Chen Date: Wed, 1 Jul 2026 22:12:06 -0400 Subject: [PATCH] Autoquant and GPTQ in support in Megatron-Core [OMNIML-3151] (#1562) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ### What does this PR do? Type of change: New Feature Autoquant and GPTQ in support in Megatron-Core - Add EP support to AutoQuantize - Register MCore support in AutoQuantize - Add decoder `output_layer` (lm head) to layerwise hook so that GPTQ can register all decoder layers & lm head - Split dataloader helper function out of megatron calibration utils so that AutoQuantize in Megatron-LM can reuse the same dataloader ### Usage See https://github.com/NVIDIA/Megatron-LM/pull/4821 for Autoquant usage in Megatron ```python # For GPTQ pick a recipe that uses gptq algorithm and run mtq.quantize # e.g. general/ptq/nvfp4_default-kv_none-gptq ``` ### Testing Tested AutoQuant on Nemotron Nano and Ultra. Tested GPTQ on Nano 3. Added unit tests for both AutoQuant and GPTQ ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A - Did you write any new necessary tests?: ✅ / ❌ / N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A - Did you get Claude approval on this PR?: ✅ / ❌ / N/A ### Additional Information ## Summary by CodeRabbit ## Summary by CodeRabbit * **New Features** * Added lazy Megatron-Core AutoQuant integration with Megatron-specific quantization hooks and better decoder-layer discovery for layerwise calibration. * Improved AutoQuantize for expert-parallel (EP) models, including consistent per-layer recipe selection across DP/TP/EP. * Extended quant-layer grouping for NemotronH MCore fused “local_experts” linear layers. * **Bug Fixes** * Prevented division-by-zero when calibration inputs are empty during Hessian updates. * **Tests** * Added/extended unit and GPU coverage for EP AutoQuant, decoder-layer calibration discovery behavior, and zero-token Hessian no-op. --------- Signed-off-by: Jennifer Chen Signed-off-by: Jenny Chen Co-authored-by: Claude Opus 4.8 (1M context) --- .github/workflows/gpu_tests.yml | 2 +- CHANGELOG.rst | 1 + .../torch/quantization/_auto_quantize_cost.py | 15 +- modelopt/torch/quantization/algorithms.py | 44 ++- modelopt/torch/quantization/model_calib.py | 15 +- .../torch/quantization/plugins/megatron.py | 57 +++- .../torch/quantization/utils/calib_utils.py | 3 + .../utils/plugins/megatron_calibration.py | 75 +++-- tests/conftest.py | 2 +- tests/gpu/torch/quantization/test_gptq.py | 83 ------ tests/gpu_megatron/conftest.py | 18 ++ .../quantization/plugins/test_megatron.py | 260 +++++++++++++++++- .../unit/torch/quantization/test_autoquant.py | 78 ++++-- tests/unit/torch/quantization/test_gptq.py | 119 ++++++++ ...tron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml | 46 ++++ tools/launcher/modules/Megatron-LM | 2 +- 16 files changed, 637 insertions(+), 183 deletions(-) create mode 100644 tests/unit/torch/quantization/test_gptq.py create mode 100644 tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml diff --git a/.github/workflows/gpu_tests.yml b/.github/workflows/gpu_tests.yml index 34dac8ba4..b79485f78 100644 --- a/.github/workflows/gpu_tests.yml +++ b/.github/workflows/gpu_tests.yml @@ -42,7 +42,7 @@ jobs: timeout: 60 container_image: nvcr.io/nvidia/pytorch:26.05-py3 - example: gpu_megatron - timeout: 60 + timeout: 75 container_image: nvcr.io/nvidia/nemo:26.06 - example: gpu_trtllm timeout: 15 diff --git a/CHANGELOG.rst b/CHANGELOG.rst index 7960dac38..500291817 100755 --- a/CHANGELOG.rst +++ b/CHANGELOG.rst @@ -77,6 +77,7 @@ Changelog - Add shared Megatron-Core calibration forward loop: ``modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`` produces the ``forward_loop`` callable expected by ``mtq.quantize`` / ``mtp.prune``. Replaces the bespoke calibration loops in Megatron-LM and Megatron-Bridge for quantization and pruning with a single canonical implementation. - Support Megatron-Core checkpoint restore and export for MSE ``NVFP4StaticQuantizer``. - Add mixed-precision FP8 + NVFP4 export for Megatron-Core: per-layer ``quant_algo`` recorded under ``quantized_layers`` in ``hf_quant_config.json``, PP-aware ``kv_cache_dtype`` gather, fused-QKV exclude split into per-HF-name ``q/k/v_proj`` entries. +- Add AutoQuant and GPTQ support for Megatron-Core models, including MCore-specific AutoQuant hooks and decoder-layer discovery for GPTQ layerwise calibration. - Add support for ``active_params`` (for MoE models) and ``memory_mb`` constraints in Minitron pruning on top of existing ``params`` constraint. You can also provide multiple constraints. See `examples/pruning/README.md `_ for more details. The underlying utility functions ``mcore_param_count``, ``mcore_memory_footprint_mb``, and ``print_mcore_model_stats`` in ``modelopt.torch.nas.plugins.megatron_model_stats`` are also available for standalone use to compute parameter counts and memory footprints (weights + KV-cache + Mamba state) for any Megatron-Core model. - Add Minitron pruning support for Megatron-Bridge Gemma3 models. - Add end-to-end optimization tutorial for Minitron pruning + two-phase distillation (80B @ 8K + 20B @ 32K long-context = 100B tokens) + FP8 PTQ + vLLM deployment for Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) → Pruned 22B/A3.0B active params, along with data blend preparation steps (with tool-calling data) and detailed pruning / data-blend / long-context ablations. See `examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md `_ for details. diff --git a/modelopt/torch/quantization/_auto_quantize_cost.py b/modelopt/torch/quantization/_auto_quantize_cost.py index 68f937708..297e8cd8d 100644 --- a/modelopt/torch/quantization/_auto_quantize_cost.py +++ b/modelopt/torch/quantization/_auto_quantize_cost.py @@ -16,7 +16,7 @@ """Cost models for AutoQuantize effective-bits accounting.""" import fnmatch -from collections.abc import Callable, Iterable, Sequence +from collections.abc import Sequence from typing import Any, Final import regex as re @@ -156,19 +156,6 @@ class AutoQuantizeCostModel: return 0.0 return 1.0 - def total_weight_size( - self, - named_modules: Iterable[tuple[str, nn.Module]], - is_auto_quantize_module: Callable[[nn.Module], bool], - cost_constraints: dict[str, Any], - ) -> float: - """Return the cost denominator for the effective-bits constraint.""" - return sum( - _get_module_weight_numel(module) * self.module_cost_weight([name], cost_constraints) - for name, module in named_modules - if is_auto_quantize_module(module) - ) - class WeightCostModel(AutoQuantizeCostModel): """Count all quantizable weights equally.""" diff --git a/modelopt/torch/quantization/algorithms.py b/modelopt/torch/quantization/algorithms.py index 2e00dc03b..a48655e1f 100644 --- a/modelopt/torch/quantization/algorithms.py +++ b/modelopt/torch/quantization/algorithms.py @@ -378,13 +378,14 @@ class QuantRecipeHparam(Hparam): total_score += importance.cpu().item() continue - if parallel_state.expert_model_parallel_group.is_initialized(): - # TODO: Support expert model parallelism for score estimation - warnings.warn("AutoQuantize does not support expert model parallelism yet.") importance = importance.cpu() importance = DistributedProcessGroup.get_dist_syncd_obj( importance, - [parallel_state.tensor_parallel_group, parallel_state.data_parallel_group], + [ + parallel_state.tensor_parallel_group, + parallel_state.data_parallel_group, + parallel_state.expert_model_parallel_group, + ], sum, ) total_score += importance.item() @@ -408,13 +409,12 @@ class QuantRecipeHparam(Hparam): cost += weight_size * recipe.compression continue - if parallel_state.expert_model_parallel_group.is_initialized(): - # TODO: Support expert model parallelism - warnings.warn("AutoQuantize does not support expert model parallelism yet.") - weight_size = DistributedProcessGroup.get_dist_syncd_obj( weight_size, - [parallel_state.tensor_parallel_group], + [ + parallel_state.tensor_parallel_group, + parallel_state.expert_model_parallel_group, + ], sum, ) @@ -466,6 +466,8 @@ class _AutoQuantizeBaseSearcher(BaseSearcher, ABC): # gate_proj, up_proj, down_proj for Qwen3 like MoE models r"^(.*?\.mlp\.experts)\.\d+\.(gate_proj|up_proj|down_proj)$", r"^(.*?\.mixer\.experts)\.\d+\.(up_proj|down_proj)$", # NemotronH MoE experts + # NemotronH MoE experts in MCore naming (linear_fc1=gate+up fused, linear_fc2=down) + r"^(.*?\.mlp\.experts\.local_experts)\.\d+\.(linear_fc1|linear_fc2)$", r"^(.*?)\.(gate_proj|up_proj)$", # gate_proj, up_proj for llama like models r"^(.*?)\.(\d+\.(w1|w2|w3))$", # mixtral experts r"^(.*?)\.((w1_linear|w2_linear|w3_linear)\.\d+)$", # dbrx experts @@ -875,6 +877,15 @@ class _AutoQuantizeBaseSearcher(BaseSearcher, ABC): for module in modules ) + @staticmethod + def _get_total_weight_size_from_candidate_stats(candidate_stats): + no_quant_recipe = QuantRecipe(quant_cfg=None) + total_weight_size = 0 + for candidate_stat in candidate_stats.values(): + no_quant_idx = candidate_stat["formats"].index(no_quant_recipe) + total_weight_size += candidate_stat["costs"][no_quant_idx] + return total_weight_size + def _get_constraints_for_search(self, max_weight_size, lower_bound=None): constraints = { "weight_size_after_compression": ( @@ -905,9 +916,10 @@ class _AutoQuantizeBaseSearcher(BaseSearcher, ABC): ) compression = self._get_formatted_weight_compression_constraint() - total_weight_size = self._cost_model.total_weight_size( - self.model.named_modules(), self._is_auto_quantize_module, self.config["cost"] + assert self.candidate_stats, ( + "candidate_stats must be populated by before_search() before run_search()" ) + total_weight_size = self._get_total_weight_size_from_candidate_stats(self.candidate_stats) self.cost_denominator = total_weight_size max_weight_size = total_weight_size * compression if verbose: @@ -928,12 +940,16 @@ class _AutoQuantizeBaseSearcher(BaseSearcher, ABC): best_recipe = {} best_constraints, best_scores = 0, 0 for name, best_hparam_recipe_info in best_recipe_info.items(): - # Solvers could give different solutions for the same layer across DP/TP groups even though - # the scores and costs are the same. Lets make sure the same recipe is selected across DP/TP + # Solvers could give different solutions for the same layer across DP/TP/EP groups even though + # the scores and costs are the same. Lets make sure the same recipe is selected across DP/TP/EP _ps = self.model.get_submodule(name.split(".quant_recipe")[0]).parallel_state best_format = DistributedProcessGroup.get_dist_syncd_obj( best_hparam_recipe_info["format"], - [_ps.data_parallel_group, _ps.tensor_parallel_group], + [ + _ps.data_parallel_group, + _ps.tensor_parallel_group, + _ps.expert_model_parallel_group, + ], lambda a: a[0], ) diff --git a/modelopt/torch/quantization/model_calib.py b/modelopt/torch/quantization/model_calib.py index 249841797..661f429b2 100644 --- a/modelopt/torch/quantization/model_calib.py +++ b/modelopt/torch/quantization/model_calib.py @@ -57,14 +57,6 @@ from .utils import ( ) from .utils.calib_utils import _GPTQ_HELPER_REGISTRY, GPTQHelper -try: - from .plugins.megatron import _check_nvfp4_static_tp_supported -except ImportError: - - def _check_nvfp4_static_tp_supported(model: nn.Module) -> None: # no-op without megatron - return - - __all__ = [ "CalibratorFactory", "awq", @@ -305,7 +297,12 @@ def max_calibrate( module.layer_sync_moe_local_experts_amax(sync_weight_amax=sync_expert_weight_amax) # Fail fast on NVFP4 static-block with TP>1 (sharded_state_dict treats _amax as replicated). - _check_nvfp4_static_tp_supported(model) + try: + from .plugins.megatron import _check_nvfp4_static_tp_supported + except ImportError: + pass + else: + _check_nvfp4_static_tp_supported(model) if not distributed_sync: # Single-process: _amax is final. diff --git a/modelopt/torch/quantization/plugins/megatron.py b/modelopt/torch/quantization/plugins/megatron.py index e18e3cd06..752dd801a 100644 --- a/modelopt/torch/quantization/plugins/megatron.py +++ b/modelopt/torch/quantization/plugins/megatron.py @@ -16,6 +16,7 @@ """Support quantization for megatron linear layers.""" import types +from contextlib import contextmanager from typing import Any import megatron.core.parallel_state as mcore_parallel @@ -39,11 +40,13 @@ from modelopt.torch.opt.plugins.megatron import ( from modelopt.torch.utils import warn_rank_0 from modelopt.torch.utils.distributed import ParallelState +from ..algorithms import AutoQuantizeGradientSearcher from ..conversion import maybe_promote_nvfp4_static_quantizer from ..nn import QuantModule, QuantModuleRegistry, SequentialQuantizer, TensorQuantizer from ..nn.modules.quant_linear import RealQuantLinear from ..qtensor import QTensorWrapper from ..utils import sync_moe_expert_amax +from ..utils.layerwise_calib import LayerActivationCollector from .custom import CUSTOM_MODEL_PLUGINS, _ParallelLinear try: @@ -612,10 +615,11 @@ class _MegatronSequentialMLP(DynamicModule): expert_model_parallel_group=mcore_parallel.get_expert_model_parallel_group(), ) - # Initialize parallel state for submodules local_experts.*.linear_fc1 and local_experts.*.linear_fc2 + # These child linears are still native MCore modules here. Seed `_parallel_state` + # directly so the later QuantModule conversion sees the intended parallel state. for expert in self.local_experts: - expert.linear_fc1.parallel_state = self.parallel_state - expert.linear_fc2.parallel_state = self.parallel_state + expert.linear_fc1._parallel_state = self.parallel_state + expert.linear_fc2._parallel_state = self.parallel_state def layer_sync_moe_local_experts_amax(self, sync_weight_amax=False): """Sync quantizer amax across local experts in a SequentialMLP. @@ -721,9 +725,10 @@ if HAS_TE: tensor_parallel_group=mcore_parallel.get_expert_tensor_parallel_group(), expert_model_parallel_group=mcore_parallel.get_expert_model_parallel_group(), ) - # initialize parallel state for submodules linear_fc1 and linear_fc2 - self.linear_fc1.parallel_state = self.parallel_state - self.linear_fc2.parallel_state = self.parallel_state + # These child linears are still native MCore modules here. Seed `_parallel_state` + # directly so the later QuantModule conversion sees the intended parallel state. + self.linear_fc1._parallel_state = self.parallel_state + self.linear_fc2._parallel_state = self.parallel_state @QuantModuleRegistry.register({TEDotProductAttention: "TEDotProductAttention"}) class _QuantTEDotProductAttention(QuantModule): @@ -779,3 +784,43 @@ if HAS_TE: # Affine KVCache Quant bias vector. state_dict = self.state_dict(prefix="", keep_vars=True) return make_sharded_tensors_for_checkpoint(state_dict, prefix, {}, sharded_offsets) + + +def _is_supported_megatron_model(model: torch.nn.Module) -> bool: + return isinstance(model, MegatronModule) + + +@contextmanager +def _megatron_grad_ckpt_context(model: torch.nn.Module): + # Megatron configures activation recompute at model build time via TransformerConfig, + # so there is no runtime flag to flip here. + yield + + +def _is_param_grad_enabled_for_megatron(pname: str, model: torch.nn.Module) -> bool: + return "weight" in pname + + +AutoQuantizeGradientSearcher.register_custom_support( + _is_supported_megatron_model, + _megatron_grad_ckpt_context, + _is_param_grad_enabled_for_megatron, +) + + +def get_mcore_layerwise_calibration_layers( + model: torch.nn.Module, +) -> list[torch.nn.Module] | torch.nn.ModuleList | None: + if not hasattr(model, "decoder") or not hasattr(model.decoder, "layers"): + return None + decoder_layers = model.decoder.layers + if getattr(model, "output_layer", None) is None: + return decoder_layers + layers = list(decoder_layers) + layers.append(model.output_layer) + return layers + + +LayerActivationCollector.register_decoder_layer_support( + _is_supported_megatron_model, get_mcore_layerwise_calibration_layers +) diff --git a/modelopt/torch/quantization/utils/calib_utils.py b/modelopt/torch/quantization/utils/calib_utils.py index aadfe40d2..2e4c3160b 100644 --- a/modelopt/torch/quantization/utils/calib_utils.py +++ b/modelopt/torch/quantization/utils/calib_utils.py @@ -63,6 +63,9 @@ def update_hessian(input, hessian, n_samples): input_flat = input.reshape(-1, input.shape[-1]).t().float() batch_size = input_flat.shape[1] + if batch_size == 0: # in MOEs some experts receive no tokens + return hessian, n_samples + # Incremental averaging: scale down old hessian hessian *= n_samples / (n_samples + batch_size) n_samples += batch_size diff --git a/modelopt/torch/utils/plugins/megatron_calibration.py b/modelopt/torch/utils/plugins/megatron_calibration.py index fd0d6ff5c..4da388582 100644 --- a/modelopt/torch/utils/plugins/megatron_calibration.py +++ b/modelopt/torch/utils/plugins/megatron_calibration.py @@ -34,11 +34,57 @@ if TYPE_CHECKING: from transformers import PreTrainedTokenizerBase, ProcessorMixin __all__ = [ + "get_megatron_calibration_dataloader", "get_megatron_calibration_forward_loop", "get_megatron_vlm_calibration_forward_loop", ] +def get_megatron_calibration_dataloader( + tokenizer: "PreTrainedTokenizerBase", + *, + dataset_name: str | list[str] = "cnn_dailymail", + batch_size: int = 1, + num_samples: int | list[int] = 512, + seq_length: int = 512, + device: torch.device | str | None = "cuda", + apply_chat_template: bool = True, + pack: bool = False, +) -> torch.utils.data.DataLoader: + """Build a DP-sharded calibration dataloader for Megatron-Core models. + + Each batch is a dict with at least ``input_ids`` and ``attention_mask`` tensors + on ``device``. The dataloader is suitable as the ``data_loader`` argument to + ``mtq.auto_quantize`` or any other API that iterates batches directly. + + All kwargs are forwarded to :func:`get_dataset_dataloader`; ``seq_length`` + maps to that function's ``max_sample_length``. + """ + # Deepcopy before mutating pad_token so the caller's tokenizer isn't silently changed. + if getattr(tokenizer, "pad_token", None) is None: + tokenizer = copy.deepcopy(tokenizer) + tokenizer.pad_token = tokenizer.eos_token + + # Shard calibration data across DP ranks; amax is max-reduced across DP inside ``mtq``. + dp_size = mpu.get_data_parallel_world_size() + return get_dataset_dataloader( + dataset_name=dataset_name, + tokenizer=tokenizer, + batch_size=batch_size, + num_samples=num_samples, + max_sample_length=seq_length, + device=device, + apply_chat_template=apply_chat_template, + pack=pack, + distributed=dp_size > 1, + sampler_kwargs={ + "num_replicas": dp_size, + "rank": mpu.get_data_parallel_rank(), + "shuffle": False, + }, + ) + + def get_megatron_calibration_forward_loop( tokenizer: "PreTrainedTokenizerBase", *, @@ -52,37 +98,24 @@ def get_megatron_calibration_forward_loop( ) -> Callable[[torch.nn.Module], None]: """Build a Megatron-Core calibration ``forward_loop(model)``. - Iterates a packed dataloader built via ``get_dataset_dataloader(pack=True)`` + Iterates a dataloader built via :func:`get_megatron_calibration_dataloader` and drives a logits-free prefill pass through the model so activation hooks - fire on every layer. All kwargs except ``seq_length`` are forwarded - 1:1 — see :func:`get_dataset_dataloader` for their semantics. ``seq_length`` - maps to that function's ``max_sample_length``. + fire on every layer. All kwargs are forwarded 1:1 — see + :func:`get_megatron_calibration_dataloader` for their semantics. Returns: - A ``forward_loop(model)`` callable to pass into ``mtq.quantize``, ``mtp.prune``, or other such APIs. + A ``forward_loop(model)`` callable to pass into ``mtq.quantize``, + ``mtp.prune``, or other such APIs. """ - # Deepcopy before mutating pad_token so the caller's tokenizer isn't silently changed. - if getattr(tokenizer, "pad_token", None) is None: - tokenizer = copy.deepcopy(tokenizer) - tokenizer.pad_token = tokenizer.eos_token - - # Shard calibration data across DP ranks; amax is max-reduced across DP inside ``mtq``. - dp_size = mpu.get_data_parallel_world_size() - dataloader = get_dataset_dataloader( + dataloader = get_megatron_calibration_dataloader( + tokenizer, dataset_name=dataset_name, - tokenizer=tokenizer, batch_size=batch_size, num_samples=num_samples, - max_sample_length=seq_length, + seq_length=seq_length, device=device, apply_chat_template=apply_chat_template, pack=pack, - distributed=dp_size > 1, - sampler_kwargs={ - "num_replicas": dp_size, - "rank": mpu.get_data_parallel_rank(), - "shuffle": False, - }, ) def _forward_loop(model: torch.nn.Module) -> None: diff --git a/tests/conftest.py b/tests/conftest.py index 16a9f3f26..370e606a0 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -54,7 +54,7 @@ _DEFAULT_TIMEOUT = { "gpu_trtllm": 60, "gpu_vllm": 60, "regression": 180, - "unit": 60, + "unit": 120 if platform.system() == "Windows" else 60, } diff --git a/tests/gpu/torch/quantization/test_gptq.py b/tests/gpu/torch/quantization/test_gptq.py index d1d8c0c23..0a1849544 100644 --- a/tests/gpu/torch/quantization/test_gptq.py +++ b/tests/gpu/torch/quantization/test_gptq.py @@ -37,89 +37,6 @@ RAND_SEED = 42 torch.manual_seed(RAND_SEED) -def test_update_hessian(): - """Test for update_hessian function with both random and known inputs.""" - # Test 1: Random input - general functionality test - torch.manual_seed(42) - batch_size = 2 - seq_len = 3 - features = 4 - input_tensor = torch.randn(batch_size, seq_len, features, dtype=torch.float32) - - hessian = torch.zeros(features, features, dtype=torch.float32) - n_samples = 0 - - updated_hessian, new_n_samples = update_hessian(input_tensor, hessian, n_samples) - - # Verify output shape - assert updated_hessian.shape == (features, features), ( - f"Expected hessian shape ({features}, {features}), got {updated_hessian.shape}" - ) - - # Verify sample count is updated correctly (incremented by total tokens = batch * seq_len) - expected_n_samples = batch_size * seq_len - assert new_n_samples == expected_n_samples, ( - f"Expected n_samples={expected_n_samples}, got {new_n_samples}" - ) - - # Verify hessian is not all zeros after update - assert not torch.allclose(updated_hessian, torch.zeros_like(updated_hessian)), ( - "Hessian should not be all zeros after update" - ) - - # Verify hessian is symmetric (should be for outer product X @ X.T) - assert torch.allclose(updated_hessian, updated_hessian.t()), "Hessian should be symmetric" - - # Test 2: Known input - verify correct hessian calculation - batch_size = 6 - seq_len = 2 - features = 2 - input_tensor = torch.ones(batch_size, seq_len, features, dtype=torch.float32) - - hessian = torch.zeros(features, features, dtype=torch.float32) - n_samples = 0 - - updated_hessian, new_n_samples = update_hessian(input_tensor, hessian, n_samples) - - # Manual calculation: - # input_flat shape: (features, batch*seq) = (2, 12), all ones - # n_samples = batch * seq = 12 (token count after flattening) - # scaled_input = sqrt(2/12) * ones(2, 12) - # outer_product = (2/12) * ones(2,12) @ ones(12,2) = [[2,2], [2,2]] - expected_n_samples = batch_size * seq_len # 12 tokens - expected_hessian = torch.ones(features, features, dtype=torch.float32) * 2.0 - - assert torch.allclose(updated_hessian, expected_hessian, atol=1e-5), ( - f"Expected hessian {expected_hessian}, got {updated_hessian}" - ) - assert new_n_samples == expected_n_samples - - # Test 3: Accumulated hessians - verify equivalence - # Processing [6,2,2] in one step should equal processing [2,2,2] three times - seq_len = 2 - features = 2 - - # Process in 3 steps of batch_size=2 (4 tokens each, 12 total) - hessian_accumulated = torch.zeros(features, features, dtype=torch.float32) - n_samples_accumulated = 0 - - for i in range(3): - input_batch = torch.ones(2, seq_len, features, dtype=torch.float32) - hessian_accumulated, n_samples_accumulated = update_hessian( - input_batch, hessian_accumulated, n_samples_accumulated - ) - - # Verify that accumulated result matches single-step result from Test 2 - assert torch.allclose(hessian_accumulated, updated_hessian, atol=1e-5), ( - f"Accumulated hessian should match single-step: expected {updated_hessian}, got {hessian_accumulated}" - ) - assert torch.allclose(hessian_accumulated, expected_hessian, atol=1e-5), ( - f"Accumulated hessian should match expected: expected {expected_hessian}, got {hessian_accumulated}" - ) - # 3 batches * 2 batch_size * 2 seq_len = 12 tokens - assert n_samples_accumulated == 12, f"Expected n_samples=12, got {n_samples_accumulated}" - - @pytest.mark.parametrize( ("block_size", "dim", "model_weight", "expect_weight_change"), [ diff --git a/tests/gpu_megatron/conftest.py b/tests/gpu_megatron/conftest.py index 405a83db0..d84ea9976 100644 --- a/tests/gpu_megatron/conftest.py +++ b/tests/gpu_megatron/conftest.py @@ -26,6 +26,24 @@ with contextlib.suppress(ImportError): from apex.transformer.parallel_state import destroy_model_parallel as apex_destroy +@pytest.fixture(scope="session", autouse=True) +def _prebuild_quant_cuda_extensions(): + """Prebuild quant CUDA extensions before per-test timeouts start. + + First-use JIT compilation can take minutes in CI, so build the base, FP8, and MX + extensions during session setup and let tests fall back to on-demand JIT if needed. + + Doing it here in session setup (``pyproject`` sets ``timeout_func_only``) keeps the + build off the per-test clock and, unlike the ``_extensions/test_torch_extensions.py`` + prebuild tests, runs regardless of test selection/ordering (e.g. ``-k`` filters) and + is not itself capped by a per-test timeout. Worker subprocesses then load the cached + .so from the shared ``TORCH_EXTENSIONS_DIR``. + """ + import modelopt.torch.quantization.extensions as ext + + ext.precompile() + + def megatron_worker_teardown(rank, world_size): """Clean up model-parallel state between tests in persistent workers.""" if dist.is_initialized(): diff --git a/tests/gpu_megatron/torch/quantization/plugins/test_megatron.py b/tests/gpu_megatron/torch/quantization/plugins/test_megatron.py index 35cb967de..ecf050abb 100644 --- a/tests/gpu_megatron/torch/quantization/plugins/test_megatron.py +++ b/tests/gpu_megatron/torch/quantization/plugins/test_megatron.py @@ -16,6 +16,7 @@ import copy from contextlib import nullcontext from functools import partial +from pathlib import Path import pytest import torch @@ -28,6 +29,7 @@ from _test_utils.torch.megatron.models import ( from _test_utils.torch.megatron.utils import ( compare_amax_sync_across_expert_parallel, copy_weights_from_grouped_to_non_grouped, + get_batch, get_forward, initialize_for_megatron, run_mcore_inference, @@ -53,8 +55,14 @@ from megatron.core.transformer.moe.router import TopKRouter import modelopt import modelopt.torch.opt as mto import modelopt.torch.quantization as mtq +from modelopt.torch.quantization.algorithms import QuantRecipe, _AutoQuantizeBaseSearcher from modelopt.torch.quantization.nn import QuantModuleRegistry -from modelopt.torch.quantization.plugins.megatron import _QuantTEMCoreRowParallelLinear +from modelopt.torch.quantization.plugins.megatron import ( + _QuantTEMCoreRowParallelLinear, + get_mcore_layerwise_calibration_layers, +) +from modelopt.torch.quantization.utils import is_quantized_linear +from modelopt.torch.quantization.utils.layerwise_calib import LayerActivationCollector try: from megatron.core.extensions.transformer_engine import TERowParallelLinear @@ -432,10 +440,47 @@ FP8_GEMM_KV_CFG["quant_cfg"].extend(mtq.FP8_KV_CFG["quant_cfg"]) mtq.NVFP4_KV_CFG, ], ) -@pytest.mark.parametrize("compress", [False, True]) @pytest.mark.parametrize("meta_device", [False, True]) @pytest.mark.parametrize("transformer_impl", ["local", "modelopt"]) def test_homogeneous_sharded_state_dict( + dist_workers, tmp_path, config, meta_device, transformer_impl +): + _run_homogeneous_sharded_state_dict( + dist_workers, + tmp_path, + config, + compress=False, + meta_device=meta_device, + transformer_impl=transformer_impl, + ) + + +@pytest.mark.parametrize( + "config", + [ + mtq.FP8_DEFAULT_CFG, + mtq.INT4_AWQ_CFG, + mtq.NVFP4_DEFAULT_CFG, + ], +) +@pytest.mark.parametrize("meta_device", [False, True]) +@pytest.mark.parametrize("transformer_impl", ["local", "modelopt"]) +@pytest.mark.timeout(240) +# Compressed state dict takes longer due to real quant conversion & saving/loading +def test_homogeneous_compressed_sharded_state_dict( + dist_workers, tmp_path, config, meta_device, transformer_impl +): + _run_homogeneous_sharded_state_dict( + dist_workers, + tmp_path, + config, + compress=True, + meta_device=meta_device, + transformer_impl=transformer_impl, + ) + + +def _run_homogeneous_sharded_state_dict( dist_workers, tmp_path, config, compress, meta_device, transformer_impl ): if compress and config is mtq.W4A8_AWQ_BETA_CFG: @@ -712,6 +757,215 @@ def test_te_grouped_vs_sequential_quantize(dist_workers_size_4, quant_cfg): ) +def _test_auto_quantize_moe_ep_helper(rank, size): + initialize_for_megatron( + tensor_model_parallel_size=1, + expert_model_parallel_size=size, + seed=SEED, + ) + model = _gpt_model_provider( + tp_size=1, + ep_size=size, + hidden_size=32, + num_moe_experts=4, + moe_grouped_gemm=False, + transformer_impl="modelopt", + ) + + def forward_step(model, batch): + input_ids, labels, position_ids, attention_mask, loss_mask = batch + return model.forward( + input_ids=input_ids, + position_ids=position_ids, + attention_mask=attention_mask, + labels=labels, + loss_mask=loss_mask, + ) + + auto_quantize_helper( + model, + data_loader=[get_batch(model, batch_size=2) for _ in range(2)], + forward_step=forward_step, + forward_backward_step=lambda m, b: forward_step(m, b).mean().backward(), + quantization_formats=[mtq.NVFP4_DEFAULT_CFG, mtq.FP8_DEFAULT_CFG], + ) + + +def test_auto_quantize_moe_ep(dist_workers): + """auto_quantize must pick a consistent recipe across EP ranks when multiple GPUs run.""" + dist_workers.run(_test_auto_quantize_moe_ep_helper) + + +def _mamba_hybrid_forward_step(model, batch): + input_ids, labels, position_ids, attention_mask, loss_mask = batch + return model.forward( + input_ids=input_ids, + position_ids=position_ids, + attention_mask=attention_mask, + labels=labels, + ) + + +@pytest.mark.skipif(not HAS_MAMBA, reason="Mamba not installed") +def test_gptq_mamba_hybrid(dist_workers_size_1): + """End-to-end GPTQ (NVFP4) on a tiny Megatron-Core NemotronH-style hybrid model.""" + dist_workers_size_1.run(_test_gptq_mamba_hybrid) + + +def _test_gptq_mamba_hybrid(rank, size): + initialize_for_megatron(tensor_model_parallel_size=1, seed=SEED) + model = get_mcore_mamba_hybrid_model( + tensor_model_parallel_size=1, + hidden_size=32, + num_attention_heads=4, + ffn_hidden_size=64, + mamba_state_dim=16, + mamba_head_dim=8, + num_moe_experts=4, + moe_grouped_gemm=False, + moe_ffn_hidden_size=32, + moe_shared_expert_intermediate_size=16, + transformer_impl="modelopt", + ).cuda() + + quant_cfg = copy.deepcopy(mtq.NVFP4_DEFAULT_CFG) + quant_cfg["algorithm"] = {"method": "gptq"} + forward = get_forward(model, batch_size=1) + model = mtq.quantize(model, quant_cfg, forward) + + for m in model.modules(): + if isinstance(m, SequentialMLP): + assert all( + is_quantized_linear(e.linear_fc1) + and e.linear_fc1.weight_quantizer.is_enabled + and e.linear_fc1.input_quantizer.is_enabled + for e in m.local_experts + ) + assert all( + is_quantized_linear(e.linear_fc2) + and e.linear_fc2.weight_quantizer.is_enabled + and e.linear_fc2.input_quantizer.is_enabled + for e in m.local_experts + ) + assert torch.isfinite(forward(model)).all() + + +def _auto_quantize_mamba_hybrid_cost_helper(rank, size, expert_model_parallel_size, result_path): + initialize_for_megatron( + tensor_model_parallel_size=1, + expert_model_parallel_size=expert_model_parallel_size, + seed=SEED, + ) + model = get_mcore_mamba_hybrid_model( + tensor_model_parallel_size=1, + expert_model_parallel_size=expert_model_parallel_size, + hidden_size=32, + num_attention_heads=4, + ffn_hidden_size=64, + mamba_state_dim=16, + mamba_head_dim=8, + num_moe_experts=4, + moe_grouped_gemm=False, + moe_ffn_hidden_size=32, + moe_shared_expert_intermediate_size=16, + transformer_impl="modelopt", + ).cuda() + + def forward_backward_step(model, batch): + _mamba_hybrid_forward_step(model, batch).mean().backward() + + _, search_state = mtq.auto_quantize( + model, + constraints={"effective_bits": 8.0}, + quantization_formats=[mtq.NVFP4_DEFAULT_CFG, mtq.FP8_DEFAULT_CFG], + data_loader=[get_batch(model, batch_size=1)], + forward_step=_mamba_hybrid_forward_step, + forward_backward_step=forward_backward_step, + num_calib_steps=1, + num_score_steps=1, + verbose=True, + ) + + no_quant = QuantRecipe(quant_cfg=None) + summed_cost = sum( + stat["costs"][stat["formats"].index(no_quant)] + for stat in search_state["candidate_stats"].values() + ) + # The per-op no-quant costs must sum to the cost denominator AutoQuantize uses, + # which is the full quantizable weight size aggregated across EP ranks. + assert summed_cost == pytest.approx(search_state["cost_denominator"], rel=1e-6) + assert search_state["best"]["is_satisfied"] + + if expert_model_parallel_size == 1: + # With EP=1, every rank has the full expert set. DP de-duplication should make the + # summed no-quant cost match the local quantizable weight size on each rank. + local_total = _AutoQuantizeBaseSearcher._get_total_weight_size(list(model.modules())) + assert summed_cost == pytest.approx(local_total, rel=1e-6) + + if rank == 0: + Path(result_path).write_text(repr(summed_cost)) + + +@pytest.mark.skipif(not HAS_MAMBA, reason="Mamba not installed") +def test_auto_quantize_mamba_hybrid_ep_cost(dist_workers, tmp_path): + """AutoQuantize cost must match for EP=1 and EP=2 when two GPUs are available.""" + ep1_path = tmp_path / "ep1_cost.txt" + dist_workers.run( + partial( + _auto_quantize_mamba_hybrid_cost_helper, + expert_model_parallel_size=1, + result_path=str(ep1_path), + ) + ) + if dist_workers.world_size < 2: + return + + ep2_path = tmp_path / "ep2_cost.txt" + dist_workers.run( + partial( + _auto_quantize_mamba_hybrid_cost_helper, + expert_model_parallel_size=2, + result_path=str(ep2_path), + ) + ) + cost_ep1 = float(ep1_path.read_text()) + cost_ep2 = float(ep2_path.read_text()) + assert cost_ep1 == pytest.approx(cost_ep2, rel=1e-6) + + +def _test_mcore_layerwise_calibration_layers_do_not_mutate_decoder(rank, size): + initialize_for_megatron(tensor_model_parallel_size=1, seed=SEED) + model = _gpt_model_provider( + tp_size=1, + hidden_size=32, + meta_device=True, + transformer_impl="modelopt", + ) + decoder_layers = model.decoder.layers + decoder_len = len(decoder_layers) + output_layer = model.output_layer + + discovered_layers = get_mcore_layerwise_calibration_layers(model) + + assert discovered_layers is not None + assert len(discovered_layers) == decoder_len + 1 + assert discovered_layers[-1] is output_layer + assert len(decoder_layers) == decoder_len + assert all(layer is not output_layer for layer in decoder_layers) + + assert LayerActivationCollector.is_supported(model) + discovered_layers = LayerActivationCollector.get_decoder_layers(model) + assert discovered_layers is not None + assert len(discovered_layers) == decoder_len + 1 + assert discovered_layers[-1] is output_layer + assert len(decoder_layers) == decoder_len + assert all(layer is not output_layer for layer in decoder_layers) + + +def test_mcore_layerwise_calibration_layers_do_not_mutate_decoder(dist_workers_size_1): + dist_workers_size_1.run(_test_mcore_layerwise_calibration_layers_do_not_mutate_decoder) + + @pytest.mark.parametrize("ep_size", [1, 2]) @pytest.mark.parametrize("sync_weight_amax", [True, False]) def test_layer_sync_moe_local_experts_amax(dist_workers, ep_size, sync_weight_amax): @@ -985,8 +1239,6 @@ def test_kv_cache_amax_sync(dist_workers): def test_convert_mcore_te_gpt_model(distributed_setup_size_1): - if not HAS_TE: - pytest.skip("Transformer Engine is not installed") initialize_for_megatron(tensor_model_parallel_size=1, seed=SEED) model = get_mcore_gpt_model(tensor_model_parallel_size=1, transformer_impl="transformer_engine") diff --git a/tests/unit/torch/quantization/test_autoquant.py b/tests/unit/torch/quantization/test_autoquant.py index 671372a8e..1978f3890 100644 --- a/tests/unit/torch/quantization/test_autoquant.py +++ b/tests/unit/torch/quantization/test_autoquant.py @@ -26,6 +26,7 @@ import modelopt.torch.opt as mto import modelopt.torch.quantization as mtq from modelopt.torch.quantization._auto_quantize_cost import ( EXCLUDED_MODULE_NAME_PATTERNS_KEY, + _get_module_weight_numel, get_auto_quantize_cost_model, infer_active_moe_expert_ratio, ) @@ -33,6 +34,7 @@ from modelopt.torch.quantization.algorithms import ( AutoQuantizeGradientSearcher, QuantRecipe, QuantRecipeHparam, + _AutoQuantizeBaseSearcher, estimate_quant_compression, ) from modelopt.torch.quantization.config import _base_disable_all, _default_disabled_quantizer_cfg @@ -172,25 +174,15 @@ def test_quant_recipe_hparam_zero_cost_weight(): def test_auto_quantize_cost_model_excludes_module_name_patterns(): - visual = torch.nn.Linear(4, 16) - mtp = torch.nn.Linear(4, 16) - lm_head = torch.nn.Linear(4, 16) cost_model = get_auto_quantize_cost_model("weight") cost_constraints = {EXCLUDED_MODULE_NAME_PATTERNS_KEY: ["*visual*", "*vision_tower*", "*mtp*"]} - total_weight_size = cost_model.total_weight_size( - [ - ("model.visual.blocks.0.attn.qkv", visual), - ("model.mtp.layers.0.mlp", mtp), - ("lm_head", lm_head), - ], - is_auto_quantize_module=lambda module: True, - cost_constraints=cost_constraints, - ) - - assert total_weight_size == pytest.approx(lm_head.weight.numel()) + # Modules whose name matches an excluded pattern contribute zero cost weight. assert cost_model.module_cost_weight(["model.visual.blocks.0.attn.qkv"], cost_constraints) == 0 assert cost_model.module_cost_weight(["model.mtp.layers.0.mlp"], cost_constraints) == 0 + # Non-excluded modules keep full weight. + assert cost_model.module_cost_weight(["lm_head"], cost_constraints) == 1.0 + # A group is only excluded when *all* of its module names match; a mixed group is not. assert ( cost_model.module_cost_weight( ["model.visual.blocks.0.attn.qkv", "lm_head"], cost_constraints @@ -203,24 +195,19 @@ def test_active_moe_cost_model_counts_fused_experts_without_weight(): fused_experts = torch.nn.Module() fused_experts.gate_up_proj = torch.nn.Parameter(torch.empty(2, 3, 5)) fused_experts.down_proj = torch.nn.Parameter(torch.empty(2, 5, 3)) - visual = torch.nn.Linear(4, 16) cost_model = get_auto_quantize_cost_model("active_moe") + cost_constraints = { + "active_moe_expert_ratio": 0.25, + EXCLUDED_MODULE_NAME_PATTERNS_KEY: ["*visual*"], + } - total_weight_size = cost_model.total_weight_size( - [ - ("layers.0.mlp.experts", fused_experts), - ("model.visual.blocks.0.attn.qkv", visual), - ], - is_auto_quantize_module=lambda module: True, - cost_constraints={ - "active_moe_expert_ratio": 0.25, - EXCLUDED_MODULE_NAME_PATTERNS_KEY: ["*visual*"], - }, - ) - - assert total_weight_size == pytest.approx( - (fused_experts.gate_up_proj.numel() + fused_experts.down_proj.numel()) * 0.25 + # Fused experts expose no `.weight`; their size is summed across all parameters. + assert _get_module_weight_numel(fused_experts) == ( + fused_experts.gate_up_proj.numel() + fused_experts.down_proj.numel() ) + # Routed MoE experts are scaled by the active-expert ratio; excluded modules drop to zero. + assert cost_model.module_cost_weight(["layers.0.mlp.experts"], cost_constraints) == 0.25 + assert cost_model.module_cost_weight(["model.visual.blocks.0.attn.qkv"], cost_constraints) == 0 @pytest.mark.parametrize("num_experts_attr", ["num_experts", "num_local_experts"]) @@ -486,6 +473,39 @@ def test_data_parallel_auto_quantize(skip_on_windows): spawn_multiprocess_job(2, _test_data_parallel_auto_quantize, backend="gloo") +def test_auto_quantize_budget_uses_no_quant_candidate_cost(monkeypatch): + class _BudgetCaptureSearcher(AutoQuantizeGradientSearcher): + def run_search_with_stats(self, max_weight_size, verbose=False): + self.max_weight_size = max_weight_size + return {}, True + + def _raise_local_total_weight_size(modules): + pytest.fail("run_search should derive total weight size from candidate costs") + + monkeypatch.setattr( + _AutoQuantizeBaseSearcher, + "_get_total_weight_size", + staticmethod(_raise_local_total_weight_size), + ) + + searcher = _BudgetCaptureSearcher() + searcher.reset_search() + searcher.model = torch.nn.Module() + searcher.config = {"verbose": False} + searcher.constraints = {"effective_bits": 8.0} + searcher.candidate_stats = { + "local_expert.quant_recipe": { + "formats": [QuantRecipe(mtq.NVFP4_DEFAULT_CFG), QuantRecipe(None)], + "scores": [1.0, 0.0], + "costs": [25.0, 100.0], + } + } + + searcher.run_search() + + assert searcher.max_weight_size == 50.0 + + def test_estimate_quant_compression(): nvfp4_affine_kv_cfg = mtq.config.QuantizeConfig(**mtq.NVFP4_AFFINE_KV_CFG) assert estimate_quant_compression(nvfp4_affine_kv_cfg) == 0.25 diff --git a/tests/unit/torch/quantization/test_gptq.py b/tests/unit/torch/quantization/test_gptq.py new file mode 100644 index 000000000..345bd1671 --- /dev/null +++ b/tests/unit/torch/quantization/test_gptq.py @@ -0,0 +1,119 @@ +# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""CPU unit tests for GPTQ utilities.""" + +import torch + +from modelopt.torch.quantization.utils.calib_utils import update_hessian + + +def test_update_hessian(): + """Test for update_hessian function with both random and known inputs.""" + # Test 1: Random input - general functionality test + torch.manual_seed(42) + batch_size = 2 + seq_len = 3 + features = 4 + input_tensor = torch.randn(batch_size, seq_len, features, dtype=torch.float32) + + hessian = torch.zeros(features, features, dtype=torch.float32) + n_samples = 0 + + updated_hessian, new_n_samples = update_hessian(input_tensor, hessian, n_samples) + + # Verify output shape + assert updated_hessian.shape == (features, features), ( + f"Expected hessian shape ({features}, {features}), got {updated_hessian.shape}" + ) + + # Verify sample count is updated correctly (incremented by total tokens = batch * seq_len) + expected_n_samples = batch_size * seq_len + assert new_n_samples == expected_n_samples, ( + f"Expected n_samples={expected_n_samples}, got {new_n_samples}" + ) + + # Verify hessian is not all zeros after update + assert not torch.allclose(updated_hessian, torch.zeros_like(updated_hessian)), ( + "Hessian should not be all zeros after update" + ) + + # Verify hessian is symmetric (should be for outer product X @ X.T) + assert torch.allclose(updated_hessian, updated_hessian.t()), "Hessian should be symmetric" + + # Test 2: Known input - verify correct hessian calculation + batch_size = 6 + seq_len = 2 + features = 2 + input_tensor = torch.ones(batch_size, seq_len, features, dtype=torch.float32) + + hessian = torch.zeros(features, features, dtype=torch.float32) + n_samples = 0 + + updated_hessian, new_n_samples = update_hessian(input_tensor, hessian, n_samples) + + # Manual calculation: + # input_flat shape: (features, batch*seq) = (2, 12), all ones + # n_samples = batch * seq = 12 (token count after flattening) + # scaled_input = sqrt(2/12) * ones(2, 12) + # outer_product = (2/12) * ones(2,12) @ ones(12,2) = [[2,2], [2,2]] + expected_n_samples = batch_size * seq_len # 12 tokens + expected_hessian = torch.ones(features, features, dtype=torch.float32) * 2.0 + + assert torch.allclose(updated_hessian, expected_hessian, atol=1e-5), ( + f"Expected hessian {expected_hessian}, got {updated_hessian}" + ) + assert new_n_samples == expected_n_samples + + # Test 3: Accumulated hessians - verify equivalence + # Processing [6,2,2] in one step should equal processing [2,2,2] three times + seq_len = 2 + features = 2 + + # Process in 3 steps of batch_size=2 (4 tokens each, 12 total) + hessian_accumulated = torch.zeros(features, features, dtype=torch.float32) + n_samples_accumulated = 0 + + for _ in range(3): + input_batch = torch.ones(2, seq_len, features, dtype=torch.float32) + hessian_accumulated, n_samples_accumulated = update_hessian( + input_batch, hessian_accumulated, n_samples_accumulated + ) + + # Verify that accumulated result matches single-step result from Test 2 + assert torch.allclose(hessian_accumulated, updated_hessian, atol=1e-5), ( + f"Accumulated hessian should match single-step: expected {updated_hessian}, got {hessian_accumulated}" + ) + assert torch.allclose(hessian_accumulated, expected_hessian, atol=1e-5), ( + f"Accumulated hessian should match expected: expected {expected_hessian}, got {hessian_accumulated}" + ) + # 3 batches * 2 batch_size * 2 seq_len = 12 tokens + assert n_samples_accumulated == 12, f"Expected n_samples=12, got {n_samples_accumulated}" + + +def test_update_hessian_zero_token_input_noops(): + """Empty activations must not mutate the Hessian or sample count.""" + features = 4 + hessian = torch.eye(features, dtype=torch.float32) + expected_hessian = hessian.clone() + n_samples = 7 + + updated_hessian, new_n_samples = update_hessian( + torch.empty(0, features, dtype=torch.float32), hessian, n_samples + ) + + assert updated_hessian is hessian + assert new_n_samples == n_samples + torch.testing.assert_close(hessian, expected_hessian) diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml new file mode 100644 index 000000000..92de27e97 --- /dev/null +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml @@ -0,0 +1,46 @@ +# Nemotron-3-Nano-30B-A3B-BF16 AutoQuantize + inline MMLU. +# Tested on one B200 node (4 x B200). +# +# Usage: +# source .env-slurm +# cd tools/launcher +# uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-AutoQuantize.yaml --yes + +job_name: Nemotron-3-Nano-30B-A3B_AutoQuantize +pipeline: + skip: false + allow_to_fail: false + note: "AutoQuantize on Nemotron-3-Nano-30B-A3B-BF16: 8192 sequence length, 512 calibration samples, inline MMLU, 1 node x 4 GPUs" + + task_0: + script: common/megatron_lm/quantize/quantize.sh + args: + - --calib-size 512 + - --calib-max-sequence-length 8192 + - --auto-quantize-bits 4.8 + - --auto-quantize-formats NVFP4_DEFAULT_CFG FP8_DEFAULT_CFG + - --auto-quantize-score-size 256 + - --auto-quantize-checkpoint /cicd/auto_quantize_8192x512_score256 + environment: + - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 + - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 + - QUANT_CFG: auto + - MMLU_DATASET: /hf-local/cais/mmlu + - MMLU_LOWER_BOUND: "0.60" + - RUN_MMLU: "true" + - RUN_EXPORT: "false" + - TP: "1" + - ETP: "1" + - EP: "4" + - PP: "1" + - CP: "1" + - DP: "1" + slurm_config: + _factory_: "slurm_factory" + container: nvcr.io/nvidia/nemo:26.04 + modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt + partition: batch + nodes: 1 + ntasks_per_node: 4 + gpus_per_node: 4 + time: "04:00:00" diff --git a/tools/launcher/modules/Megatron-LM b/tools/launcher/modules/Megatron-LM index c69697d05..0429f9fe6 160000 --- a/tools/launcher/modules/Megatron-LM +++ b/tools/launcher/modules/Megatron-LM @@ -1 +1 @@ -Subproject commit c69697d0510198acddf124694dcfd5057021092f +Subproject commit 0429f9fe6243c92e494895012f7078d7d3d4e351