mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix + new tests (test-suite only; changes are confined to `tests/`) - **Fix `test_parallel_load_and_export` hang on 2 GPUs.** The temp paths were built from `os.getpid()` *inside* the workers, so each rank got a different `ckpt_dir`; rank 1 then took `_resolve_checkpoint_dir`'s Hub branch and blocked on an extra barrier while rank 0 ran the loader's broadcasts. The checkpoint is now built once in the parent under `tmp_path` and passed in, so all ranks agree on the path. - **Imports moved to module top** across `tests/gpu*` and `tests/examples`; function-local imports kept only where guarded (`importorskip`/`try`), where the import *is* the test (JIT compile), or where it must follow `sys.path` setup. - **Reuse `_test_utils` instead of local copies:** added `get_tiny_mixtral`; deduped `assert_nodes_are_quantized` (5 copies), the accelerate-offload/layerwise config helpers, `make_quant_attention`, `get_dflash_config`, the NVFP4 amax assertions, and 3 copies of the `tiny_wan22_path` fixture. - **Shared model-dir fixtures assert they were not modified** (`assert_unmodified_tree`): a file manifest is compared on teardown, so a test that writes into a session-scoped fixture directory fails instead of silently changing what later tests see. - **Dropped dependency guards the CI env already guarantees** (diffusers, tensorrt_llm in `gpu_trtllm`, transformer_engine in `gpu_megatron`, transformers in examples) so a missing dep fails loudly instead of skipping. - **`test_heterogenous_sharded_state_dict` is skipped on Blackwell** (sm_120), matching the existing marker and its TE/CUDA-13 rationale — same tracking issue as #1901. ### Testing Local, 2x RTX 6000 Ada: `tests/gpu/torch/utils/test_model_load_utils.py` passes on 2 GPUs and on 1 GPU (previously hung on 2). Also ran the touched files in `tests/unit` (608 passed) and `tests/gpu` (~340 passed). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — test-only PR - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ — not yet run <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added reusable validation utilities for generated files, quantization behavior, attention modules, model fixtures, offloading, and speculative decoding. * Expanded coverage for tiny Wan, Mixtral, Llama, and related model scenarios. * Consolidated duplicated setup and assertions across ONNX, GPU, quantization, export, and sparsity tests. * Improved fixture integrity checks, NVFP4 validation, and handling of identity inputs. * Reduced unnecessary dependency-based skips and isolated known platform-specific flakiness. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
58 lines
2.4 KiB
Python
58 lines
2.4 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
"""Filesystem helpers for tests."""
|
|
|
|
from collections.abc import Iterator
|
|
from contextlib import contextmanager
|
|
from pathlib import Path
|
|
|
|
|
|
def _manifest(root: Path) -> dict[str, tuple[int, int]]:
|
|
"""``relative path -> (size, mtime_ns)`` for every file below ``root``."""
|
|
return {
|
|
str(p.relative_to(root)): (p.stat().st_size, p.stat().st_mtime_ns)
|
|
for p in root.rglob("*")
|
|
if p.is_file() and not p.is_symlink()
|
|
}
|
|
|
|
|
|
@contextmanager
|
|
def assert_unmodified_tree(path: Path | str) -> Iterator[Path]:
|
|
"""Fail if anything under ``path`` is added, removed, or rewritten inside the ``with``.
|
|
|
|
For session/module-scoped model-directory fixtures: a test that writes into a shared
|
|
directory silently changes what every later test sees. Comparing a file manifest on
|
|
teardown catches that. ``chmod``-ing the tree read-only would report at the write rather
|
|
than at teardown, but it only works for an unprivileged user -- root has
|
|
``CAP_DAC_OVERRIDE`` and writes straight through the permission bits, and the CI
|
|
containers run as root.
|
|
"""
|
|
path = Path(path)
|
|
before = _manifest(path)
|
|
yield path
|
|
if not path.exists():
|
|
raise AssertionError(f"shared fixture directory {path} was deleted by a test")
|
|
after = _manifest(path)
|
|
added = sorted(after.keys() - before.keys())
|
|
removed = sorted(before.keys() - after.keys())
|
|
changed = sorted(k for k in before.keys() & after.keys() if before[k] != after[k])
|
|
if added or removed or changed:
|
|
raise AssertionError(
|
|
f"shared fixture directory {path} was modified by a test "
|
|
f"(added={added}, removed={removed}, changed={changed}); "
|
|
"copy it into the test's own tmp_path instead of writing into the shared tree"
|
|
)
|