mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? **Type of change**: New feature On DLA, the whole DLA-eligible region is compiled as one node, which runs in INT8 or FP16, and it expects scales to be present throughout. A tensor without a usable scale typically forces either that region to run in FP16 or a GPU fallback (if enabled) — otherwise the build fails. With IQ (implicit quantization) being deprecated in TensorRT, users are migrating to ModelOpt for quantization/calibration. However, this breaks the DLA workflow since DLA still only supports IQ. The suggested workflow is then to: 1. Use ModelOpt to obtain the EQ (explicitly quantized) model; 2. Use [NVIDIA's Q/DQ Translator Toolkit](https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/tree/main/tools/qdq-translator) to obtain the `calib.cache` and `layer_arg.txt` files, which can be used with the non-quantized model to generate a DLA loadable. A [study on Yolov5](https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes) has shown that EQ can achieve perf parity with IQ on DLA if Q/DQ nodes are inserted at every layer, making sure all tensors have INT8 scales. From the study: _"With this option, all layers’ scales can be obtained during model fine-tuning. However, this method may potentially disrupt TensorRT fusion strategy with Q/DQ layers when running inference on GPU and lead to higher latency on the GPU. For DLA, on the other hand, the rule of thumb with PTQ scales is, “The more available scales, the lower the latency.” "_ This PR aims to enable a quantization path targeting DLA. ### Usage ```python $ python -m modelopt.onnx.quantization --onnx=model.onnx --target_dla ``` ### Testing - Two new parametrized tests (target_dla=False/True) cover both the Conv/Mul quantization expansion and the GEMV (MatMul m=1) exclusion bypass, with dedicated model builders. - Internal test: 6241485@10 I ran the following experiments on various `timm` models: | Exp | ModelOpt flag | QDQ-Translator flag | |-------|-----------------------|------------------------------| | 1 | `--high_precision_dtype=fp32` | default | | 2 | `--high_precision_dtype=fp32 --target_dla` | default | | 3 | `--high_precision_dtype=fp32` | `--addtl_ops_to_infer_adjacent_scales` [1] | | 4 | `--high_precision_dtype=fp32 --target_dla` | `--addtl_ops_to_infer_adjacent_scales` [1] | > [1] See https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/pull/35 Results (DOS Orin Linux with TRT 10.15.3.2): | Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 | |-------|------|----|------|----| | resnet50 | 5.09 | 1.25 | 1.26 | 1.22 | | mobilenetv2_100 | 4.07 | 3.78 | 0.80 | 0.77 | | efficientnet_lite0 | 6.10 | 5.65 | 1.07 | 1.07 | | inception_v3 | 11.43 | 1.57 | 1.56 | 1.57 | | res2net50_14w_8s | 17.39 | 3.06 | 3.82 | 3.02 | Observations: 1. Exp 1 vs 2: `--target_dla` is essential to recover performance. 2. Exp 3 vs 4: `--target_dla` is necessary for perf parity or improved perf compared to the default ModelOpt behavior. This is demonstrated in the `res2net50_14w_8s`, which benefits from this new flag due to its architecture containing 8 Convs operating on 14-channel tensors (below the 16-channel minimum check in `int8.py / find_nodes_from_convs_to_exclude()`. Accuracy evaluation also shows no degradation for any of the experiments. Top-1 with 1,000 ImageNet samples (%): | Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 | |-------|------|----|------|----| | resnet50 | 75.6 | 76.0 | 75.1 | 76.0 | | mobilenetv2_100 | 72.1 | 72.0 | 72.1 | 72.3 | | efficientnet_lite0 | 75.1 | 75.2 | 75.1 | 75.2 | | inception_v3 | 76.4 | 75.5 | 76.2 | 75.5 | | res2net50_14w_8s | 75.5 | 75.6 | 75.8 | 75.6 | ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ❌ <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional info Related blogpost: https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--target_dla` option for INT8 quantization to enable optimized Q/DQ placement for DLA. * **Behavior Changes** * Adjusts quantization pre-processing rules when DLA targeting (or autotune) is enabled, and defaults to quantizing all op types when none are specified. * **Examples** * Added deterministic `--seed` for evaluation; enhanced ImageNet dataset and calibration image loading/preprocessing (local or dataset-based). * **Tests** * Added coverage to verify Q/DQ placement differences for Conv and MatMul with `target_dla`. * **Documentation** * Updated the changelog to highlight the new DLA targeting option. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
234 lines
7.8 KiB
Python
234 lines
7.8 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
"""Module to evaluate a device model for the specified task."""
|
|
|
|
import os
|
|
import random
|
|
from pathlib import Path
|
|
from typing import Final
|
|
|
|
import torch
|
|
from datasets import load_dataset
|
|
from PIL import Image
|
|
from tqdm import tqdm
|
|
|
|
from modelopt.torch._deploy.device_model import DeviceModel
|
|
|
|
ACCURACY: Final[str] = "accuracy"
|
|
|
|
|
|
class ImageNetWrapper(torch.utils.data.Dataset):
|
|
"""Wrapper for the ILSVRC/imagenet-1k Hugging Face dataset."""
|
|
|
|
def __init__(self, hf_dataset, transform=None):
|
|
"""Initialize the wrapper.
|
|
|
|
Args:
|
|
hf_dataset: The Hugging Face dataset object.
|
|
transform: Optional transform to apply to images.
|
|
"""
|
|
self.dataset = hf_dataset
|
|
self.transform = transform
|
|
|
|
def __len__(self):
|
|
return len(self.dataset)
|
|
|
|
def __getitem__(self, idx):
|
|
item = self.dataset[idx]
|
|
image = item["image"]
|
|
|
|
# Convert to RGB if needed
|
|
if image.mode != "RGB":
|
|
image = image.convert("RGB")
|
|
|
|
if self.transform:
|
|
image = self.transform(image)
|
|
|
|
label = item["label"]
|
|
return image, label
|
|
|
|
|
|
class LocalImageNetDataset(torch.utils.data.Dataset):
|
|
"""Local ImageNet validation set from a flat directory and a label file.
|
|
|
|
Expects:
|
|
<root>/validation/ILSVRC2012_val_XXXXXXXX.JPEG (50k images, flat)
|
|
<root>/val.txt (one line per image: "<filename> <class_idx>")
|
|
"""
|
|
|
|
def __init__(self, root, transform=None):
|
|
"""Initialize the dataset.
|
|
|
|
Args:
|
|
root: Path to the ImageNet root directory.
|
|
transform: Optional transform to apply to images.
|
|
"""
|
|
img_dir = Path(root) / "validation"
|
|
with open(Path(root) / "val.txt") as f:
|
|
entries = [line.strip().split() for line in f]
|
|
self.samples = [(img_dir / name, int(label)) for name, label in entries]
|
|
self.transform = transform
|
|
|
|
def __len__(self):
|
|
return len(self.samples)
|
|
|
|
def __getitem__(self, idx):
|
|
path, label = self.samples[idx]
|
|
with Image.open(path) as img:
|
|
image = img.convert("RGB")
|
|
if self.transform:
|
|
image = self.transform(image)
|
|
return image, label
|
|
|
|
|
|
def _load_dataset(dataset_path: str):
|
|
"""Load an ImageNet-style dataset from a local directory or HF Hub.
|
|
|
|
Supports three path types:
|
|
- HF Hub dataset card name (e.g. ILSVRC/imagenet-1k)
|
|
- Local HF dataset mirror with data/validation* shards
|
|
- Local ImageNet root with flat validation dir + val.txt
|
|
"""
|
|
if os.path.isfile(os.path.join(dataset_path, "val.txt")):
|
|
# Local ImageNet: flat validation/ dir + val.txt label file
|
|
return None, dataset_path
|
|
return (
|
|
load_dataset(
|
|
dataset_path,
|
|
split="validation",
|
|
data_files={"validation": "data/validation*"},
|
|
verification_mode="no_checks",
|
|
),
|
|
None,
|
|
)
|
|
|
|
|
|
def evaluate(
|
|
model: torch.nn.Module | DeviceModel,
|
|
transform,
|
|
evaluation_type: str = ACCURACY,
|
|
batch_size=1,
|
|
num_examples=None,
|
|
device="cuda",
|
|
dataset_path="ILSVRC/imagenet-1k",
|
|
seed=0,
|
|
):
|
|
"""Evaluate a model for the given dataset.
|
|
|
|
Args:
|
|
model: PyTorch model or DeviceModel to evaluate.
|
|
transform: Transform to apply to the dataset images.
|
|
evaluation_type: Type of evaluation to perform. Currently only accuracy is supported.
|
|
batch_size: Batch size to use for evaluation. Currently only batch_size=1 is supported.
|
|
num_examples: Number of examples to evaluate on. If None, evaluate on the entire dataset.
|
|
device: Device to run evaluation on. Supported devices: "cpu" and "cuda". Defaults to "cuda".
|
|
dataset_path: HF dataset card (e.g. "ILSVRC/imagenet-1k"), local HF mirror with
|
|
data/validation* shards, or local ImageNet root dir containing val.txt and
|
|
a flat validation/ directory. Defaults to "ILSVRC/imagenet-1k".
|
|
seed: Random seed for the DataLoader shuffle, ensuring reproducible image sampling across
|
|
runs. Defaults to 0.
|
|
Returns:
|
|
The evaluation result.
|
|
"""
|
|
hf_dataset, local_root = _load_dataset(dataset_path)
|
|
if local_root is not None:
|
|
val_dataset = LocalImageNetDataset(local_root, transform=transform)
|
|
else:
|
|
val_dataset = ImageNetWrapper(hf_dataset, transform=transform)
|
|
|
|
generator = torch.Generator()
|
|
generator.manual_seed(seed)
|
|
val_loader = torch.utils.data.DataLoader(
|
|
val_dataset, batch_size=batch_size, shuffle=True, num_workers=4, generator=generator
|
|
)
|
|
|
|
# TODO: Add support for segmentation tasks.
|
|
if evaluation_type == ACCURACY:
|
|
return evaluate_accuracy(
|
|
model,
|
|
val_loader,
|
|
num_examples,
|
|
batch_size,
|
|
topk=(1, 5),
|
|
random_seed=seed,
|
|
device=device,
|
|
)
|
|
else:
|
|
raise ValueError(f"Unsupported evaluation type: {evaluation_type}")
|
|
|
|
|
|
def evaluate_accuracy(
|
|
model, val_loader, num_examples, batch_size, topk=(1,), random_seed=None, device="cuda"
|
|
):
|
|
"""Evaluate the accuracy of the model on the validation dataset.
|
|
|
|
Args:
|
|
model: Model to evaluate.
|
|
val_loader: DataLoader for the validation dataset.
|
|
num_examples: Number of examples to evaluate on. If None, evaluate on the entire dataset.
|
|
batch_size: Batch size to use for evaluation.
|
|
topk: function support topk accuracy. Return list of accuracy equal to topk length.
|
|
example of usage `top1, top5 = evaluate_accuracy(..., topk=(1,5))`
|
|
`top1, top5, top10 = evaluate_accuracy(..., topk=(1,5,10))`
|
|
random_seed: Random seed to use for evaluation.
|
|
device: Device to run evaluation on. Supported devices: "cpu" and "cuda". Defaults to "cuda".
|
|
|
|
Returns:
|
|
The accuracy of the model on the validation dataset.
|
|
"""
|
|
|
|
if random_seed is not None:
|
|
torch.manual_seed(random_seed)
|
|
torch.cuda.manual_seed_all(random_seed)
|
|
random.seed(random_seed)
|
|
torch.backends.cudnn.deterministic = True
|
|
torch.backends.cudnn.benchmark = False
|
|
|
|
if isinstance(model, torch.nn.Module):
|
|
model.eval()
|
|
model = model.to(device)
|
|
total = 0
|
|
corrects = [0] * len(topk)
|
|
for _, (inputs, labels) in tqdm(
|
|
enumerate(val_loader),
|
|
total=num_examples // batch_size if num_examples is not None else len(val_loader),
|
|
desc="Evaluation progress: ",
|
|
):
|
|
if num_examples is not None and total >= num_examples:
|
|
break
|
|
# Forward pass
|
|
if not isinstance(model, torch.nn.Module):
|
|
inputs = [inputs]
|
|
else:
|
|
inputs = inputs.to(device)
|
|
outputs = model(inputs)
|
|
|
|
# Calculate accuracy
|
|
outputs = outputs[0] if isinstance(outputs, list) else outputs.data
|
|
labels_size = labels.size(0)
|
|
outputs = outputs[:labels_size]
|
|
|
|
total += labels_size
|
|
|
|
labels = labels.to(outputs.device)
|
|
|
|
for ind, k in enumerate(topk):
|
|
_, predicted = torch.topk(outputs, k, dim=1)
|
|
corrects[ind] += (predicted == labels.unsqueeze(1)).any(dim=1).sum().item()
|
|
|
|
res = [100 * corr / total for corr in corrects]
|
|
return res
|