mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What does this PR do? **Type of change:** refactor / deprecation (examples) Follow-up to #1705 (which consolidated `examples/vlm_ptq` into `examples/llm_ptq`). Since that example now covers Hugging Face **LLM and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the directory to `examples/hf_ptq` and leaves a relative symlink `examples/llm_ptq → hf_ptq` so existing paths/commands keep working during a deprecation window. Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat approach), targeted for the **same 0.46 release** as the consolidation. ### Changes - `git mv examples/llm_ptq → examples/hf_ptq` and `tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the matrix name to both `examples/<name>` and `tests/examples/<name>`). - Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`. - Update CI matrices and all repo **path references** (docs, READMEs, agent skills, launcher/debugger tools, tests) from `llm_ptq` to `hf_ptq`. - Keep Python identifiers / test-util module names (`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task, not the directory. - Preserve the CODEOWNERS team slug (`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG entries; add a CHANGELOG deprecation note. ### Back-compat caveats (inherent to git directory symlinks) - ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work through the symlink. - ⚠️ Windows git checkouts don't materialize symlinks by default (low impact — this example is Linux-only in practice). - ⚠️ GitHub web doesn't follow directory symlinks, so legacy external deep-links to `examples/llm_ptq/...` won't navigate in. All **internal** references are repointed to `hf_ptq`, so the symlink is only for legacy external/CLI use. ### Usage (unchanged via symlink) ```bash # New canonical path cd examples/hf_ptq scripts/huggingface_example.sh --model <hf_model> --quant fp8 # Old path still works (forwards via symlink) cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8 ``` ### Testing - `bash -n` on moved/edited shell scripts (new path + via symlink). - `py_compile` on moved/edited Python; test re-export shim repointed to `examples/hf_ptq/example_utils`. - Verified git tracks `examples/llm_ptq` as a single symlink (mode 120000), not a duplicated tree (no pre-commit / pytest double-processing). - `pre-commit run` on all changed files passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (relative symlink keeps `examples/llm_ptq` paths valid; see caveats above) - Did you write any new necessary tests?: N/A (pure rename; existing tests moved with the dir) - Did you update Changelog?: ✅ ### Additional Information Follow-up (later release): remove the `examples/llm_ptq` symlink once external references have migrated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PTQ guidance now directs to the unified Hugging Face PTQ flow, including VLM quantization via the shared `--vlm` entry point. * **Documentation** * Updated README and guide links, references, and command snippets to use `hf_ptq` (replacing `llm_ptq`). * Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed VILA/NVILA coverage from the Hugging Face PTQ examples. * **Bug Fixes** * Improved detection and routing so local/manual setup uses the correct PTQ source. * **Tests / Chores** * CI and example tests updated to run the `hf_ptq` variants. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
4.7 KiB
4.7 KiB
File-Based Command Relay (Debugger)
A lightweight client/server system for running commands inside a Docker container from the host, using only a shared filesystem — no networking required.
Overview
Host (Claude Code) Docker Container
┌─────────────┐ ┌─────────────────┐
│ client.sh │ writes cmd file │ server.sh │
│ run "X" │ ───────────────────► │ detects cmd │
│ │ │ executes X │
│ reads │ writes result file │ writes result │
│ result │ ◄─────────────────── │ │
└─────────────┘ └─────────────────┘
└──── shared filesystem (.relay/) ────┘
Assumptions
- The ModelOpt repo is accessible from both host and container (e.g., bind-mounted)
- HuggingFace models are mounted at
/hf-local - The server auto-detects the repo root from the location of
server.sh
Quick Start
1. Start the server (inside Docker)
# The server auto-detects the repo root (two levels up from tools/debugger/)
bash /path/to/modelopt/tools/debugger/server.sh
The server automatically sets the working directory to the repo root. You can override with --workdir.
2. Connect from the host
bash tools/debugger/client.sh handshake
3. Run commands
# Run a simple command
bash tools/debugger/client.sh run "echo hello"
# Run a test script
bash tools/debugger/client.sh run "bash hf_ptq/scripts/huggingface_example.sh"
# Run with a long timeout (default is 600s)
bash tools/debugger/client.sh --timeout 1800 run "python my_long_test.py"
# Cancel a running command
bash tools/debugger/client.sh cancel
# Check status
bash tools/debugger/client.sh status
Protocol
The relay uses a directory at tools/debugger/.relay/ with this structure:
.relay/
├── server.ready # Written by server on startup
├── owner # Current relay owner id (host:pid:nanos); newest server wins
├── client.ready # Written by client during handshake
├── handshake.done # Written by server to confirm handshake
├── running # Written by server while a command is executing (cmd_id:pid)
├── cancel # Written by client to request cancellation of the running command
├── cmd/ # Client writes command .sh files here
│ └── <id>.sh # Command to execute
└── result/ # Server writes results here
├── <id>.log # stdout + stderr
└── <id>.exit # Exit code
Handshake
- Server starts, creates
.relay/server.ready - Client writes
.relay/client.ready - Server detects it, writes
.relay/handshake.done - Both sides are now connected
Command Execution
- Client writes a command to
.relay/cmd/<id>.sh - Server detects the file, reads the command content, and removes the
.shfile - Server runs
bash -c <content>in a new process group, writes.relay/running - Server writes
.relay/result/<id>.exitand.relay/result/<id>.log, then removes.relay/running - Client reads results and cleans up
Cancellation
- Client writes the target
cmd_idto.relay/cancel - Server verifies the
cmd_idmatches, then kills the command's process group - Server writes exit code 130 and removes
.relay/runningand.relay/cancel - Client-side timeout also triggers cancellation automatically
Options
Server
| Flag | Default | Description |
|---|---|---|
--relay-dir |
<script_dir>/.relay |
Relay directory path |
--workdir |
Auto-detected repo root | Working directory for commands |
Client
| Flag | Default | Description |
|---|---|---|
--relay-dir |
<script_dir>/.relay |
Relay directory path |
--timeout |
600 |
Seconds to wait for command result |
Notes
- The
.relay/directory is in.gitignore— it is not checked in. - Only one server serves at a time. A newly started server claims
.relay/owner; any older server (even on another host sharing this NFS relay) sees the changed owner on its next poll and exits cleanly, rather than competing for commands. - Commands run sequentially in the order the server discovers them.
- A running command can be cancelled via
client.sh cancel. Cancelled commands exit with code 130. - Client-side timeouts automatically cancel the running command on the server.