mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary - Adds a lightweight file-based client/server relay (`tools/debugger/`) that enables Claude Code (or any host-side automation) to execute commands inside a remote Docker container using only a shared filesystem — no networking setup required. - The server auto-detects the repo root, installs modelopt (`pip install -e .[dev]`), sets `PYTHONPATH`, and listens for commands. - The client supports `handshake`, `run`, `status`, and `flush` subcommands. - Includes `README.md` (full protocol docs) and `CLAUDE.md` (quick reference for Claude Code). ## Tested - Ran Qwen3.5-35B-A3B MoE PTQ with `nvfp4_experts_only` quantization via the relay: ``` bash examples/llm_ptq/scripts/huggingface_example.sh --model /hf-local/Qwen/Qwen3.5-35B-A3B/ --quant nvfp4_experts_only ``` - 42,140 quantizers inserted, MTP layers correctly excluded - Quantized checkpoint exported successfully (~208s, 147.61 GB peak GPU memory) ## Test plan - [x] Start `server.sh` inside a Docker container with the repo mounted - [x] Run `client.sh handshake` from the host - [x] Run `client.sh run "echo hello"` and verify output - [x] Run `client.sh flush` and verify `.relay/` is cleared - [x] Run a real PTQ workload end-to-end 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a file-based command relay system with host↔container client and server CLIs, supporting handshake, run, status and flush workflows to execute commands inside containers. * **Documentation** * Added guides describing the relay protocol, usage examples, CLI options, lifecycle, and operational notes (workdir, timeouts, sequential execution). * **Chores** * Updated ignore rules to exclude ephemeral relay artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
3.7 KiB
3.7 KiB
File-Based Command Relay (Debugger)
A lightweight client/server system for running commands inside a Docker container from the host, using only a shared filesystem — no networking required.
Overview
Host (Claude Code) Docker Container
┌─────────────┐ ┌─────────────────┐
│ client.sh │ writes cmd file │ server.sh │
│ run "X" │ ───────────────────► │ detects cmd │
│ │ │ executes X │
│ reads │ writes result file │ writes result │
│ result │ ◄─────────────────── │ │
└─────────────┘ └─────────────────┘
└──── shared filesystem (.relay/) ────┘
Assumptions
- The ModelOpt repo is accessible from both host and container (e.g., bind-mounted)
- HuggingFace models are mounted at
/hf-local - The server auto-detects the repo root from the location of
server.sh
Quick Start
1. Start the server (inside Docker)
# The server auto-detects the repo root (two levels up from tools/debugger/)
bash /path/to/modelopt/tools/debugger/server.sh
The server automatically sets the working directory to the repo root. You can override with --workdir.
2. Connect from the host
bash tools/debugger/client.sh handshake
3. Run commands
# Run a simple command
bash tools/debugger/client.sh run "echo hello"
# Run a test script
bash tools/debugger/client.sh run "bash llm_ptq/scripts/huggingface_example.sh"
# Run with a long timeout (default is 600s)
bash tools/debugger/client.sh --timeout 1800 run "python my_long_test.py"
# Check status
bash tools/debugger/client.sh status
Protocol
The relay uses a directory at tools/debugger/.relay/ with this structure:
.relay/
├── server.ready # Written by server on startup
├── client.ready # Written by client during handshake
├── handshake.done # Written by server to confirm handshake
├── cmd/ # Client writes command .sh files here
│ └── <id>.sh # Command to execute
└── result/ # Server writes results here
├── <id>.log # stdout + stderr
└── <id>.exit # Exit code
Handshake
- Server starts, creates
.relay/server.ready - Client writes
.relay/client.ready - Server detects it, writes
.relay/handshake.done - Both sides are now connected
Command Execution
- Client writes a command to
.relay/cmd/<timestamp>.sh - Server detects the file, runs
bash <file>in the workdir, captures output - Server writes
.relay/result/<timestamp>.logand.relay/result/<timestamp>.exit - Server removes the
.shfile; client reads results and cleans up
Options
Server
| Flag | Default | Description |
|---|---|---|
--relay-dir |
<script_dir>/.relay |
Relay directory path |
--workdir |
Auto-detected repo root | Working directory for commands |
Client
| Flag | Default | Description |
|---|---|---|
--relay-dir |
<script_dir>/.relay |
Relay directory path |
--timeout |
600 |
Seconds to wait for command result |
Notes
- The
.relay/directory is in.gitignore— it is not checked in. - Only one server should run at a time (startup clears the relay directory).
- Commands run sequentially in the order the server discovers them.