mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix The `tools/debugger` file-based relay assumes a single server but never enforced it. Because the relay lives on shared NFS (the repo is often the same checkout mounted on multiple hosts), a forgotten `server.sh` on another host kept polling the same `.relay/` and could **steal commands** (executing them on the wrong host), and killing one server's cleanup could **wipe the active server's markers**. This adds a `.relay/owner` ownership token (`host:pid:nanos`): - Each server writes `owner` atomically at startup and **takes over** instead of refusing when a stale `server.ready` exists (the old `kill -0 <pid>` guard was host-local and meaningless across hosts). - The handshake and main loops exit cleanly if `owner` changes (`[server] Superseded by <id> — exiting.`), so a freshly started server **evicts** any stale one — even on another host. - `cleanup()` only clears shared markers if we still own them, so a stepping-down server never clobbers its successor's `server.ready`/`owner`. Also gitignores `tools/debugger/logs/` and documents the `owner` file in the README. ### Usage ```bash # Inside the container; a previously-running server elsewhere that shares this # NFS .relay/ steps down automatically once this one claims ownership: bash tools/debugger/server.sh # [server] Note: existing server.ready found (<host:pid:ts>); taking over. # (the stale server logs: "[server] Superseded by <id> — exiting.") ``` ### Testing Verified live on computelab: a forgotten `server.sh` on another host was evicted when a new server started, after which `client.sh run` executed on the correct (new) host; confirmed the stepping-down server's cleanup does not remove the successor's `server.ready`/`owner`. `server.sh` passes `bash -n`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- additive: new .relay/owner file; client.sh and the wire protocol are unchanged --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A <!-- the file-based relay tool has no test harness; behavior verified manually --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- internal dev tooling, not a shipped feature/API --> - Did you get Claude approval on this PR?: N/A <!-- can run /claude review --> ### Additional Information Scope is limited to `tools/debugger/` (`server.sh`, `README.md`, `.gitignore`). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated relay protocol documentation to clarify ownership-based server coordination. * **Bug Fixes** * Improved reliability of multi-server coordination in shared relay environments. * **Chores** * Updated ignore patterns for logging files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
261 lines
10 KiB
Bash
Executable File
261 lines
10 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
# File-based command relay server.
|
|
# Run this inside the Docker container. It watches for command files from the
|
|
# client, executes them, and writes results back.
|
|
#
|
|
# Usage: bash server.sh [--relay-dir <path>] [--workdir <path>]
|
|
|
|
set -euo pipefail
|
|
|
|
RELAY_DIR=""
|
|
WORKDIR=""
|
|
POLL_INTERVAL=1
|
|
|
|
# Derive the modelopt repo root from the location of this script (tools/debugger/server.sh)
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
DEFAULT_WORKDIR="$(cd "$SCRIPT_DIR/../.." && pwd)"
|
|
|
|
usage() {
|
|
echo "Usage: $0 [--relay-dir <path>] [--workdir <path>]"
|
|
echo ""
|
|
echo "Options:"
|
|
echo " --relay-dir Path to relay directory (default: <script_dir>/.relay)"
|
|
echo " --workdir Working directory for commands (default: auto-detected repo root)"
|
|
exit 1
|
|
}
|
|
|
|
while [[ $# -gt 0 ]]; do
|
|
case "$1" in
|
|
--relay-dir) RELAY_DIR="$2"; shift 2 ;;
|
|
--workdir) WORKDIR="$2"; shift 2 ;;
|
|
-h|--help) usage ;;
|
|
*) echo "Unknown option: $1"; usage ;;
|
|
esac
|
|
done
|
|
|
|
# Default relay dir is .relay next to this script
|
|
if [[ -z "$RELAY_DIR" ]]; then
|
|
RELAY_DIR="$SCRIPT_DIR/.relay"
|
|
fi
|
|
|
|
# Default workdir is the repo root (two levels up from tools/debugger/)
|
|
if [[ -z "$WORKDIR" ]]; then
|
|
WORKDIR="$DEFAULT_WORKDIR"
|
|
fi
|
|
|
|
CMD_DIR="$RELAY_DIR/cmd"
|
|
RESULT_DIR="$RELAY_DIR/result"
|
|
|
|
echo "[server] Workdir: $WORKDIR"
|
|
|
|
# Unique id for this server instance. Servers can share one relay dir over NFS
|
|
# (e.g. the same repo mounted on multiple hosts), so single-owner is enforced by
|
|
# this token rather than by host-local PIDs: whoever writes .relay/owner last wins,
|
|
# and the others step down on their next poll (see below).
|
|
SERVER_ID="$(hostname):$$:$(date +%s%N)"
|
|
|
|
cleanup() {
|
|
echo "[server] Shutting down..."
|
|
# Kill any running command (guard all reads with || true to prevent set -e
|
|
# from aborting the trap and leaving stale marker files)
|
|
running_pid=$(cut -d: -f2 "$RELAY_DIR/running" 2>/dev/null) || true
|
|
if [[ -n "$running_pid" ]]; then
|
|
kill -- -"$running_pid" 2>/dev/null || kill "$running_pid" 2>/dev/null || true
|
|
fi
|
|
# Kill any child processes in our process group
|
|
pkill -P $$ 2>/dev/null || true
|
|
# Only clear shared markers if we still own the relay — never clobber a
|
|
# successor server that has taken over (servers share one NFS relay dir).
|
|
if [[ "$(cat "$RELAY_DIR/owner" 2>/dev/null)" == "$SERVER_ID" ]]; then
|
|
rm -f "$RELAY_DIR/server.ready" "$RELAY_DIR/handshake.done" \
|
|
"$RELAY_DIR/running" "$RELAY_DIR/cancel" "$RELAY_DIR/owner"
|
|
fi
|
|
exit 0
|
|
}
|
|
trap cleanup SIGINT SIGTERM
|
|
|
|
# Set environment
|
|
export PYTHONPATH="$WORKDIR"
|
|
|
|
# A previously-running server (possibly on another host sharing this NFS relay)
|
|
# steps down on its next poll once we claim ownership below, so take over rather
|
|
# than refuse to start.
|
|
if [[ -f "$RELAY_DIR/server.ready" ]]; then
|
|
echo "[server] Note: existing server.ready found ($(cat "$RELAY_DIR/server.ready" 2>/dev/null)); taking over."
|
|
fi
|
|
|
|
# Initialize relay directories
|
|
rm -rf "$RELAY_DIR"
|
|
mkdir -p "$CMD_DIR" "$RESULT_DIR"
|
|
|
|
# Claim ownership of the relay (single-owner enforcement; see main loop).
|
|
echo "$SERVER_ID" > "$RELAY_DIR/owner.tmp" && mv "$RELAY_DIR/owner.tmp" "$RELAY_DIR/owner"
|
|
|
|
# Ensure modelopt is editable-installed from WORKDIR
|
|
check_modelopt_local() {
|
|
python3 -c "
|
|
import modelopt, os, sys
|
|
actual = os.path.realpath(modelopt.__path__[0])
|
|
expected = os.path.realpath('$WORKDIR')
|
|
if os.path.commonpath([actual, expected]) != expected:
|
|
print(f'modelopt loaded from {actual}, expected under {expected}', file=sys.stderr)
|
|
sys.exit(1)
|
|
" 2>&1
|
|
}
|
|
|
|
if check_modelopt_local >/dev/null 2>&1; then
|
|
echo "[server] modelopt already editable-installed from $WORKDIR, skipping pip install."
|
|
else
|
|
echo "[server] Installing modelopt (python3 -m pip install -e .[dev]) ..."
|
|
(cd "$WORKDIR" && python3 -m pip install -e ".[dev]")
|
|
if ! check_modelopt_local; then
|
|
echo "[server] ERROR: modelopt is not running from the local folder ($WORKDIR)."
|
|
echo "[server] Try: python3 -m pip install -e '.[dev]' inside the container, then restart the server."
|
|
exit 1
|
|
fi
|
|
echo "[server] Install done."
|
|
fi
|
|
|
|
# Signal that server is ready
|
|
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/server.ready"
|
|
echo "[server] Ready. Relay dir: $RELAY_DIR"
|
|
echo "[server] Waiting for client handshake..."
|
|
|
|
# Wait for client handshake
|
|
while [[ ! -f "$RELAY_DIR/client.ready" ]]; do
|
|
current_owner="$(cat "$RELAY_DIR/owner" 2>/dev/null || true)"
|
|
if [[ -n "$current_owner" && "$current_owner" != "$SERVER_ID" ]]; then
|
|
echo "[server] Superseded by $current_owner before handshake — exiting."
|
|
exit 0
|
|
fi
|
|
sleep "$POLL_INTERVAL"
|
|
done
|
|
|
|
CLIENT_INFO=$(cat "$RELAY_DIR/client.ready")
|
|
echo "[server] Client connected: $CLIENT_INFO"
|
|
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/handshake.done"
|
|
echo "[server] Handshake complete. Listening for commands..."
|
|
|
|
# Main loop: watch for command files and re-handshake requests
|
|
shopt -s nullglob
|
|
while true; do
|
|
# Step down if a newer server has claimed the relay (single-owner across hosts).
|
|
current_owner="$(cat "$RELAY_DIR/owner" 2>/dev/null || true)"
|
|
if [[ -n "$current_owner" && "$current_owner" != "$SERVER_ID" ]]; then
|
|
echo "[server] Superseded by $current_owner — exiting."
|
|
exit 0
|
|
fi
|
|
|
|
# Detect re-handshake (client flushed and reconnected)
|
|
if [[ -f "$RELAY_DIR/client.ready" && ! -f "$RELAY_DIR/handshake.done" ]]; then
|
|
CLIENT_INFO=$(cat "$RELAY_DIR/client.ready")
|
|
echo "[server] Client re-connected: $CLIENT_INFO"
|
|
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/handshake.done"
|
|
echo "[server] Re-handshake complete."
|
|
fi
|
|
|
|
for cmd_file in "$CMD_DIR"/*.sh; do
|
|
# Guard against command files deleted by the client between glob expansion
|
|
# and processing (e.g., client timeout on a queued command)
|
|
[[ -f "$cmd_file" ]] || continue
|
|
|
|
cmd_id="$(basename "$cmd_file" .sh)"
|
|
# Tolerate file disappearing between guard and read (TOCTOU with client timeout)
|
|
cmd_content=$(cat "$cmd_file" 2>/dev/null) || continue
|
|
# Remove command file immediately after reading to prevent re-execution
|
|
# and to avoid TOCTOU with client timeout deleting it during execution
|
|
rm -f "$cmd_file"
|
|
echo "[server] Executing command $cmd_id: $cmd_content"
|
|
|
|
# Clear any stale cancel file from a previous timed-out client
|
|
rm -f "$RELAY_DIR/cancel"
|
|
|
|
# Create log file and stream output to server console via tail
|
|
: > "$RESULT_DIR/$cmd_id.log"
|
|
tail -f "$RESULT_DIR/$cmd_id.log" &
|
|
tail_pid=$!
|
|
|
|
# Run in a new process group (setsid) for clean cancellation of entire process tree
|
|
(cd "$WORKDIR" && exec setsid bash -c "$cmd_content") >> "$RESULT_DIR/$cmd_id.log" 2>&1 &
|
|
cmd_pid=$!
|
|
|
|
# Track the running command (ID and PID) — atomic write to prevent partial reads
|
|
echo "$cmd_id:$cmd_pid" > "$RELAY_DIR/running.tmp"
|
|
mv "$RELAY_DIR/running.tmp" "$RELAY_DIR/running"
|
|
|
|
# Wait for completion or cancellation
|
|
cancelled=""
|
|
while kill -0 "$cmd_pid" 2>/dev/null; do
|
|
if [[ -f "$RELAY_DIR/cancel" ]]; then
|
|
# Verify cancel targets this command (reject empty or mismatched signals)
|
|
cancel_target=$(cat "$RELAY_DIR/cancel" 2>/dev/null) || true
|
|
if [[ "$cancel_target" != "$cmd_id" ]]; then
|
|
rm -f "$RELAY_DIR/cancel"
|
|
sleep "$POLL_INTERVAL"
|
|
continue
|
|
fi
|
|
echo "[server] Cancelling command $cmd_id (PID $cmd_pid)..."
|
|
# Kill entire process group (negative PID) for full tree cleanup
|
|
kill -- -"$cmd_pid" 2>/dev/null || kill "$cmd_pid" 2>/dev/null || true
|
|
# Wait up to 5s for graceful exit, then escalate to SIGKILL
|
|
for _ in $(seq 1 5); do
|
|
kill -0 "$cmd_pid" 2>/dev/null || break
|
|
sleep 1
|
|
done
|
|
if kill -0 "$cmd_pid" 2>/dev/null; then
|
|
echo "[server] Process $cmd_pid did not exit, sending SIGKILL..."
|
|
kill -9 -- -"$cmd_pid" 2>/dev/null || kill -9 "$cmd_pid" 2>/dev/null || true
|
|
fi
|
|
wait "$cmd_pid" 2>/dev/null || true
|
|
cancelled="true"
|
|
rm -f "$RELAY_DIR/cancel"
|
|
echo "[cancelled]" >> "$RESULT_DIR/$cmd_id.log"
|
|
echo "[server] Command $cmd_id cancelled."
|
|
break
|
|
fi
|
|
sleep "$POLL_INTERVAL"
|
|
done
|
|
|
|
# Determine exit code (|| exit_code=$? prevents set -e from killing the
|
|
# server when the command exits non-zero)
|
|
if [[ -n "$cancelled" ]]; then
|
|
exit_code=130
|
|
else
|
|
exit_code=0
|
|
wait "$cmd_pid" 2>/dev/null || exit_code=$?
|
|
fi
|
|
|
|
# Stop console streaming
|
|
kill "$tail_pid" 2>/dev/null || true
|
|
wait "$tail_pid" 2>/dev/null || true
|
|
|
|
# Write exit code BEFORE removing the running marker, so any observer
|
|
# that sees running disappear can immediately find the result
|
|
echo "$exit_code" > "$RESULT_DIR/$cmd_id.exit.tmp"
|
|
mv "$RESULT_DIR/$cmd_id.exit.tmp" "$RESULT_DIR/$cmd_id.exit"
|
|
|
|
# Now safe to remove markers
|
|
rm -f "$RELAY_DIR/running"
|
|
rm -f "$RELAY_DIR/cancel"
|
|
|
|
echo "[server] Command $cmd_id finished (exit=$exit_code)"
|
|
done
|
|
|
|
sleep "$POLL_INTERVAL"
|
|
done
|