Files
Chenjie LuoandClaude Opus 4.8 d7df14d12a [tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do?

Type of change: Bug fix

The `tools/debugger` file-based relay assumes a single server but never
enforced it.
Because the relay lives on shared NFS (the repo is often the same
checkout mounted on
multiple hosts), a forgotten `server.sh` on another host kept polling
the same
`.relay/` and could **steal commands** (executing them on the wrong
host), and killing
one server's cleanup could **wipe the active server's markers**.

This adds a `.relay/owner` ownership token (`host:pid:nanos`):

- Each server writes `owner` atomically at startup and **takes over**
instead of
refusing when a stale `server.ready` exists (the old `kill -0 <pid>`
guard was
  host-local and meaningless across hosts).
- The handshake and main loops exit cleanly if `owner` changes
(`[server] Superseded by <id> — exiting.`), so a freshly started server
**evicts**
  any stale one — even on another host.
- `cleanup()` only clears shared markers if we still own them, so a
stepping-down
  server never clobbers its successor's `server.ready`/`owner`.

Also gitignores `tools/debugger/logs/` and documents the `owner` file in
the README.

### Usage

```bash
# Inside the container; a previously-running server elsewhere that shares this
# NFS .relay/ steps down automatically once this one claims ownership:
bash tools/debugger/server.sh
# [server] Note: existing server.ready found (<host:pid:ts>); taking over.
#   (the stale server logs: "[server] Superseded by <id> — exiting.")
```

### Testing

Verified live on computelab: a forgotten `server.sh` on another host was
evicted when
a new server started, after which `client.sh run` executed on the
correct (new) host;
confirmed the stepping-down server's cleanup does not remove the
successor's
`server.ready`/`owner`. `server.sh` passes `bash -n`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- additive: new .relay/owner
file; client.sh and the wire protocol are unchanged -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- the file-based relay
tool has no test harness; behavior verified manually -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- internal dev tooling, not a shipped feature/API -->
- Did you get Claude approval on this PR?: N/A <!-- can run /claude
review -->

### Additional Information

Scope is limited to `tools/debugger/` (`server.sh`, `README.md`,
`.gitignore`).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated relay protocol documentation to clarify ownership-based server
coordination.

* **Bug Fixes**
* Improved reliability of multi-server coordination in shared relay
environments.

* **Chores**
  * Updated ignore patterns for logging files.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:27:52 +00:00

261 lines
10 KiB
Bash
Executable File

#!/usr/bin/env bash
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# File-based command relay server.
# Run this inside the Docker container. It watches for command files from the
# client, executes them, and writes results back.
#
# Usage: bash server.sh [--relay-dir <path>] [--workdir <path>]
set -euo pipefail
RELAY_DIR=""
WORKDIR=""
POLL_INTERVAL=1
# Derive the modelopt repo root from the location of this script (tools/debugger/server.sh)
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
DEFAULT_WORKDIR="$(cd "$SCRIPT_DIR/../.." && pwd)"
usage() {
echo "Usage: $0 [--relay-dir <path>] [--workdir <path>]"
echo ""
echo "Options:"
echo " --relay-dir Path to relay directory (default: <script_dir>/.relay)"
echo " --workdir Working directory for commands (default: auto-detected repo root)"
exit 1
}
while [[ $# -gt 0 ]]; do
case "$1" in
--relay-dir) RELAY_DIR="$2"; shift 2 ;;
--workdir) WORKDIR="$2"; shift 2 ;;
-h|--help) usage ;;
*) echo "Unknown option: $1"; usage ;;
esac
done
# Default relay dir is .relay next to this script
if [[ -z "$RELAY_DIR" ]]; then
RELAY_DIR="$SCRIPT_DIR/.relay"
fi
# Default workdir is the repo root (two levels up from tools/debugger/)
if [[ -z "$WORKDIR" ]]; then
WORKDIR="$DEFAULT_WORKDIR"
fi
CMD_DIR="$RELAY_DIR/cmd"
RESULT_DIR="$RELAY_DIR/result"
echo "[server] Workdir: $WORKDIR"
# Unique id for this server instance. Servers can share one relay dir over NFS
# (e.g. the same repo mounted on multiple hosts), so single-owner is enforced by
# this token rather than by host-local PIDs: whoever writes .relay/owner last wins,
# and the others step down on their next poll (see below).
SERVER_ID="$(hostname):$$:$(date +%s%N)"
cleanup() {
echo "[server] Shutting down..."
# Kill any running command (guard all reads with || true to prevent set -e
# from aborting the trap and leaving stale marker files)
running_pid=$(cut -d: -f2 "$RELAY_DIR/running" 2>/dev/null) || true
if [[ -n "$running_pid" ]]; then
kill -- -"$running_pid" 2>/dev/null || kill "$running_pid" 2>/dev/null || true
fi
# Kill any child processes in our process group
pkill -P $$ 2>/dev/null || true
# Only clear shared markers if we still own the relay — never clobber a
# successor server that has taken over (servers share one NFS relay dir).
if [[ "$(cat "$RELAY_DIR/owner" 2>/dev/null)" == "$SERVER_ID" ]]; then
rm -f "$RELAY_DIR/server.ready" "$RELAY_DIR/handshake.done" \
"$RELAY_DIR/running" "$RELAY_DIR/cancel" "$RELAY_DIR/owner"
fi
exit 0
}
trap cleanup SIGINT SIGTERM
# Set environment
export PYTHONPATH="$WORKDIR"
# A previously-running server (possibly on another host sharing this NFS relay)
# steps down on its next poll once we claim ownership below, so take over rather
# than refuse to start.
if [[ -f "$RELAY_DIR/server.ready" ]]; then
echo "[server] Note: existing server.ready found ($(cat "$RELAY_DIR/server.ready" 2>/dev/null)); taking over."
fi
# Initialize relay directories
rm -rf "$RELAY_DIR"
mkdir -p "$CMD_DIR" "$RESULT_DIR"
# Claim ownership of the relay (single-owner enforcement; see main loop).
echo "$SERVER_ID" > "$RELAY_DIR/owner.tmp" && mv "$RELAY_DIR/owner.tmp" "$RELAY_DIR/owner"
# Ensure modelopt is editable-installed from WORKDIR
check_modelopt_local() {
python3 -c "
import modelopt, os, sys
actual = os.path.realpath(modelopt.__path__[0])
expected = os.path.realpath('$WORKDIR')
if os.path.commonpath([actual, expected]) != expected:
print(f'modelopt loaded from {actual}, expected under {expected}', file=sys.stderr)
sys.exit(1)
" 2>&1
}
if check_modelopt_local >/dev/null 2>&1; then
echo "[server] modelopt already editable-installed from $WORKDIR, skipping pip install."
else
echo "[server] Installing modelopt (python3 -m pip install -e .[dev]) ..."
(cd "$WORKDIR" && python3 -m pip install -e ".[dev]")
if ! check_modelopt_local; then
echo "[server] ERROR: modelopt is not running from the local folder ($WORKDIR)."
echo "[server] Try: python3 -m pip install -e '.[dev]' inside the container, then restart the server."
exit 1
fi
echo "[server] Install done."
fi
# Signal that server is ready
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/server.ready"
echo "[server] Ready. Relay dir: $RELAY_DIR"
echo "[server] Waiting for client handshake..."
# Wait for client handshake
while [[ ! -f "$RELAY_DIR/client.ready" ]]; do
current_owner="$(cat "$RELAY_DIR/owner" 2>/dev/null || true)"
if [[ -n "$current_owner" && "$current_owner" != "$SERVER_ID" ]]; then
echo "[server] Superseded by $current_owner before handshake — exiting."
exit 0
fi
sleep "$POLL_INTERVAL"
done
CLIENT_INFO=$(cat "$RELAY_DIR/client.ready")
echo "[server] Client connected: $CLIENT_INFO"
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/handshake.done"
echo "[server] Handshake complete. Listening for commands..."
# Main loop: watch for command files and re-handshake requests
shopt -s nullglob
while true; do
# Step down if a newer server has claimed the relay (single-owner across hosts).
current_owner="$(cat "$RELAY_DIR/owner" 2>/dev/null || true)"
if [[ -n "$current_owner" && "$current_owner" != "$SERVER_ID" ]]; then
echo "[server] Superseded by $current_owner — exiting."
exit 0
fi
# Detect re-handshake (client flushed and reconnected)
if [[ -f "$RELAY_DIR/client.ready" && ! -f "$RELAY_DIR/handshake.done" ]]; then
CLIENT_INFO=$(cat "$RELAY_DIR/client.ready")
echo "[server] Client re-connected: $CLIENT_INFO"
echo "$(hostname):$$:$(date -Iseconds)" > "$RELAY_DIR/handshake.done"
echo "[server] Re-handshake complete."
fi
for cmd_file in "$CMD_DIR"/*.sh; do
# Guard against command files deleted by the client between glob expansion
# and processing (e.g., client timeout on a queued command)
[[ -f "$cmd_file" ]] || continue
cmd_id="$(basename "$cmd_file" .sh)"
# Tolerate file disappearing between guard and read (TOCTOU with client timeout)
cmd_content=$(cat "$cmd_file" 2>/dev/null) || continue
# Remove command file immediately after reading to prevent re-execution
# and to avoid TOCTOU with client timeout deleting it during execution
rm -f "$cmd_file"
echo "[server] Executing command $cmd_id: $cmd_content"
# Clear any stale cancel file from a previous timed-out client
rm -f "$RELAY_DIR/cancel"
# Create log file and stream output to server console via tail
: > "$RESULT_DIR/$cmd_id.log"
tail -f "$RESULT_DIR/$cmd_id.log" &
tail_pid=$!
# Run in a new process group (setsid) for clean cancellation of entire process tree
(cd "$WORKDIR" && exec setsid bash -c "$cmd_content") >> "$RESULT_DIR/$cmd_id.log" 2>&1 &
cmd_pid=$!
# Track the running command (ID and PID) — atomic write to prevent partial reads
echo "$cmd_id:$cmd_pid" > "$RELAY_DIR/running.tmp"
mv "$RELAY_DIR/running.tmp" "$RELAY_DIR/running"
# Wait for completion or cancellation
cancelled=""
while kill -0 "$cmd_pid" 2>/dev/null; do
if [[ -f "$RELAY_DIR/cancel" ]]; then
# Verify cancel targets this command (reject empty or mismatched signals)
cancel_target=$(cat "$RELAY_DIR/cancel" 2>/dev/null) || true
if [[ "$cancel_target" != "$cmd_id" ]]; then
rm -f "$RELAY_DIR/cancel"
sleep "$POLL_INTERVAL"
continue
fi
echo "[server] Cancelling command $cmd_id (PID $cmd_pid)..."
# Kill entire process group (negative PID) for full tree cleanup
kill -- -"$cmd_pid" 2>/dev/null || kill "$cmd_pid" 2>/dev/null || true
# Wait up to 5s for graceful exit, then escalate to SIGKILL
for _ in $(seq 1 5); do
kill -0 "$cmd_pid" 2>/dev/null || break
sleep 1
done
if kill -0 "$cmd_pid" 2>/dev/null; then
echo "[server] Process $cmd_pid did not exit, sending SIGKILL..."
kill -9 -- -"$cmd_pid" 2>/dev/null || kill -9 "$cmd_pid" 2>/dev/null || true
fi
wait "$cmd_pid" 2>/dev/null || true
cancelled="true"
rm -f "$RELAY_DIR/cancel"
echo "[cancelled]" >> "$RESULT_DIR/$cmd_id.log"
echo "[server] Command $cmd_id cancelled."
break
fi
sleep "$POLL_INTERVAL"
done
# Determine exit code (|| exit_code=$? prevents set -e from killing the
# server when the command exits non-zero)
if [[ -n "$cancelled" ]]; then
exit_code=130
else
exit_code=0
wait "$cmd_pid" 2>/dev/null || exit_code=$?
fi
# Stop console streaming
kill "$tail_pid" 2>/dev/null || true
wait "$tail_pid" 2>/dev/null || true
# Write exit code BEFORE removing the running marker, so any observer
# that sees running disappear can immediately find the result
echo "$exit_code" > "$RESULT_DIR/$cmd_id.exit.tmp"
mv "$RESULT_DIR/$cmd_id.exit.tmp" "$RESULT_DIR/$cmd_id.exit"
# Now safe to remove markers
rm -f "$RELAY_DIR/running"
rm -f "$RELAY_DIR/cancel"
echo "[server] Command $cmd_id finished (exit=$exit_code)"
done
sleep "$POLL_INTERVAL"
done