fix(supervisor): restore canonical stdin after connection loss (#3852)

* fix(supervisor): restore canonical stdin after connection loss

Probe idle SSH peers and enforce a receive deadline during transport I/O,
including writes blocked by a stalled relay. Release the dead attachment's
stdin lease through existing handler cleanup.

Retry denied write intent on later ordinary input without displacing a
healthy owner. Preserve explicit read-only, EOF and detach behavior, and
discard control bytes retained while input ownership was denied.

Cover half-open forwarding, blocked writes, healthy idle peers and competing
reconnects through the production supervisor frame bridge and real SSH.

Fixes #3648

Signed-off-by: Shiju <shiju@nvidia.com>

* docs(skills): describe read-only reconnect input retry

Explain what an openshell-cli user sees when automatic recovery reattaches before the supervisor closes the dead connection: the attachment reports read-only, later ordinary input retries stdin acquisition and prints `input enabled`, input typed while read-only is discarded, exit keys still detach, and an explicitly read-only viewer or a healthy owner is never affected.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
This commit is contained in:
Shiju
2026-09-29 20:28:55 +00:00
committed by GitHub
parent 33a8eac196
commit 0ea0d31020
7 changed files with 787 additions and 11 deletions
@@ -39,6 +39,7 @@ libc = "0.2"
[dev-dependencies]
openshell-ocsf = { path = "../openshell-ocsf", features = ["test-support"] }
tempfile = "3"
tokio = { workspace = true, features = ["test-util"] }
[lints]
workspace = true
+61 -10
View File
@@ -26,6 +26,14 @@ const NO_LOGIN_SHELL_ENV: (&str, &str) = ("OPENSHELL_NO_LOGIN_SHELL", "1");
const MAIN_DETACH_PREFIX: u8 = 0x10;
const MAIN_DETACH_KEY: u8 = 0x11;
const MAIN_DETACH_EOF: u8 = 0x04;
const SSH_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(15);
const SSH_PEER_TIMEOUT: Duration = Duration::from_mins(1);
mod peer_stream;
#[cfg(test)]
#[path = "ssh/reconnect_tests.rs"]
mod reconnect_tests;
fn filter_main_detach_sequence(prefix_pending: &mut bool, data: &[u8]) -> (Vec<u8>, bool) {
let mut forward = Vec::with_capacity(data.len() + usize::from(*prefix_pending));
@@ -70,6 +78,9 @@ fn ssh_server_init(
let mut config = russh::server::Config {
server_id: russh::SshId::Standard(Cow::Owned(format!("SSH-2.0-OpenShell_{VERSION}"))),
auth_rejection_time: Duration::from_secs(1),
// Peer replies keep idle attachments alive without relying on shell I/O.
keepalive_interval: Some(SSH_KEEPALIVE_INTERVAL),
keepalive_max: 3,
..Default::default()
};
config.keys.push(host_key);
@@ -159,9 +170,15 @@ pub async fn run_ssh_server(
let main_session = main_session.clone();
tokio::spawn(async move {
if let Err(err) =
handle_connection(stream, config, port_forward, boundary_exec, main_session)
.await
if let Err(err) = handle_connection(
stream,
config,
port_forward,
boundary_exec,
main_session,
SSH_PEER_TIMEOUT,
)
.await
{
ocsf_emit!(
SshActivityBuilder::new(openshell_ocsf::ctx::ctx())
@@ -300,6 +317,7 @@ async fn handle_connection(
port_forward: Arc<dyn openshell_isolation_interface::contract::BoundaryLoopbackConnector>,
boundary_exec: Arc<dyn openshell_isolation_interface::contract::BoundaryExec>,
main_session: Option<Arc<MainSession>>,
peer_timeout: Duration,
) -> Result<()> {
// Access is gated by the Unix-socket filesystem permissions (root-only),
// not by an application-level preface. The supervisor bridges the
@@ -316,9 +334,12 @@ async fn handle_connection(
);
let handler = SshHandler::new(port_forward, boundary_exec, main_session);
let stream = peer_stream::PeerStream::new(stream, peer_timeout);
russh::server::run_stream(config, stream, handler)
.await
.map_err(|err| miette::miette!("ssh stream error: {err}"))?;
.map_err(|err| miette::miette!("ssh stream error: {err}"))?
.await
.map_err(|err| miette::miette!("ssh session error: {err}"))?;
Ok(())
}
@@ -339,6 +360,9 @@ struct ChannelState {
main_input_owner: Option<u64>,
main_attached: bool,
main_read_only: bool,
/// A writable attachment denied stdin may retry on later input. Explicit
/// viewers and channels that sent EOF must never acquire a new lease.
main_input_pending: bool,
main_detach_prefix_pending: bool,
main_output_task: Option<tokio::task::AbortHandle>,
}
@@ -704,6 +728,7 @@ impl russh::server::Handler for SshHandler {
Err(error) => (None, Some(error)),
}
};
state.main_input_pending = warning.is_some();
state.input_sender = input;
state.main_detach_prefix_pending = false;
let mut output = main_session.subscribe();
@@ -715,7 +740,7 @@ impl russh::server::Handler for SshHandler {
.extended_data(
channel,
1,
format!("openshell: {error}; attached read-only; press Ctrl-C or Ctrl-D to exit{line_ending}").into_bytes(),
format!("openshell: {error}; attached read-only; retry input after the owner disconnects; press Ctrl-C or Ctrl-D to exit{line_ending}").into_bytes(),
)
.await;
}
@@ -831,11 +856,35 @@ impl russh::server::Handler for SshHandler {
.await;
return Ok(());
}
let (forward, detach) = if state.main_attached {
let denied_prefix = state.main_input_pending && state.main_detach_prefix_pending;
let (mut forward, detach) = if state.main_attached {
filter_main_detach_sequence(&mut state.main_detach_prefix_pending, data)
} else {
(data.to_vec(), false)
};
// Remember a denied viewer's prefix only to recognize split detach
// sequences. It must not become process input when a later frame wins
// the lease: that earlier keystroke was sent without write ownership.
if denied_prefix && forward.first() == Some(&MAIN_DETACH_PREFIX) {
forward.remove(0);
}
// A reconnect can precede detection of the old connection's failure.
// Retry only on new input, using the same exclusive acquisition as a
// fresh attachment. Detach keys retain their read-only behavior.
if state.main_attached
&& state.main_input_pending
&& !state.main_read_only
&& !detach
&& !forward.is_empty()
&& let Some(main_session) = self.main_session.as_ref()
&& !main_session.finished()
&& let Ok((owner, input)) = main_session.acquire_input()
{
state.main_input_owner = Some(owner);
state.input_sender = Some(InputSender::Main(input));
state.main_input_pending = false;
session.extended_data(channel, 1, b"openshell: input enabled\r\n".to_vec())?;
}
let error = (!forward.is_empty())
.then(|| state.input_sender.as_ref()?.send(forward).err())
.flatten();
@@ -866,6 +915,7 @@ impl russh::server::Handler for SshHandler {
main_session.release_input(owner);
}
state.input_sender.take();
state.main_input_pending = false;
state.main_detach_prefix_pending = false;
} else {
warn!("channel_eof on unknown channel {channel:?}");
@@ -1086,6 +1136,7 @@ impl SshHandler {
main_session.release_input(owner);
}
state.input_sender.take();
state.main_input_pending = false;
state.main_detach_prefix_pending = false;
if let Some(task) = state.main_output_task.take() {
task.abort();
@@ -1221,7 +1272,7 @@ mod tests {
use std::io::Write as _;
use std::process::{Command, Stdio};
struct AcceptAnyServerKey;
pub(super) struct AcceptAnyServerKey;
impl russh::client::Handler for AcceptAnyServerKey {
type Error = russh::Error;
@@ -1234,7 +1285,7 @@ mod tests {
}
}
struct TestLoopbackConnector;
pub(super) struct TestLoopbackConnector;
#[async_trait::async_trait]
impl openshell_isolation_interface::contract::BoundaryLoopbackConnector for TestLoopbackConnector {
@@ -1256,7 +1307,7 @@ mod tests {
}
}
struct RejectingExec;
pub(super) struct RejectingExec;
#[async_trait::async_trait]
impl openshell_isolation_interface::contract::BoundaryExec for RejectingExec {
@@ -1546,7 +1597,7 @@ mod tests {
assert_eq!(
String::from_utf8_lossy(&data),
format!(
"openshell: canonical main process already has an input owner; attached read-only; press Ctrl-C or Ctrl-D to exit{line_ending}"
"openshell: canonical main process already has an input owner; attached read-only; retry input after the owner disconnects; press Ctrl-C or Ctrl-D to exit{line_ending}"
)
);
}
@@ -0,0 +1,118 @@
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Bound SSH peer silence even while the session is blocked writing to its relay.
use std::future::Future;
use std::io;
use std::pin::Pin;
use std::task::{Context, Poll};
use std::time::Duration;
use tokio::io::{AsyncRead, AsyncWrite, ReadBuf};
use tokio::time::{Instant, Sleep};
/// Only inbound bytes extend the deadline. SSH keepalive replies allow healthy
/// idle clients to retain their connections without sending application input.
pub(super) struct PeerStream<S> {
stream: S,
timeout: Duration,
deadline: Pin<Box<Sleep>>,
}
impl<S> PeerStream<S> {
pub(super) fn new(stream: S, timeout: Duration) -> Self {
Self {
stream,
timeout,
deadline: Box::pin(tokio::time::sleep(timeout)),
}
}
fn poll_deadline(&mut self, cx: &mut Context<'_>) -> io::Result<()> {
if self.deadline.as_mut().poll(cx).is_ready() {
return Err(io::Error::new(
io::ErrorKind::TimedOut,
"SSH peer did not respond before the receive deadline",
));
}
Ok(())
}
}
impl<S: AsyncRead + Unpin> AsyncRead for PeerStream<S> {
fn poll_read(
mut self: Pin<&mut Self>,
cx: &mut Context<'_>,
buf: &mut ReadBuf<'_>,
) -> Poll<io::Result<()>> {
self.poll_deadline(cx)?;
let before = buf.filled().len();
match Pin::new(&mut self.stream).poll_read(cx, buf) {
Poll::Ready(Ok(())) => {
if buf.filled().len() > before {
let next = Instant::now() + self.timeout;
self.deadline.as_mut().reset(next);
}
Poll::Ready(Ok(()))
}
other => other,
}
}
}
impl<S: AsyncWrite + Unpin> AsyncWrite for PeerStream<S> {
fn poll_write(
mut self: Pin<&mut Self>,
cx: &mut Context<'_>,
buf: &[u8],
) -> Poll<io::Result<usize>> {
// Russh awaits packet writes outside its keepalive select. Polling here
// makes an expired peer close even when the relay stops draining output.
self.poll_deadline(cx)?;
Pin::new(&mut self.stream).poll_write(cx, buf)
}
fn poll_flush(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<io::Result<()>> {
self.poll_deadline(cx)?;
Pin::new(&mut self.stream).poll_flush(cx)
}
fn poll_shutdown(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<io::Result<()>> {
self.poll_deadline(cx)?;
Pin::new(&mut self.stream).poll_shutdown(cx)
}
}
#[cfg(test)]
mod tests {
use super::*;
use tokio::io::{AsyncReadExt, AsyncWriteExt};
#[tokio::test(start_paused = true)]
async fn blocked_write_expires_at_receive_deadline() {
let (stream, _peer) = tokio::io::duplex(1);
let mut stream = PeerStream::new(stream, Duration::from_mins(1));
let started = Instant::now();
let error = stream.write_all(b"ab").await.unwrap_err();
assert_eq!(error.kind(), io::ErrorKind::TimedOut);
assert_eq!(started.elapsed(), Duration::from_mins(1));
}
#[tokio::test(start_paused = true)]
async fn only_received_bytes_extend_the_deadline() {
let (stream, mut peer) = tokio::io::duplex(1);
let mut stream = PeerStream::new(stream, Duration::from_mins(1));
let started = Instant::now();
tokio::time::advance(Duration::from_secs(40)).await;
peer.write_all(b"r").await.unwrap();
let mut reply = [0];
stream.read_exact(&mut reply).await.unwrap();
assert_eq!(&reply, b"r");
tokio::time::advance(Duration::from_secs(40)).await;
// A successful write must not renew peer liveness either.
stream.write_all(b"a").await.unwrap();
let error = stream.write_all(b"b").await.unwrap_err();
assert_eq!(error.kind(), io::ErrorKind::TimedOut);
assert_eq!(started.elapsed(), Duration::from_secs(100));
}
}
@@ -0,0 +1,568 @@
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0
//! Exercise a half-open gateway path through the production supervisor frame bridge.
use super::tests::{AcceptAnyServerKey, RejectingExec, TestLoopbackConnector};
use super::*;
use openshell_core::proto::{RelayFrame, relay_frame};
use openshell_isolation_interface::contract::{
BackendError, BoundaryExitStatus, BoundaryProcess, BoundarySignal, ProcessAttachment,
};
use std::sync::atomic::{AtomicUsize, Ordering};
use tokio::io::{AsyncReadExt, AsyncWriteExt};
use tokio::sync::{mpsc, watch};
const TEST_INTERVAL: Duration = Duration::from_millis(50);
const WAIT: Duration = Duration::from_secs(3);
#[derive(Clone, Copy, PartialEq)]
enum RelayMode {
Forward,
Discard,
Stalled,
}
#[derive(Default)]
struct RetainedProcess(AtomicUsize);
#[async_trait::async_trait]
impl BoundaryProcess for RetainedProcess {
async fn wait(&self) -> Result<BoundaryExitStatus, BackendError> {
std::future::pending().await
}
async fn signal(&self, _: BoundarySignal) -> Result<(), BackendError> {
self.0.fetch_add(1, Ordering::SeqCst);
Ok(())
}
async fn terminate(&self) -> Result<(), BackendError> {
self.0.fetch_add(1, Ordering::SeqCst);
Ok(())
}
}
struct CanonicalProcess {
session: Arc<MainSession>,
process: Arc<RetainedProcess>,
input: tokio::io::DuplexStream,
output: Option<tokio::io::DuplexStream>,
}
impl CanonicalProcess {
fn new() -> Self {
let (stdin, input) = tokio::io::duplex(64 * 1024);
let (stdout, output) = tokio::io::duplex(64 * 1024);
let process = Arc::new(RetainedProcess::default());
let session = MainSession::from_boundary(
ProcessAttachment {
stdin: Box::new(stdin),
stdout: Box::new(stdout),
stderr: None,
terminal: None,
},
process.clone(),
);
Self {
session,
process,
input,
output: Some(output),
}
}
async fn assert_input(&mut self, channel: &russh::Channel<russh::client::Msg>) {
channel.data(&b"typing restored\n"[..]).await.unwrap();
self.expect_input().await;
}
async fn expect_input(&mut self) {
let mut received = [0; 16];
tokio::time::timeout(WAIT, self.input.read_exact(&mut received))
.await
.expect("attachment must deliver input to the retained process")
.unwrap();
assert_eq!(&received, b"typing restored\n");
assert!(!self.session.finished());
assert_eq!(self.process.0.load(Ordering::SeqCst), 0);
}
}
struct Connection {
client: russh::client::Handle<AcceptAnyServerKey>,
mode: watch::Sender<RelayMode>,
tasks: Vec<tokio::task::JoinHandle<()>>,
}
impl Drop for Connection {
fn drop(&mut self) {
for task in &self.tasks {
task.abort();
}
}
}
impl Connection {
async fn new(main: Arc<MainSession>) -> Self {
Self::with_window(main, None).await
}
async fn with_window(main: Arc<MainSession>, window: Option<u32>) -> Self {
let dir = tempfile::tempdir().unwrap();
let socket = dir.path().join("ssh.sock");
let (listener, mut config, _) = ssh_server_init(&socket, &None, false).unwrap();
// Preserve whether production enables probes. Only shorten the interval.
let config_mut = Arc::get_mut(&mut config).unwrap();
config_mut.keepalive_interval = config_mut.keepalive_interval.map(|_| TEST_INTERVAL);
let server_task = tokio::spawn(async move {
let (stream, _) = listener.accept().await.unwrap();
let _ = handle_connection(
stream,
config,
Arc::new(TestLoopbackConnector),
Arc::new(RejectingExec),
Some(main),
TEST_INTERVAL * 4,
)
.await;
});
let target = tokio::net::UnixStream::connect(&socket).await.unwrap();
let (to_supervisor, inbound) = mpsc::channel(16);
let (outbound, mut from_supervisor) = mpsc::channel::<RelayFrame>(16);
let relay = tokio::spawn(crate::supervisor_session::test_bridge_ssh_relay(
target, inbound, outbound,
));
let (client_stream, gateway_stream) = tokio::io::duplex(64 * 1024);
let (mut gateway_read, mut gateway_write) = tokio::io::split(gateway_stream);
let (mode, mut input_mode) = watch::channel(RelayMode::Forward);
let mut output_mode = input_mode.clone();
let inbound_pump = tokio::spawn(async move {
let mut bytes = [0; 16 * 1024];
loop {
if *input_mode.borrow() == RelayMode::Stalled {
if input_mode.changed().await.is_err() {
break;
}
continue;
}
tokio::select! {
changed = input_mode.changed() => {
if changed.is_err() { break; }
}
read = gateway_read.read(&mut bytes) => {
let Ok(size) = read else { break; };
if size == 0 { break; }
if *input_mode.borrow() == RelayMode::Forward {
let frame = RelayFrame {
payload: Some(relay_frame::Payload::Data(bytes[..size].to_vec())),
};
if to_supervisor.send(Ok(frame)).await.is_err() { break; }
}
}
}
}
});
let outbound_pump = tokio::spawn(async move {
loop {
if *output_mode.borrow() == RelayMode::Stalled {
if output_mode.changed().await.is_err() {
break;
}
continue;
}
tokio::select! {
changed = output_mode.changed() => {
if changed.is_err() { break; }
}
frame = from_supervisor.recv() => {
let Some(frame) = frame else { break; };
if let Some(relay_frame::Payload::Data(data)) = frame.payload
&& *output_mode.borrow() == RelayMode::Forward
&& gateway_write.write_all(&data).await.is_err()
{
break;
}
}
}
}
});
let mut client_config = russh::client::Config::default();
if let Some(window) = window {
client_config.window_size = window;
}
let mut client = russh::client::connect_stream(
Arc::new(client_config),
client_stream,
AcceptAnyServerKey,
)
.await
.unwrap();
assert!(matches!(
client.authenticate_none("sandbox").await.unwrap(),
russh::client::AuthResult::Success
));
Self {
client,
mode,
tasks: vec![server_task, relay, inbound_pump, outbound_pump],
}
}
async fn attach(&self) -> russh::Channel<russh::client::Msg> {
let mut channel = self.client.channel_open_session().await.unwrap();
channel
.request_subsystem(true, "openshell-main")
.await
.unwrap();
assert!(matches!(
next_event(&mut channel).await,
russh::ChannelMsg::Success
));
channel
}
}
async fn next_event(channel: &mut russh::Channel<russh::client::Msg>) -> russh::ChannelMsg {
tokio::time::timeout(WAIT, channel.wait())
.await
.unwrap()
.unwrap()
}
async fn assert_read_only(channel: &mut russh::Channel<russh::client::Msg>) {
loop {
match next_event(channel).await {
russh::ChannelMsg::ExtendedData { data, ext: 1 } => {
assert!(String::from_utf8_lossy(&data).contains("attached read-only"));
break;
}
russh::ChannelMsg::Data { .. } => {}
other => panic!("expected read-only diagnostic, got {other:?}"),
}
}
}
async fn wait_for_release(main: &MainSession) {
tokio::time::timeout(WAIT, async {
loop {
if let Ok((owner, _)) = main.acquire_input() {
main.release_input(owner);
break;
}
tokio::time::sleep(Duration::from_millis(10)).await;
}
})
.await
.expect("dead relay must release stdin within the liveness bound");
}
async fn reconnect_after_cut(mode: RelayMode, producing: bool) {
let mut main = CanonicalProcess::new();
let original = Connection::new(main.session.clone()).await;
let original_channel = original.attach().await;
main.assert_input(&original_channel).await;
let mut early = Connection::new(main.session.clone()).await;
let mut early_channel = early.attach().await;
assert_read_only(&mut early_channel).await;
let (mut early_read, early_write) = early_channel.split();
let (notice_tx, mut notice_rx) = mpsc::channel(1);
early.tasks.push(tokio::spawn(async move {
// The replacement remains a healthy output consumer while the old
// path is cut. Its own SSH receive window must not become the fault.
while let Some(event) = early_read.wait().await {
if let russh::ChannelMsg::ExtendedData { data, ext: 1 } = event {
let _ = notice_tx.send(data).await;
}
}
}));
original.mode.send_replace(mode);
let output_task = producing.then(|| {
let mut output = main.output.take().unwrap();
tokio::spawn(async move {
// Faster than the stalled relay can drain, but below the retained
// output budget until the SSH event queue has applied backpressure.
let chunk = [b'x'; 16 * 1024];
loop {
if output.write_all(&chunk).await.is_err() {
break;
}
tokio::time::sleep(Duration::from_millis(2)).await;
}
})
});
wait_for_release(&main.session).await;
if let Some(task) = output_task {
task.abort();
}
// The already reconnected writer can acquire the freed lease on new input.
// Keystrokes ignored while the previous owner was alive are never replayed.
early_write.data(&b"typing restored\n"[..]).await.unwrap();
main.expect_input().await;
let notice = tokio::time::timeout(WAIT, notice_rx.recv())
.await
.unwrap()
.unwrap();
assert_eq!(&notice[..], b"openshell: input enabled\r\n");
early_write.close().await.unwrap();
wait_for_release(&main.session).await;
// A fresh connection after cleanup can also attach to that same process.
let mut recovered = Connection::new(main.session.clone()).await;
let recovered_channel = recovered.attach().await;
let (mut recovered_read, recovered_write) = recovered_channel.split();
recovered.tasks.push(tokio::spawn(async move {
while recovered_read.wait().await.is_some() {}
}));
recovered_write
.data(&b"typing restored\n"[..])
.await
.unwrap();
main.expect_input().await;
assert!(main.session.acquire_input().is_err());
}
#[tokio::test]
async fn silent_half_open_relay_releases_canonical_input() {
reconnect_after_cut(RelayMode::Discard, false).await;
}
#[tokio::test]
async fn output_does_not_keep_a_half_open_relay_alive() {
reconnect_after_cut(RelayMode::Discard, true).await;
}
#[tokio::test]
async fn stalled_relay_writes_release_canonical_input() {
reconnect_after_cut(RelayMode::Stalled, true).await;
}
#[tokio::test]
async fn healthy_idle_owner_survives_probes_and_cannot_be_displaced() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let channel = owner.attach().await;
main.assert_input(&channel).await;
tokio::time::sleep(TEST_INTERVAL * 12).await;
let viewer = Connection::new(main.session.clone()).await;
let mut viewer_channel = viewer.attach().await;
assert_read_only(&mut viewer_channel).await;
viewer_channel.data(&b"ignored\n"[..]).await.unwrap();
viewer_channel.eof().await.unwrap();
viewer_channel.close().await.unwrap();
main.assert_input(&channel).await;
assert!(main.session.acquire_input().is_err());
}
#[tokio::test]
async fn explicit_viewer_does_not_acquire_released_input() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let viewer = Connection::new(main.session.clone()).await;
let mut channel = viewer.client.channel_open_session().await.unwrap();
channel
.set_env(true, "OPENSHELL_MAIN_READ_ONLY", "1")
.await
.unwrap();
assert!(matches!(
next_event(&mut channel).await,
russh::ChannelMsg::Success
));
channel
.request_subsystem(true, "openshell-main")
.await
.unwrap();
assert!(matches!(
next_event(&mut channel).await,
russh::ChannelMsg::Success
));
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(&b"ignored\n"[..]).await.unwrap();
tokio::time::sleep(Duration::from_millis(30)).await;
let fresh = Connection::new(main.session.clone()).await;
let fresh_channel = fresh.attach().await;
main.assert_input(&fresh_channel).await;
}
#[tokio::test]
async fn waiting_writer_eof_cancels_acquisition() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let waiting = Connection::new(main.session.clone()).await;
let mut channel = waiting.attach().await;
assert_read_only(&mut channel).await;
channel.eof().await.unwrap();
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(&b"ignored\n"[..]).await.unwrap();
tokio::time::sleep(Duration::from_millis(30)).await;
let fresh = Connection::new(main.session.clone()).await;
let fresh_channel = fresh.attach().await;
main.assert_input(&fresh_channel).await;
}
#[tokio::test]
async fn production_config_probes_before_the_receive_deadline() {
let dir = tempfile::tempdir().unwrap();
let (_, config, _) = ssh_server_init(&dir.path().join("ssh.sock"), &None, false).unwrap();
assert_eq!(config.keepalive_interval, Some(Duration::from_secs(15)));
assert_eq!(config.keepalive_max, 3);
assert_eq!(SSH_PEER_TIMEOUT, Duration::from_mins(1));
}
#[tokio::test]
async fn waiting_writer_detach_keys_do_not_acquire_released_input() {
for key in [b"\x03".as_slice(), b"\x04", b"\x10\x11"] {
let main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let waiting = Connection::new(main.session.clone()).await;
let mut channel = waiting.attach().await;
assert_read_only(&mut channel).await;
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(key).await.unwrap();
loop {
if matches!(next_event(&mut channel).await, russh::ChannelMsg::Close) {
break;
}
}
wait_for_release(&main.session).await;
assert_eq!(main.process.0.load(Ordering::SeqCst), 0);
}
}
#[tokio::test]
async fn competing_waiting_writers_preserve_exclusive_ownership() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let first = Connection::new(main.session.clone()).await;
let mut first_channel = first.attach().await;
assert_read_only(&mut first_channel).await;
let second = Connection::new(main.session.clone()).await;
let mut second_channel = second.attach().await;
assert_read_only(&mut second_channel).await;
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
let (a, b) = tokio::join!(
first_channel.data(&b"a"[..]),
second_channel.data(&b"b"[..])
);
a.unwrap();
b.unwrap();
let mut winner = [0];
tokio::time::timeout(WAIT, main.input.read_exact(&mut winner))
.await
.unwrap()
.unwrap();
let mut extra = [0];
assert!(
tokio::time::timeout(Duration::from_millis(30), main.input.read(&mut extra))
.await
.is_err()
);
let winning_channel = match winner[0] {
b'a' => {
second_channel.close().await.unwrap();
first_channel
}
b'b' => {
first_channel.close().await.unwrap();
second_channel
}
other => panic!("unexpected input: {other}"),
};
tokio::time::sleep(Duration::from_millis(30)).await;
assert!(main.session.acquire_input().is_err());
main.assert_input(&winning_channel).await;
}
#[tokio::test]
async fn detached_waiting_writer_cannot_reacquire_with_pending_output() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
// The read-only diagnostic exhausts this channel's output window, so
// russh retains the channel while the detach close waits behind output.
let waiting = Connection::with_window(main.session.clone(), Some(0)).await;
let channel = waiting.attach().await;
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(&b"\x04"[..]).await.unwrap();
channel.data(&b"after detach"[..]).await.unwrap();
let mut received = [0];
assert!(
tokio::time::timeout(Duration::from_millis(100), main.input.read(&mut received))
.await
.is_err(),
"detached channel forwarded input: {received:?}"
);
let (lease, _) = main
.session
.acquire_input()
.expect("detach must cancel pending ownership");
main.session.release_input(lease);
channel.close().await.unwrap();
wait_for_release(&main.session).await;
}
#[tokio::test]
async fn retry_does_not_forward_prefix_typed_before_owner_release() {
let mut main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let waiting = Connection::new(main.session.clone()).await;
let mut channel = waiting.attach().await;
assert_read_only(&mut channel).await;
channel.data(&b"\x10"[..]).await.unwrap();
// A request response on the same SSH channel proves the prefix was
// processed before the original owner releases its lease.
channel.set_env(true, "REVIEW_BARRIER", "1").await.unwrap();
assert!(matches!(
next_event(&mut channel).await,
russh::ChannelMsg::Success
));
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(&b"z"[..]).await.unwrap();
let mut received = [0];
tokio::time::timeout(WAIT, main.input.read_exact(&mut received))
.await
.unwrap()
.unwrap();
assert_eq!(
received,
[b'z'],
"retry replayed a prefix typed while stdin was denied"
);
}
#[tokio::test]
async fn waiting_writer_split_detach_survives_owner_release() {
let main = CanonicalProcess::new();
let owner = Connection::new(main.session.clone()).await;
let owner_channel = owner.attach().await;
let waiting = Connection::new(main.session.clone()).await;
let mut channel = waiting.attach().await;
assert_read_only(&mut channel).await;
channel.data(&b"\x10"[..]).await.unwrap();
channel.set_env(true, "TEST_BARRIER", "1").await.unwrap();
assert!(matches!(
next_event(&mut channel).await,
russh::ChannelMsg::Success
));
owner_channel.close().await.unwrap();
wait_for_release(&main.session).await;
channel.data(&b"\x11"[..]).await.unwrap();
loop {
if matches!(next_event(&mut channel).await, russh::ChannelMsg::Close) {
break;
}
}
wait_for_release(&main.session).await;
assert_eq!(main.process.0.load(Ordering::SeqCst), 0);
}
@@ -522,6 +522,22 @@ pub async fn report_main_process_exit(
Ok(())
}
#[cfg(test)]
pub(crate) async fn test_bridge_ssh_relay(
target: tokio::net::UnixStream,
inbound: mpsc::Receiver<Result<RelayFrame, tonic::Status>>,
out_tx: mpsc::Sender<RelayFrame>,
) {
let _ = bridge_relay(
Box::new(target),
tokio_stream::wrappers::ReceiverStream::new(inbound),
out_tx,
"half-open-test".into(),
Arc::new(AtomicBool::new(false)),
)
.await;
}
/// Confirm terminal delivery and permit ephemeral cleanup.
pub async fn finalize_main_process_exit(
endpoint: &str,
@@ -687,8 +703,24 @@ async fn handle_relay_open(
}
Err(e) => return Err(format!("relay_stream RPC failed: {e}").into()),
};
let mut inbound = response.into_inner();
bridge_relay(
target,
response.into_inner(),
out_tx,
channel_id,
terminating,
)
.await
}
/// Forward the relay's data frames without interpreting the target protocol.
async fn bridge_relay(
target: Box<dyn TargetStream>,
mut inbound: impl tokio_stream::Stream<Item = Result<RelayFrame, tonic::Status>> + Unpin,
out_tx: mpsc::Sender<RelayFrame>,
channel_id: String,
terminating: Arc<AtomicBool>,
) -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
// Connect to the local SSH daemon on its Unix socket.
let (mut target_r, mut target_w) = tokio::io::split(target);
+4
View File
@@ -278,6 +278,10 @@ process. It retries transient transport failures for up to 60 seconds. Initial
authentication failures, sandbox lifecycle changes, and clean SSH exits are
not retried.
The supervisor sends SSH keepalive probes after 15 seconds without inbound traffic and closes a connection after 60 seconds without receiving peer bytes, including when relay writes are stalled. Healthy idle clients answer these probes and stay attached. The timeout starts when the supervisor last reads peer bytes; relay buffering and scheduling can delay detection after physical network loss. The canonical process keeps running when its dead attachment closes.
If recovery reaches the sandbox before the old connection releases stdin, the replacement reports `attached read-only`. After the old connection times out, type again: the replacement acquires stdin if it is available, reports `input enabled`, and forwards that new input. Keystrokes sent while another attachment owned stdin are discarded. An explicitly read-only attachment stays read-only, and recovery never displaces a healthy input owner.
`Ctrl-C` retains its normal terminal behavior and interrupts the foreground
process. For read-only attachments, `Ctrl-C` or `Ctrl-D` exits only the current
viewer.
+2
View File
@@ -377,6 +377,8 @@ attachments running. Configure VS Code Remote-SSH with:
openshell sandbox ssh-config my-sandbox >> ~/.ssh/config
```
A writable attachment that finds another attachment holding stdin reports `attached read-only; retry input after the owner disconnects`. Automatic recovery can hit this when it reattaches before the supervisor closes the dead connection. The supervisor closes a connection 60 seconds after it last received bytes from it, which can be later than 60 seconds after the network failed if the relay buffered data. After the old owner disconnects or times out, send the input you meant to type next. If stdin is free, the attachment prints `input enabled` and forwards that input to the process, so do not probe with Enter or a prompt answer such as `y`. If nothing prints, the old connection still holds stdin. Input sent while the attachment was read-only never reaches the process, so send it again later. `Ctrl-C`, `Ctrl-D`, and `Ctrl-P` then `Ctrl-Q` still exit a read-only attachment instead of enabling input; enable input first if you need `Ctrl-C` to interrupt the process. Recovery never takes stdin from a healthy owner, and an explicitly read-only attachment stays read-only.
If `connect` reports `canonical main process already finished`, inspect the
result with `sandbox get`. A pending foreground attachment can still retrieve
retained output in `Completed` or `Error`; phase alone does not determine