Container isolation: fix stdio stall + per-exec connection leak; add guest image pipeline

Host-side (ships with a normal swift build):
- LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking)
  so a wedged stream can't head-of-line-block sibling execs' relays; atomic
  stdio-or-abort start (patches #5, #6).
- Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged
  channel (patch #7).
- ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on
  force-close — fixes a per-turn leak (per-exec vsock/gRPC connection +
  runConnections() task) in the long-lived shared control container. Likely
  the "degrades until app restart" root cause.
- ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so
  a failed launch settles as retryable instead of locking the composer.

Guest-side (rides the custom vminitd initfs; inert until the image is built):
- ManagedProcess: offload the blocking start off the gRPC event loop (patch #8).
- Per-exec cgroups (patch #9) recorded as design only — cross-cutting.

Pipeline:
- .github/workflows/vminit-image.yml builds vminitd from the vendored source
  and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed
  at the custom image.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
2026-07-13 18:46:48 -07:00
co-authored by Claude Opus 4.8
parent 4bcf34d91a
commit 7972dfb02d
4 changed files with 188 additions and 23 deletions
@@ -147,7 +147,27 @@ final class ManagedProcess: ContainerProcess, Sendable {
}
extension ManagedProcess {
/// [Nucleic vendored patch] Run the blocking start sequence OFF the cooperative executor / gRPC
/// event loop. `startBlocking()` does synchronous, potentially slow pipe reads (waiting for
/// `vmexec` to hand back the pid, then for the error pipe to close) while holding `state`'s Mutex.
/// Upstream ran that directly on the calling task, so a slow exec start parked the event loop and
/// head-of-line-blocked sibling execs' control RPCs multiplexed on the same loop (each host `exec`
/// dials its own connection, but NIO pins several connections per loop). Dispatching to a worker
/// keeps the loop responsive; the body has no `await` and `ManagedProcess` is `Sendable`, so it is
/// safe off-actor, and the per-exec Mutex still serializes only this exec's own operations.
func start() async throws -> Int32 {
try await withCheckedThrowingContinuation { (cont: CheckedContinuation<Int32, Error>) in
DispatchQueue.global(qos: .userInitiated).async {
do {
cont.resume(returning: try self.startBlocking())
} catch {
cont.resume(throwing: error)
}
}
}
}
private func startBlocking() throws -> Int32 {
do {
return try self.state.withLock {
log.info(