Container isolation: fix stdio stall + per-exec connection leak; add guest image pipeline
Host-side (ships with a normal swift build): - LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking) so a wedged stream can't head-of-line-block sibling execs' relays; atomic stdio-or-abort start (patches #5, #6). - Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged channel (patch #7). - ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on force-close — fixes a per-turn leak (per-exec vsock/gRPC connection + runConnections() task) in the long-lived shared control container. Likely the "degrades until app restart" root cause. - ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so a failed launch settles as retryable instead of locking the composer. Guest-side (rides the custom vminitd initfs; inert until the image is built): - ManagedProcess: offload the blocking start off the gRPC event loop (patch #8). - Per-exec cgroups (patch #9) recorded as design only — cross-cutting. Pipeline: - .github/workflows/vminit-image.yml builds vminitd from the vendored source and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed at the custom image. Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
@@ -147,7 +147,27 @@ final class ManagedProcess: ContainerProcess, Sendable {
|
||||
}
|
||||
|
||||
extension ManagedProcess {
|
||||
/// [Nucleic vendored patch] Run the blocking start sequence OFF the cooperative executor / gRPC
|
||||
/// event loop. `startBlocking()` does synchronous, potentially slow pipe reads (waiting for
|
||||
/// `vmexec` to hand back the pid, then for the error pipe to close) while holding `state`'s Mutex.
|
||||
/// Upstream ran that directly on the calling task, so a slow exec start parked the event loop and
|
||||
/// head-of-line-blocked sibling execs' control RPCs multiplexed on the same loop (each host `exec`
|
||||
/// dials its own connection, but NIO pins several connections per loop). Dispatching to a worker
|
||||
/// keeps the loop responsive; the body has no `await` and `ManagedProcess` is `Sendable`, so it is
|
||||
/// safe off-actor, and the per-exec Mutex still serializes only this exec's own operations.
|
||||
func start() async throws -> Int32 {
|
||||
try await withCheckedThrowingContinuation { (cont: CheckedContinuation<Int32, Error>) in
|
||||
DispatchQueue.global(qos: .userInitiated).async {
|
||||
do {
|
||||
cont.resume(returning: try self.startBlocking())
|
||||
} catch {
|
||||
cont.resume(throwing: error)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func startBlocking() throws -> Int32 {
|
||||
do {
|
||||
return try self.state.withLock {
|
||||
log.info(
|
||||
|
||||
Reference in New Issue
Block a user