Nucleic Control: spawn watchdog + reliable Stop + stdio diagnostics
Containerized control sessions could hang with no output ("Working…" forever): the agent's exec
stdio over the vminitd vsock channel intermittently failed to carry bytes, so claude ran and
reached the approval server but its stdin/stdout never connected — it idled in interactive
stream-json mode and never exited. Forensics (nucleic.sqlite + per-session claude-home MCP logs)
showed every stdio layer byte-identical to a working state, i.e. a flaky framework race, not a
regression in our code. Make the failure recoverable and visible instead of an eternal spinner:
- ClaudeCodeBackend: a 60s spawn watchdog on containerized runs — no first stdout → emit a
recoverable error (with guest stderr + container probe), SIGKILL the wedged process, and finish
the run errored, instead of awaiting stdoutLines forever.
- Stop reliability: ProcessHandle.forceCloseStreams() (ContainerizedProcessHandle finishes its
line streams host-side; default no-op for the host pipe handle), wired into every kill
escalation (terminate/interruptThenKill/killGroupAfter) so a force-killed run always settles
even when the guest wait/stdio RPC wedges — the real cause of "Stop is inconsistent".
- Cap MCP_TIMEOUT (connection) to 60s so an unreachable approval server can't wedge startup for
~24.8 days; MCP_TOOL_TIMEOUT stays unbounded for human-answered approvals.
- LinuxProcess.setupIO logs which stdio stream fails to connect (os.Logger, com.nucleic /
container-io) so a stall pinpoints the failing stream. Vendored patch #3.
- Tests: stop-escalation force-close, responsive-process no-op, MCP timeout asymmetry.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
050605b13e
commit
4bcf34d91a
+13
-4
@@ -29,7 +29,15 @@ in-tree means the patch can't be lost to a dependency re-resolve.
|
||||
which let a forked child survive a Stop in a long-lived shared container. Marked with
|
||||
`[Nucleic vendored patch]`; used by `ContainerizedProcessHandle.sendSignal` in NucleicCore.
|
||||
|
||||
3. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
|
||||
3. **`Sources/Containerization/LinuxProcess.swift` — stdio-connection diagnostics (log-only).**
|
||||
`setupIO` logs (`os.Logger`, subsystem `com.nucleic`, category `container-io`) when a *configured*
|
||||
stdio stream's guest side never connects — which leaves its host `FileHandle` nil, so the relay /
|
||||
readability handler is never wired and the agent's stdin is never delivered (it hangs) or its
|
||||
stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers).
|
||||
Behavior is unchanged; it only surfaces the failing stream. Marked `[Nucleic vendored patch]`
|
||||
(the `import os`, the `nucleicIOLog` static, and the per-stream check in `setupIO`).
|
||||
|
||||
4. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
|
||||
were dropped, and the corresponding `.testTarget(...)` entries removed from `Package.swift`. The
|
||||
library/executable targets we build are untouched.
|
||||
|
||||
@@ -38,9 +46,10 @@ in-tree means the patch can't be lost to a dependency re-resolve.
|
||||
1. `git clone` upstream (or copy `.build/checkouts/containerization` after bumping the URL pin
|
||||
temporarily), check out the desired commit.
|
||||
2. `rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \
|
||||
--exclude=examples/ --exclude=images/ <upstream>/ third_party/containerization/`
|
||||
--exclude=images/ <upstream>/ third_party/containerization/`
|
||||
3. Remove the `.testTarget(...)` blocks from `third_party/containerization/Package.swift`.
|
||||
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward) and patch
|
||||
#2 (`LinuxProcess.killProcessGroup(_:)`). Grep for `[Nucleic vendored patch]` to find every site.
|
||||
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward), patch #2
|
||||
(`LinuxProcess.killProcessGroup(_:)`), and patch #3 (the `setupIO` stdio-connection log + its
|
||||
`import os` / `nucleicIOLog`). Grep for `[Nucleic vendored patch]` to find every site.
|
||||
5. Update the commit hash above and in the root `Package.swift` comment.
|
||||
6. `swift build` and run the balloon tests.
|
||||
|
||||
@@ -21,10 +21,14 @@ import ContainerizationOS
|
||||
import Foundation
|
||||
import Logging
|
||||
import Synchronization
|
||||
import os // [Nucleic vendored patch] stdio-connection diagnostics
|
||||
|
||||
/// `LinuxProcess` represents a Linux process and is used to
|
||||
/// setup and control the full lifecycle for the process.
|
||||
public final class LinuxProcess: Sendable {
|
||||
/// [Nucleic vendored patch] Diagnostic log for stdio stream-connection failures (see `setupIO`).
|
||||
static let nucleicIOLog = os.Logger(subsystem: "com.nucleic", category: "container-io")
|
||||
|
||||
/// The ID of the process. This is purely metadata for the caller.
|
||||
public let id: String
|
||||
|
||||
@@ -97,7 +101,7 @@ public final class LinuxProcess: Sendable {
|
||||
private let agent: any VirtualMachineAgent
|
||||
private let vm: any VirtualMachineInstance
|
||||
private let ociRuntimePath: String?
|
||||
private let logger: Logger?
|
||||
private let logger: Logging.Logger? // [Nucleic vendored patch] disambiguated from os.Logger
|
||||
private let onDelete: (@Sendable () async -> Void)?
|
||||
|
||||
init(
|
||||
@@ -108,7 +112,7 @@ public final class LinuxProcess: Sendable {
|
||||
ociRuntimePath: String?,
|
||||
agent: any VirtualMachineAgent,
|
||||
vm: any VirtualMachineInstance,
|
||||
logger: Logger?,
|
||||
logger: Logging.Logger?, // [Nucleic vendored patch] disambiguated from os.Logger
|
||||
onDelete: (@Sendable () async -> Void)? = nil
|
||||
) {
|
||||
self.id = id
|
||||
@@ -146,6 +150,18 @@ extension LinuxProcess {
|
||||
}
|
||||
}
|
||||
|
||||
// [Nucleic vendored patch] Diagnostics: a configured stdio stream whose guest side never
|
||||
// connected leaves its host FileHandle `nil`, so the relay / readability handler below is
|
||||
// never wired — the agent's stdin is then never delivered (it hangs waiting for input) or
|
||||
// its stdout is never read ("no output, just a spinner"). Log that specific failure (Console
|
||||
// / `log show`, subsystem com.nucleic, category container-io) so a stall pinpoints the stream
|
||||
// instead of proceeding silently. Log-only; behavior is unchanged.
|
||||
let configured = [self.ioSetup.stdin != nil, self.ioSetup.stdout != nil, self.ioSetup.stderr != nil]
|
||||
for (index, label) in [(0, "stdin"), (1, "stdout"), (2, "stderr")] where configured[index] && handles[index] == nil {
|
||||
Self.nucleicIOLog.error(
|
||||
"setupIO[\(self.id, privacy: .public)]: \(label, privacy: .public) stream never connected from the guest — agent stdio will stall")
|
||||
}
|
||||
|
||||
// Note: stdin relay is started separately via startStdinRelay() after
|
||||
// the process has started, to avoid a deadlock where closeStdin is
|
||||
// called before the process is consuming from the pipe.
|
||||
|
||||
Reference in New Issue
Block a user