Files
containerization/PATCHES.md
T
NucleicandClaude Opus 4.8 4bcf34d91a Nucleic Control: spawn watchdog + reliable Stop + stdio diagnostics
Containerized control sessions could hang with no output ("Working…" forever): the agent's exec
stdio over the vminitd vsock channel intermittently failed to carry bytes, so claude ran and
reached the approval server but its stdin/stdout never connected — it idled in interactive
stream-json mode and never exited. Forensics (nucleic.sqlite + per-session claude-home MCP logs)
showed every stdio layer byte-identical to a working state, i.e. a flaky framework race, not a
regression in our code. Make the failure recoverable and visible instead of an eternal spinner:

- ClaudeCodeBackend: a 60s spawn watchdog on containerized runs — no first stdout → emit a
  recoverable error (with guest stderr + container probe), SIGKILL the wedged process, and finish
  the run errored, instead of awaiting stdoutLines forever.
- Stop reliability: ProcessHandle.forceCloseStreams() (ContainerizedProcessHandle finishes its
  line streams host-side; default no-op for the host pipe handle), wired into every kill
  escalation (terminate/interruptThenKill/killGroupAfter) so a force-killed run always settles
  even when the guest wait/stdio RPC wedges — the real cause of "Stop is inconsistent".
- Cap MCP_TIMEOUT (connection) to 60s so an unreachable approval server can't wedge startup for
  ~24.8 days; MCP_TOOL_TIMEOUT stays unbounded for human-answered approvals.
- LinuxProcess.setupIO logs which stdio stream fails to connect (os.Logger, com.nucleic /
  container-io) so a stall pinpoints the failing stream. Vendored patch #3.
- Tests: stop-escalation force-close, responsive-process no-op, MCP timeout asymmetry.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-25 17:37:46 -07:00

3.7 KiB

Vendored containerization — Nucleic patches

This is a vendored copy of apple/containerization at upstream commit 6b7b42ca3efeee8c706070e4355e6a807c5336ae, referenced by the root Package.swift via .package(path: "third_party/containerization") instead of the github URL.

It is vendored (not pulled) because we carry a local patch upstream doesn't have. Keeping it in-tree means the patch can't be lost to a dependency re-resolve.

What's changed vs. upstream

  1. Sources/Containerization/LinuxContainer.swift — forward VM extensions. LinuxContainer.Configuration gains a vmExtensions: [any Sendable] field, and LinuxContainer assigns it into VMConfiguration.extensions when it builds the VM config. Upstream already supports VMConfiguration.extensions + the VZInstanceExtension hook (configureVZ/didCreate), but LinuxContainer — the only entry point we use — never forwarded it, so there was no way to attach a device (e.g. a virtio memory balloon) to a container's VM. Search for the marker comment [Nucleic vendored patch] to find both edit sites.

    Nucleic uses this to attach a VZVirtioTraditionalMemoryBalloonDeviceConfiguration and drive its target at runtime for automatic VM memory reclamation — see MemoryBalloon.swift / ContainerEngine in NucleicCore.

  2. Sources/Containerization/LinuxProcess.swift — process-group kill. LinuxProcess gains killProcessGroup(_:), which signals the negative pid (-pid) so the guest's kill(2) targets the exec'd process's whole process group, not just the leader. Every exec is setsid()'d by vmexec, so the process is its own group leader (pgid == pid) and a group signal reaches the children it forked. Upstream only exposes the leader-only kill(_:), which let a forked child survive a Stop in a long-lived shared container. Marked with [Nucleic vendored patch]; used by ContainerizedProcessHandle.sendSignal in NucleicCore.

  3. Sources/Containerization/LinuxProcess.swift — stdio-connection diagnostics (log-only). setupIO logs (os.Logger, subsystem com.nucleic, category container-io) when a configured stdio stream's guest side never connects — which leaves its host FileHandle nil, so the relay / readability handler is never wired and the agent's stdin is never delivered (it hangs) or its stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers). Behavior is unchanged; it only surfaces the failing stream. Marked [Nucleic vendored patch] (the import os, the nucleicIOLog static, and the per-stream check in setupIO).

  4. Trimmed for footprint (no behavior change). Tests/, docs/, examples/, and images/ were dropped, and the corresponding .testTarget(...) entries removed from Package.swift. The library/executable targets we build are untouched.

Re-vendoring a newer upstream commit

  1. git clone upstream (or copy .build/checkouts/containerization after bumping the URL pin temporarily), check out the desired commit.
  2. rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \ --exclude=images/ <upstream>/ third_party/containerization/
  3. Remove the .testTarget(...) blocks from third_party/containerization/Package.swift.
  4. Re-apply patch #1 (the vmExtensions field + the vmConfig.extensions = … forward), patch #2 (LinuxProcess.killProcessGroup(_:)), and patch #3 (the setupIO stdio-connection log + its import os / nucleicIOLog). Grep for [Nucleic vendored patch] to find every site.
  5. Update the commit hash above and in the root Package.swift comment.
  6. swift build and run the balloon tests.