Host-side (ships with a normal swift build): - LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking) so a wedged stream can't head-of-line-block sibling execs' relays; atomic stdio-or-abort start (patches #5, #6). - Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged channel (patch #7). - ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on force-close — fixes a per-turn leak (per-exec vsock/gRPC connection + runConnections() task) in the long-lived shared control container. Likely the "degrades until app restart" root cause. - ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so a failed launch settles as retryable instead of locking the composer. Guest-side (rides the custom vminitd initfs; inert until the image is built): - ManagedProcess: offload the blocking start off the gRPC event loop (patch #8). - Per-exec cgroups (patch #9) recorded as design only — cross-cutting. Pipeline: - .github/workflows/vminit-image.yml builds vminitd from the vendored source and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed at the custom image. Co-Authored-By: Claude Opus 4.8 <[email protected]>
9.7 KiB
Vendored containerization — Nucleic patches
This is a vendored copy of apple/containerization
at upstream commit 6b7b42ca3efeee8c706070e4355e6a807c5336ae, referenced by the root Package.swift
via .package(path: "third_party/containerization") instead of the github URL.
It is vendored (not pulled) because we carry a local patch upstream doesn't have. Keeping it in-tree means the patch can't be lost to a dependency re-resolve.
What's changed vs. upstream
-
Sources/Containerization/LinuxContainer.swift— forward VM extensions.LinuxContainer.Configurationgains avmExtensions: [any Sendable]field, andLinuxContainerassigns it intoVMConfiguration.extensionswhen it builds the VM config. Upstream already supportsVMConfiguration.extensions+ theVZInstanceExtensionhook (configureVZ/didCreate), butLinuxContainer— the only entry point we use — never forwarded it, so there was no way to attach a device (e.g. a virtio memory balloon) to a container's VM. Search for the marker comment[Nucleic vendored patch]to find both edit sites.Nucleic uses this to attach a
VZVirtioTraditionalMemoryBalloonDeviceConfigurationand drive its target at runtime for automatic VM memory reclamation — seeMemoryBalloon.swift/ContainerEnginein NucleicCore. -
Sources/Containerization/LinuxProcess.swift— process-group kill.LinuxProcessgainskillProcessGroup(_:), which signals the negative pid (-pid) so the guest'skill(2)targets the exec'd process's whole process group, not just the leader. Every exec issetsid()'d byvmexec, so the process is its own group leader (pgid == pid) and a group signal reaches the children it forked. Upstream only exposes the leader-onlykill(_:), which let a forked child survive a Stop in a long-lived shared container. Marked with[Nucleic vendored patch]; used byContainerizedProcessHandle.sendSignalin NucleicCore. -
Sources/Containerization/LinuxProcess.swift— stdio-connection diagnostics (log-only).setupIOlogs (os.Logger, subsystemcom.nucleic, categorycontainer-io) when a configured stdio stream's guest side never connects — which leaves its hostFileHandlenil, so the relay / readability handler is never wired and the agent's stdin is never delivered (it hangs) or its stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers). Behavior is unchanged; it only surfaces the failing stream. Marked[Nucleic vendored patch](theimport os, thenucleicIOLogstatic, and the per-stream check insetupIO). -
Trimmed for footprint (no behavior change).
Tests/,docs/,examples/, andimages/were dropped, and the corresponding.testTarget(...)entries removed fromPackage.swift. The library/executable targets we build are untouched. -
Sources/Containerization/LinuxProcess.swift— non-blocking stdio relay. Upstream'ssetupIOrelays guest stdout/stderr withFileHandle.availableData, a blocking read, from inside areadabilityHandler. Those handlers run on Foundation's shared readability queue, so if one exec's guest stdout wedged mid-stream that blocking read parked the shared thread and head-of-line-blocked every other exec's stdout/stderr relay across all containers — one stuck session froze the others. The patch marks each connected fdO_NONBLOCKand drains it via a newnucleicDrainNonBlocking(returns bytes + EOF, never blocks; EAGAIN just waits for the next readable event). A wedged stream is now contained to its own exec. Marked[Nucleic vendored patch](the two static helpersnucleicSetNonBlocking/nucleicDrainNonBlockingand the two rewrittenreadabilityHandlerblocks). Requires host-side POSIXread/fcntl/errno. -
Sources/Containerization/LinuxProcess.swift— atomic stdio-or-abort start. Instart(), aftersetupIOreturns, if a configured stdio stream never connected from the guest (itsFileHandleis nil — patch #3's logged failure), the patch tears the just-created exec back down (agent.deleteProcess) and throws instead of callingstartProcess. Upstream proceeds and runs a process with a dead stream (stdin never delivered → hangs; stdout never read → the "no output, just a spinner" 60s stall in Nucleic Control). Now that permanent silent stall surfaces as a clean, retryable start error. Marked[Nucleic vendored patch](the guard block beforestartProcess). -
Sources/Containerization/Vminitd.swift— bounded teardown RPC.deleteProcessnow sends a 30sCallOptions.timeout(upstream sends none, so it can block forever on a wedged agent channel). Nucleic callsLinuxProcess.delete()after every turn to reclaim the per-exec vsock/gRPC connectionexec()dials; an unboundeddeleteProcesswould let that reclaim hang and the connection leak. On the thrown deadline,performDeletionstill closes the agent connection. Marked[Nucleic vendored patch](thecallOptsblock indeleteProcess). NOTE: this pairs with a Nucleic-side change inContainerizedProcessHandle(calldelete()after the exec exits / on force-close) — without that caller, upstream never deletes execs at all and the shared control container leaks a connection +runConnections()task per turn.
GUEST-side patches (require rebuilding the initfs — see below)
Patches #1–#7 are host-side (the Containerization library), shipped by a normal swift build.
Patches #8+ live in vminitd/ (the guest agent), which rides in the initfs OCI image. They are INERT
until that image is rebuilt from this source and published, and ContainerEngine.vminitReference
points at it. That is now automated: .github/workflows/vminit-image.yml builds vminitd from this
vendored tree and pushes ghcr.io/abkslm/vminit:<tag>; vminitReference is pinned to that custom
image. Bump the -nucleicN tag suffix and re-run the workflow whenever a guest patch changes.
vminitd/Sources/VminitdCore/ManagedProcess.swift— offload the blocking start off the event loop.ManagedProcess.start()did synchronous, potentially slow pipe reads (waiting forvmexecto return the pid, then for the error pipe to close) while holdingstate's Mutex, ON the calling task — which is the gRPC handler's event-loop thread. A slow start therefore parked the loop and head-of-line-blocked sibling execs' control RPCs sharing it. The patch splits the body into a synchronousstartBlocking()and an asyncstart()that runs it onDispatchQueue.globalvia a checked continuation, keeping the loop responsive. Safe because the body has noawaitandManagedProcessisSendable. Marked[Nucleic vendored patch].
PLANNED guest patch (design recorded; NOT yet implemented)
- Per-exec cgroups (memory/cpu/pids isolation). Today the whole container shares ONE cgroup
(
/container/<id>):vmexec runplaces the init there via the OCIcgroupsPath+applyResources(RunCommand.swift), and each exec joins it vialoadFromPid(init.pid).addProcessinManagedProcess.start. So one session's runaway RSS trips the VM OOM-killer against a random sibling. Target layout (cgroup v2): make/container/<id>an intermediary (enablecgroup.subtree_control—Cgroup2Manager.toggleSubtreeControllersalready skips the leaf so this composes), move init to a leaf/container/<id>/init, and place each exec in its own leaf/container/<id>/<execID>with generousmemory.high/memory.max/cpu.max/pids.maxso a runaway session is throttled/OOM-killed within its own cgroup, siblings untouched — WITHOUT hard-partitioning RAM (soft limits preserve burst). This is CROSS-CUTTING, not a one-file patch: the per-exec limits must be carried on the exec RPC (theCreateProcess/exec OCI spec has no resources field today), which means a protobuf field (SandboxContext) + host-side plumbing (Vminitd.createProcess/ContainerEngine.exec) in addition to the vminitd cgroup restructure (ManagedContainer,ManagedProcess,vmexec/RunCommand). Sequence it after #8 lands via CI, and validate in a real container (a wrong v2 hierarchy fails at runtime, not at compile).
Re-vendoring a newer upstream commit
git cloneupstream (or copy.build/checkouts/containerizationafter bumping the URL pin temporarily), check out the desired commit.rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \ --exclude=images/ <upstream>/ third_party/containerization/- Remove the
.testTarget(...)blocks fromthird_party/containerization/Package.swift. - Re-apply patch #1 (the
vmExtensionsfield + thevmConfig.extensions = …forward), patch #2 (LinuxProcess.killProcessGroup(_:)), patch #3 (thesetupIOstdio-connection log + itsimport os/nucleicIOLog), patch #5 (the non-blocking stdio relay:nucleicSetNonBlocking/nucleicDrainNonBlocking+ the rewrittenreadabilityHandlerblocks), and patch #6 (the atomic stdio-or-abort guard instart()), patch #7 (the boundeddeleteProcesstimeout inVminitd.swift), and patch #8 (theManagedProcess.startevent-loop offload invminitd/). Grep for[Nucleic vendored patch]to find every site. Patch #9 (per-exec cgroups) is design-only so far — see its entry. After re-applying anyvminitd/patch, re-run.github/workflows/vminit-image.ymlto rebuild + publish the custom init image, and bumpContainerEngine.vminitReference. - Update the commit hash above and in the root
Package.swiftcomment. swift buildand run the balloon tests.