Host-side (ships with a normal swift build): - LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking) so a wedged stream can't head-of-line-block sibling execs' relays; atomic stdio-or-abort start (patches #5, #6). - Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged channel (patch #7). - ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on force-close — fixes a per-turn leak (per-exec vsock/gRPC connection + runConnections() task) in the long-lived shared control container. Likely the "degrades until app restart" root cause. - ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so a failed launch settles as retryable instead of locking the composer. Guest-side (rides the custom vminitd initfs; inert until the image is built): - ManagedProcess: offload the blocking start off the gRPC event loop (patch #8). - Per-exec cgroups (patch #9) recorded as design only — cross-cutting. Pipeline: - .github/workflows/vminit-image.yml builds vminitd from the vendored source and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed at the custom image. Co-Authored-By: Claude Opus 4.8 <[email protected]>
128 lines
9.7 KiB
Markdown
128 lines
9.7 KiB
Markdown
# Vendored `containerization` — Nucleic patches
|
||
|
||
This is a **vendored copy** of [apple/containerization](https://github.com/apple/containerization)
|
||
at upstream commit `6b7b42ca3efeee8c706070e4355e6a807c5336ae`, referenced by the root `Package.swift`
|
||
via `.package(path: "third_party/containerization")` instead of the github URL.
|
||
|
||
It is vendored (not pulled) because we carry a local patch upstream doesn't have. Keeping it
|
||
in-tree means the patch can't be lost to a dependency re-resolve.
|
||
|
||
## What's changed vs. upstream
|
||
|
||
1. **`Sources/Containerization/LinuxContainer.swift` — forward VM extensions.**
|
||
`LinuxContainer.Configuration` gains a `vmExtensions: [any Sendable]` field, and
|
||
`LinuxContainer` assigns it into `VMConfiguration.extensions` when it builds the VM config.
|
||
Upstream already supports `VMConfiguration.extensions` + the `VZInstanceExtension` hook
|
||
(`configureVZ`/`didCreate`), but `LinuxContainer` — the only entry point we use — never forwarded
|
||
it, so there was no way to attach a device (e.g. a virtio memory balloon) to a container's VM.
|
||
Search for the marker comment `[Nucleic vendored patch]` to find both edit sites.
|
||
|
||
Nucleic uses this to attach a `VZVirtioTraditionalMemoryBalloonDeviceConfiguration` and drive its
|
||
target at runtime for automatic VM memory reclamation — see `MemoryBalloon.swift` /
|
||
`ContainerEngine` in NucleicCore.
|
||
|
||
2. **`Sources/Containerization/LinuxProcess.swift` — process-group kill.**
|
||
`LinuxProcess` gains `killProcessGroup(_:)`, which signals the negative pid (`-pid`) so the
|
||
guest's `kill(2)` targets the exec'd process's whole **process group**, not just the leader.
|
||
Every exec is `setsid()`'d by `vmexec`, so the process is its own group leader (pgid == pid) and
|
||
a group signal reaches the children it forked. Upstream only exposes the leader-only `kill(_:)`,
|
||
which let a forked child survive a Stop in a long-lived shared container. Marked with
|
||
`[Nucleic vendored patch]`; used by `ContainerizedProcessHandle.sendSignal` in NucleicCore.
|
||
|
||
3. **`Sources/Containerization/LinuxProcess.swift` — stdio-connection diagnostics (log-only).**
|
||
`setupIO` logs (`os.Logger`, subsystem `com.nucleic`, category `container-io`) when a *configured*
|
||
stdio stream's guest side never connects — which leaves its host `FileHandle` nil, so the relay /
|
||
readability handler is never wired and the agent's stdin is never delivered (it hangs) or its
|
||
stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers).
|
||
Behavior is unchanged; it only surfaces the failing stream. Marked `[Nucleic vendored patch]`
|
||
(the `import os`, the `nucleicIOLog` static, and the per-stream check in `setupIO`).
|
||
|
||
4. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
|
||
were dropped, and the corresponding `.testTarget(...)` entries removed from `Package.swift`. The
|
||
library/executable targets we build are untouched.
|
||
|
||
5. **`Sources/Containerization/LinuxProcess.swift` — non-blocking stdio relay.**
|
||
Upstream's `setupIO` relays guest stdout/stderr with `FileHandle.availableData`, a **blocking**
|
||
read, from inside a `readabilityHandler`. Those handlers run on Foundation's shared readability
|
||
queue, so if one exec's guest stdout wedged mid-stream that blocking read parked the shared thread
|
||
and head-of-line-blocked **every** other exec's stdout/stderr relay across all containers — one
|
||
stuck session froze the others. The patch marks each connected fd `O_NONBLOCK` and drains it via a
|
||
new `nucleicDrainNonBlocking` (returns bytes + EOF, never blocks; EAGAIN just waits for the next
|
||
readable event). A wedged stream is now contained to its own exec. Marked `[Nucleic vendored patch]`
|
||
(the two static helpers `nucleicSetNonBlocking`/`nucleicDrainNonBlocking` and the two rewritten
|
||
`readabilityHandler` blocks). Requires host-side POSIX `read`/`fcntl`/`errno`.
|
||
|
||
6. **`Sources/Containerization/LinuxProcess.swift` — atomic stdio-or-abort start.**
|
||
In `start()`, after `setupIO` returns, if a *configured* stdio stream never connected from the
|
||
guest (its `FileHandle` is nil — patch #3's logged failure), the patch tears the just-created exec
|
||
back down (`agent.deleteProcess`) and throws instead of calling `startProcess`. Upstream proceeds
|
||
and runs a process with a dead stream (stdin never delivered → hangs; stdout never read → the "no
|
||
output, just a spinner" 60s stall in Nucleic Control). Now that permanent silent stall surfaces as
|
||
a clean, retryable start error. Marked `[Nucleic vendored patch]` (the guard block before
|
||
`startProcess`).
|
||
|
||
7. **`Sources/Containerization/Vminitd.swift` — bounded teardown RPC.**
|
||
`deleteProcess` now sends a 30s `CallOptions.timeout` (upstream sends none, so it can block
|
||
forever on a wedged agent channel). Nucleic calls `LinuxProcess.delete()` after every turn to
|
||
reclaim the per-exec vsock/gRPC connection `exec()` dials; an unbounded `deleteProcess` would let
|
||
that reclaim hang and the connection leak. On the thrown deadline, `performDeletion` still closes
|
||
the agent connection. Marked `[Nucleic vendored patch]` (the `callOpts` block in `deleteProcess`).
|
||
NOTE: this pairs with a Nucleic-side change in `ContainerizedProcessHandle` (call `delete()` after
|
||
the exec exits / on force-close) — without that caller, upstream never deletes execs at all and
|
||
the shared control container leaks a connection + `runConnections()` task per turn.
|
||
|
||
### GUEST-side patches (require rebuilding the initfs — see below)
|
||
|
||
Patches #1–#7 are host-side (the `Containerization` library), shipped by a normal `swift build`.
|
||
Patches #8+ live in `vminitd/` (the guest agent), which rides in the initfs OCI image. They are INERT
|
||
until that image is rebuilt from this source and published, and `ContainerEngine.vminitReference`
|
||
points at it. That is now automated: **`.github/workflows/vminit-image.yml`** builds vminitd from this
|
||
vendored tree and pushes `ghcr.io/abkslm/vminit:<tag>`; `vminitReference` is pinned to that custom
|
||
image. Bump the `-nucleicN` tag suffix and re-run the workflow whenever a guest patch changes.
|
||
|
||
8. **`vminitd/Sources/VminitdCore/ManagedProcess.swift` — offload the blocking start off the event loop.**
|
||
`ManagedProcess.start()` did synchronous, potentially slow pipe reads (waiting for `vmexec` to
|
||
return the pid, then for the error pipe to close) while holding `state`'s Mutex, ON the calling
|
||
task — which is the gRPC handler's event-loop thread. A slow start therefore parked the loop and
|
||
head-of-line-blocked sibling execs' control RPCs sharing it. The patch splits the body into a
|
||
synchronous `startBlocking()` and an async `start()` that runs it on `DispatchQueue.global` via a
|
||
checked continuation, keeping the loop responsive. Safe because the body has no `await` and
|
||
`ManagedProcess` is `Sendable`. Marked `[Nucleic vendored patch]`.
|
||
|
||
### PLANNED guest patch (design recorded; NOT yet implemented)
|
||
|
||
9. **Per-exec cgroups (memory/cpu/pids isolation).** Today the whole container shares ONE cgroup
|
||
(`/container/<id>`): `vmexec run` places the init there via the OCI `cgroupsPath` + `applyResources`
|
||
(`RunCommand.swift`), and each exec joins it via `loadFromPid(init.pid).addProcess` in
|
||
`ManagedProcess.start`. So one session's runaway RSS trips the VM OOM-killer against a *random*
|
||
sibling. Target layout (cgroup v2): make `/container/<id>` an intermediary (enable
|
||
`cgroup.subtree_control` — `Cgroup2Manager.toggleSubtreeControllers` already skips the leaf so this
|
||
composes), move init to a leaf `/container/<id>/init`, and place each exec in its own leaf
|
||
`/container/<id>/<execID>` with generous `memory.high`/`memory.max`/`cpu.max`/`pids.max` so a
|
||
runaway session is throttled/OOM-killed *within its own cgroup*, siblings untouched — WITHOUT
|
||
hard-partitioning RAM (soft limits preserve burst). This is CROSS-CUTTING, not a one-file patch:
|
||
the per-exec limits must be carried on the exec RPC (the `CreateProcess`/exec OCI spec has no
|
||
resources field today), which means a protobuf field (`SandboxContext`) + host-side plumbing
|
||
(`Vminitd.createProcess` / `ContainerEngine.exec`) in addition to the vminitd cgroup restructure
|
||
(`ManagedContainer`, `ManagedProcess`, `vmexec/RunCommand`). Sequence it after #8 lands via CI, and
|
||
validate in a real container (a wrong v2 hierarchy fails at runtime, not at compile).
|
||
|
||
## Re-vendoring a newer upstream commit
|
||
|
||
1. `git clone` upstream (or copy `.build/checkouts/containerization` after bumping the URL pin
|
||
temporarily), check out the desired commit.
|
||
2. `rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \
|
||
--exclude=images/ <upstream>/ third_party/containerization/`
|
||
3. Remove the `.testTarget(...)` blocks from `third_party/containerization/Package.swift`.
|
||
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward), patch #2
|
||
(`LinuxProcess.killProcessGroup(_:)`), patch #3 (the `setupIO` stdio-connection log + its
|
||
`import os` / `nucleicIOLog`), patch #5 (the non-blocking stdio relay: `nucleicSetNonBlocking` /
|
||
`nucleicDrainNonBlocking` + the rewritten `readabilityHandler` blocks), and patch #6 (the atomic
|
||
stdio-or-abort guard in `start()`), patch #7 (the bounded `deleteProcess` timeout in
|
||
`Vminitd.swift`), and patch #8 (the `ManagedProcess.start` event-loop offload in `vminitd/`). Grep
|
||
for `[Nucleic vendored patch]` to find every site. Patch #9 (per-exec cgroups) is design-only so
|
||
far — see its entry. After re-applying any `vminitd/` patch, re-run `.github/workflows/vminit-image.yml`
|
||
to rebuild + publish the custom init image, and bump `ContainerEngine.vminitReference`.
|
||
5. Update the commit hash above and in the root `Package.swift` comment.
|
||
6. `swift build` and run the balloon tests.
|