# Nucleic — Container session isolation > **Status (2026-07-15):** host-side fixes are **live** (any `dev` build) — now including the > control-plane accept-loop hardening (#10/#11) behind the recurring all-sessions "produced no output > within 60s" stall. The custom vminit image **`-nucleic3` is built, pushed, and DEPLOYED** > (`ContainerEngine.vminitReference` points at it): it carries the offload (#8), per-exec cgroups > (#9, with the graceful flat-layout fallback), and the guest halves of the control-plane hardening > (#10/#12 — the VsockProxy fd-leak and black-hole-listener fixes). #9/#10/#12 are compile-verified > but were deployed ahead of the §5 runtime validation by explicit decision — if containers > misbehave, revert `vminitReference` to `-nucleic1` (the previous validated tag); the > reference-keyed initfs cache re-pulls on next launch either way. The spawn-stall watchdog's probe + > container auto-recovery (below) additionally converts any control-plane wedge from "restart the > app" into "retry the turn". Why this exists: Nucleic Control runs several agent sessions concurrently, and to keep that affordable on an average Mac it runs them as separate `exec`s inside **one shared control container** (per agent family) rather than one VM per session — see [RUNTIME_ARCHITECTURE](RUNTIME_ARCHITECTURE.md) (`ContainerEngine` / `ContainerManager`) and [VSOCK_CONTROL_PLANE](VSOCK_CONTROL_PLANE.md). Sharing a container is a memory optimization, but it created a blast-radius problem: one session could wedge, leak, OOM-kill, starve, or fork-bomb the others. This document is the map of how each of those vectors is now contained **without** giving up the shared-container memory efficiency (project/coding sessions already get one container each — `ContainerManager.containerName(for:)` — so this is specifically about the *shared control* container). The low-level per-patch detail lives in [`third_party/containerization/PATCHES.md`](../third_party/containerization/PATCHES.md) (the vendored Apple `containerization` framework carries Nucleic patches #1–#12). This doc is the "why + where + how to build/validate" overview. ## 1. The vectors and the fixes | "One session takes down another" via… | Fix | Where | State | | --- | --- | --- | --- | | **stdio wedge** — a stuck stream freezes *sibling* sessions' relays | Non-blocking per-exec stdio relay (`O_NONBLOCK` + `nucleicDrainNonBlocking`) replacing blocking `FileHandle.availableData`; **atomic stdio-or-abort start** (fail fast + retryable instead of a silent 60s stall) | vendored `LinuxProcess.swift` (patches #5, #6); recoverable-error mapping in `ClaudeCodeBackend` | **live** (host build) | | **per-turn connection leak** — the "degrades until you quit the app" driver | `exec()` already dials a DEDICATED per-exec vsock/gRPC agent connection; the leak was that **Nucleic never called `LinuxProcess.delete()`**, so each turn leaked a connection + `runConnections()` task + guest exec record. Now `ContainerizedProcessHandle` fires `delete()` on exit and on force-close; `deleteProcess` is bounded by a 30s deadline (patch #7) | `ContainerizedProcessHandle.swift`; vendored `Vminitd.swift` | **live** (host build) | | **control-plane head-of-line block** — a slow exec start stalls sibling RPCs on the shared event loop | Offload `ManagedProcess.start`'s blocking pipe reads off the gRPC event loop (`DispatchQueue.global` + continuation) | vendored `vminitd/ManagedProcess.swift` (patch #8) | in deployed **`-nucleic3`** | | **OOM cross-kill** — a runaway session's OOM kills a *random* sibling | Per-exec cgroup with `memory.oom.group=1` so an OOM kills only the offending session's process tree | vendored `vminitd/…` (patch #9) | deployed in **`-nucleic3`** (runtime validation pending) | | **CPU starvation / fork bomb** | Per-exec `cpu.weight` (fair share) + `pids.max` (backstop) | vendored `vminitd/…` (patch #9) | deployed in **`-nucleic3`** (runtime validation pending) | | **one session eating all memory** before its own OOM | Opt-in host-configured per-exec `memory.max` (rides the exec's OCI spec — no protobuf change) | vendored host + guest + `ContainerManager`/`ContainerEngine`; setting `controlPerSessionMemoryGiB` | deployed in **`-nucleic3`** (runtime validation pending) | | **control-plane black hole** — one bad accept/connect permanently kills the shared control socket's relay while it stays bound; every session then hangs in MCP init ("produced no output within 60s") until an app restart | Accept loops tolerate transient `accept(2)` errors (#10); host relay contains per-connection failures + closes both ends on failure (#11); guest `VsockProxy` cleanup is leak-proof/once-guarded, registrations aren't `try!`, and a dead accept loop closes its listener so peers fail fast (#12). Backstops: `listen` backlog 16→`SOMAXCONN`, the in-guest bridge is supervised (restart loop), and the spawn-stall watchdog now **probes the control plane from inside the guest** and **auto-recreates the container** when the probe confirms it dead | vendored `Socket.swift` / `UnixSocketRelay.swift` / `vminitd/VsockProxy.swift`; `MCPApprovalServer` / `ContainerEngine` / `ContainerManager` / `ClaudeCodeBackend` | host halves + probe/recovery **live** (host build); guest halves deployed in **`-nucleic3`** (runtime validation pending) | Note the connection-leak finding also settled a design question: **per-exec control channels already exist** — every `LinuxContainer.exec()` dials its own agent connection. The "one shared channel per container" concern applies only to the container's init/lifecycle agent, not to session runs. ## 2. Two shipping surfaces: host vs guest - **Host-side** patches compile into the app via a normal `swift build` (the `Containerization` library) — they take effect the moment the code is in `dev`, no image work. This is the stdio fixes, the connection-leak fix, the bounded teardown, and the recoverable-error mapping. - **Guest-side** patches live in `third_party/containerization/vminitd/` (the guest agent, PID 1). They ride the **vminit initfs OCI image** and are **inert until that image is rebuilt and published** and `ContainerEngine.vminitReference` points at it. This is the offload (#8) and the cgroups work (#9). The materialized initfs (`…/Nucleic/containers/initfs/vminit.ext4`) is cached **keyed on the image reference** (a `vminit.ext4.reference` sidecar): repointing `vminitReference` re-pulls on next launch. (Before the sidecar existed the first-ever materialized initfs was trusted forever, so a repoint silently kept booting the old guest — verify a rollout actually landed by checking the sidecar's contents.) The pull authenticates with the app's GitHub token when the image is on ghcr.io, so the package can stay private; if the pull fails and a cached initfs exists, the cache is used and the pull retries next launch. ## 3. The per-exec cgroup layout (#9) cgroup v2's "no internal process" rule means a cgroup can hold processes **or** delegate controllers to children, not both. So the container cgroup is restructured into an intermediary: ``` /container/ ← intermediary: resource ceiling (memory.max) + subtree_control delegated ├── init ← leaf: the container init (sleep infinity + control-bridge) ├── exec-aaaa ← leaf: session A's agent — memory.oom.group=1, cpu.weight, pids.max, [memory.max] ├── exec-bbbb ← leaf: session B's agent — … └── … ``` - `ManagedContainer.init` builds this: creates `/container/`, creates `/container//init`, enables `cgroup.subtree_control` up the chain (by enabling controllers from the init leaf), applies the container ceiling to the parent, and repoints the init's `cgroupsPath` at the leaf. - `ManagedProcess.start` places each exec in its own `/container//` and sets `memory.oom.group=1`, `cpu.weight=100`, `pids.max=4096`, and (if a host limit was passed) `memory.max`. - **Graceful fallback:** if *any* per-exec setup step fails, `ManagedContainer` wipes the partial state and reverts to the upstream **flat** layout (init + all execs share `/container/`), and each exec independently falls back to the container cgroup. `execCgroupParent == nil` marks flat mode. A cgroup hiccup therefore degrades to today's behavior — it can never block a container from starting. This is the key safety property given the change is compile-verified but not yet runtime-validated. New `Cgroup2Manager` helpers: `setOomGroup` / `setCpuWeight` / `setPidsMax` / `setMemoryMax` / `remove`. ## 4. Hard per-exec `memory.max` (opt-in follow-up) `memory.oom.group` scopes an OOM to the offending session, but without a cap a session can still consume the whole container before that OOM. The hard cap closes that — and needs **no protobuf change**, because the exec already ships the full OCI `Spec` and the guest was just ignoring `linux.resources`: - **Host:** `LinuxProcessConfiguration.memoryLimitInBytes` → `LinuxContainer.exec` stamps it onto `spec.linux.resources.memory.limit`. - **Guest:** `Server+GRPC.createProcess` reads it back and applies it as the exec cgroup's `memory.max` (`ManagedProcess` / `Cgroup2Manager.setMemoryMax`). - **Nucleic:** `ContainerServiceSettings.controlPerSessionMemoryGiB` — **default `0` = off**. When > 0 it applies (via `ContainerManager.exec`) **only to the shared control container** (per-session sandbox containers already run one session each). Off by default preserves #9's generous scoped-OOM behavior. ## 5. Building & validating the custom vminit image The guest image is built **locally**, not in CI: `make containerization` compiles the host framework, which uses macOS 26+ `Virtualization`/vmnet symbols that GitHub-hosted runners' SDKs lack (and the app targets macOS 27). The content/image store is **user-global** (`~/Library/Application Support/com.apple.containerization/`), so the built image is reachable from any checkout. Root-`Makefile` targets: ```bash make vminit-image-prep # ONE-TIME per machine: swiftly + the aarch64 musl static SDK (~1 GB) make vminit-image # build cctl (Xcode Swift) + cross-build vminitd/vmexec (musl) → package the image make vminit-image-login # ONE-TIME: cctl login → GHCR token stored in the macOS Keychain (no env vars) make vminit-image-push # push to ghcr.io/abkslm/vminit (auth from the Keychain) ``` (The targets force `WARNINGS_AS_ERRORS=false` — Xcode's Swift 6.4 rejects `-warnings-as-errors` alongside SwiftPM's `-suppress-warnings` on dependencies. `cctl` is built on demand so the targets work from a fresh checkout.) **Tag scheme & the current-vs-next invariant:** - `-nucleicN` is our guest-patch revision (`0.34.0` = the upstream vminit version we replace). - `Makefile: NUCLEIC_VMINIT_REF` = the tag you **build** (the *next* image). - `ContainerEngine.vminitReference` = the **currently deployed, validated** tag. - On a new guest patch: bump `NUCLEIC_VMINIT_REF`, `make vminit-image` + `make vminit-image-push`, **validate**, then set `vminitReference` to match. So `NUCLEIC_VMINIT_REF` may be one revision ahead while a build is validated. Currently both are on `-nucleic3` (#9 + the #10/#12 control-plane hardening), deployed ahead of full validation by explicit decision — revert `vminitReference` to `-nucleic1` if it misbehaves. **Validation gate (before pointing a shipping build at a new guest image):** boot a real Control session and confirm (a) sessions start normally, (b) `/sys/fs/cgroup/container//` exists per session, and (c) — with `controlPerSessionMemoryGiB` set — a memory hog is capped/killed within its own session and siblings survive. If good, bump `vminitReference`; if not, the graceful fallback means it's still running (just flat) and you stay on the previous tag. ## 6. Re-vendoring note The Apple `containerization` framework is vendored at a pinned commit with Nucleic patches #1–#12. When re-vendoring a newer upstream, re-apply every `[Nucleic vendored patch]` site (grep for that marker) per [`PATCHES.md`](../third_party/containerization/PATCHES.md), then rebuild + publish the image and bump `vminitReference`. ## 7. Open follow-up - A host-tunable per-exec `cpu.max` (hard CPU quota) and `pids.max` would be a small extension of the same OCI-spec-riding mechanism used for `memory.max` in §4. - The `controlPerSessionMemoryGiB` policy is a flat GiB cap; a fraction-of-container or concurrency-aware policy could be layered on if static caps prove too blunt.