Per-exec cgroups (guest patch #9): scope a session's OOM/CPU/fork-bomb to itself
Restructures the guest cgroup layout so each exec gets its OWN child cgroup (/container/<id>/<execID>) with memory.oom.group=1, a fair cpu.weight, and a pids.max backstop — so one control session can't OOM-kill, starve, or fork-bomb its siblings in the shared container. The container init moves to its own leaf so the container cgroup can delegate controllers to children (cgroup v2 no-internal-process rule). New Cgroup2Manager helpers: setOomGroup/setCpuWeight/ setPidsMax/remove. Best-effort with graceful fallback: any failure in the per-exec setup wipes the partial state and reverts to today's flat layout, and each exec falls back to the container cgroup — a cgroup hiccup degrades to current behavior, never a failed start. COMPILE-VERIFIED via the musl cross-build; NOT yet runtime-validated. Built as image tag -nucleic2; vminitReference stays on the validated -nucleic1 until -nucleic2 is checked in a real container. A hard host-configured per-exec memory.max (exec-RPC resources field) remains a follow-up. Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
+25
-19
@@ -98,23 +98,28 @@ rebuild whenever a guest patch changes. Built locally, not in CI: the host frame
|
||||
checked continuation, keeping the loop responsive. Safe because the body has no `await` and
|
||||
`ManagedProcess` is `Sendable`. Marked `[Nucleic vendored patch]`.
|
||||
|
||||
### PLANNED guest patch (design recorded; NOT yet implemented)
|
||||
|
||||
9. **Per-exec cgroups (memory/cpu/pids isolation).** Today the whole container shares ONE cgroup
|
||||
(`/container/<id>`): `vmexec run` places the init there via the OCI `cgroupsPath` + `applyResources`
|
||||
(`RunCommand.swift`), and each exec joins it via `loadFromPid(init.pid).addProcess` in
|
||||
`ManagedProcess.start`. So one session's runaway RSS trips the VM OOM-killer against a *random*
|
||||
sibling. Target layout (cgroup v2): make `/container/<id>` an intermediary (enable
|
||||
`cgroup.subtree_control` — `Cgroup2Manager.toggleSubtreeControllers` already skips the leaf so this
|
||||
composes), move init to a leaf `/container/<id>/init`, and place each exec in its own leaf
|
||||
`/container/<id>/<execID>` with generous `memory.high`/`memory.max`/`cpu.max`/`pids.max` so a
|
||||
runaway session is throttled/OOM-killed *within its own cgroup*, siblings untouched — WITHOUT
|
||||
hard-partitioning RAM (soft limits preserve burst). This is CROSS-CUTTING, not a one-file patch:
|
||||
the per-exec limits must be carried on the exec RPC (the `CreateProcess`/exec OCI spec has no
|
||||
resources field today), which means a protobuf field (`SandboxContext`) + host-side plumbing
|
||||
(`Vminitd.createProcess` / `ContainerEngine.exec`) in addition to the vminitd cgroup restructure
|
||||
(`ManagedContainer`, `ManagedProcess`, `vmexec/RunCommand`). Sequence it after #8 lands via CI, and
|
||||
validate in a real container (a wrong v2 hierarchy fails at runtime, not at compile).
|
||||
9. **Per-exec cgroups (OOM/CPU/pids isolation).** Upstream puts the container init AND every exec in
|
||||
ONE cgroup (`/container/<id>`), so one session's runaway RSS trips the in-VM OOM-killer against a
|
||||
*random* sibling, and a fork bomb / CPU hog hits the whole box. This patch makes `/container/<id>`
|
||||
an intermediary: the resource ceiling stays on it, `ManagedContainer.init` moves the init into its
|
||||
own leaf (`/container/<id>/init`) which enables `cgroup.subtree_control` up the chain, and
|
||||
`ManagedProcess.start` places each exec in its OWN child (`/container/<id>/<execID>`) with
|
||||
`memory.oom.group=1` (a runaway session's OOM kills only *its* tree), a fair `cpu.weight`, and a
|
||||
`pids.max` fork-bomb backstop. New `Cgroup2Manager` helpers: `setOomGroup`/`setCpuWeight`/
|
||||
`setPidsMax`/`remove`. **Best-effort with a graceful fallback**: if any step of the per-exec setup
|
||||
fails it wipes the partial state and reverts to the flat layout, and `ManagedProcess` falls back to
|
||||
the container cgroup per exec — so a cgroup hiccup degrades to today's behavior, never a failed
|
||||
start. `ManagedContainer.execCgroupParent == nil` marks flat mode. NOTE: this delivers *scoped-OOM*
|
||||
containment without host-configured limits; a hard per-exec `memory.max` (host-chosen, so a session
|
||||
can't consume the whole box before its own OOM) still wants the exec-RPC resources field
|
||||
(protobuf + `Vminitd.createProcess`/`ContainerEngine.exec` plumbing) — a follow-up. Marked
|
||||
`[Nucleic vendored patch]` across `Cgroup2Manager.swift`, `ManagedContainer.swift`,
|
||||
`ManagedProcess.swift`. **COMPILE-VERIFIED ONLY (musl cross-build); NOT yet runtime-validated** — a
|
||||
wrong cgroup-v2 hierarchy fails at runtime, so boot a container with the new image and confirm
|
||||
sessions start, `/sys/fs/cgroup/container/<id>/<execID>` exists per session, and a hog is contained,
|
||||
before pointing a shipping build at it. `vmexec/RunCommand` is unchanged: it still applies
|
||||
`linux.resources` at `linux.cgroupsPath`, which the patch repoints (init leaf) and clears
|
||||
accordingly.
|
||||
|
||||
## Re-vendoring a newer upstream commit
|
||||
|
||||
@@ -129,8 +134,9 @@ rebuild whenever a guest patch changes. Built locally, not in CI: the host frame
|
||||
`nucleicDrainNonBlocking` + the rewritten `readabilityHandler` blocks), and patch #6 (the atomic
|
||||
stdio-or-abort guard in `start()`), patch #7 (the bounded `deleteProcess` timeout in
|
||||
`Vminitd.swift`), and patch #8 (the `ManagedProcess.start` event-loop offload in `vminitd/`). Grep
|
||||
for `[Nucleic vendored patch]` to find every site. Patch #9 (per-exec cgroups) is design-only so
|
||||
far — see its entry. After re-applying any `vminitd/` patch, rebuild + publish the custom init image
|
||||
for `[Nucleic vendored patch]` to find every site, and patch #9 (per-exec cgroups) across
|
||||
`Cgroup2Manager.swift` / `ManagedContainer.swift` / `ManagedProcess.swift`. After re-applying any
|
||||
`vminitd/` patch, rebuild + publish the custom init image
|
||||
with `make vminit-image` + `make vminit-image-push`, and bump `ContainerEngine.vminitReference`.
|
||||
5. Update the commit hash above and in the root `Package.swift` comment.
|
||||
6. `swift build` and run the balloon tests.
|
||||
|
||||
Reference in New Issue
Block a user