Per-exec cgroups follow-up: host-configured hard memory.max (no protobuf)
Adds an opt-in hard per-session memory ceiling on top of patch #9's scoped-OOM. The exec already ships the full OCI Spec, so the limit rides spec.linux.resources.memory.limit — no RPC/protobuf change: - host framework: LinuxProcessConfiguration.memoryLimitInBytes; LinuxContainer.exec stamps it onto the exec spec. - guest: Server+GRPC.createProcess reads it back and applies it as the exec cgroup's memory.max (new Cgroup2Manager.setMemoryMax) via createExec/ManagedProcess. - Nucleic: ContainerServiceSettings.controlPerSessionMemoryGiB (default 0 = off), applied only to the shared control container (ContainerManager.exec); wired through ContainerEngine.exec. So one session can't consume the whole shared container's memory before its own (oom.group-scoped) OOM. Default off preserves #9's behavior. Compile-verified host + musl guest; rides the pending -nucleic2 image, still runtime-pending. Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
+14
-6
@@ -109,12 +109,20 @@ rebuild whenever a guest patch changes. Built locally, not in CI: the host frame
|
||||
`setPidsMax`/`remove`. **Best-effort with a graceful fallback**: if any step of the per-exec setup
|
||||
fails it wipes the partial state and reverts to the flat layout, and `ManagedProcess` falls back to
|
||||
the container cgroup per exec — so a cgroup hiccup degrades to today's behavior, never a failed
|
||||
start. `ManagedContainer.execCgroupParent == nil` marks flat mode. NOTE: this delivers *scoped-OOM*
|
||||
containment without host-configured limits; a hard per-exec `memory.max` (host-chosen, so a session
|
||||
can't consume the whole box before its own OOM) still wants the exec-RPC resources field
|
||||
(protobuf + `Vminitd.createProcess`/`ContainerEngine.exec` plumbing) — a follow-up. Marked
|
||||
`[Nucleic vendored patch]` across `Cgroup2Manager.swift`, `ManagedContainer.swift`,
|
||||
`ManagedProcess.swift`. **COMPILE-VERIFIED ONLY (musl cross-build); NOT yet runtime-validated** — a
|
||||
start. `ManagedContainer.execCgroupParent == nil` marks flat mode.
|
||||
|
||||
Beyond scoped-OOM, a **hard host-configured per-exec `memory.max`** is also wired — WITHOUT a
|
||||
protobuf change, because the exec already ships the full OCI `Spec` and the guest just ignored
|
||||
`linux.resources`. Host: `LinuxProcessConfiguration.memoryLimitInBytes` → `LinuxContainer.exec`
|
||||
stamps it onto `spec.linux.resources.memory.limit`. Guest: `Server+GRPC.createProcess` reads that
|
||||
back and passes it to `createExec`/`ManagedProcess`, which sets `memory.max` (new
|
||||
`Cgroup2Manager.setMemoryMax`) on the exec's cgroup — so a session can't consume the whole box
|
||||
before its own OOM. Driven by Nucleic's `ContainerServiceSettings.controlPerSessionMemoryGiB`
|
||||
(default 0 = off; applied only to the shared control container, via `ContainerManager.exec`), so the
|
||||
default stays scoped-OOM-only. Marked `[Nucleic vendored patch]` across `Cgroup2Manager.swift`,
|
||||
`ManagedContainer.swift`, `ManagedProcess.swift`, `Server+GRPC.swift` (guest) and
|
||||
`LinuxProcessConfiguration.swift`, `LinuxContainer.swift` (host). **COMPILE-VERIFIED (host + musl
|
||||
cross-build); NOT yet runtime-validated** — a
|
||||
wrong cgroup-v2 hierarchy fails at runtime, so boot a container with the new image and confirm
|
||||
sessions start, `/sys/fs/cgroup/container/<id>/<execID>` exists per session, and a hog is contained,
|
||||
before pointing a shipping build at it. `vmexec/RunCommand` is unchanged: it still applies
|
||||
|
||||
Reference in New Issue
Block a user