Per-exec cgroups follow-up: host-configured hard memory.max (no protobuf)

Adds an opt-in hard per-session memory ceiling on top of patch #9's scoped-OOM.
The exec already ships the full OCI Spec, so the limit rides
spec.linux.resources.memory.limit — no RPC/protobuf change:

- host framework: LinuxProcessConfiguration.memoryLimitInBytes; LinuxContainer.exec
  stamps it onto the exec spec.
- guest: Server+GRPC.createProcess reads it back and applies it as the exec
  cgroup's memory.max (new Cgroup2Manager.setMemoryMax) via createExec/ManagedProcess.
- Nucleic: ContainerServiceSettings.controlPerSessionMemoryGiB (default 0 = off),
  applied only to the shared control container (ContainerManager.exec); wired
  through ContainerEngine.exec.

So one session can't consume the whole shared container's memory before its own
(oom.group-scoped) OOM. Default off preserves #9's behavior. Compile-verified host
+ musl guest; rides the pending -nucleic2 image, still runtime-pending.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
2026-07-13 20:12:42 -07:00
co-authored by Claude Opus 4.8
parent 2eb563c90c
commit 831d19c3a6
7 changed files with 53 additions and 8 deletions
@@ -288,6 +288,12 @@ public struct Cgroup2Manager: Sendable {
try Self.writeValue(path: self.path, value: String(max), fileName: "pids.max")
}
/// [Nucleic vendored patch] Hard memory ceiling (`memory.max`) — a host-configured per-exec cap so
/// one session can't consume the whole container's memory before its own (oom.group-scoped) OOM.
package func setMemoryMax(bytes: UInt64) throws {
try Self.writeValue(path: self.path, value: String(bytes), fileName: "memory.max")
}
/// [Nucleic vendored patch] Remove this cgroup directory (rmdir). The cgroup must already be empty
/// of processes and child cgroups. Best-effort partial-setup cleanup for the per-exec layout.
package func remove() throws {
@@ -201,7 +201,8 @@ extension ManagedContainer {
func createExec(
id: String,
stdio: HostStdio,
process: ContainerizationOCI.Process
process: ContainerizationOCI.Process,
memoryLimitBytes: UInt64? = nil // [Nucleic vendored patch] hard per-exec memory.max
) throws {
log.debug("creating exec process with \(process)")
@@ -217,6 +218,7 @@ extension ManagedContainer {
bundle: self.bundle,
owningPid: self.initProcess.pid,
execCgroupParent: self.execCgroupParent, // [Nucleic vendored patch] per-exec cgroup
execMemoryLimitBytes: memoryLimitBytes, // [Nucleic vendored patch]
log: self.log
)
self.execs[id] = process
@@ -60,6 +60,8 @@ final class ManagedProcess: ContainerProcess, Sendable {
// [Nucleic vendored patch] Parent cgroup for this exec's OWN per-exec child (`<parent>/<id>`);
// nil means the legacy flat layout (join the container/init cgroup via `owningPid`).
private let execCgroupParent: String?
// [Nucleic vendored patch] Hard per-exec memory.max (bytes) for this exec's cgroup; nil = none.
private let execMemoryLimitBytes: UInt64?
private let ackPipe: Pipe
private let syncPipe: Pipe
private let errorPipe: Pipe
@@ -78,6 +80,7 @@ final class ManagedProcess: ContainerProcess, Sendable {
bundle: ContainerizationOCI.Bundle,
owningPid: Int32? = nil,
execCgroupParent: String? = nil, // [Nucleic vendored patch]
execMemoryLimitBytes: UInt64? = nil, // [Nucleic vendored patch]
log: Logger
) throws {
self.id = id
@@ -86,6 +89,7 @@ final class ManagedProcess: ContainerProcess, Sendable {
self.log = log
self.owningPid = owningPid
self.execCgroupParent = execCgroupParent
self.execMemoryLimitBytes = execMemoryLimitBytes
let syncPipe = Pipe()
try syncPipe.setCloexec()
@@ -233,6 +237,9 @@ extension ManagedProcess {
try? execCg.setOomGroup(true)
try? execCg.setCpuWeight(100)
try? execCg.setPidsMax(4096)
if let limit = execMemoryLimitBytes, limit > 0 {
try? execCg.setMemoryMax(bytes: limit) // host-configured hard per-exec ceiling
}
try execCg.addProcess(pid: pid)
} catch {
log.error("per-exec cgroup for \(id) failed; joining the container cgroup: \(error)")
@@ -895,10 +895,18 @@ extension Initd: Com_Apple_Containerization_Sandbox_V3_SandboxContext.SimpleServ
// This is an exec.
if let container = await self.state.containers[request.containerID] {
// [Nucleic vendored patch] A per-exec memory ceiling rides the exec's OCI
// resources (set host-side by LinuxContainer.exec); apply it as this exec's
// own memory.max (patch #9). Only positive limits count.
let execMemoryLimit: UInt64? = {
guard let limit = ociSpec.linux?.resources?.memory?.limit, limit > 0 else { return nil }
return UInt64(limit)
}()
try await container.createExec(
id: request.id,
stdio: stdioPorts,
process: process
process: process,
memoryLimitBytes: execMemoryLimit
)
} else {
// We need to make our new fangled container.