Per-exec cgroups (guest patch #9): scope a session's OOM/CPU/fork-bomb to itself
Restructures the guest cgroup layout so each exec gets its OWN child cgroup (/container/<id>/<execID>) with memory.oom.group=1, a fair cpu.weight, and a pids.max backstop — so one control session can't OOM-kill, starve, or fork-bomb its siblings in the shared container. The container init moves to its own leaf so the container cgroup can delegate controllers to children (cgroup v2 no-internal-process rule). New Cgroup2Manager helpers: setOomGroup/setCpuWeight/ setPidsMax/remove. Best-effort with graceful fallback: any failure in the per-exec setup wipes the partial state and reverts to today's flat layout, and each exec falls back to the container cgroup — a cgroup hiccup degrades to current behavior, never a failed start. COMPILE-VERIFIED via the musl cross-build; NOT yet runtime-validated. Built as image tag -nucleic2; vminitReference stays on the validated -nucleic1 until -nucleic2 is checked in a real container. A hard host-configured per-exec memory.max (exec-RPC resources field) remains a follow-up. Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
@@ -267,6 +267,33 @@ public struct Cgroup2Manager: Sendable {
|
||||
fileName: "memory.low")
|
||||
}
|
||||
|
||||
/// [Nucleic vendored patch] Make the kernel OOM-killer treat this cgroup as an atomic unit: when a
|
||||
/// memory limit (this cgroup's or an ancestor's) forces an OOM, the whole cgroup's process tree is
|
||||
/// killed together rather than one victim. Used to scope a runaway exec's OOM to that exec so
|
||||
/// sibling execs in the same container survive.
|
||||
package func setOomGroup(_ enabled: Bool) throws {
|
||||
try Self.writeValue(path: self.path, value: enabled ? "1" : "0", fileName: "memory.oom.group")
|
||||
}
|
||||
|
||||
/// [Nucleic vendored patch] Relative CPU share under contention (cgroup v2 `cpu.weight`, 1…10000,
|
||||
/// default 100). Equal weights give each exec a fair slice so one busy session can't starve
|
||||
/// siblings of CPU.
|
||||
package func setCpuWeight(_ weight: UInt64) throws {
|
||||
try Self.writeValue(path: self.path, value: String(weight), fileName: "cpu.weight")
|
||||
}
|
||||
|
||||
/// [Nucleic vendored patch] Cap the pids in this cgroup (`pids.max`) — a fork-bomb backstop so one
|
||||
/// exec can't exhaust the pid space and wedge its siblings.
|
||||
package func setPidsMax(_ max: UInt64) throws {
|
||||
try Self.writeValue(path: self.path, value: String(max), fileName: "pids.max")
|
||||
}
|
||||
|
||||
/// [Nucleic vendored patch] Remove this cgroup directory (rmdir). The cgroup must already be empty
|
||||
/// of processes and child cgroups. Best-effort partial-setup cleanup for the per-exec layout.
|
||||
package func remove() throws {
|
||||
try FileManager.default.removeItem(at: self.path)
|
||||
}
|
||||
|
||||
package func getMemoryEvents() throws -> MemoryEvents {
|
||||
let content = try readFileContent(fileName: "memory.events")
|
||||
let values = parseKeyValuePairs(content)
|
||||
|
||||
Reference in New Issue
Block a user