5 Commits
Author SHA1 Message Date
abkslmandClaude Fable 5 17e84c9573 Deep-sweep fixes: cooperative-pool starvation, stdio tail loss, epoll integrity, UI pins
Host — the two remaining app-wide stall mechanisms plus main-thread pins
found by mining all nine hang reports:
- LinuxProcess.startStdinRelay wrote to a BLOCKING stdin fd on a
  width-limited cooperative-pool thread, non-cancellably; wedged guests
  starved the whole concurrency runtime (decode loops, watchdogs — an
  app-wide freeze surviving the reconcile fix). Writes now offload to a
  per-process GCD queue (vendored patch #18).
- TranscriptWriter (actor) did blocking write/fsync on the cooperative
  pool; it now runs on its own DispatchSerialQueue executor.
- UserMessageBubble's truncation probe typeset entire pasted-log-sized
  messages through CoreText per layout pass (100% main-thread pins in
  the 07-21 hang reports); certainly-long messages now skip the probe
  and render a prefix while collapsed.
- toolGroupSignature JSON-encoded every tool input in the transcript up
  to 12.5x/s on the MainActor; now a structural hash. The summary pass
  is trailing-throttled to 0.4s, and flatItems joins streaming chunks
  once instead of re-copying the prefix per delta.
- StatusFeedFetcher.parseDate allocated three formatters per call (86%
  of a pool thread in the 07-26 report); now shared statics.

Guest (vminitd) — teardown data loss and epoll registration hazards:
- IOPair no longer closes on a bare EPOLLHUP with a backpressure flush
  in flight (dropped the CLI's final output line); EPOLLOUT finishes the
  flush, then EOF closes loss-free. ManagedProcess.setExit closes only
  stdin, letting stdout/stderr self-close on EOF, with an 8s grace pass
  (patch #16).
- Epoll events carry a registration generation; the supervisor ignores
  stale events for recycled fd numbers. registerFd refuses EEXIST
  instead of clobbering the existing handler. TerminalIO's stdin relay
  writes a dup of the terminal fd so its backpressure registration
  can't collide with the stdout relay's (patch #17).
- VsockProxy flushes bytes parked toward the surviving peer on hangup,
  closes the dialing socket on a failed backend connect, and
  StandardIO/TerminalIO clean up partially-created pairs on setup
  failure (patch #16).

Full suite: 1451+292+74+20 tests, two failures — both pre-existing
environmental (MacVM base image absent on this machine; a load-flaky
liveness test that passes 3/3 in isolation).

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-07-28 15:35:35 -07:00
abkslmandClaude Opus 4.8 831d19c3a6 Per-exec cgroups follow-up: host-configured hard memory.max (no protobuf)
Adds an opt-in hard per-session memory ceiling on top of patch #9's scoped-OOM.
The exec already ships the full OCI Spec, so the limit rides
spec.linux.resources.memory.limit — no RPC/protobuf change:

- host framework: LinuxProcessConfiguration.memoryLimitInBytes; LinuxContainer.exec
  stamps it onto the exec spec.
- guest: Server+GRPC.createProcess reads it back and applies it as the exec
  cgroup's memory.max (new Cgroup2Manager.setMemoryMax) via createExec/ManagedProcess.
- Nucleic: ContainerServiceSettings.controlPerSessionMemoryGiB (default 0 = off),
  applied only to the shared control container (ContainerManager.exec); wired
  through ContainerEngine.exec.

So one session can't consume the whole shared container's memory before its own
(oom.group-scoped) OOM. Default off preserves #9's behavior. Compile-verified host
+ musl guest; rides the pending -nucleic2 image, still runtime-pending.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-07-13 20:12:42 -07:00
abkslmandClaude Opus 4.8 2eb563c90c Per-exec cgroups (guest patch #9): scope a session's OOM/CPU/fork-bomb to itself
Restructures the guest cgroup layout so each exec gets its OWN child cgroup
(/container/<id>/<execID>) with memory.oom.group=1, a fair cpu.weight, and a
pids.max backstop — so one control session can't OOM-kill, starve, or fork-bomb
its siblings in the shared container. The container init moves to its own leaf
so the container cgroup can delegate controllers to children (cgroup v2
no-internal-process rule). New Cgroup2Manager helpers: setOomGroup/setCpuWeight/
setPidsMax/remove.

Best-effort with graceful fallback: any failure in the per-exec setup wipes the
partial state and reverts to today's flat layout, and each exec falls back to the
container cgroup — a cgroup hiccup degrades to current behavior, never a failed
start.

COMPILE-VERIFIED via the musl cross-build; NOT yet runtime-validated. Built as
image tag -nucleic2; vminitReference stays on the validated -nucleic1 until
-nucleic2 is checked in a real container. A hard host-configured per-exec
memory.max (exec-RPC resources field) remains a follow-up.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-07-13 20:03:20 -07:00
abkslmandClaude Opus 4.8 7972dfb02d Container isolation: fix stdio stall + per-exec connection leak; add guest image pipeline
Host-side (ships with a normal swift build):
- LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking)
  so a wedged stream can't head-of-line-block sibling execs' relays; atomic
  stdio-or-abort start (patches #5, #6).
- Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged
  channel (patch #7).
- ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on
  force-close — fixes a per-turn leak (per-exec vsock/gRPC connection +
  runConnections() task) in the long-lived shared control container. Likely
  the "degrades until app restart" root cause.
- ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so
  a failed launch settles as retryable instead of locking the composer.

Guest-side (rides the custom vminitd initfs; inert until the image is built):
- ManagedProcess: offload the blocking start off the gRPC event loop (patch #8).
- Per-exec cgroups (patch #9) recorded as design only — cross-cutting.

Pipeline:
- .github/workflows/vminit-image.yml builds vminitd from the vendored source
  and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed
  at the custom image.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-07-13 18:46:48 -07:00
NucleicandClaude Opus 4.8 11b9825e09 Vendor apple/containerization with a VM-extensions forwarding patch
Switch the containerization dependency from the github URL to a vendored copy
(third_party/containerization, upstream commit 6b7b42ca) referenced by path, so
we can carry a small local patch that upstream lacks: LinuxContainer.Configuration
gains a `vmExtensions` field forwarded into VMConfiguration.extensions. Upstream
already supports VMConfiguration.extensions + the VZInstanceExtension hook, but
LinuxContainer — our only entry point — never forwarded them, so there was no way
to attach a device (e.g. a memory balloon) to a container's VM.

Tests/, docs/, examples/, images/ and the corresponding test targets are trimmed
for footprint (we never build the dependency's tests). See PATCHES.md for the full
diff vs. upstream and the re-vendoring procedure. Also adds the ContainerizationExtras
product to NucleicCore (AddressAllocator, named in the configureVZ signature).

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-21 20:22:21 -07:00