Deep-sweep fixes: cooperative-pool starvation, stdio tail loss, epoll integrity, UI pins

Host — the two remaining app-wide stall mechanisms plus main-thread pins
found by mining all nine hang reports:
- LinuxProcess.startStdinRelay wrote to a BLOCKING stdin fd on a
  width-limited cooperative-pool thread, non-cancellably; wedged guests
  starved the whole concurrency runtime (decode loops, watchdogs — an
  app-wide freeze surviving the reconcile fix). Writes now offload to a
  per-process GCD queue (vendored patch #18).
- TranscriptWriter (actor) did blocking write/fsync on the cooperative
  pool; it now runs on its own DispatchSerialQueue executor.
- UserMessageBubble's truncation probe typeset entire pasted-log-sized
  messages through CoreText per layout pass (100% main-thread pins in
  the 07-21 hang reports); certainly-long messages now skip the probe
  and render a prefix while collapsed.
- toolGroupSignature JSON-encoded every tool input in the transcript up
  to 12.5x/s on the MainActor; now a structural hash. The summary pass
  is trailing-throttled to 0.4s, and flatItems joins streaming chunks
  once instead of re-copying the prefix per delta.
- StatusFeedFetcher.parseDate allocated three formatters per call (86%
  of a pool thread in the 07-26 report); now shared statics.

Guest (vminitd) — teardown data loss and epoll registration hazards:
- IOPair no longer closes on a bare EPOLLHUP with a backpressure flush
  in flight (dropped the CLI's final output line); EPOLLOUT finishes the
  flush, then EOF closes loss-free. ManagedProcess.setExit closes only
  stdin, letting stdout/stderr self-close on EOF, with an 8s grace pass
  (patch #16).
- Epoll events carry a registration generation; the supervisor ignores
  stale events for recycled fd numbers. registerFd refuses EEXIST
  instead of clobbering the existing handler. TerminalIO's stdin relay
  writes a dup of the terminal fd so its backpressure registration
  can't collide with the stdout relay's (patch #17).
- VsockProxy flushes bytes parked toward the surviving peer on hangup,
  closes the dialing socket on a failed backend connect, and
  StandardIO/TerminalIO clean up partially-created pairs on setup
  failure (patch #16).

Full suite: 1451+292+74+20 tests, two failures — both pre-existing
environmental (MacVM base image absent on this machine; a load-flaky
liveness test that passes 3/3 in isolation).

Co-Authored-By: Claude Fable 5 <[email protected]>
This commit is contained in:
2026-07-28 15:35:35 -07:00
co-authored by Claude Fable 5
parent 6f950d859b
commit 17e84c9573
10 changed files with 358 additions and 104 deletions
@@ -23,7 +23,18 @@ import Synchronization
final class ProcessSupervisor: Sendable {
private let poller: Epoll
private let handlers = Mutex<[Int32: @Sendable (Epoll.Mask) -> Void]>([:])
/// [Nucleic vendored patch] Handler table keyed by fd, each entry stamped with the
/// registration generation echoed back through epoll (`Epoll.Event.generation`).
/// Dispatch compares the event's generation against the live entry's, so an event
/// queued for a CLOSED registration of a recycled fd number can never fire the new
/// registration's handler (a stale EPOLLHUP used to be able to tear down a brand-new
/// healthy connection that inherited the number mid-batch).
private struct HandlerTable {
var nextGeneration: UInt32 = 1
var entries: [Int32: (generation: UInt32, handler: @Sendable (Epoll.Mask) -> Void)] = [:]
}
private let handlers = Mutex<HandlerTable>(HandlerTable())
private let queue: DispatchQueue
// `DispatchSourceSignal` is thread-safe.
@@ -58,8 +69,13 @@ final class ProcessSupervisor: Sendable {
return
}
for event in events {
let handler = self.handlers.withLock { $0[event.fd] }
handler?(event.mask)
// [Nucleic vendored patch] Dispatch only when the event belongs to the
// CURRENT registration of this fd number — an earlier handler in this
// batch may have closed the fd and something else re-registered the
// recycled number already (see HandlerTable).
let entry = self.handlers.withLock { $0.entries[event.fd] }
guard let entry, entry.generation == event.generation else { continue }
entry.handler(event.mask)
}
}
}
@@ -70,23 +86,36 @@ final class ProcessSupervisor: Sendable {
///
/// The handler is stored before the fd is added to epoll, ensuring no
/// events are missed.
///
/// [Nucleic vendored patch] Refuses (EEXIST) an fd that is already registered rather
/// than clobbering its handler: the old overwrite-then-fail-EEXIST path destroyed the
/// existing registration's handler AND removed the map entry, leaving the fd armed in
/// epoll with no handler — a silently dead relay (the TerminalIO shared-fd case).
func registerFd(
_ fd: Int32,
mask: Epoll.Mask = [.input, .output],
handler: @escaping @Sendable (Epoll.Mask) -> Void
) throws {
self.handlers.withLock { $0[fd] = handler }
let generation: UInt32 = try self.handlers.withLock { table in
guard table.entries[fd] == nil else { throw POSIXError(.EEXIST) }
let generation = table.nextGeneration
// 0 is reserved for the poller's internal eventFD registration.
table.nextGeneration = table.nextGeneration &+ 1
if table.nextGeneration == 0 { table.nextGeneration = 1 }
table.entries[fd] = (generation, handler)
return generation
}
do {
try self.poller.add(fd, mask: mask)
try self.poller.add(fd, mask: mask, generation: generation)
} catch {
self.handlers.withLock { _ = $0.removeValue(forKey: fd) }
self.handlers.withLock { _ = $0.entries.removeValue(forKey: fd) }
throw error
}
}
/// Remove a file descriptor from epoll monitoring and discard its handler.
func unregisterFd(_ fd: Int32) throws {
self.handlers.withLock { _ = $0.removeValue(forKey: fd) }
self.handlers.withLock { _ = $0.entries.removeValue(forKey: fd) }
try self.poller.delete(fd)
}