Files
gitea-macos-vm-orchestrator/docs/DESIGN.md
T

41 KiB
Raw Blame History

gitea-macos-runner — Design

1. Overview

gitea-macos-runner is a single-host daemon for an Apple Silicon Mac. It watches a Gitea instance for queued Actions jobs that require macOS, and for each one it boots a fresh, ephemeral macOS VM on Apple's Virtualization.framework, registers a single-use runner inside it, lets the job run, and then destroys the VM.

The design goal is that no state survives a job. Not a checkout, not a keychain entry, not a ~/Library mutation, not a leftover process. The guest that runs job N+1 is a byte-identical copy-on-write clone of the same base image that job N started from. This is the property that a persistent self-hosted Mac runner cannot offer, and it is the whole reason this tool exists.

Three constraints shape everything below:

  1. Apple's kernel allows at most two concurrent macOS guests per host. Not a policy, not a licence term we chose — a hard limit that surfaces as VZError.virtualMachineLimitExceeded from start(). Concurrency is therefore 2, permanently, and the config value is clamped rather than trusted.
  2. Virtualization needs a GUI session and a signed bundle. The daemon runs as a LaunchAgent in a logged-in user session, from inside a signed .app carrying com.apple.security.virtualization — Developer ID when a certificate is available, ad-hoc otherwise (see "Verified facts", item 10).
  3. Gitea decides which job a runner claims, not us. We supply capacity; the server matches. Trying to pin a specific job to a specific VM would mean reimplementing Gitea's matching rules, and would be wrong the moment they change.

Non-goals

  • Multi-host scheduling. One daemon, one Mac, two slots.
  • Container-based execution. Gitea's host schema runs jobs directly on the guest; that is the point of having a real macOS VM.
  • Bridged networking. NAT only — see §6.
  • Guest reuse or warm pools. See §9 for why save/restore is deferred rather than rejected.

2. Component diagram

┌──────────────────────────────── Host (Apple Silicon Mac, macOS 26+) ─────────────────────────────┐
│                                                                                                  │
│  LaunchAgent (user session, auto-login, login.keychain unlocked)                                 │
│    └── GiteaMacosRunner.app  (signed, com.apple.security.virtualization, LSUIElement)            │
│          │                                                                                       │
│          │  NSApplication(.prohibited).run()  ── main thread, required by Virtualization         │
│          │                                                                                       │
│    ┌─────▼──────────────────────────── Orchestrator (actor) ────────────────────────────────┐    │
│    │                                                                                        │    │
│    │   poll loop ──► GiteaClient.listQueuedJobs()  ──► [WorkflowJob]                         │    │
│    │        │                                                                                │   │
│    │        ├──────► SchedulerCore.plan(...)  ── PURE, no I/O ──► [SchedulerAction]          │    │
│    │        │                                                                                │   │
│    │        ├──► bootVM(slot,jobHint) ──► VMStore.cloneImage ──► VMInstance.start            │    │
│    │        │                                  │                       │                     │   │
│    │        │                                  │                       ▼                     │   │
│    │        │                                  │            DHCPLeaseParser(/var/db/…)       │   │
│    │        │                                  │                       │                     │   │
│    │        │                                  │                       ▼                     │   │
│    │        │                                  │            SSHExecutor ──► gitea-runner     │   │
│    │        │                                  │                            register+daemon  │   │
│    │        └──► teardownVM(slot,reason) ──► VMInstance.requestStopThenForce ──► deleteClone │   │
│    │                                                                                        │    │
│    │   reconcile loop ──► GiteaClient.listRunners / deleteRunner  (sweep orphaned rows)      │    │
│    └────────────────────────────────────────────────────────────────────────────────────────┘    │
│                                                                                                  │
│  <storeDir>/                                                                                     │
│    images/default/{disk.asif, nvram.bin, config.json}     ← built once, provisioned, read-only    │
│    vms/<uuid>/{disk.asif, nvram.bin, config.json}         ← APFS CoW clones, destroyed per job    │
│    ipsw/                                                  ← downloaded restore images            │
│    state.json                                             ← the two persistent per-slot MACs     │
│                                                                                                  │
│  ┌──── VM slot 0 (MAC A) ────┐   ┌──── VM slot 1 (MAC B) ────┐    ← at most 2, kernel-enforced   │
│  │ macOS guest               │   │ macOS guest               │                                   │
│  │  gitea-runner --ephemeral │   │  gitea-runner --ephemeral │                                   │
│  │  node, git, bash          │   │  node, git, bash          │                                   │
│  └───────────┬───────────────┘   └───────────┬───────────────┘                                   │
└──────────────┼───────────────────────────────┼───────────────────────────────────────────────────┘
               │        NAT (vmenet, bootpd)   │
               └───────────────┬───────────────┘
                               ▼
                    ┌─────────────────────┐
                    │  Gitea 1.25+        │
                    │  /api/v1/admin/…    │
                    └─────────────────────┘

Module boundaries

Target Contains Constraint
RunnerCore Config, Gitea models + client, LabelSet, DHCPLeaseParser, SSHExec, SchedulerCore No import Virtualization. Builds on Linux, so scheduling and parsing logic can be unit-tested anywhere.
RunnerHost VMBundle, VMStore, VZConfigFactory, VMInstance, IPSW, ImageBuilder, GuestProvisioner, Orchestrator, LaunchdService, Doctor macOS-only. Everything that touches the framework.
gitea-macos-runner CLI + daemon entry point Depends on both.

The split is not cosmetic: SchedulerCore being pure and portable is what makes the scheduling policy — the part most likely to have subtle bugs — testable without a Mac, a VM, or a Gitea instance.


3. Job lifecycle

Gitea                    Orchestrator                VMStore / VMInstance          Guest
  │                           │                             │                        │
  │◄── listQueuedJobs ────────┤  (every pollIntervalSeconds) │                        │
  ├─── [job 4711, labels ─────►                              │                        │
  │     ["macos-arm64"]]      │                              │                        │
  │                           ├── LabelSet.matches? ─────────┤                        │
  │                           ├── SchedulerCore.plan ────────┤                        │
  │                           │   → .bootVM(slot: 0,         │                        │
  │                           │             jobHint: 4711)   │                        │
  │                           │                              │                        │
  │                           ├── ensureFreeSpace(minGB) ───►│                        │
  │                           ├── cloneImage("default",  ───►│  APFS CoW copy         │
  │                           │     slotMAC: MAC-A)          │  + rewrite config.json │
  │                           ├── VMInstance.start() ───────►│  ──── boot ───────────►│
  │                           │                              │                        │
  │                           ├── poll /var/db/dhcpd_leases ─┤◄─── DHCP request ──────┤
  │                           │   until MAC-A has an IP      │                        │
  │                           ├── waitForSSH(ip) ────────────┼───────────────────────►│
  │                           │                              │                        │
  │                           ├── uploadData(token, 0600) ───┼───────────────────────►│
  │                           ├── ssh: gitea-runner register ┼───────────────────────►│
  │◄────────────────────── register (name=macos-vm-<uuid>, --ephemeral) ──────────────┤
  │                           ├── ssh: rm -f <tokenfile>     │                        │
  │                           ├── ssh: gitea-runner daemon ──┼───────────────────────►│
  │                           │                              │                        │
  │◄────────────────────── poll for task ─────────────────────────────────────────────┤
  ├─── assign job 4711 ───────────────────────────────────────────────────────────────►
  │                           │                              │        ...running...   │
  │◄────────────────────── job result, logs ──────────────────────────────────────────┤
  ├─── auto-deregister runner (server-enforced --ephemeral) ──────────────────────────►
  │                           │                              │  daemon exits          │
  │                           │◄─ SSH command returns ───────┼────────────────────────┤
  │                           ├── teardownVM(slot: 0) ──────►│                        │
  │                           │   requestStopThenForce ──────┼───────────────────────►│ (halt)
  │                           │   deleteClone ──────────────►│  rm -rf vms/<uuid>     │
  │                           ├── markIdle(slot: 0)          │                        │

Two details in that sequence carry more weight than their size suggests.

The token goes through a file, not an argument. gitea-runner register is invoked with --token-file <f>, where <f> was written by uploadData with mode 0600 and is rm -f'd in the same shell command. Passing --token would put a fleet-wide credential into the guest's process table, visible to any process the job spawns — and the job is arbitrary code from a repository.

--ephemeral, not --once. --ephemeral (Gitea 1.24+) is enforced by the server: it hands this runner exactly one task and then deletes the registration. --once is a runner-side convention only — the server still considers the runner live, and a misbehaving or patched runner could claim more work. Since the whole security story here rests on "one VM, one job", the enforcement has to live on the side we don't hand to the job.


4. Image build pipeline

Base images are built once with image build, and every job clones one. The build is slow (most of an hour, mostly a ~15 GB download); the clone is milliseconds.

  image build --name default [--ipsw PATH]
        │
        ├─ 1. IPSWProvider.latestSupported()      → CDN url + buildVersion
        │       (VZMacOSRestoreImage.latestSupported returns a NETWORK url —
        │        it cannot be handed to the installer)
        ├─ 2. IPSWProvider.download()             → <storeDir>/ipsw/*.ipsw
        ├─ 3. IPSWProvider.load(localPath:)        → VZMacOSRestoreImage
        │       (resolveSymlinksInPath first; the framework rejects symlinks)
        │
        ├─ 4. restoreImage.mostFeaturefulSupportedConfiguration
        │       nil  ⇒  this host cannot run this image. Fail loudly; do not guess.
        │
        ├─ 5. createBundle()
        │       hardwareModel.dataRepresentation      → config.json
        │       VZMacMachineIdentifier() (fresh)      → config.json
        │       VZMacAuxiliaryStorage(creatingStorageAt:hardwareModel:) → nvram.bin
        │       disk:  diskutil image create blank --fs none --format ASIF --size <N>G
        │              └─ fallback: sparse RAW file (format recorded in config.json)
        │
        ├─ 6. VZMacOSInstaller(virtualMachine:restoringFromImageAt:) on a STOPPED vm
        │       KVO on installer.progress → percentage
        │
        ├─ 7. first boot with Setup Assistant automation
        │       #available(macOS 27.0, *):
        │         VZMacGuestProvisioningOptions(username/password/fullName,
        │                                       logsInAutomatically: true,
        │                                       enablesRemoteLogin: true)
        │         → VZMacOSVirtualMachineStartOptions.setGuestProvisioning(_:)
        │       ⚠ an OLDER GUEST SILENTLY IGNORES THIS — no error, no account, no SSH
        │
        ├─ 8. wait for DHCP lease (by MAC) → wait for SSH → GuestProvisioner
        │       provision.sh   (sudoers, no-sleep, no-Spotlight, maxfiles, known_hosts)
        │       Node.js        (official arm64 .pkg → installer -pkg)   ← REQUIRED
        │       verify git / bash / node
        │       gitea-runner   (host downloads asset → upload → chmod +x)
        │       [optional] Xcode from a .xip
        │
        └─ 9. clean shutdown → config.provisioned = true   ← only now is it clonable

On step 7 and its failure mode

VZMacGuestProvisioningOptions needs macOS 27 or newer on both the host and the guest. The host side is a compile/availability check we control. The guest side is not: an older guest accepts the boot and simply ignores the options. There is no error to catch. The observable symptom is that the VM boots, sits at Setup Assistant forever, never requests a DHCP lease with a usable hostname, and never answers SSH — so the build fails at step 8 with a timeout that says nothing useful.

firstBootAndProvision therefore detects the timeout and reports it as an explicit "guest is too old for unattended setup; supply a macOS 27+ IPSW" failure. A --manual-setup flow that opens a window and lets a human click through Setup Assistant once is out of scope for v1 — deliberately, because a GUI step in a tool whose whole purpose is unattended operation is a trap. It is noted here so the omission is a decision rather than an oversight.

On step 5's disk format

ASIF is preferred because it is sparse: a 64 GB nominal disk costs what the guest actually writes, and it CoW-clones cleanly on APFS. diskutil image create is shelled out to because there is no framework API for it. If that call fails for any reason — older diskutil, unusual volume — a sparse RAW file is created instead and the format is recorded in config.json, so VZConfigFactory attaches the right file without re-probing.


5. Scheduling semantics

SchedulerCore.plan is a pure function: (state, queuedJobs, labels, maxVMs, now, jobTimeout, bootTimeout) → (state', [action]). It performs no I/O, reads no clock, and is fully deterministic — which is what allows the entire scheduling policy to be tested with a fixed now and a synthetic job list.

Capacity, not assignment

This is the central idea and the easiest thing to get wrong.

A booted VM is capacity. It is not a promise to run a particular job. We see job 4711 queued, we boot a VM, we register an ephemeral runner — and the server then decides which queued job that runner claims. It may well claim job 4712 instead. That is fine and in fact preferable: Gitea's matching rules (labels, repo permissions, ordering, priority) are its business, and any attempt to predict them here would be a reimplementation that drifts out of sync.

The jobHint threaded through SchedulerAction.bootVM and SlotState.running exists for exactly two purposes: log messages, and the dedup ledger below. Nothing else may depend on it.

Dedup by job id

SchedulerState.dispatchedJobIDs is a Set<Int64> of jobs that have already caused a boot.

Without it, the loop is pathological. A VM takes tens of seconds to boot, provision, and register. The poll interval is 5 seconds. So a single queued job would still be queued on the next poll, and the next, and the next — triggering a second boot, then exhausting the slot budget, all for one job.

The ledger is expired against reality rather than against a timer: any id no longer appearing in the queued set is dropped. That way a slot freed by a completed job can be re-earned by a genuinely new job, but a job that is still waiting does not double-book.

The cap

maxVMs is clamped to 2 in plan, and again in RunnerConfig.validated(). Both places, because the kernel limit is not something a config file gets to negotiate: a third start() raises VZError.virtualMachineLimitExceeded, which VMInstance.mapVZError translates into CoreError.vmLimitExceeded and the scheduler treats as transient back-pressure rather than a failure.

Timeouts

  • A slot in .provisioning(since:) longer than bootTimeoutSeconds (default 900) is torn down. Covers a guest that never gets a lease, never starts sshd, or hangs in Setup Assistant. The default is deliberately generous: several Virtualization guests sharing one host push a boot from tens of seconds into minutes, and a limit below the worst case does not time out a bad boot, it livelocks — each replacement clone starts from zero and adds load, so the next boot is slower still and no runner ever registers.
  • The lifecycle's own waitForLease and waitForSSH budgets are derived from what is left of that deadline, not from a fresh copy of it. Given the full bootTimeoutSeconds their deadlines would fall after the planner's, so the planner would always cancel first and the specific error — which host, how many attempts, what the last one said — would be discarded in favour of a bare cancellation.
  • A slot in .running(jobHint:since:) longer than jobTimeoutMinutes (default 120) is torn down. Covers a job that hangs. This is comfortably below Gitea's own ABANDONED_JOB_TIMEOUT (24 h), so our teardown always happens first and the server sees a clean deregistration rather than an abandonment.

Teardown actions are emitted before boot actions in the returned list, so a slot freed in one pass can be reused in that same pass.

Reconcile

Every reconcileIntervalSeconds (default 300), the orchestrator lists runners and deletes any that are:

  • ephemeral == true, and
  • busy == false, and
  • name starts with our configured namePrefix, and
  • not backed by a live VM in this process.

All four conditions, because deleting a live runner fails a running job. The loop is deliberately conservative: a row we are unsure about is left alone, and will be revisited in five minutes.

This loop is not optional housekeeping — it is load-bearing. See Verified Fact 7.


6. Security model

The threat

A CI job is arbitrary code from a repository, running with the privileges of the account it executes under. On a persistent self-hosted Mac runner, that code can read every previous job's checkout, poison caches, install launch agents, and harvest whatever credentials the machine has accumulated. Every subsequent job on that host inherits the compromise.

The mitigation: genuinely ephemeral guests

  • One job per VM, enforced server-side. --ephemeral means Gitea hands the runner exactly one task and then deletes the registration. A patched or hijacked runner binary cannot ask for more work, because the server will not give it any.
  • The VM is destroyed after that job. Not reset, not cleaned — the clone directory is rm -rf'd and the next job clones the base image afresh. There is no path by which job N influences job N+1 short of compromising the host.
  • The guest holds nothing worth stealing. Its account credentials (admin/admin by default) are meaningful only on a host-private NAT link to a machine that is about to be deleted.

The shared registration token

Registration tokens in Gitea are reusable and scope-wide, and minting a new one for a scope invalidates all prior tokens of that scope. That makes per-VM tokens actively harmful: generating one for each VM would break every other runner registered against that scope, including ones on other hosts.

So the fleet shares one token. The mitigations are:

  • It is written into the guest as a file with mode 0600, never as a command argument (arguments are world-readable via ps).
  • It is deleted immediately after gitea-runner register consumes it, in the same && chain, before gitea-runner daemon starts and long before any job code runs.
  • It is a registration token, not an API token: it grants the ability to register a runner, not to read repositories or act as a user.

The residual risk is real but bounded — a job that wins a race against rm -f could register additional runners for that scope. The recommended deployment seeds a fixed token server-side via GITEA_RUNNER_REGISTRATION_TOKEN so that rotating it is a deliberate, coordinated act rather than an API call side effect.

:host schema risk

Jobs run in host schema: directly on the guest OS, not in a container. That is the point — a macOS job needs real macOS. But it means the job has full user-level access to the guest, including sudo (which provision.sh makes passwordless, because Xcode and installer need it). Everything above rests on the guest being disposable and isolated, not on the job being constrained inside it.

Host-side posture

  • The daemon runs as a LaunchAgent in a user session, not as root. The entitlement it carries (com.apple.security.virtualization) grants VM creation and nothing else.
  • Networking is NAT, not bridged. Guests can reach the LAN and Gitea, but are not first-class hosts on it. Bridged networking would require the restricted com.apple.vm.networking entitlement, which needs an Apple-approved provisioning profile and which ad-hoc signing cannot grant at all — a constraint that happens to align with what we want anyway.
  • SSH host keys are not verified. The peer is a VM this process booted moments ago on a link no other machine shares; pinning would break on every clone and add nothing.

7. Failure modes and recovery

Failure Detection Recovery
Guest never gets a DHCP lease bootTimeout in waitForLease Teardown, slot recycled, retried next poll
Guest never answers SSH bootTimeout in waitForSSH Same
Job hangs jobTimeoutMinutes Teardown; Gitea reaps the task via its zombie sweep (~10–15 min)
VM dies uncleanly Runner row left behind, task stuck Running Reconcile loop deletes the row (§5); Gitea's zombie sweep handles the task
Daemon crashes with VMs live Clones orphaned on disk purgeClones() at startup, then one immediate reconcile pass
Third VM requested VZError.virtualMachineLimitExceeded Mapped to CoreError.vmLimitExceeded, treated as back-pressure
Disk fills ensureFreeSpace(minGB:) before each clone Boot refused, logged; jobs stay queued (safe — Gitea holds them ~24 h)
Gitea unreachable Request error in the poll loop Logged, retried next tick; no state change

8. Configuration and operations summary

Config lives at ~/.config/gitea-macos-runner/config.json; see Resources/config.example.json. Order of operations for a new host:

gitea-macos-runner doctor          # verify arch, macOS, entitlement, keychain, Gitea
gitea-macos-runner config init     # write an annotated config
gitea-macos-runner image build     # ~1 hour, mostly IPSW download
gitea-macos-runner vm boot         # optional smoke test: boot a clone, print its IP
gitea-macos-runner service install # LaunchAgent, RunAtLoad + KeepAlive
gitea-macos-runner doctor          # again, now that it runs from the signed .app

doctor exists because every one of its checks corresponds to a failure that otherwise appears as an opaque error deep inside a VM boot. The most common by far: running from .build/ instead of the signed .app, so the entitlement is absent.


9. Future work

  • Save/restore for warm boots. VZVirtualMachine.saveMachineStateTo (macOS 14+) could cut per-job boot from ~60 s to near-instant by restoring a snapshot taken just after gitea-runner is ready. The blocker is that restore forbids changing the MAC address or ECID, which collides with our per-slot MAC scheme (§ Verified Fact 12) — a restored state would have to be captured per slot, and the interaction with DHCP lease reuse needs care. Deferred, not rejected.
  • vsock control channel. VZVirtioSocketDeviceConfiguration is already in the VM configuration. Replacing SSH with a vsock agent would remove password auth, the waitForSSH poll, and the Local Network privacy prompt entirely.
  • --manual-setup for pre-macOS-27 guests (§4).
  • Image versioning / garbage collection for multiple base images.

Appendix: Verified Facts

Researched facts this design depends on, each with its consequence for the implementation. Anyone changing the corresponding code should re-verify the fact first.

1. Job discovery is GET /api/v1/admin/actions/jobs?status=queued (Gitea 1.25+). The labels field on a returned job is the workflow's runs-on: value. The external status string queued maps to Gitea's internal StatusWaiting, meaning "ready, waiting for a matching runner". → Consequence: the external string waiting means something entirely different — the job is blocked on a dependency — and must never be treated as schedulable. WorkflowJob.isQueued checks status == "queued" and nothing else.

2. A queued job waits for a matching runner up to ABANDONED_JOB_TIMEOUT (default 24 h, swept every 6 h). → Consequence: there is no urgency in the poll loop. A 5-second interval is for responsiveness, not correctness; a daemon that is down for an hour loses nothing. It also sets the ceiling that jobTimeoutMinutes (default 120) must stay well under, so our teardown always precedes the server's abandonment.

3. The runner binary is gitea-runner v3.x, renamed from act_runner and published from gitea.com/gitea/runner. The register-time flag --ephemeral (Gitea 1.24+) is server-enforced: single job, then automatic deregistration. --once is weaker — runner-side only. → Consequence: --ephemeral is mandatory and --once is never used. The "one VM, one job" guarantee in §6 rests on the server enforcing it, not on the runner cooperating. Download URLs and the binary name must reference gitea-runner, not act_runner.

4. Registration tokens are REUSABLE, and minting a new token for a scope INVALIDATES prior tokens of that scope. POST /api/v1/admin/actions/runners/registration-token in practice returns the existing active token. Seeding server-side via GITEA_RUNNER_REGISTRATION_TOKEN is the recommended deployment. → Consequence: never pre-generate a token per VM — doing so would break every other runner registered against that scope. One token is resolved once, cached for the process lifetime, and shared by the fleet; fetchRegistrationTokenViaAPI defaults to false.

5. Label syntax is name:schema, schema defaults to host, and only BARE names are stored server-side. If the guest's runner config.yaml sets runner.labels, it silently overrides --labels passed at registration. → Consequence: LabelSet matches bare names (macos-arm64) and only appends :host when building the register --labels argument. The guest must never ship a config.yaml containing labels — noted in both GuestProvisioner and Resources/provision.sh, because the failure is silent: the runner registers successfully and is simply never matched.

6. Guest requirements: the gitea-runner binary, node, git, bash, and a writable $HOME. JavaScript actions such as actions/checkout spawn node directly. → Consequence: Node.js installation is not optional and not a convenience — without it essentially every real workflow fails at its first step. GuestProvisioner.verifyToolchain fails the build rather than shipping an image that will break at job time.

7. An unclean VM death leaves both a runner row and a Running task behind. The task is reaped by Gitea's zombie sweep in roughly 10–15 minutes. The runner row is swept only at midnight — and never at all if that runner claimed no task. Rows are removed with DELETE /api/v1/admin/actions/runners/{id}. → Consequence: the reconcile loop (§5) is load-bearing, not housekeeping. Without it, every crashed boot leaves a permanent phantom runner. It is also why every VM registers under a globally unique name (namePrefix + UUID): that uniqueness is what lets us look at a row and decide with certainty that it is ours and unbacked.

8. macOS guests are capped at 2 concurrent per host, enforced by the kernel. A third start() raises VZError.virtualMachineLimitExceeded. → Consequence: maxConcurrentVMs is hard-clamped to 2 in both RunnerConfig.validated() and SchedulerCore.plan; the slot table is fixed-size; and the error is mapped to CoreError.vmLimitExceeded and treated as transient back-pressure rather than a failure.

9. VZMacGuestProvisioningOptions (macOS 27+ host AND guest) automates Setup Assistant, including enablesRemoteLogin (SSH). Older guests silently ignore it. → Consequence: the API is gated at #available(macOS 27.0, *), and because the guest-side failure produces no error, firstBootAndProvision must translate its lease/SSH timeout into an explicit "guest too old" message rather than a bare timeout. The --manual-setup fallback is out of v1 scope (§4).

10. Headless Virtualization requires an NSApplication run loop with .prohibited activation policy, inside a signed .app bundle carrying com.apple.security.virtualization. That entitlement is unrestricted: ad-hoc signing (codesign -s -) grants it, and a Developer ID certificate grants it with no provisioning profile. Bridged networking would additionally need a restricted entitlement; NAT does not. → Consequence: CommandDaemon starts NSApplication and runs the orchestrator in a detached Task. This applies to every command that starts a VM, not just the daemon: vm boot, image build, and image provision all go through the same VZAppRuntime.run host, since VZMacOSInstaller and the first-boot provisioning pass need the run loop exactly as much as a job VM does. The Makefile has bundle and sign targets and install deliberately installs the bundle rather than the bare binary; Info.plist sets LSUIElement; Doctor checks the entitlement on the running binary because running from .build/ is the most common setup failure. Two packaging constraints follow from the same fact. The entitlements plist must contain no XML comments: plutil -lint accepts them, but codesign hands the file to AMFI's stricter parser, which fails with AMFIUnserializeXML: syntax error and then signs the bundle with zero entitlements — a silent downgrade that only surfaces as a failed VM start. And bundle must copy provision.sh, launchd.plist.template, and config.example.json into Contents/Resources, since GuestProvisioner, LaunchdService, and config init look there before falling back to repo-relative paths; an installed .app without them is a working binary with a broken image build, service install, and config init.

10a. Which signature is used decides whether the app's code identity is stable across rebuilds. A Developer ID signature's designated requirement is anchored to the team (… and certificate leaf[subject.OU] = <TEAM_ID>), so every build is the same program to macOS. An ad-hoc signature has no anchor, so identity falls back to the main executable's Mach-O UUID, which the linker regenerates on essentially every link. → Consequence: this is not cosmetic, because macOS Local Network privacy is keyed on exactly that UUID (Fact 16). Under ad-hoc signing a grant is withdrawn by the next make install; under Developer ID it persists. make sign therefore selects a Developer ID Application identity matching TEAM_ID when the keychain has one and falls back to ad-hoc with a warning when it does not — the fallback is required because CI builds inside a throwaway guest with neither keychain nor certificate. The Developer ID path also passes --options runtime --timestamp, so the bundle is notarizable later without re-signing; notarization itself is skipped, since it governs distribution to other Macs and this app is built and installed in place. Doctor.checkCodeSignature reports which path was taken and warns on ad-hoc.

11. macOS 15+ requires an unlocked login.keychain to start a VM. → Consequence: the service must be a LaunchAgent in the auto-logged-in user's session, never a LaunchDaemon (which has no session and no unlocked keychain). LaunchdService only ever writes to ~/Library/LaunchAgents, and Doctor probes with security show-keychain-info login.keychain.

12. Guest IPs come from parsing /var/db/dhcpd_leases, keyed by MAC. hw_address lines carry a 1, hardware-type prefix and octets that may lack zero-padding (aa:bb:c:dd:ee:ff). Duplicate MACs occur; the newest lease wins. macOS's DHCP lease time is 24 hours. → Consequence: DHCPLeaseParser.normalizeMAC must strip the prefix and zero-pad, or lookups silently fail against VZMACAddress.string. And ephemeral fleets must not randomize MACs per clone — a day of dead leases would accumulate and exhaust the NAT subnet. Hence exactly two persistent per-slot MACs, generated once with VZMACAddress.randomLocallyAdministered() and stored in state.json.

13. APFS copy-on-write cloning via FileManager.copyItem requires source and destination on the same volume. Clones grow as the guest writes. → Consequence: base images and ephemeral clones both live under storeDir, and cloning is done per file rather than by copying a directory wholesale. ensureFreeSpace(minGB:) runs before every clone, with a floor (default 20 GB) well above one clone's nominal cost, because the apparent size and the real cost diverge over a job's lifetime.

14. The ASIF sparse disk format is created via diskutil image create (macOS 26+); RAW is the fallback. → Consequence: ImageBuilder.createDisk shells out to /usr/sbin/diskutil image create blank --fs none --format ASIF --size <N>G <path> and falls back to a sparse RAW file, recording which was used in VMBundleConfig.diskFormat so VZConfigFactory attaches the right file without re-probing. This is also the reason the package's minimum platform is macOS 26.

15. Save/restore (macOS 14+) could give near-instant warm boots, but forbids changing the MAC address or ECID. → Consequence: documented as future work (§9) rather than implemented. The prohibition collides directly with the per-slot MAC scheme from Fact 12, so adopting it would require per-slot saved states and a careful look at DHCP lease reuse — not a drop-in optimization.

16. macOS 15+ Local Network privacy can block host→guest connections, and granting it interactively takes deliberate work. Per TN3179 it is not TCC — the check is a Network Extension packet filter, so it is absent from TCC.db, cannot be queried, cannot be reset, and it "uses your main executable UUID as part of its implementation". A denial returns EHOSTUNREACH (errno 65), indistinguishable from a genuinely unreachable host. Three things then conspire against the interactive grant: a LaunchAgent has no UI to show the prompt in; a run started from a shell is attributed to the responsible process, so the prompt and the System Settings row belong to Terminal rather than to this app, and granting it to Terminal does not carry to the agent; and under ad-hoc signing the UUID keying (Fact 10a) withdraws the grant on the next rebuild. → Consequence: the deterministic fix is the subnet allowlist (com.apple.network.local-network, keys AllowedEthernetLocalNetworkAddresses and AllowedWiFiLocalNetworkAddresses), which is keyed on the network rather than the app and is read at boot — so it needs a reboot, not a service restart. LocalNetworkPolicy owns the arithmetic and requires coverage of 192.168.64.0/18, not a single /24, because the NAT subnet is chosen at runtime and slides to the next free /24; Doctor.localNetworkNote and SSHExec.localNetworkHint both report against it, since errno 65 gives the operator nothing to go on by itself.

→ Consequence: both routes are commands rather than documentation. permissions grant writes the allowlist and verifies it read back (sudo defaults write lands in root's or the invoking user's preferences depending on whether sudo preserved HOME, so where it went is not assumable). permissions grant --method prompt addresses the attribution problem head-on: launching the installed bundle through LaunchServices (open -n -b …) makes the app its own responsible process, so the prompt and the Settings row belong to it rather than to Terminal — and because the LaunchAgent runs the same signed identity, the grant carries. That only became worth building once the bundle was Developer ID signed; under ad-hoc signing the UUID churn withdraws it on the next rebuild, which is why permissions status reports code identity alongside the allowlist.

→ Consequence: the prompt route is best-effort and says so. Observed on a host where the decision was already recorded: UserEventAgent resolves the flow to the bundle ID on every attempt — so the attribution works — but presents no alert, because macOS asks once per app identity and then answers from that record, silently, forever. There is no supported reset. So --method prompt verifies by probing rather than by trusting the launch, and on a denial says plainly that it did not take and points at the allowlist, which is not subject to the per-app check at all. The allowlist stays the recommendation.