39 KiB
gitea-macos-runner — Design
1. Overview
gitea-macos-runner is a single-host daemon for an Apple Silicon Mac. It watches
a Gitea instance for queued Actions jobs that require macOS, and for each one it
boots a fresh, ephemeral macOS VM on Apple's Virtualization.framework,
registers a single-use runner inside it, lets the job run, and then destroys the
VM.
The design goal is that no state survives a job. Not a checkout, not a
keychain entry, not a ~/Library mutation, not a leftover process. The guest
that runs job N+1 is a byte-identical copy-on-write clone of the same base
image that job N started from. This is the property that a persistent
self-hosted Mac runner cannot offer, and it is the whole reason this tool exists.
Three constraints shape everything below:
- Apple's kernel allows at most two concurrent macOS guests per host. Not a
policy, not a licence term we chose — a hard limit that surfaces as
VZError.virtualMachineLimitExceededfromstart(). Concurrency is therefore 2, permanently, and the config value is clamped rather than trusted. - Virtualization needs a GUI session and a signed bundle. The daemon runs as
a LaunchAgent in a logged-in user session, from inside a signed
.appcarryingcom.apple.security.virtualization— Developer ID when a certificate is available, ad-hoc otherwise (see "Verified facts", item 10). - Gitea decides which job a runner claims, not us. We supply capacity; the server matches. Trying to pin a specific job to a specific VM would mean reimplementing Gitea's matching rules, and would be wrong the moment they change.
Non-goals
- Multi-host scheduling. One daemon, one Mac, two slots.
- Container-based execution. Gitea's
hostschema runs jobs directly on the guest; that is the point of having a real macOS VM. - Bridged networking. NAT only — see §6.
- Guest reuse or warm pools. See §9 for why save/restore is deferred rather than rejected.
2. Component diagram
┌──────────────────────────────── Host (Apple Silicon Mac, macOS 26+) ─────────────────────────────┐
│ │
│ LaunchAgent (user session, auto-login, login.keychain unlocked) │
│ └── GiteaMacosRunner.app (signed, com.apple.security.virtualization, LSUIElement) │
│ │ │
│ │ NSApplication(.prohibited).run() ── main thread, required by Virtualization │
│ │ │
│ ┌─────▼──────────────────────────── Orchestrator (actor) ────────────────────────────────┐ │
│ │ │ │
│ │ poll loop ──► GiteaClient.listQueuedJobs() ──► [WorkflowJob] │ │
│ │ │ │ │
│ │ ├──────► SchedulerCore.plan(...) ── PURE, no I/O ──► [SchedulerAction] │ │
│ │ │ │ │
│ │ ├──► bootVM(slot,jobHint) ──► VMStore.cloneImage ──► VMInstance.start │ │
│ │ │ │ │ │ │
│ │ │ │ ▼ │ │
│ │ │ │ DHCPLeaseParser(/var/db/…) │ │
│ │ │ │ │ │ │
│ │ │ │ ▼ │ │
│ │ │ │ SSHExecutor ──► gitea-runner │ │
│ │ │ │ register+daemon │ │
│ │ └──► teardownVM(slot,reason) ──► VMInstance.requestStopThenForce ──► deleteClone │ │
│ │ │ │
│ │ reconcile loop ──► GiteaClient.listRunners / deleteRunner (sweep orphaned rows) │ │
│ └────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ <storeDir>/ │
│ images/default/{disk.asif, nvram.bin, config.json} ← built once, provisioned, read-only │
│ vms/<uuid>/{disk.asif, nvram.bin, config.json} ← APFS CoW clones, destroyed per job │
│ ipsw/ ← downloaded restore images │
│ state.json ← the two persistent per-slot MACs │
│ │
│ ┌──── VM slot 0 (MAC A) ────┐ ┌──── VM slot 1 (MAC B) ────┐ ← at most 2, kernel-enforced │
│ │ macOS guest │ │ macOS guest │ │
│ │ gitea-runner --ephemeral │ │ gitea-runner --ephemeral │ │
│ │ node, git, bash │ │ node, git, bash │ │
│ └───────────┬───────────────┘ └───────────┬───────────────┘ │
└──────────────┼───────────────────────────────┼───────────────────────────────────────────────────┘
│ NAT (vmenet, bootpd) │
└───────────────┬───────────────┘
▼
┌─────────────────────┐
│ Gitea 1.25+ │
│ /api/v1/admin/… │
└─────────────────────┘
Module boundaries
| Target | Contains | Constraint |
|---|---|---|
RunnerCore |
Config, Gitea models + client, LabelSet, DHCPLeaseParser, SSHExec, SchedulerCore |
No import Virtualization. Builds on Linux, so scheduling and parsing logic can be unit-tested anywhere. |
RunnerHost |
VMBundle, VMStore, VZConfigFactory, VMInstance, IPSW, ImageBuilder, GuestProvisioner, Orchestrator, LaunchdService, Doctor |
macOS-only. Everything that touches the framework. |
gitea-macos-runner |
CLI + daemon entry point | Depends on both. |
The split is not cosmetic: SchedulerCore being pure and portable is what makes
the scheduling policy — the part most likely to have subtle bugs — testable
without a Mac, a VM, or a Gitea instance.
3. Job lifecycle
Gitea Orchestrator VMStore / VMInstance Guest
│ │ │ │
│◄── listQueuedJobs ────────┤ (every pollIntervalSeconds) │ │
├─── [job 4711, labels ─────► │ │
│ ["macos-arm64"]] │ │ │
│ ├── LabelSet.matches? ─────────┤ │
│ ├── SchedulerCore.plan ────────┤ │
│ │ → .bootVM(slot: 0, │ │
│ │ jobHint: 4711) │ │
│ │ │ │
│ ├── ensureFreeSpace(minGB) ───►│ │
│ ├── cloneImage("default", ───►│ APFS CoW copy │
│ │ slotMAC: MAC-A) │ + rewrite config.json │
│ ├── VMInstance.start() ───────►│ ──── boot ───────────►│
│ │ │ │
│ ├── poll /var/db/dhcpd_leases ─┤◄─── DHCP request ──────┤
│ │ until MAC-A has an IP │ │
│ ├── waitForSSH(ip) ────────────┼───────────────────────►│
│ │ │ │
│ ├── uploadData(token, 0600) ───┼───────────────────────►│
│ ├── ssh: gitea-runner register ┼───────────────────────►│
│◄────────────────────── register (name=macos-vm-<uuid>, --ephemeral) ──────────────┤
│ ├── ssh: rm -f <tokenfile> │ │
│ ├── ssh: gitea-runner daemon ──┼───────────────────────►│
│ │ │ │
│◄────────────────────── poll for task ─────────────────────────────────────────────┤
├─── assign job 4711 ───────────────────────────────────────────────────────────────►
│ │ │ ...running... │
│◄────────────────────── job result, logs ──────────────────────────────────────────┤
├─── auto-deregister runner (server-enforced --ephemeral) ──────────────────────────►
│ │ │ daemon exits │
│ │◄─ SSH command returns ───────┼────────────────────────┤
│ ├── teardownVM(slot: 0) ──────►│ │
│ │ requestStopThenForce ──────┼───────────────────────►│ (halt)
│ │ deleteClone ──────────────►│ rm -rf vms/<uuid> │
│ ├── markIdle(slot: 0) │ │
Two details in that sequence carry more weight than their size suggests.
The token goes through a file, not an argument. gitea-runner register is
invoked with --token-file <f>, where <f> was written by uploadData with
mode 0600 and is rm -f'd in the same shell command. Passing --token would
put a fleet-wide credential into the guest's process table, visible to any
process the job spawns — and the job is arbitrary code from a repository.
--ephemeral, not --once. --ephemeral (Gitea 1.24+) is enforced by the
server: it hands this runner exactly one task and then deletes the registration.
--once is a runner-side convention only — the server still considers the runner
live, and a misbehaving or patched runner could claim more work. Since the whole
security story here rests on "one VM, one job", the enforcement has to live on
the side we don't hand to the job.
4. Image build pipeline
Base images are built once with image build, and every job clones one. The
build is slow (most of an hour, mostly a ~15 GB download); the clone is
milliseconds.
image build --name default [--ipsw PATH]
│
├─ 1. IPSWProvider.latestSupported() → CDN url + buildVersion
│ (VZMacOSRestoreImage.latestSupported returns a NETWORK url —
│ it cannot be handed to the installer)
├─ 2. IPSWProvider.download() → <storeDir>/ipsw/*.ipsw
├─ 3. IPSWProvider.load(localPath:) → VZMacOSRestoreImage
│ (resolveSymlinksInPath first; the framework rejects symlinks)
│
├─ 4. restoreImage.mostFeaturefulSupportedConfiguration
│ nil ⇒ this host cannot run this image. Fail loudly; do not guess.
│
├─ 5. createBundle()
│ hardwareModel.dataRepresentation → config.json
│ VZMacMachineIdentifier() (fresh) → config.json
│ VZMacAuxiliaryStorage(creatingStorageAt:hardwareModel:) → nvram.bin
│ disk: diskutil image create blank --fs none --format ASIF --size <N>G
│ └─ fallback: sparse RAW file (format recorded in config.json)
│
├─ 6. VZMacOSInstaller(virtualMachine:restoringFromImageAt:) on a STOPPED vm
│ KVO on installer.progress → percentage
│
├─ 7. first boot with Setup Assistant automation
│ #available(macOS 27.0, *):
│ VZMacGuestProvisioningOptions(username/password/fullName,
│ logsInAutomatically: true,
│ enablesRemoteLogin: true)
│ → VZMacOSVirtualMachineStartOptions.setGuestProvisioning(_:)
│ ⚠ an OLDER GUEST SILENTLY IGNORES THIS — no error, no account, no SSH
│
├─ 8. wait for DHCP lease (by MAC) → wait for SSH → GuestProvisioner
│ provision.sh (sudoers, no-sleep, no-Spotlight, maxfiles, known_hosts)
│ Node.js (official arm64 .pkg → installer -pkg) ← REQUIRED
│ verify git / bash / node
│ gitea-runner (host downloads asset → upload → chmod +x)
│ [optional] Xcode from a .xip
│
└─ 9. clean shutdown → config.provisioned = true ← only now is it clonable
On step 7 and its failure mode
VZMacGuestProvisioningOptions needs macOS 27 or newer on both the host and
the guest. The host side is a compile/availability check we control. The guest
side is not: an older guest accepts the boot and simply ignores the options.
There is no error to catch. The observable symptom is that the VM boots, sits at
Setup Assistant forever, never requests a DHCP lease with a usable hostname, and
never answers SSH — so the build fails at step 8 with a timeout that says nothing
useful.
firstBootAndProvision therefore detects the timeout and reports it as an
explicit "guest is too old for unattended setup; supply a macOS 27+ IPSW"
failure. A --manual-setup flow that opens a window and lets a human click
through Setup Assistant once is out of scope for v1 — deliberately, because a
GUI step in a tool whose whole purpose is unattended operation is a trap. It is
noted here so the omission is a decision rather than an oversight.
On step 5's disk format
ASIF is preferred because it is sparse: a 64 GB nominal disk costs what the guest
actually writes, and it CoW-clones cleanly on APFS. diskutil image create is
shelled out to because there is no framework API for it. If that call fails for
any reason — older diskutil, unusual volume — a sparse RAW file is created
instead and the format is recorded in config.json, so VZConfigFactory attaches
the right file without re-probing.
5. Scheduling semantics
SchedulerCore.plan is a pure function: (state, queuedJobs, labels, maxVMs, now, jobTimeout, bootTimeout) → (state', [action]). It performs no I/O, reads no
clock, and is fully deterministic — which is what allows the entire scheduling
policy to be tested with a fixed now and a synthetic job list.
Capacity, not assignment
This is the central idea and the easiest thing to get wrong.
A booted VM is capacity. It is not a promise to run a particular job. We see job 4711 queued, we boot a VM, we register an ephemeral runner — and the server then decides which queued job that runner claims. It may well claim job 4712 instead. That is fine and in fact preferable: Gitea's matching rules (labels, repo permissions, ordering, priority) are its business, and any attempt to predict them here would be a reimplementation that drifts out of sync.
The jobHint threaded through SchedulerAction.bootVM and
SlotState.running exists for exactly two purposes: log messages, and the dedup
ledger below. Nothing else may depend on it.
Dedup by job id
SchedulerState.dispatchedJobIDs is a Set<Int64> of jobs that have already
caused a boot.
Without it, the loop is pathological. A VM takes tens of seconds to boot, provision, and register. The poll interval is 5 seconds. So a single queued job would still be queued on the next poll, and the next, and the next — triggering a second boot, then exhausting the slot budget, all for one job.
The ledger is expired against reality rather than against a timer: any id no longer appearing in the queued set is dropped. That way a slot freed by a completed job can be re-earned by a genuinely new job, but a job that is still waiting does not double-book.
The cap
maxVMs is clamped to 2 in plan, and again in RunnerConfig.validated(). Both
places, because the kernel limit is not something a config file gets to
negotiate: a third start() raises VZError.virtualMachineLimitExceeded, which
VMInstance.mapVZError translates into CoreError.vmLimitExceeded and the
scheduler treats as transient back-pressure rather than a failure.
Timeouts
- A slot in
.provisioning(since:)longer thanbootTimeoutSeconds(default 900) is torn down. Covers a guest that never gets a lease, never startssshd, or hangs in Setup Assistant. The default is deliberately generous: several Virtualization guests sharing one host push a boot from tens of seconds into minutes, and a limit below the worst case does not time out a bad boot, it livelocks — each replacement clone starts from zero and adds load, so the next boot is slower still and no runner ever registers. - The lifecycle's own
waitForLeaseandwaitForSSHbudgets are derived from what is left of that deadline, not from a fresh copy of it. Given the fullbootTimeoutSecondstheir deadlines would fall after the planner's, so the planner would always cancel first and the specific error — which host, how many attempts, what the last one said — would be discarded in favour of a bare cancellation. - A slot in
.running(jobHint:since:)longer thanjobTimeoutMinutes(default 120) is torn down. Covers a job that hangs. This is comfortably below Gitea's ownABANDONED_JOB_TIMEOUT(24 h), so our teardown always happens first and the server sees a clean deregistration rather than an abandonment.
Teardown actions are emitted before boot actions in the returned list, so a slot freed in one pass can be reused in that same pass.
Reconcile
Every reconcileIntervalSeconds (default 300), the orchestrator lists runners and
deletes any that are:
ephemeral == true, andbusy == false, andnamestarts with our configurednamePrefix, and- not backed by a live VM in this process.
All four conditions, because deleting a live runner fails a running job. The loop is deliberately conservative: a row we are unsure about is left alone, and will be revisited in five minutes.
This loop is not optional housekeeping — it is load-bearing. See Verified Fact 7.
6. Security model
The threat
A CI job is arbitrary code from a repository, running with the privileges of the account it executes under. On a persistent self-hosted Mac runner, that code can read every previous job's checkout, poison caches, install launch agents, and harvest whatever credentials the machine has accumulated. Every subsequent job on that host inherits the compromise.
The mitigation: genuinely ephemeral guests
- One job per VM, enforced server-side.
--ephemeralmeans Gitea hands the runner exactly one task and then deletes the registration. A patched or hijacked runner binary cannot ask for more work, because the server will not give it any. - The VM is destroyed after that job. Not reset, not cleaned — the clone
directory is
rm -rf'd and the next job clones the base image afresh. There is no path by which job N influences job N+1 short of compromising the host. - The guest holds nothing worth stealing. Its account credentials
(
admin/adminby default) are meaningful only on a host-private NAT link to a machine that is about to be deleted.
The shared registration token
Registration tokens in Gitea are reusable and scope-wide, and minting a new one for a scope invalidates all prior tokens of that scope. That makes per-VM tokens actively harmful: generating one for each VM would break every other runner registered against that scope, including ones on other hosts.
So the fleet shares one token. The mitigations are:
- It is written into the guest as a file with mode
0600, never as a command argument (arguments are world-readable viaps). - It is deleted immediately after
gitea-runner registerconsumes it, in the same&&chain, beforegitea-runner daemonstarts and long before any job code runs. - It is a registration token, not an API token: it grants the ability to register a runner, not to read repositories or act as a user.
The residual risk is real but bounded — a job that wins a race against rm -f
could register additional runners for that scope. The recommended deployment
seeds a fixed token server-side via GITEA_RUNNER_REGISTRATION_TOKEN so that
rotating it is a deliberate, coordinated act rather than an API call side effect.
:host schema risk
Jobs run in host schema: directly on the guest OS, not in a container. That is
the point — a macOS job needs real macOS. But it means the job has full user-level
access to the guest, including sudo (which provision.sh makes passwordless,
because Xcode and installer need it). Everything above rests on the guest being
disposable and isolated, not on the job being constrained inside it.
Host-side posture
- The daemon runs as a LaunchAgent in a user session, not as root. The
entitlement it carries (
com.apple.security.virtualization) grants VM creation and nothing else. - Networking is NAT, not bridged. Guests can reach the LAN and Gitea, but are
not first-class hosts on it. Bridged networking would require the restricted
com.apple.vm.networkingentitlement, which needs an Apple-approved provisioning profile and which ad-hoc signing cannot grant at all — a constraint that happens to align with what we want anyway. - SSH host keys are not verified. The peer is a VM this process booted moments ago on a link no other machine shares; pinning would break on every clone and add nothing.
7. Failure modes and recovery
| Failure | Detection | Recovery |
|---|---|---|
| Guest never gets a DHCP lease | bootTimeout in waitForLease |
Teardown, slot recycled, retried next poll |
| Guest never answers SSH | bootTimeout in waitForSSH |
Same |
| Job hangs | jobTimeoutMinutes |
Teardown; Gitea reaps the task via its zombie sweep (~10–15 min) |
| VM dies uncleanly | Runner row left behind, task stuck Running | Reconcile loop deletes the row (§5); Gitea's zombie sweep handles the task |
| Daemon crashes with VMs live | Clones orphaned on disk | purgeClones() at startup, then one immediate reconcile pass |
| Third VM requested | VZError.virtualMachineLimitExceeded |
Mapped to CoreError.vmLimitExceeded, treated as back-pressure |
| Disk fills | ensureFreeSpace(minGB:) before each clone |
Boot refused, logged; jobs stay queued (safe — Gitea holds them ~24 h) |
| Gitea unreachable | Request error in the poll loop | Logged, retried next tick; no state change |
8. Configuration and operations summary
Config lives at ~/.config/gitea-macos-runner/config.json; see
Resources/config.example.json. Order of operations for a new host:
gitea-macos-runner doctor # verify arch, macOS, entitlement, keychain, Gitea
gitea-macos-runner config init # write an annotated config
gitea-macos-runner image build # ~1 hour, mostly IPSW download
gitea-macos-runner vm boot # optional smoke test: boot a clone, print its IP
gitea-macos-runner service install # LaunchAgent, RunAtLoad + KeepAlive
gitea-macos-runner doctor # again, now that it runs from the signed .app
doctor exists because every one of its checks corresponds to a failure that
otherwise appears as an opaque error deep inside a VM boot. The most common by
far: running from .build/ instead of the signed .app, so the entitlement is
absent.
9. Future work
- Save/restore for warm boots.
VZVirtualMachine.saveMachineStateTo(macOS 14+) could cut per-job boot from ~60 s to near-instant by restoring a snapshot taken just aftergitea-runneris ready. The blocker is that restore forbids changing the MAC address or ECID, which collides with our per-slot MAC scheme (§ Verified Fact 12) — a restored state would have to be captured per slot, and the interaction with DHCP lease reuse needs care. Deferred, not rejected. - vsock control channel.
VZVirtioSocketDeviceConfigurationis already in the VM configuration. Replacing SSH with a vsock agent would remove password auth, thewaitForSSHpoll, and the Local Network privacy prompt entirely. --manual-setupfor pre-macOS-27 guests (§4).- Image versioning / garbage collection for multiple base images.
Appendix: Verified Facts
Researched facts this design depends on, each with its consequence for the implementation. Anyone changing the corresponding code should re-verify the fact first.
1. Job discovery is GET /api/v1/admin/actions/jobs?status=queued (Gitea
1.25+). The labels field on a returned job is the workflow's runs-on:
value. The external status string queued maps to Gitea's internal
StatusWaiting, meaning "ready, waiting for a matching runner".
→ Consequence: the external string waiting means something entirely
different — the job is blocked on a dependency — and must never be
treated as schedulable. WorkflowJob.isQueued checks status == "queued" and
nothing else.
2. A queued job waits for a matching runner up to ABANDONED_JOB_TIMEOUT
(default 24 h, swept every 6 h).
→ Consequence: there is no urgency in the poll loop. A 5-second interval is for
responsiveness, not correctness; a daemon that is down for an hour loses nothing.
It also sets the ceiling that jobTimeoutMinutes (default 120) must stay well
under, so our teardown always precedes the server's abandonment.
3. The runner binary is gitea-runner v3.x, renamed from act_runner and
published from gitea.com/gitea/runner. The register-time flag --ephemeral
(Gitea 1.24+) is server-enforced: single job, then automatic deregistration.
--once is weaker — runner-side only.
→ Consequence: --ephemeral is mandatory and --once is never used. The
"one VM, one job" guarantee in §6 rests on the server enforcing it, not on the
runner cooperating. Download URLs and the binary name must reference
gitea-runner, not act_runner.
4. Registration tokens are REUSABLE, and minting a new token for a scope
INVALIDATES prior tokens of that scope. POST /api/v1/admin/actions/runners/registration-token in practice returns the
existing active token. Seeding server-side via GITEA_RUNNER_REGISTRATION_TOKEN
is the recommended deployment.
→ Consequence: never pre-generate a token per VM — doing so would break
every other runner registered against that scope. One token is resolved once,
cached for the process lifetime, and shared by the fleet;
fetchRegistrationTokenViaAPI defaults to false.
5. Label syntax is name:schema, schema defaults to host, and only BARE
names are stored server-side. If the guest's runner config.yaml sets
runner.labels, it silently overrides --labels passed at registration.
→ Consequence: LabelSet matches bare names (macos-arm64) and only appends
:host when building the register --labels argument. The guest must never
ship a config.yaml containing labels — noted in both GuestProvisioner and
Resources/provision.sh, because the failure is silent: the runner registers
successfully and is simply never matched.
6. Guest requirements: the gitea-runner binary, node, git, bash, and a
writable $HOME. JavaScript actions such as actions/checkout spawn node
directly.
→ Consequence: Node.js installation is not optional and not a convenience —
without it essentially every real workflow fails at its first step.
GuestProvisioner.verifyToolchain fails the build rather than shipping an image
that will break at job time.
7. An unclean VM death leaves both a runner row and a Running task behind.
The task is reaped by Gitea's zombie sweep in roughly 10–15 minutes. The runner
row is swept only at midnight — and never at all if that runner claimed no
task. Rows are removed with DELETE /api/v1/admin/actions/runners/{id}.
→ Consequence: the reconcile loop (§5) is load-bearing, not housekeeping.
Without it, every crashed boot leaves a permanent phantom runner. It is also why
every VM registers under a globally unique name (namePrefix + UUID): that
uniqueness is what lets us look at a row and decide with certainty that it is
ours and unbacked.
8. macOS guests are capped at 2 concurrent per host, enforced by the kernel.
A third start() raises VZError.virtualMachineLimitExceeded.
→ Consequence: maxConcurrentVMs is hard-clamped to 2 in both
RunnerConfig.validated() and SchedulerCore.plan; the slot table is fixed-size;
and the error is mapped to CoreError.vmLimitExceeded and treated as transient
back-pressure rather than a failure.
9. VZMacGuestProvisioningOptions (macOS 27+ host AND guest) automates Setup
Assistant, including enablesRemoteLogin (SSH). Older guests silently
ignore it.
→ Consequence: the API is gated at #available(macOS 27.0, *), and because the
guest-side failure produces no error, firstBootAndProvision must translate its
lease/SSH timeout into an explicit "guest too old" message rather than a bare
timeout. The --manual-setup fallback is out of v1 scope (§4).
10. Headless Virtualization requires an NSApplication run loop with
.prohibited activation policy, inside a signed .app bundle carrying
com.apple.security.virtualization. That entitlement is unrestricted: ad-hoc
signing (codesign -s -) grants it, and a Developer ID certificate grants it
with no provisioning profile. Bridged networking would additionally need a
restricted entitlement; NAT does not.
→ Consequence: CommandDaemon starts NSApplication and runs the orchestrator
in a detached Task. This applies to every command that starts a VM, not
just the daemon: vm boot, image build, and image provision all go through
the same VZAppRuntime.run host, since VZMacOSInstaller and the first-boot
provisioning pass need the run loop exactly as much as a job VM does. The
Makefile has bundle and sign targets and
install deliberately installs the bundle rather than the bare binary;
Info.plist sets LSUIElement; Doctor checks the entitlement on the running
binary because running from .build/ is the most common setup failure.
Two packaging constraints follow from the same fact. The entitlements plist must
contain no XML comments: plutil -lint accepts them, but codesign hands
the file to AMFI's stricter parser, which fails with AMFIUnserializeXML: syntax error and then signs the bundle with zero entitlements — a silent
downgrade that only surfaces as a failed VM start. And bundle must copy
provision.sh, launchd.plist.template, and config.example.json into
Contents/Resources, since GuestProvisioner, LaunchdService, and
config init look there before falling back to repo-relative paths; an installed
.app without them is a working binary with a broken image build,
service install, and config init.
10a. Which signature is used decides whether the app's code identity is stable
across rebuilds. A Developer ID signature's designated requirement is anchored
to the team (… and certificate leaf[subject.OU] = <TEAM_ID>), so every build
is the same program to macOS. An ad-hoc signature has no anchor, so identity
falls back to the main executable's Mach-O UUID, which the linker regenerates on
essentially every link.
→ Consequence: this is not cosmetic, because macOS Local Network privacy is
keyed on exactly that UUID (Fact 16). Under ad-hoc signing a grant is
withdrawn by the next make install; under Developer ID it persists. make sign
therefore selects a Developer ID Application identity matching TEAM_ID when
the keychain has one and falls back to ad-hoc with a warning when it does not —
the fallback is required because CI builds inside a throwaway guest with neither
keychain nor certificate. The Developer ID path also passes --options runtime --timestamp, so the bundle is notarizable later without re-signing; notarization
itself is skipped, since it governs distribution to other Macs and this app is
built and installed in place. Doctor.checkCodeSignature reports which path was
taken and warns on ad-hoc.
11. macOS 15+ requires an unlocked login.keychain to start a VM.
→ Consequence: the service must be a LaunchAgent in the auto-logged-in
user's session, never a LaunchDaemon (which has no session and no unlocked
keychain). LaunchdService only ever writes to ~/Library/LaunchAgents, and
Doctor probes with security show-keychain-info login.keychain.
12. Guest IPs come from parsing /var/db/dhcpd_leases, keyed by MAC.
hw_address lines carry a 1, hardware-type prefix and octets that may lack
zero-padding (aa:bb:c:dd:ee:ff). Duplicate MACs occur; the newest lease wins.
macOS's DHCP lease time is 24 hours.
→ Consequence: DHCPLeaseParser.normalizeMAC must strip the prefix and
zero-pad, or lookups silently fail against VZMACAddress.string. And ephemeral
fleets must not randomize MACs per clone — a day of dead leases would
accumulate and exhaust the NAT subnet. Hence exactly two persistent per-slot
MACs, generated once with VZMACAddress.randomLocallyAdministered() and stored
in state.json.
13. APFS copy-on-write cloning via FileManager.copyItem requires source and
destination on the same volume. Clones grow as the guest writes.
→ Consequence: base images and ephemeral clones both live under storeDir, and
cloning is done per file rather than by copying a directory wholesale.
ensureFreeSpace(minGB:) runs before every clone, with a floor (default 20 GB)
well above one clone's nominal cost, because the apparent size and the real cost
diverge over a job's lifetime.
14. The ASIF sparse disk format is created via diskutil image create (macOS
26+); RAW is the fallback.
→ Consequence: ImageBuilder.createDisk shells out to
/usr/sbin/diskutil image create blank --fs none --format ASIF --size <N>G <path>
and falls back to a sparse RAW file, recording which was used in
VMBundleConfig.diskFormat so VZConfigFactory attaches the right file without
re-probing. This is also the reason the package's minimum platform is macOS 26.
15. Save/restore (macOS 14+) could give near-instant warm boots, but forbids changing the MAC address or ECID. → Consequence: documented as future work (§9) rather than implemented. The prohibition collides directly with the per-slot MAC scheme from Fact 12, so adopting it would require per-slot saved states and a careful look at DHCP lease reuse — not a drop-in optimization.
16. macOS 15+ Local Network privacy can block host→guest connections, and
there is no reliable way to grant it interactively here. Per
TN3179
it is not TCC — the check is a Network Extension packet filter, so it is absent
from TCC.db, cannot be queried, cannot be reset, and it "uses your main
executable UUID as part of its implementation". A denial returns EHOSTUNREACH
(errno 65), indistinguishable from a genuinely unreachable host. Three things
then conspire against the interactive grant: a LaunchAgent has no UI to show the
prompt in; a run started from a shell is attributed to the responsible
process, so the prompt and the System Settings row belong to Terminal rather
than to this app, and granting it to Terminal does not carry to the agent; and
under ad-hoc signing the UUID keying (Fact 10a) withdraws the grant on the next
rebuild.
→ Consequence: the deterministic fix is the subnet allowlist
(com.apple.network.local-network, keys AllowedEthernetLocalNetworkAddresses
and AllowedWiFiLocalNetworkAddresses), which is keyed on the network rather
than the app and is read at boot — so it needs a reboot, not a service restart.
Doctor.checkLocalNetwork reads those keys and requires coverage of
192.168.64.0/18, not a single /24, because the NAT subnet is chosen at runtime
and slides to the next free /24. SSHExec.localNetworkHint appends the same
guidance to connection failures, since errno 65 gives the operator nothing to go
on by itself.