582 lines
35 KiB
Markdown
582 lines
35 KiB
Markdown
# gitea-macos-runner — Design
|
||
|
||
## 1. Overview
|
||
|
||
`gitea-macos-runner` is a single-host daemon for an Apple Silicon Mac. It watches
|
||
a Gitea instance for queued Actions jobs that require macOS, and for each one it
|
||
boots a **fresh, ephemeral macOS VM** on Apple's Virtualization.framework,
|
||
registers a single-use runner inside it, lets the job run, and then destroys the
|
||
VM.
|
||
|
||
The design goal is that **no state survives a job**. Not a checkout, not a
|
||
keychain entry, not a `~/Library` mutation, not a leftover process. The guest
|
||
that runs job *N+1* is a byte-identical copy-on-write clone of the same base
|
||
image that job *N* started from. This is the property that a persistent
|
||
self-hosted Mac runner cannot offer, and it is the whole reason this tool exists.
|
||
|
||
Three constraints shape everything below:
|
||
|
||
1. **Apple's kernel allows at most two concurrent macOS guests per host.** Not a
|
||
policy, not a licence term we chose — a hard limit that surfaces as
|
||
`VZError.virtualMachineLimitExceeded` from `start()`. Concurrency is therefore
|
||
2, permanently, and the config value is clamped rather than trusted.
|
||
2. **Virtualization needs a GUI session and a signed bundle.** The daemon runs as
|
||
a LaunchAgent in a logged-in user session, from inside an ad-hoc-signed `.app`
|
||
carrying `com.apple.security.virtualization`.
|
||
3. **Gitea decides which job a runner claims, not us.** We supply capacity; the
|
||
server matches. Trying to pin a specific job to a specific VM would mean
|
||
reimplementing Gitea's matching rules, and would be wrong the moment they
|
||
change.
|
||
|
||
### Non-goals
|
||
|
||
* Multi-host scheduling. One daemon, one Mac, two slots.
|
||
* Container-based execution. Gitea's `host` schema runs jobs directly on the
|
||
guest; that is the point of having a real macOS VM.
|
||
* Bridged networking. NAT only — see §6.
|
||
* Guest reuse or warm pools. See §9 for why save/restore is deferred rather than
|
||
rejected.
|
||
|
||
---
|
||
|
||
## 2. Component diagram
|
||
|
||
```
|
||
┌──────────────────────────────── Host (Apple Silicon Mac, macOS 26+) ─────────────────────────────┐
|
||
│ │
|
||
│ LaunchAgent (user session, auto-login, login.keychain unlocked) │
|
||
│ └── GiteaMacosRunner.app (ad-hoc signed, com.apple.security.virtualization, LSUIElement) │
|
||
│ │ │
|
||
│ │ NSApplication(.prohibited).run() ── main thread, required by Virtualization │
|
||
│ │ │
|
||
│ ┌─────▼──────────────────────────── Orchestrator (actor) ────────────────────────────────┐ │
|
||
│ │ │ │
|
||
│ │ poll loop ──► GiteaClient.listQueuedJobs() ──► [WorkflowJob] │ │
|
||
│ │ │ │ │
|
||
│ │ ├──────► SchedulerCore.plan(...) ── PURE, no I/O ──► [SchedulerAction] │ │
|
||
│ │ │ │ │
|
||
│ │ ├──► bootVM(slot,jobHint) ──► VMStore.cloneImage ──► VMInstance.start │ │
|
||
│ │ │ │ │ │ │
|
||
│ │ │ │ ▼ │ │
|
||
│ │ │ │ DHCPLeaseParser(/var/db/…) │ │
|
||
│ │ │ │ │ │ │
|
||
│ │ │ │ ▼ │ │
|
||
│ │ │ │ SSHExecutor ──► gitea-runner │ │
|
||
│ │ │ │ register+daemon │ │
|
||
│ │ └──► teardownVM(slot,reason) ──► VMInstance.requestStopThenForce ──► deleteClone │ │
|
||
│ │ │ │
|
||
│ │ reconcile loop ──► GiteaClient.listRunners / deleteRunner (sweep orphaned rows) │ │
|
||
│ └────────────────────────────────────────────────────────────────────────────────────────┘ │
|
||
│ │
|
||
│ <storeDir>/ │
|
||
│ images/default/{disk.asif, nvram.bin, config.json} ← built once, provisioned, read-only │
|
||
│ vms/<uuid>/{disk.asif, nvram.bin, config.json} ← APFS CoW clones, destroyed per job │
|
||
│ ipsw/ ← downloaded restore images │
|
||
│ state.json ← the two persistent per-slot MACs │
|
||
│ │
|
||
│ ┌──── VM slot 0 (MAC A) ────┐ ┌──── VM slot 1 (MAC B) ────┐ ← at most 2, kernel-enforced │
|
||
│ │ macOS guest │ │ macOS guest │ │
|
||
│ │ gitea-runner --ephemeral │ │ gitea-runner --ephemeral │ │
|
||
│ │ node, git, bash │ │ node, git, bash │ │
|
||
│ └───────────┬───────────────┘ └───────────┬───────────────┘ │
|
||
└──────────────┼───────────────────────────────┼───────────────────────────────────────────────────┘
|
||
│ NAT (vmenet, bootpd) │
|
||
└───────────────┬───────────────┘
|
||
▼
|
||
┌─────────────────────┐
|
||
│ Gitea 1.25+ │
|
||
│ /api/v1/admin/… │
|
||
└─────────────────────┘
|
||
```
|
||
|
||
### Module boundaries
|
||
|
||
| Target | Contains | Constraint |
|
||
|---|---|---|
|
||
| `RunnerCore` | Config, Gitea models + client, `LabelSet`, `DHCPLeaseParser`, `SSHExec`, `SchedulerCore` | **No `import Virtualization`.** Builds on Linux, so scheduling and parsing logic can be unit-tested anywhere. |
|
||
| `RunnerHost` | `VMBundle`, `VMStore`, `VZConfigFactory`, `VMInstance`, `IPSW`, `ImageBuilder`, `GuestProvisioner`, `Orchestrator`, `LaunchdService`, `Doctor` | macOS-only. Everything that touches the framework. |
|
||
| `gitea-macos-runner` | CLI + daemon entry point | Depends on both. |
|
||
|
||
The split is not cosmetic: `SchedulerCore` being pure and portable is what makes
|
||
the scheduling policy — the part most likely to have subtle bugs — testable
|
||
without a Mac, a VM, or a Gitea instance.
|
||
|
||
---
|
||
|
||
## 3. Job lifecycle
|
||
|
||
```
|
||
Gitea Orchestrator VMStore / VMInstance Guest
|
||
│ │ │ │
|
||
│◄── listQueuedJobs ────────┤ (every pollIntervalSeconds) │ │
|
||
├─── [job 4711, labels ─────► │ │
|
||
│ ["macos-arm64"]] │ │ │
|
||
│ ├── LabelSet.matches? ─────────┤ │
|
||
│ ├── SchedulerCore.plan ────────┤ │
|
||
│ │ → .bootVM(slot: 0, │ │
|
||
│ │ jobHint: 4711) │ │
|
||
│ │ │ │
|
||
│ ├── ensureFreeSpace(minGB) ───►│ │
|
||
│ ├── cloneImage("default", ───►│ APFS CoW copy │
|
||
│ │ slotMAC: MAC-A) │ + rewrite config.json │
|
||
│ ├── VMInstance.start() ───────►│ ──── boot ───────────►│
|
||
│ │ │ │
|
||
│ ├── poll /var/db/dhcpd_leases ─┤◄─── DHCP request ──────┤
|
||
│ │ until MAC-A has an IP │ │
|
||
│ ├── waitForSSH(ip) ────────────┼───────────────────────►│
|
||
│ │ │ │
|
||
│ ├── uploadData(token, 0600) ───┼───────────────────────►│
|
||
│ ├── ssh: gitea-runner register ┼───────────────────────►│
|
||
│◄────────────────────── register (name=macos-vm-<uuid>, --ephemeral) ──────────────┤
|
||
│ ├── ssh: rm -f <tokenfile> │ │
|
||
│ ├── ssh: gitea-runner daemon ──┼───────────────────────►│
|
||
│ │ │ │
|
||
│◄────────────────────── poll for task ─────────────────────────────────────────────┤
|
||
├─── assign job 4711 ───────────────────────────────────────────────────────────────►
|
||
│ │ │ ...running... │
|
||
│◄────────────────────── job result, logs ──────────────────────────────────────────┤
|
||
├─── auto-deregister runner (server-enforced --ephemeral) ──────────────────────────►
|
||
│ │ │ daemon exits │
|
||
│ │◄─ SSH command returns ───────┼────────────────────────┤
|
||
│ ├── teardownVM(slot: 0) ──────►│ │
|
||
│ │ requestStopThenForce ──────┼───────────────────────►│ (halt)
|
||
│ │ deleteClone ──────────────►│ rm -rf vms/<uuid> │
|
||
│ ├── markIdle(slot: 0) │ │
|
||
```
|
||
|
||
Two details in that sequence carry more weight than their size suggests.
|
||
|
||
**The token goes through a file, not an argument.** `gitea-runner register` is
|
||
invoked with `--token-file <f>`, where `<f>` was written by `uploadData` with
|
||
mode `0600` and is `rm -f`'d in the same shell command. Passing `--token` would
|
||
put a fleet-wide credential into the guest's process table, visible to any
|
||
process the job spawns — and the job is arbitrary code from a repository.
|
||
|
||
**`--ephemeral`, not `--once`.** `--ephemeral` (Gitea 1.24+) is enforced *by the
|
||
server*: it hands this runner exactly one task and then deletes the registration.
|
||
`--once` is a runner-side convention only — the server still considers the runner
|
||
live, and a misbehaving or patched runner could claim more work. Since the whole
|
||
security story here rests on "one VM, one job", the enforcement has to live on
|
||
the side we don't hand to the job.
|
||
|
||
---
|
||
|
||
## 4. Image build pipeline
|
||
|
||
Base images are built once with `image build`, and every job clones one. The
|
||
build is slow (most of an hour, mostly a ~15 GB download); the clone is
|
||
milliseconds.
|
||
|
||
```
|
||
image build --name default [--ipsw PATH]
|
||
│
|
||
├─ 1. IPSWProvider.latestSupported() → CDN url + buildVersion
|
||
│ (VZMacOSRestoreImage.latestSupported returns a NETWORK url —
|
||
│ it cannot be handed to the installer)
|
||
├─ 2. IPSWProvider.download() → <storeDir>/ipsw/*.ipsw
|
||
├─ 3. IPSWProvider.load(localPath:) → VZMacOSRestoreImage
|
||
│ (resolveSymlinksInPath first; the framework rejects symlinks)
|
||
│
|
||
├─ 4. restoreImage.mostFeaturefulSupportedConfiguration
|
||
│ nil ⇒ this host cannot run this image. Fail loudly; do not guess.
|
||
│
|
||
├─ 5. createBundle()
|
||
│ hardwareModel.dataRepresentation → config.json
|
||
│ VZMacMachineIdentifier() (fresh) → config.json
|
||
│ VZMacAuxiliaryStorage(creatingStorageAt:hardwareModel:) → nvram.bin
|
||
│ disk: diskutil image create blank --fs none --format ASIF --size <N>G
|
||
│ └─ fallback: sparse RAW file (format recorded in config.json)
|
||
│
|
||
├─ 6. VZMacOSInstaller(virtualMachine:restoringFromImageAt:) on a STOPPED vm
|
||
│ KVO on installer.progress → percentage
|
||
│
|
||
├─ 7. first boot with Setup Assistant automation
|
||
│ #available(macOS 27.0, *):
|
||
│ VZMacGuestProvisioningOptions(username/password/fullName,
|
||
│ logsInAutomatically: true,
|
||
│ enablesRemoteLogin: true)
|
||
│ → VZMacOSVirtualMachineStartOptions.setGuestProvisioning(_:)
|
||
│ ⚠ an OLDER GUEST SILENTLY IGNORES THIS — no error, no account, no SSH
|
||
│
|
||
├─ 8. wait for DHCP lease (by MAC) → wait for SSH → GuestProvisioner
|
||
│ provision.sh (sudoers, no-sleep, no-Spotlight, maxfiles, known_hosts)
|
||
│ Node.js (official arm64 .pkg → installer -pkg) ← REQUIRED
|
||
│ verify git / bash / node
|
||
│ gitea-runner (host downloads asset → upload → chmod +x)
|
||
│ [optional] Xcode from a .xip
|
||
│
|
||
└─ 9. clean shutdown → config.provisioned = true ← only now is it clonable
|
||
```
|
||
|
||
### On step 7 and its failure mode
|
||
|
||
`VZMacGuestProvisioningOptions` needs **macOS 27 or newer on both the host and
|
||
the guest**. The host side is a compile/availability check we control. The guest
|
||
side is not: an older guest accepts the boot and simply ignores the options.
|
||
There is no error to catch. The observable symptom is that the VM boots, sits at
|
||
Setup Assistant forever, never requests a DHCP lease with a usable hostname, and
|
||
never answers SSH — so the build fails at step 8 with a timeout that says nothing
|
||
useful.
|
||
|
||
`firstBootAndProvision` therefore detects the timeout and reports it as an
|
||
explicit "guest is too old for unattended setup; supply a macOS 27+ IPSW"
|
||
failure. A `--manual-setup` flow that opens a window and lets a human click
|
||
through Setup Assistant once is **out of scope for v1** — deliberately, because a
|
||
GUI step in a tool whose whole purpose is unattended operation is a trap. It is
|
||
noted here so the omission is a decision rather than an oversight.
|
||
|
||
### On step 5's disk format
|
||
|
||
ASIF is preferred because it is sparse: a 64 GB nominal disk costs what the guest
|
||
actually writes, and it CoW-clones cleanly on APFS. `diskutil image create` is
|
||
shelled out to because there is no framework API for it. If that call fails for
|
||
any reason — older `diskutil`, unusual volume — a sparse RAW file is created
|
||
instead and the format is recorded in `config.json`, so `VZConfigFactory` attaches
|
||
the right file without re-probing.
|
||
|
||
---
|
||
|
||
## 5. Scheduling semantics
|
||
|
||
`SchedulerCore.plan` is a pure function: `(state, queuedJobs, labels, maxVMs,
|
||
now, jobTimeout, bootTimeout) → (state', [action])`. It performs no I/O, reads no
|
||
clock, and is fully deterministic — which is what allows the entire scheduling
|
||
policy to be tested with a fixed `now` and a synthetic job list.
|
||
|
||
### Capacity, not assignment
|
||
|
||
This is the central idea and the easiest thing to get wrong.
|
||
|
||
A booted VM is **capacity**. It is not a promise to run a particular job. We see
|
||
job 4711 queued, we boot a VM, we register an ephemeral runner — and the *server*
|
||
then decides which queued job that runner claims. It may well claim job 4712
|
||
instead. That is fine and in fact preferable: Gitea's matching rules (labels,
|
||
repo permissions, ordering, priority) are its business, and any attempt to
|
||
predict them here would be a reimplementation that drifts out of sync.
|
||
|
||
The `jobHint` threaded through `SchedulerAction.bootVM` and
|
||
`SlotState.running` exists for exactly two purposes: log messages, and the dedup
|
||
ledger below. Nothing else may depend on it.
|
||
|
||
### Dedup by job id
|
||
|
||
`SchedulerState.dispatchedJobIDs` is a `Set<Int64>` of jobs that have already
|
||
caused a boot.
|
||
|
||
Without it, the loop is pathological. A VM takes tens of seconds to boot,
|
||
provision, and register. The poll interval is 5 seconds. So a single queued job
|
||
would still be queued on the next poll, and the next, and the next — triggering a
|
||
second boot, then exhausting the slot budget, all for one job.
|
||
|
||
The ledger is expired against reality rather than against a timer: any id no
|
||
longer appearing in the queued set is dropped. That way a slot freed by a
|
||
completed job can be re-earned by a genuinely new job, but a job that is *still*
|
||
waiting does not double-book.
|
||
|
||
### The cap
|
||
|
||
`maxVMs` is clamped to 2 in `plan`, and again in `RunnerConfig.validated()`. Both
|
||
places, because the kernel limit is not something a config file gets to
|
||
negotiate: a third `start()` raises `VZError.virtualMachineLimitExceeded`, which
|
||
`VMInstance.mapVZError` translates into `CoreError.vmLimitExceeded` and the
|
||
scheduler treats as transient back-pressure rather than a failure.
|
||
|
||
### Timeouts
|
||
|
||
* A slot in `.provisioning(since:)` longer than `bootTimeoutSeconds` (default
|
||
300) is torn down. Covers a guest that never gets a lease, never starts `sshd`,
|
||
or hangs in Setup Assistant.
|
||
* A slot in `.running(jobHint:since:)` longer than `jobTimeoutMinutes` (default
|
||
120) is torn down. Covers a job that hangs. This is comfortably below Gitea's
|
||
own `ABANDONED_JOB_TIMEOUT` (24 h), so our teardown always happens first and
|
||
the server sees a clean deregistration rather than an abandonment.
|
||
|
||
Teardown actions are emitted **before** boot actions in the returned list, so a
|
||
slot freed in one pass can be reused in that same pass.
|
||
|
||
### Reconcile
|
||
|
||
Every `reconcileIntervalSeconds` (default 300), the orchestrator lists runners and
|
||
deletes any that are:
|
||
|
||
* `ephemeral == true`, **and**
|
||
* `busy == false`, **and**
|
||
* `name` starts with our configured `namePrefix`, **and**
|
||
* not backed by a live VM in this process.
|
||
|
||
All four conditions, because deleting a live runner fails a running job. The loop
|
||
is deliberately conservative: a row we are unsure about is left alone, and will be
|
||
revisited in five minutes.
|
||
|
||
This loop is not optional housekeeping — it is load-bearing. See Verified Fact 7.
|
||
|
||
---
|
||
|
||
## 6. Security model
|
||
|
||
### The threat
|
||
|
||
A CI job is arbitrary code from a repository, running with the privileges of the
|
||
account it executes under. On a persistent self-hosted Mac runner, that code can
|
||
read every previous job's checkout, poison caches, install launch agents, and
|
||
harvest whatever credentials the machine has accumulated. Every subsequent job on
|
||
that host inherits the compromise.
|
||
|
||
### The mitigation: genuinely ephemeral guests
|
||
|
||
* **One job per VM, enforced server-side.** `--ephemeral` means Gitea hands the
|
||
runner exactly one task and then deletes the registration. A patched or
|
||
hijacked runner binary cannot ask for more work, because the server will not
|
||
give it any.
|
||
* **The VM is destroyed after that job.** Not reset, not cleaned — the clone
|
||
directory is `rm -rf`'d and the next job clones the base image afresh. There is
|
||
no path by which job *N* influences job *N+1* short of compromising the host.
|
||
* **The guest holds nothing worth stealing.** Its account credentials
|
||
(`admin`/`admin` by default) are meaningful only on a host-private NAT link to a
|
||
machine that is about to be deleted.
|
||
|
||
### The shared registration token
|
||
|
||
Registration tokens in Gitea are **reusable and scope-wide**, and minting a new
|
||
one for a scope **invalidates all prior tokens of that scope**. That makes
|
||
per-VM tokens actively harmful: generating one for each VM would break every
|
||
other runner registered against that scope, including ones on other hosts.
|
||
|
||
So the fleet shares one token. The mitigations are:
|
||
|
||
* It is written into the guest as a **file with mode `0600`**, never as a command
|
||
argument (arguments are world-readable via `ps`).
|
||
* It is **deleted immediately** after `gitea-runner register` consumes it, in the
|
||
same `&&` chain, before `gitea-runner daemon` starts and long before any job
|
||
code runs.
|
||
* It is a *registration* token, not an API token: it grants the ability to
|
||
register a runner, not to read repositories or act as a user.
|
||
|
||
The residual risk is real but bounded — a job that wins a race against `rm -f`
|
||
could register additional runners for that scope. The recommended deployment
|
||
seeds a fixed token server-side via `GITEA_RUNNER_REGISTRATION_TOKEN` so that
|
||
rotating it is a deliberate, coordinated act rather than an API call side effect.
|
||
|
||
### `:host` schema risk
|
||
|
||
Jobs run in `host` schema: directly on the guest OS, not in a container. That is
|
||
the point — a macOS job needs real macOS. But it means the job has full user-level
|
||
access to the guest, including `sudo` (which `provision.sh` makes passwordless,
|
||
because Xcode and `installer` need it). Everything above rests on the guest being
|
||
disposable and isolated, not on the job being constrained inside it.
|
||
|
||
### Host-side posture
|
||
|
||
* The daemon runs as a **LaunchAgent in a user session**, not as root. The
|
||
entitlement it carries (`com.apple.security.virtualization`) grants VM creation
|
||
and nothing else.
|
||
* Networking is **NAT**, not bridged. Guests can reach the LAN and Gitea, but are
|
||
not first-class hosts on it. Bridged networking would require the restricted
|
||
`com.apple.vm.networking` entitlement, which ad-hoc signing cannot grant — a
|
||
constraint that happens to align with what we want anyway.
|
||
* **SSH host keys are not verified.** The peer is a VM this process booted
|
||
moments ago on a link no other machine shares; pinning would break on every
|
||
clone and add nothing.
|
||
|
||
---
|
||
|
||
## 7. Failure modes and recovery
|
||
|
||
| Failure | Detection | Recovery |
|
||
|---|---|---|
|
||
| Guest never gets a DHCP lease | `bootTimeout` in `waitForLease` | Teardown, slot recycled, retried next poll |
|
||
| Guest never answers SSH | `bootTimeout` in `waitForSSH` | Same |
|
||
| Job hangs | `jobTimeoutMinutes` | Teardown; Gitea reaps the task via its zombie sweep (~10–15 min) |
|
||
| VM dies uncleanly | Runner row left behind, task stuck Running | Reconcile loop deletes the row (§5); Gitea's zombie sweep handles the task |
|
||
| Daemon crashes with VMs live | Clones orphaned on disk | `purgeClones()` at startup, then one immediate reconcile pass |
|
||
| Third VM requested | `VZError.virtualMachineLimitExceeded` | Mapped to `CoreError.vmLimitExceeded`, treated as back-pressure |
|
||
| Disk fills | `ensureFreeSpace(minGB:)` before each clone | Boot refused, logged; jobs stay queued (safe — Gitea holds them ~24 h) |
|
||
| Gitea unreachable | Request error in the poll loop | Logged, retried next tick; no state change |
|
||
|
||
---
|
||
|
||
## 8. Configuration and operations summary
|
||
|
||
Config lives at `~/.config/gitea-macos-runner/config.json`; see
|
||
`Resources/config.example.json`. Order of operations for a new host:
|
||
|
||
```
|
||
gitea-macos-runner doctor # verify arch, macOS, entitlement, keychain, Gitea
|
||
gitea-macos-runner config init # write an annotated config
|
||
gitea-macos-runner image build # ~1 hour, mostly IPSW download
|
||
gitea-macos-runner vm boot # optional smoke test: boot a clone, print its IP
|
||
gitea-macos-runner service install # LaunchAgent, RunAtLoad + KeepAlive
|
||
gitea-macos-runner doctor # again, now that it runs from the signed .app
|
||
```
|
||
|
||
`doctor` exists because every one of its checks corresponds to a failure that
|
||
otherwise appears as an opaque error deep inside a VM boot. The most common by
|
||
far: running from `.build/` instead of the signed `.app`, so the entitlement is
|
||
absent.
|
||
|
||
---
|
||
|
||
## 9. Future work
|
||
|
||
* **Save/restore for warm boots.** `VZVirtualMachine.saveMachineStateTo` (macOS
|
||
14+) could cut per-job boot from ~60 s to near-instant by restoring a snapshot
|
||
taken just after `gitea-runner` is ready. The blocker is that restore forbids
|
||
changing the MAC address or ECID, which collides with our per-slot MAC scheme
|
||
(§ Verified Fact 12) — a restored state would have to be captured per slot, and
|
||
the interaction with DHCP lease reuse needs care. Deferred, not rejected.
|
||
* **vsock control channel.** `VZVirtioSocketDeviceConfiguration` is already in the
|
||
VM configuration. Replacing SSH with a vsock agent would remove password auth,
|
||
the `waitForSSH` poll, and the Local Network privacy prompt entirely.
|
||
* **`--manual-setup`** for pre-macOS-27 guests (§4).
|
||
* **Image versioning / garbage collection** for multiple base images.
|
||
|
||
---
|
||
|
||
## Appendix: Verified Facts
|
||
|
||
Researched facts this design depends on, each with its consequence for the
|
||
implementation. Anyone changing the corresponding code should re-verify the fact
|
||
first.
|
||
|
||
**1. Job discovery is `GET /api/v1/admin/actions/jobs?status=queued` (Gitea
|
||
1.25+).** The `labels` field on a returned job is the workflow's `runs-on:`
|
||
value. The external status string `queued` maps to Gitea's internal
|
||
`StatusWaiting`, meaning "ready, waiting for a matching runner".
|
||
→ *Consequence:* the external string `waiting` means something entirely
|
||
different — the job is **blocked** on a dependency — and must **never** be
|
||
treated as schedulable. `WorkflowJob.isQueued` checks `status == "queued"` and
|
||
nothing else.
|
||
|
||
**2. A queued job waits for a matching runner up to `ABANDONED_JOB_TIMEOUT`
|
||
(default 24 h, swept every 6 h).**
|
||
→ *Consequence:* there is no urgency in the poll loop. A 5-second interval is for
|
||
responsiveness, not correctness; a daemon that is down for an hour loses nothing.
|
||
It also sets the ceiling that `jobTimeoutMinutes` (default 120) must stay well
|
||
under, so our teardown always precedes the server's abandonment.
|
||
|
||
**3. The runner binary is `gitea-runner` v3.x**, renamed from `act_runner` and
|
||
published from `gitea.com/gitea/runner`. The register-time flag `--ephemeral`
|
||
(Gitea 1.24+) is **server-enforced**: single job, then automatic deregistration.
|
||
`--once` is weaker — runner-side only.
|
||
→ *Consequence:* `--ephemeral` is mandatory and `--once` is never used. The
|
||
"one VM, one job" guarantee in §6 rests on the server enforcing it, not on the
|
||
runner cooperating. Download URLs and the binary name must reference
|
||
`gitea-runner`, not `act_runner`.
|
||
|
||
**4. Registration tokens are REUSABLE, and minting a new token for a scope
|
||
INVALIDATES prior tokens of that scope.** `POST
|
||
/api/v1/admin/actions/runners/registration-token` in practice returns the
|
||
existing active token. Seeding server-side via `GITEA_RUNNER_REGISTRATION_TOKEN`
|
||
is the recommended deployment.
|
||
→ *Consequence:* **never pre-generate a token per VM** — doing so would break
|
||
every other runner registered against that scope. One token is resolved once,
|
||
cached for the process lifetime, and shared by the fleet;
|
||
`fetchRegistrationTokenViaAPI` defaults to `false`.
|
||
|
||
**5. Label syntax is `name:schema`, schema defaults to `host`, and only BARE
|
||
names are stored server-side.** If the guest's runner `config.yaml` sets
|
||
`runner.labels`, it **silently overrides** `--labels` passed at registration.
|
||
→ *Consequence:* `LabelSet` matches bare names (`macos-arm64`) and only appends
|
||
`:host` when building the `register --labels` argument. The guest must **never**
|
||
ship a `config.yaml` containing labels — noted in both `GuestProvisioner` and
|
||
`Resources/provision.sh`, because the failure is silent: the runner registers
|
||
successfully and is simply never matched.
|
||
|
||
**6. Guest requirements: the `gitea-runner` binary, `node`, `git`, `bash`, and a
|
||
writable `$HOME`.** JavaScript actions such as `actions/checkout` spawn `node`
|
||
**directly**.
|
||
→ *Consequence:* Node.js installation is **not optional** and not a convenience —
|
||
without it essentially every real workflow fails at its first step.
|
||
`GuestProvisioner.verifyToolchain` fails the build rather than shipping an image
|
||
that will break at job time.
|
||
|
||
**7. An unclean VM death leaves both a runner row and a Running task behind.**
|
||
The task is reaped by Gitea's zombie sweep in roughly 10–15 minutes. The runner
|
||
row is swept only at midnight — and **never at all** if that runner claimed no
|
||
task. Rows are removed with `DELETE /api/v1/admin/actions/runners/{id}`.
|
||
→ *Consequence:* the reconcile loop (§5) is load-bearing, not housekeeping.
|
||
Without it, every crashed boot leaves a permanent phantom runner. It is also why
|
||
every VM registers under a **globally unique** name (`namePrefix` + UUID): that
|
||
uniqueness is what lets us look at a row and decide with certainty that it is
|
||
ours and unbacked.
|
||
|
||
**8. macOS guests are capped at 2 concurrent per host, enforced by the kernel.**
|
||
A third `start()` raises `VZError.virtualMachineLimitExceeded`.
|
||
→ *Consequence:* `maxConcurrentVMs` is hard-clamped to 2 in both
|
||
`RunnerConfig.validated()` and `SchedulerCore.plan`; the slot table is fixed-size;
|
||
and the error is mapped to `CoreError.vmLimitExceeded` and treated as transient
|
||
back-pressure rather than a failure.
|
||
|
||
**9. `VZMacGuestProvisioningOptions` (macOS 27+ host AND guest) automates Setup
|
||
Assistant**, including `enablesRemoteLogin` (SSH). Older guests **silently
|
||
ignore** it.
|
||
→ *Consequence:* the API is gated at `#available(macOS 27.0, *)`, and because the
|
||
guest-side failure produces no error, `firstBootAndProvision` must translate its
|
||
lease/SSH timeout into an explicit "guest too old" message rather than a bare
|
||
timeout. The `--manual-setup` fallback is out of v1 scope (§4).
|
||
|
||
**10. Headless Virtualization requires an `NSApplication` run loop with
|
||
`.prohibited` activation policy, inside a signed `.app` bundle** carrying
|
||
`com.apple.security.virtualization`. Ad-hoc signing (`codesign -s -`) suffices.
|
||
Bridged networking would additionally need a restricted entitlement; NAT does
|
||
not.
|
||
→ *Consequence:* `CommandDaemon` starts `NSApplication` and runs the orchestrator
|
||
in a detached `Task`. This applies to **every** command that starts a VM, not
|
||
just the daemon: `vm boot`, `image build`, and `image provision` all go through
|
||
the same `VZAppRuntime.run` host, since `VZMacOSInstaller` and the first-boot
|
||
provisioning pass need the run loop exactly as much as a job VM does. The
|
||
`Makefile` has `bundle` and `sign` targets and
|
||
`install` deliberately installs the bundle rather than the bare binary;
|
||
`Info.plist` sets `LSUIElement`; `Doctor` checks the entitlement on the running
|
||
binary because running from `.build/` is the most common setup failure.
|
||
Two packaging constraints follow from the same fact. The entitlements plist must
|
||
contain **no XML comments**: `plutil -lint` accepts them, but `codesign` hands
|
||
the file to AMFI's stricter parser, which fails with `AMFIUnserializeXML: syntax
|
||
error` and then signs the bundle with *zero* entitlements — a silent
|
||
downgrade that only surfaces as a failed VM start. And `bundle` must copy
|
||
`provision.sh`, `launchd.plist.template`, and `config.example.json` into
|
||
`Contents/Resources`, since `GuestProvisioner`, `LaunchdService`, and
|
||
`config init` look there before falling back to repo-relative paths; an installed
|
||
`.app` without them is a working binary with a broken `image build`,
|
||
`service install`, and `config init`.
|
||
|
||
**11. macOS 15+ requires an unlocked `login.keychain` to start a VM.**
|
||
→ *Consequence:* the service **must** be a LaunchAgent in the auto-logged-in
|
||
user's session, never a LaunchDaemon (which has no session and no unlocked
|
||
keychain). `LaunchdService` only ever writes to `~/Library/LaunchAgents`, and
|
||
`Doctor` probes with `security show-keychain-info login.keychain`.
|
||
|
||
**12. Guest IPs come from parsing `/var/db/dhcpd_leases`, keyed by MAC.**
|
||
`hw_address` lines carry a `1,` hardware-type prefix and octets that may lack
|
||
zero-padding (`aa:bb:c:dd:ee:ff`). Duplicate MACs occur; the newest lease wins.
|
||
macOS's DHCP lease time is 24 hours.
|
||
→ *Consequence:* `DHCPLeaseParser.normalizeMAC` must strip the prefix and
|
||
zero-pad, or lookups silently fail against `VZMACAddress.string`. And ephemeral
|
||
fleets must **not** randomize MACs per clone — a day of dead leases would
|
||
accumulate and exhaust the NAT subnet. Hence exactly two **persistent per-slot
|
||
MACs**, generated once with `VZMACAddress.randomLocallyAdministered()` and stored
|
||
in `state.json`.
|
||
|
||
**13. APFS copy-on-write cloning via `FileManager.copyItem` requires source and
|
||
destination on the same volume.** Clones grow as the guest writes.
|
||
→ *Consequence:* base images and ephemeral clones both live under `storeDir`, and
|
||
cloning is done **per file** rather than by copying a directory wholesale.
|
||
`ensureFreeSpace(minGB:)` runs before every clone, with a floor (default 20 GB)
|
||
well above one clone's nominal cost, because the apparent size and the real cost
|
||
diverge over a job's lifetime.
|
||
|
||
**14. The ASIF sparse disk format is created via `diskutil image create` (macOS
|
||
26+); RAW is the fallback.**
|
||
→ *Consequence:* `ImageBuilder.createDisk` shells out to
|
||
`/usr/sbin/diskutil image create blank --fs none --format ASIF --size <N>G <path>`
|
||
and falls back to a sparse RAW file, recording which was used in
|
||
`VMBundleConfig.diskFormat` so `VZConfigFactory` attaches the right file without
|
||
re-probing. This is also the reason the package's minimum platform is macOS 26.
|
||
|
||
**15. Save/restore (macOS 14+) could give near-instant warm boots, but forbids
|
||
changing the MAC address or ECID.**
|
||
→ *Consequence:* documented as future work (§9) rather than implemented. The
|
||
prohibition collides directly with the per-slot MAC scheme from Fact 12, so
|
||
adopting it would require per-slot saved states and a careful look at DHCP lease
|
||
reuse — not a drop-in optimization.
|