664 lines
34 KiB
Markdown
664 lines
34 KiB
Markdown
# Troubleshooting
|
||
|
||
Start with `gitea-macos-runner doctor` — it catches most misconfiguration before you go
|
||
symptom-hunting. Then find your symptom below.
|
||
|
||
Useful log commands throughout:
|
||
|
||
```sh
|
||
# Daemon logs, live
|
||
log stream --predicate 'process == "gitea-macos-runner"' --info
|
||
|
||
# Daemon logs, last hour
|
||
log show --predicate 'process == "gitea-macos-runner"' --info --last 1h
|
||
|
||
gitea-macos-runner service status
|
||
```
|
||
|
||
---
|
||
|
||
## Quick reference
|
||
|
||
| Symptom | Cause | Fix |
|
||
| --- | --- | --- |
|
||
| VM won't start; entitlement / `com.apple.security.virtualization` error | Running an unsigned binary, or one outside the signed `.app` bundle | `make sign` (or re-run `make install`); invoke the installed bundle, never `.build/release/…` |
|
||
| `virtualMachineLimitExceeded` at boot | macOS allows at most **2** concurrent macOS VMs | Set `scheduler.maxConcurrentVMs` ≤ 2; kill stray VMs from earlier runs |
|
||
| VM boots but never gets an IP | DHCP lease not yet written, or Local Network privacy denial (macOS 15+) | Check `/var/db/dhcpd_leases`; grant Local Network permission or pre-authorize the subnet |
|
||
| Runner not listed under Privacy & Security → Local Network | Expected — the list is populated only after the app first attempts a local connection, and a run started from a shell is attributed to Terminal, not to the app | Allowlist the subnet with `defaults write com.apple.network.local-network` and reboot; the interactive grant does not reach the LaunchAgent |
|
||
| SSH times out on a freshly built image | Guest macOS < 27, so provisioning options were ignored and Setup Assistant is waiting | Rebuild the image from a macOS **27+** IPSW |
|
||
| VMs boot in a loop; every teardown says `reason=cancelled` and nothing is logged between the lease and the teardown | `scheduler.bootTimeoutSeconds` is below the guest's *worst-case* boot on a contended host, so each clone is killed while still starting — and each replacement makes the next one slower | Raise `scheduler.bootTimeoutSeconds` (default 900) and reduce the number of concurrent guests; see [The daemon boots VMs forever](#the-daemon-boots-vms-forever-and-every-teardown-says-reasoncancelled) |
|
||
| `ssh failed: cannot connect … No route to host) (errno: 65)` part-way through provisioning | macOS 15+ Local Network privacy blocking the app — the grant is keyed on the executable's UUID, so `make install` withdraws it | Allowlist the subnet (`192.168.64.0/18`) and **reboot**; see [SSH fails with "No route to host" mid-run](#ssh-fails-with-no-route-to-host-errno-65-mid-run) |
|
||
| Allowlist is set but guests are still unreachable | It names `192.168.64.0/24` while the NAT has moved to `192.168.65.x` | Widen it to `192.168.64.0/18` and reboot; `doctor` now warns about too-narrow allowlists |
|
||
| `SecKeyCreateRandomKey` / "Interaction is not allowed" | `login.keychain` is locked — no GUI session | Run as a LaunchAgent in an unlocked GUI session; enable auto-login |
|
||
| Job stays queued forever | Label mismatch, or the daemon isn't running/reaching Gitea | Use bare label names in `runs-on`; match `runner.labels`; check daemon logs |
|
||
| `actions/checkout` fails instantly | Node.js missing from the guest image | `gitea-macos-runner image provision <name>` |
|
||
| Runner rows piling up in the Gitea UI | VMs killed uncleanly; registrations orphaned | Reconcile loop cleans them; force it by restarting the daemon; delete manually if needed |
|
||
| `doctor` says `runner download url → 403` but the URL works in a browser | Old build: the check used `HEAD`, and the presigned redirect target is signed per method | Upgrade — the check now uses a ranged `GET`. If it persists, the asset really is missing |
|
||
| Disk filling up | Copy-on-write clones grow as jobs write | Raise `storage.minFreeDiskGB`; delete stale clones in `storeDir/vms` |
|
||
| `image build` appears to hang during install | Normal — macOS install is slow | Wait. **Do not stop the VM mid-install**; if you did, delete the image and rebuild |
|
||
| `image build` prints its banner and then nothing, at 0% CPU | Old build: the build task was queued behind the `NSApplication` run loop and never started | Upgrade — the task is now detached. Stage lines should appear within seconds |
|
||
| `image build` stuck at `loading restore image metadata…` | A truncated or partial `.ipsw` — the framework blocks rather than failing | Upgrade (the file is now size- and magic-checked first); re-download the IPSW |
|
||
| `provisioning failed: … Failed to lock auxiliary storage` after install | The installer's VM had not yet released `nvram.bin` when first boot started | Upgrade — the install now drains its queue and first boot retries for 30 s. Re-run `image build`; it resumes |
|
||
| `image build` says `image 'default' already exists` after a failed first boot | Old build: an installed-but-unprovisioned bundle was treated as a finished image | Upgrade — `image build` now resumes it instead of refusing |
|
||
|
||
---
|
||
|
||
## VM won't start — entitlement error
|
||
|
||
**Symptom.** Any VM operation fails immediately with an error naming
|
||
`com.apple.security.virtualization`, or a generic "operation not permitted" from
|
||
Virtualization.framework.
|
||
|
||
**Cause.** Virtualization.framework checks the entitlement on the calling binary, and entitlements
|
||
are only honoured on a signed binary inside a proper `.app` bundle. The raw product of
|
||
`swift build` has neither.
|
||
|
||
**Fix.**
|
||
|
||
```sh
|
||
make install # build → bundle → sign → install (make sign alone re-signs in place)
|
||
codesign -d --entitlements - ~/Applications/GiteaMacosRunner.app # verify
|
||
```
|
||
|
||
The output must list `com.apple.security.virtualization`. Then confirm the command you're running
|
||
resolves to the installed bundle's binary — `which -a gitea-macos-runner` should point at
|
||
`~/Applications/GiteaMacosRunner.app/Contents/MacOS/gitea-macos-runner`, not at
|
||
`.build/release/gitea-macos-runner`. Ad-hoc signing is sufficient for *this* error; you do not need
|
||
a developer account to start VMs. (A Developer ID certificate solves a different problem — Local
|
||
Network grants evaporating on rebuild. See [setup.md](setup.md#code-signing).)
|
||
|
||
---
|
||
|
||
## `virtualMachineLimitExceeded`
|
||
|
||
**Symptom.** The first VM boots fine; a second or third fails with `virtualMachineLimitExceeded`.
|
||
|
||
**Cause.** macOS permits **two** concurrent macOS guests per host. This is an Apple kernel and
|
||
licensing limit, not a resource constraint — more RAM will not raise it.
|
||
|
||
**Fix.** Set `scheduler.maxConcurrentVMs` to 2 or less. If you're already at 2 and still hitting
|
||
the limit, a VM from a previous run is still alive — check for stray processes and for leftover
|
||
directories under `storeDir/vms`, then restart the daemon so it starts from a clean state.
|
||
|
||
To handle more macOS jobs in parallel, add another Mac.
|
||
|
||
---
|
||
|
||
## VM starts but never gets an IP
|
||
|
||
**Symptom.** The VM boots (you can see it progress if you use `vm boot`), but the daemon reports
|
||
that it could not resolve the guest address, or gives up at `scheduler.bootTimeoutSeconds`.
|
||
|
||
**Cause.** The daemon resolves the guest's NAT address from the host's DHCP lease file, which is
|
||
only written once the guest requests a lease — several seconds after the boot screen appears. If the
|
||
address never appears at all, the usual culprit on macOS 15+ is the **Local Network privacy
|
||
prompt**: a LaunchAgent that was never granted permission (or was denied) cannot talk to the guest.
|
||
|
||
**Fix.**
|
||
|
||
1. Check the lease file after the guest has been up for ~30 seconds:
|
||
|
||
```sh
|
||
cat /var/db/dhcpd_leases
|
||
```
|
||
|
||
An entry with a recent `lease` timestamp and the guest's MAC means networking is fine and the
|
||
problem is timing — raise `scheduler.bootTimeoutSeconds`.
|
||
|
||
2. No entry at all: pre-authorize the VM subnet, then **reboot** (these are read at boot):
|
||
|
||
```sh
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"
|
||
```
|
||
|
||
The `/18` is deliberate: the NAT subnet is chosen at runtime and slides to the next free /24
|
||
(192.168.65.x, .66.x, …) when one is taken, so a pinned `192.168.64.0/24` breaks the day it
|
||
moves. `doctor` reports `local network access` as a pass once it sees an allowlist covering that
|
||
span. This is the deterministic fix for an unattended host —
|
||
see [Local Network: the app is not listed in System Settings](#local-network-the-app-is-not-listed-in-system-settings)
|
||
for why the interactive grant is not.
|
||
|
||
Do not reach for the interactive grant instead. Booting a VM by hand from a Terminal makes the
|
||
prompt (and the **System Settings → Privacy & Security → Local Network** row) belong to
|
||
*Terminal* rather than to this app, and approving it there does not carry over to the
|
||
LaunchAgent.
|
||
|
||
---
|
||
|
||
## Local Network: the app is not listed in System Settings
|
||
|
||
**Symptom.** `doctor` prints the `local network access` note telling you to approve the app under
|
||
**System Settings → Privacy & Security → Local Network**, but the runner is nowhere in that list —
|
||
so there is nothing to switch on.
|
||
|
||
**Cause.** This is expected, not a bug. The Local Network list is populated lazily: an app appears
|
||
there only *after* it has actually attempted a local-network connection and been evaluated. It
|
||
cannot be pre-approved. A freshly installed runner that has not yet reached a guest has no entry.
|
||
|
||
There is also nothing to query, so `doctor` cannot tell you the grant's state. Unlike most macOS
|
||
privacy controls, Local Network privacy is not stored in TCC — per
|
||
[TN3179](https://developer.apple.com/documentation/technotes/tn3179-understanding-local-network-privacy)
|
||
the checks live "deep in the networking stack" as a Network Extension packet filter, so the
|
||
permission is absent from `TCC.db` and `tccutil reset` does not apply to it.
|
||
|
||
**Why you will probably never see the app's own row.** macOS assigns a privacy decision to the
|
||
*responsible process*, not necessarily to the binary that opened the socket. Launching the runner
|
||
the obvious way —
|
||
|
||
```sh
|
||
gitea-macos-runner vm boot --image default
|
||
```
|
||
|
||
— execs the bundle's binary from a shell, so the system holds **Terminal** responsible. Both the
|
||
prompt and the Settings row belong to Terminal, and allowing it there does **not** carry over to the
|
||
LaunchAgent, which is the process that actually needs it. Meanwhile the LaunchAgent itself has no UI
|
||
to show a prompt in, so from it the connection is denied outright and surfaces as `No route to host`
|
||
(errno 65) rather than as a permission error. Between the two, there is no reliable way to grant
|
||
this interactively on an unattended host.
|
||
|
||
**Fix — deterministic, and what to use on any host running the LaunchAgent.** Allowlist the VM
|
||
subnet instead. It is keyed on the network rather than on the app, so no prompt is involved, and
|
||
nothing needs redoing:
|
||
|
||
```sh
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"
|
||
```
|
||
|
||
Reboot afterwards. `doctor` then reports `local network access` as a **pass**. Both keys are
|
||
Apple's, documented in TN3179; the same pair is what [Tart's FAQ](https://tart.run/faq/) recommends
|
||
for this exact problem on CI hosts.
|
||
|
||
> **And an ad-hoc signed bundle cannot hold the grant anyway.** TN3179 notes that "local network
|
||
> privacy uses your main executable UUID as part of its implementation". An ad-hoc signature has no
|
||
> team anchor, so that UUID *is* the app's identity — and the linker writes a new `LC_UUID` on
|
||
> essentially every rebuild, so a rebuilt-and-reinstalled runner reads as a *different* program and
|
||
> prompts again, while the old row lingers (macOS provides no way to reset a Local Network decision
|
||
> to undetermined). Expect duplicate entries after a few upgrades.
|
||
>
|
||
> Signing with a **Developer ID Application** certificate fixes the churn: its designated
|
||
> requirement is anchored to your team, so every build is recognised as the same program. `make sign`
|
||
> uses one automatically when it is in the keychain — check with `codesign -dvv` or `doctor`'s
|
||
> `code identity` line, and see [setup.md](setup.md#code-signing). The subnet allowlist avoids the
|
||
> mechanism entirely and is still the right answer for an unattended host.
|
||
|
||
---
|
||
|
||
## SSH times out on a freshly built image
|
||
|
||
**Symptom.** The VM boots and gets an IP, but SSH never connects. Attaching a display to the guest
|
||
shows **Setup Assistant** — the region/Apple ID welcome flow — rather than a login window.
|
||
|
||
**Cause.** The image builder uses `VZMacGuestProvisioningOptions` to create the admin account, enable
|
||
Remote Login, and skip Setup Assistant. That API requires **macOS 27 or newer in the guest as well
|
||
as the host**. An older guest **silently ignores** the options: no error, no account, no SSH server
|
||
— it just sits at first-run setup forever.
|
||
|
||
**Fix.** Rebuild the base image from a macOS 27+ IPSW:
|
||
|
||
```sh
|
||
gitea-macos-runner image delete default
|
||
gitea-macos-runner image build --ipsw ~/Downloads/UniversalMac_27.0_XXXXX_Restore.ipsw
|
||
```
|
||
|
||
Verify the host is also 27+ (`sw_vers`). There is no way to make a pre-27 guest work unattended
|
||
with this builder.
|
||
|
||
---
|
||
|
||
## The daemon boots VMs forever and every teardown says `reason=cancelled`
|
||
|
||
**Symptom.** A job is queued, the daemon is running, and the log repeats the same three lines with a
|
||
new runner name each time — but no runner ever appears in Gitea:
|
||
|
||
```
|
||
info orchestrator: job=1 runner=macos-vm-1cd8e83f… slot=0 booting VM
|
||
info orchestrator: ip=192.168.65.233 slot=0 guest leased address
|
||
info orchestrator: reason=cancelled slot=0 tearing down slot
|
||
```
|
||
|
||
Note what is missing: nothing between the lease and the teardown, and a teardown reason that names
|
||
no cause.
|
||
|
||
**Cause.** The guest takes longer to reach `sshd` than `scheduler.bootTimeoutSeconds` allows, so the
|
||
scheduler tears the slot down while it is still coming up — usually seconds before it would have
|
||
succeeded. This is not a timeout that fires once; it is a **livelock**. The replacement clone starts
|
||
from zero *and* adds load to an already contended host, so the next boot is slower still and the
|
||
loop never converges.
|
||
|
||
Several Virtualization guests on one Mac is enough to cause it: a guest that reaches SSH in 40
|
||
seconds on an idle host can take four or five minutes when it is sharing the machine, and each slot
|
||
holds 4 vCPU and 8 GB for the whole attempt. Check with `uptime` inside a guest — a load average in
|
||
the tens means the guest is starved, not broken.
|
||
|
||
**Fix.**
|
||
|
||
1. Raise `scheduler.bootTimeoutSeconds`. The default is 900; treat it as a ceiling on the guest's
|
||
*worst* case, not its typical one. Timing out too early costs far more than noticing a genuinely
|
||
wedged guest late.
|
||
|
||
2. Reduce contention. Count what is actually running:
|
||
|
||
```sh
|
||
ps -Ao pid,rss,etime,comm | grep -i -e virtual -e vmnet
|
||
```
|
||
|
||
Virtualization guests belonging to *other* tools compete for the same cores and the same two-VM
|
||
macOS limit. Shut down what you are not using, or lower `scheduler.maxConcurrentVMs`.
|
||
|
||
3. Confirm the guest itself is fine, independently of the daemon, with `doctor` — its `guest ssh`
|
||
check authenticates against whichever slot currently holds a lease:
|
||
|
||
```sh
|
||
gitea-macos-runner doctor
|
||
```
|
||
|
||
**If you are on an older build**, upgrade: the empty gap in that log was three bugs, all now fixed.
|
||
`waitForSSH` was handed the full `bootTimeoutSeconds` even though the scheduler's clock had started
|
||
before the clone — so the scheduler always fired first and `waitForSSH`'s own error was unreachable;
|
||
its per-attempt failures were only ever reported in that unreachable error; and the teardown
|
||
overwrote the planner's reason with `cancelled`. Current builds log `waiting for guest ssh` with the
|
||
attempt count and the last error while it is happening, and report
|
||
`reason="boot timeout: provisioning for 312s (limit 300s)"`.
|
||
|
||
---
|
||
|
||
## `SecKeyCreateRandomKey` / "Interaction is not allowed"
|
||
|
||
**Symptom.** The daemon starts but fails during VM setup with a Security-framework error mentioning
|
||
`SecKeyCreateRandomKey`, `errSecInteractionNotAllowed`, or "Interaction is not allowed". Often it
|
||
works when you run the daemon by hand in Terminal and fails under launchd.
|
||
|
||
**Cause.** The **`login.keychain` is locked.** macOS 15+ requires it unlocked for key operations the
|
||
VM lifecycle performs, and it is only unlocked inside a live, logged-in GUI session. A LaunchDaemon,
|
||
an SSH-only session, or a Mac sitting at the login window all fail this.
|
||
|
||
**Fix.**
|
||
|
||
1. Confirm the service is installed as a **LaunchAgent**, not a LaunchDaemon:
|
||
`gitea-macos-runner service install` does the right thing; a hand-written plist in
|
||
`/Library/LaunchDaemons` does not.
|
||
2. Ensure the runner user is actually logged in with the desktop loaded. Enable auto-login:
|
||
**System Settings → Users & Groups → Automatically log in as**.
|
||
3. Prevent the machine from returning to a locked state:
|
||
|
||
```sh
|
||
sudo pmset -a sleep 0 disablesleep 1
|
||
```
|
||
|
||
and disable "Require password after screen saver begins" for the runner user.
|
||
|
||
Connecting over Screen Sharing to a Mac at the login window does not unlock `login.keychain` for
|
||
launchd's session — auto-login is the reliable answer on a dedicated CI Mac.
|
||
|
||
---
|
||
|
||
## Job stays queued and no VM boots
|
||
|
||
**Symptom.** The workflow shows as queued in Gitea indefinitely. Nothing appears in the daemon logs
|
||
about it.
|
||
|
||
**Causes and fixes, in the order worth checking:**
|
||
|
||
1. **Label mismatch.** `runs-on` must use **bare label names** (`macos-arm64`), and every label
|
||
listed must appear in the host config's `runner.labels`. The `:host` suffix used at registration
|
||
is runner-side only and must never appear in workflow YAML. A single typo produces exactly this
|
||
symptom with no error anywhere.
|
||
|
||
2. **Daemon not running or not polling.**
|
||
|
||
```sh
|
||
gitea-macos-runner service status
|
||
log show --predicate 'process == "gitea-macos-runner"' --info --last 15m
|
||
```
|
||
|
||
You should see a poll every `scheduler.pollIntervalSeconds`.
|
||
|
||
3. **Gitea too old.** The queued-jobs API with the `labels` field requires **Gitea ≥ 1.25**. On an
|
||
older instance the daemon can never see jobs. `doctor` reports the server version.
|
||
|
||
4. **Admin PAT wrong or under-scoped.** The token must belong to a **site admin** and carry
|
||
`read:admin` + `write:admin`. Test it:
|
||
|
||
```sh
|
||
curl -H "Authorization: token $TOKEN" \
|
||
"https://gitea.example.com/api/v1/admin/actions/jobs?status=queued"
|
||
```
|
||
|
||
A 403 means scope or admin status; a 404 usually means the Gitea version predates the endpoint.
|
||
|
||
5. **A VM booted but its runner never came online.** Then the job is queued *and* you see VM
|
||
activity in the logs. Look at registration failures — most often an invalid or invalidated
|
||
registration token (see below).
|
||
|
||
6. **Job expired.** Gitea abandons a job after `ABANDONED_JOB_TIMEOUT` (default 24h). If the daemon
|
||
was down longer than that, the job is gone; re-run it.
|
||
|
||
### Registration fails with an invalid token
|
||
|
||
Registration tokens are reusable, but **creating a new token for a scope invalidates the previous
|
||
one**. If someone clicked "create new registration token" in the Gitea UI, the token in your
|
||
`registrationTokenFile` is now dead. Re-seed
|
||
`GITEA_RUNNER_REGISTRATION_TOKEN` on the server (and restart Gitea), or switch to
|
||
`fetchRegistrationTokenViaAPI: true`. See
|
||
[setup.md §1.3](setup.md#13-choose-a-registration-token-strategy).
|
||
|
||
---
|
||
|
||
## `actions/checkout` fails instantly
|
||
|
||
**Symptom.** The job starts, the runner connects, and the very first step fails immediately —
|
||
typically a spawn error naming `node`, or an unhelpful non-zero exit before any output.
|
||
|
||
**Cause.** Gitea Actions' JavaScript actions (`actions/checkout` and most of the ecosystem) run by
|
||
spawning `node` inside the guest. **Node.js is required in the image**, and if provisioning was
|
||
interrupted it may be absent.
|
||
|
||
**Fix.** Confirm, then reprovision:
|
||
|
||
```sh
|
||
gitea-macos-runner vm boot
|
||
ssh admin@<guest-ip> 'node --version && git --version'
|
||
|
||
gitea-macos-runner image provision default
|
||
```
|
||
|
||
If `git` is also missing, the provisioning step failed early — check the build log and rerun
|
||
provisioning.
|
||
|
||
---
|
||
|
||
## Runner rows piling up in the Gitea UI
|
||
|
||
**Symptom.** **Site Administration → Actions → Runners** accumulates offline `macos-vm-…` entries.
|
||
|
||
**Cause.** Gitea deletes an ephemeral registration when its job completes normally. A VM that is
|
||
killed uncleanly — daemon crash, host power loss, `jobTimeoutMinutes` kill — never reaches that
|
||
point, so the row is orphaned.
|
||
|
||
**Fix.** Usually nothing: the daemon's reconcile loop sweeps orphaned registrations every
|
||
`scheduler.reconcileIntervalSeconds` (default 300) via
|
||
`DELETE /api/v1/admin/actions/runners/{id}`, plus a daily midnight sweep. Orphans should clear
|
||
within a few minutes.
|
||
|
||
If they persist, the daemon's admin PAT probably lacks `write:admin` — check the logs for delete
|
||
failures. To clear them by hand, delete the rows in the Gitea UI; they are inert (offline
|
||
registrations that have already been spent cannot receive jobs).
|
||
|
||
---
|
||
|
||
## `doctor` warns "runner download url → 403"
|
||
|
||
**Symptom.** `doctor` reports the runner download URL as a 403, but pasting the same URL into a
|
||
browser downloads the binary fine:
|
||
|
||
```
|
||
! runner download url https://gitea.com/gitea/runner/releases/download/v3.0.2/… → 403
|
||
```
|
||
|
||
**Cause.** gitea.com does not serve release assets itself. It answers with a `303 See Other`
|
||
pointing at a presigned object-storage URL, and that signature covers the HTTP method of the
|
||
request that minted it. A `HEAD` probe therefore gets a HEAD-signed URL — and then the HTTP client,
|
||
following the 303, rewrites the method to `GET` (RFC 7231 §6.4.4) and replays the signed URL with
|
||
the one verb it was not signed for. The store answers `403 SignatureDoesNotMatch`. Nothing is
|
||
actually wrong with the asset.
|
||
|
||
**Fix.** Upgrade — the check now probes with a one-byte ranged `GET` (`Range: bytes=0-0`), the same
|
||
verb `image provision` uses for the real download, and falls back to a `HEAD` only if the range is
|
||
rejected. If you still see a 403 after upgrading, the asset really is gone: check `runner.version`
|
||
and `runner.runnerDownloadURL` for a `darwin-arm64` build.
|
||
|
||
---
|
||
|
||
## Disk filling up
|
||
|
||
**Symptom.** Free space falls steadily; the daemon starts refusing to launch VMs, citing
|
||
`storage.minFreeDiskGB`.
|
||
|
||
**Cause.** Each VM is an APFS copy-on-write clone of the base image. The clone is free at creation
|
||
but **grows as the job writes** — dependency caches, build outputs, Xcode's derived data. Clones
|
||
from uncleanly-killed VMs are not reclaimed automatically.
|
||
|
||
**Fix.**
|
||
|
||
```sh
|
||
# See what's there.
|
||
du -sh ~/Library/Application\ Support/gitea-macos-runner/*
|
||
ls -la ~/Library/Application\ Support/gitea-macos-runner/vms
|
||
|
||
# With the daemon stopped, remove stale clones.
|
||
gitea-macos-runner service uninstall # or stop the daemon
|
||
rm -rf ~/Library/Application\ Support/gitea-macos-runner/vms/<stale-clone>
|
||
gitea-macos-runner service install
|
||
```
|
||
|
||
Only delete entries under `vms/` — `images/` holds the base images you'd otherwise have to rebuild.
|
||
Longer term: raise `storage.minFreeDiskGB` so the guard trips earlier, delete unused base images
|
||
with `image delete`, and remember an Xcode image needs 140 GB+ of headroom, more with two
|
||
concurrent clones diverging.
|
||
|
||
---
|
||
|
||
## `image build` prints the banner and then nothing at all
|
||
|
||
**Symptom.** `image build` prints
|
||
|
||
```
|
||
building image 'default' (this takes a while; the IPSW alone is ~15 GB)
|
||
```
|
||
|
||
and then stops — no stage lines, no progress bar, no error. Activity Monitor shows the process
|
||
using no CPU and no VM services running. Ctrl-C does nothing.
|
||
|
||
**Cause.** A bug in releases before this fix. `image build` hosts an `NSApplication` run loop
|
||
(Virtualization.framework requires one), and the work was started with a `Task` that inherited the
|
||
main actor. Because `NSApplication.run()` is itself reached from Swift's async `main`, the main
|
||
dispatch queue already had a block in flight and would not re-enter — so the build task was queued
|
||
behind a run loop that never yields and never got a first tick. Nothing ran, including the code
|
||
that would have reported the error. `SIGINT`/`SIGTERM` handling was stuck the same way, which is
|
||
why Ctrl-C did not work either.
|
||
|
||
**Fix.** Upgrade. The build task is now detached and signals are handled off the main queue. You
|
||
should see stage lines within a second or two:
|
||
|
||
```
|
||
resolving restore image…
|
||
loading restore image metadata…
|
||
creating VM bundle (disk 64 GB)…
|
||
installing macOS [########----------------------] 27%
|
||
```
|
||
|
||
If a build still goes quiet, the line last printed tells you which stage owns the silence — see
|
||
below.
|
||
|
||
---
|
||
|
||
## `image build` sits at "loading restore image metadata…"
|
||
|
||
**Symptom.** The build reaches `loading restore image metadata…` and stays there. After a minute it
|
||
adds:
|
||
|
||
```
|
||
still loading — a truncated or partially downloaded .ipsw can block here; verify the download completed
|
||
```
|
||
|
||
**Cause.** `VZMacOSRestoreImage.image(from:)` reads the whole archive's metadata and reports no
|
||
progress while it does. On a healthy ~21 GB IPSW this takes seconds to a minute or two. On a
|
||
*partial* download it can block for a very long time instead of failing.
|
||
|
||
**Fix.** The obvious checks are now made before the framework is handed the file — it must exist, be
|
||
a regular file, be at least 1 GB, and start with the zip magic `PK` — so an incomplete download now
|
||
fails immediately with its actual size rather than hanging. If you are on an older build, check by
|
||
hand:
|
||
|
||
```sh
|
||
ls -la ~/Downloads/UniversalMac_*.ipsw # ~15-22 GB, no .download/.crdownload sibling
|
||
xxd -l 2 -p ~/Downloads/UniversalMac_*.ipsw # must print 504b
|
||
```
|
||
|
||
Anything smaller, or not starting `504b`, is an incomplete or wrong file: delete it and download it
|
||
again.
|
||
|
||
> A pattern that matches **more than one** IPSW is refused outright, listing the matches — for
|
||
> example `~/Downloads/UniversalMac_27.0_*.ipsw` when two 27.0 builds are sitting in `~/Downloads`.
|
||
> Name exactly one of them.
|
||
|
||
---
|
||
|
||
## `image build` hangs at install
|
||
|
||
**Symptom.** `image build` sits for a long time at the macOS install phase with little visible
|
||
progress.
|
||
|
||
**Cause.** Usually none — installing macOS from an IPSW genuinely takes a long time (tens of
|
||
minutes, longer on slower storage). The install phase is largely silent.
|
||
|
||
**Fix.** **Wait, and do not stop the VM mid-install.** Interrupting the installer leaves the disk
|
||
image in an undefined state; the resulting image may boot and then fail in confusing ways later.
|
||
There is no resume from a *partial* install — only from a complete one (see the next section).
|
||
|
||
If you did interrupt it, or the build genuinely failed:
|
||
|
||
```sh
|
||
gitea-macos-runner image delete <name>
|
||
gitea-macos-runner image build --ipsw <path> --name <name>
|
||
```
|
||
|
||
Before rebuilding, verify the IPSW is complete and matches your host architecture (Apple Silicon)
|
||
and version requirement (macOS 27+ for unattended provisioning), and that you have enough free disk
|
||
for the IPSW plus the target disk size.
|
||
|
||
---
|
||
|
||
## `Failed to lock auxiliary storage` right after the install finishes
|
||
|
||
**Symptom.** The macOS install runs to 100%, then:
|
||
|
||
```
|
||
first boot + guest provisioning…
|
||
error: provisioning failed: could not boot image 'default': provisioning failed: Invalid
|
||
virtual machine configuration. Failed to lock auxiliary storage.
|
||
```
|
||
|
||
**Cause.** A `VZVirtualMachine` holds an exclusive lock on its bundle's `nvram.bin` for its entire
|
||
lifetime and releases it in `dealloc`. `VZMacOSInstaller` owns a VM of its own, and older builds
|
||
resumed the caller from inside the installer's completion handler — before the framework's frame
|
||
had unwound and dropped the last reference. First boot then constructed a *second* VM over the same
|
||
bundle and lost the race.
|
||
|
||
**Fix.** Upgrade. Two changes address it:
|
||
|
||
- The install now tears down its VM on its own serial queue and resumes the caller only from a
|
||
later block on that queue, so the installer's VM is deallocated before `install()` returns.
|
||
- First boot retries specifically on this failure for up to 30 s (2 s apart), printing
|
||
`waiting for installer to release the VM bundle…`. Any other configuration error still fails
|
||
immediately — an invalid configuration does not become valid by waiting.
|
||
|
||
Nothing is lost when it does happen: the install is complete, so re-running `image build` resumes.
|
||
|
||
---
|
||
|
||
## `image build` resumes an installed-but-unprovisioned image
|
||
|
||
**Symptom.** A previous `image build` finished installing macOS and then failed at first boot or
|
||
provisioning. Re-running it prints:
|
||
|
||
```
|
||
image default already installed — resuming first boot + provisioning
|
||
```
|
||
|
||
**This is intended.** An `image build` that fails after the install has left an hour of work on
|
||
disk, and throwing it away to redo an identical install is not a reasonable default. When the
|
||
bundle exists, is complete, and is not yet marked provisioned, the install phase is skipped and the
|
||
build goes straight to first boot.
|
||
|
||
It is safe because provisioning is exactly what did *not* happen: the guest has never been booted,
|
||
so its first boot is still ahead of it and `VZMacGuestProvisioningOptions` — which macOS evaluates
|
||
only on the first boot after a restore — still applies.
|
||
|
||
**The one ambiguous case.** The bundle on disk cannot say whether an earlier run got far enough to
|
||
*boot* the guest. If it did, that single chance at automated Setup Assistant is spent. Rather than
|
||
guess, the resume boots and waits for SSH: if the guest answers, provisioning was never applied and
|
||
the build continues normally. Only if SSH times out does it stop, and it then tells you so
|
||
explicitly — that state is unrecoverable, and the fix is `image delete` followed by a fresh build.
|
||
|
||
An image that is already `provisioned` is untouched; `image build` still refuses with
|
||
`image '<name>' already exists`. To re-run provisioning on a finished image, use
|
||
`gitea-macos-runner image provision <name>` instead.
|
||
|
||
---
|
||
|
||
## SSH fails with "No route to host" (errno 65) mid-run
|
||
|
||
**Symptom.** A command that had *just* talked to the guest successfully suddenly cannot reach it —
|
||
most visibly `image provision`, which waits for SSH, reports the first provisioning step, and then
|
||
dies on the upload:
|
||
|
||
```
|
||
first boot + guest provisioning…
|
||
provisioning: system configuration (provision.sh)
|
||
error: provisioning failed: could not upload provision.sh from …/provision.sh:
|
||
ssh failed: cannot connect to 192.168.65.232:22: … No route to host) (errno: 65)
|
||
```
|
||
|
||
**First, check whether it is actually failing.** A single errno 65 at `attempt=1 elapsed=0s`,
|
||
seconds after `guest leased address`, is normal and not worth chasing. `waitForSSH` logs its first
|
||
probe unconditionally and then heartbeats every 30 s, and that first probe usually lands before the
|
||
host has an ARP entry for the guest — which is also `EHOSTUNREACH`. What matters is whether the
|
||
message *repeats* at `elapsed=30s`, `60s`, `90s`, … If it does not, boot is proceeding normally. If
|
||
it does, read on.
|
||
|
||
**Cause.** On macOS 15 and newer, an app that has not been granted Local Network access does not get
|
||
a "permission denied": the packet filter answers **`EHOSTUNREACH` — errno 65, "No route to host"**,
|
||
which is indistinguishable from a guest that is genuinely off the network. Guests here live on a
|
||
host-private NAT link that is reachable whenever the VM is up, so on this path a *persistent* errno
|
||
65 is far more often the privacy filter than a routing problem.
|
||
|
||
Two details make it look intermittent rather than like a permission problem:
|
||
|
||
- **Under an ad-hoc signature the grant is keyed on the executable's UUID.** Per
|
||
[TN3179](https://developer.apple.com/documentation/technotes/tn3179-understanding-local-network-privacy),
|
||
Local Network privacy "uses your main executable UUID as part of its implementation", and the
|
||
linker mints a fresh `LC_UUID` on essentially every build. A `make install` after a code change
|
||
therefore presents a program macOS has never seen, whose permission is undetermined again — even
|
||
though the binary you ran ten minutes ago worked. A Developer ID signature anchors identity to the
|
||
team instead and does not drift; `doctor`'s `code identity` check tells you which you have, and
|
||
[setup.md](setup.md#code-signing) covers switching.
|
||
- **Processes started over SSH are exempt.** Running the same command through `ssh you@host …`
|
||
succeeds while running it from a GUI Terminal fails. A remote-shell success proves nothing about
|
||
the interactive path.
|
||
- **A shell-launched run is attributed to Terminal.** macOS charges the privacy decision to the
|
||
responsible process, so the app's own identity is not what is being evaluated when you launch it
|
||
by hand — and a grant given to Terminal does nothing for the LaunchAgent.
|
||
|
||
**Fix.** Allowlist the subnet — it is keyed on the network, not on the app, so no rebuild can
|
||
withdraw it and no prompt has to be answered:
|
||
|
||
```sh
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
|
||
sudo defaults write com.apple.network.local-network \
|
||
AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"
|
||
sudo reboot
|
||
```
|
||
|
||
The values are only read at boot, so **the reboot is not optional** — until it happens, `defaults
|
||
read com.apple.network.local-network` shows the new setting while the filter still behaves as
|
||
before.
|
||
|
||
Use `/18`, not `/24`. Virtualization.framework's NAT starts at `192.168.64.0/24` but chooses the
|
||
subnet at runtime and steps to the next free /24 when that one is in use, so hosts drift to
|
||
`192.168.65.x` and beyond. An allowlist naming a single /24 that the NAT has since moved off is the
|
||
worst case: it reads as configured, `doctor` used to call it a pass, and every guest connection
|
||
still fails with errno 65. `doctor` now warns instead when the allowlist does not cover
|
||
`192.168.64.0`–`192.168.127.255`.
|
||
|
||
**Verifying.** After the reboot, `gitea-macos-runner doctor` should show `local network access` as a
|
||
pass naming the range. Re-run the command that failed; nothing else needs redoing, and
|
||
`image provision` is idempotent, so a partially completed run is safe to repeat.
|