Files
gitea-macos-vm-orchestrator/docs/troubleshooting.md
T

456 lines
21 KiB
Markdown

# Troubleshooting
Start with `gitea-macos-runner doctor` — it catches most misconfiguration before you go
symptom-hunting. Then find your symptom below.
Useful log commands throughout:
```sh
# Daemon logs, live
log stream --predicate 'process == "gitea-macos-runner"' --info
# Daemon logs, last hour
log show --predicate 'process == "gitea-macos-runner"' --info --last 1h
gitea-macos-runner service status
```
---
## Quick reference
| Symptom | Cause | Fix |
| --- | --- | --- |
| VM won't start; entitlement / `com.apple.security.virtualization` error | Running an unsigned binary, or one outside the signed `.app` bundle | `make sign` (or re-run `make install`); invoke the installed bundle, never `.build/release/…` |
| `virtualMachineLimitExceeded` at boot | macOS allows at most **2** concurrent macOS VMs | Set `scheduler.maxConcurrentVMs` ≤ 2; kill stray VMs from earlier runs |
| VM boots but never gets an IP | DHCP lease not yet written, or Local Network privacy denial (macOS 15+) | Check `/var/db/dhcpd_leases`; grant Local Network permission or pre-authorize the subnet |
| Runner not listed under Privacy & Security → Local Network | Expected — the list is populated only after the app first attempts a local connection; it cannot be pre-approved | Boot one VM by hand from a GUI Terminal to create the entry, or (better on CI) allowlist the subnet with `defaults write com.apple.network.local-network` |
| SSH times out on a freshly built image | Guest macOS < 27, so provisioning options were ignored and Setup Assistant is waiting | Rebuild the image from a macOS **27+** IPSW |
| `SecKeyCreateRandomKey` / "Interaction is not allowed" | `login.keychain` is locked — no GUI session | Run as a LaunchAgent in an unlocked GUI session; enable auto-login |
| Job stays queued forever | Label mismatch, or the daemon isn't running/reaching Gitea | Use bare label names in `runs-on`; match `runner.labels`; check daemon logs |
| `actions/checkout` fails instantly | Node.js missing from the guest image | `gitea-macos-runner image provision <name>` |
| Runner rows piling up in the Gitea UI | VMs killed uncleanly; registrations orphaned | Reconcile loop cleans them; force it by restarting the daemon; delete manually if needed |
| `doctor` says `runner download url → 403` but the URL works in a browser | Old build: the check used `HEAD`, and the presigned redirect target is signed per method | Upgrade — the check now uses a ranged `GET`. If it persists, the asset really is missing |
| Disk filling up | Copy-on-write clones grow as jobs write | Raise `storage.minFreeDiskGB`; delete stale clones in `storeDir/vms` |
| `image build` appears to hang during install | Normal — macOS install is slow | Wait. **Do not stop the VM mid-install**; if you did, delete the image and rebuild |
| `image build` prints its banner and then nothing, at 0% CPU | Old build: the build task was queued behind the `NSApplication` run loop and never started | Upgrade — the task is now detached. Stage lines should appear within seconds |
| `image build` stuck at `loading restore image metadata…` | A truncated or partial `.ipsw` — the framework blocks rather than failing | Upgrade (the file is now size- and magic-checked first); re-download the IPSW |
---
## VM won't start — entitlement error
**Symptom.** Any VM operation fails immediately with an error naming
`com.apple.security.virtualization`, or a generic "operation not permitted" from
Virtualization.framework.
**Cause.** Virtualization.framework checks the entitlement on the calling binary, and entitlements
are only honoured on a signed binary inside a proper `.app` bundle. The raw product of
`swift build` has neither.
**Fix.**
```sh
make install # build → bundle → sign → install (make sign alone re-signs in place)
codesign -d --entitlements - ~/Applications/GiteaMacosRunner.app # verify
```
The output must list `com.apple.security.virtualization`. Then confirm the command you're running
resolves to the installed bundle's binary — `which -a gitea-macos-runner` should point at
`~/Applications/GiteaMacosRunner.app/Contents/MacOS/gitea-macos-runner`, not at
`.build/release/gitea-macos-runner`. Ad-hoc signing is sufficient; you do not need a paid developer
account.
---
## `virtualMachineLimitExceeded`
**Symptom.** The first VM boots fine; a second or third fails with `virtualMachineLimitExceeded`.
**Cause.** macOS permits **two** concurrent macOS guests per host. This is an Apple kernel and
licensing limit, not a resource constraint — more RAM will not raise it.
**Fix.** Set `scheduler.maxConcurrentVMs` to 2 or less. If you're already at 2 and still hitting
the limit, a VM from a previous run is still alive — check for stray processes and for leftover
directories under `storeDir/vms`, then restart the daemon so it starts from a clean state.
To handle more macOS jobs in parallel, add another Mac.
---
## VM starts but never gets an IP
**Symptom.** The VM boots (you can see it progress if you use `vm boot`), but the daemon reports
that it could not resolve the guest address, or gives up at `scheduler.bootTimeoutSeconds`.
**Cause.** The daemon resolves the guest's NAT address from the host's DHCP lease file, which is
only written once the guest requests a lease — several seconds after the boot screen appears. If the
address never appears at all, the usual culprit on macOS 15+ is the **Local Network privacy
prompt**: a LaunchAgent that was never granted permission (or was denied) cannot talk to the guest.
**Fix.**
1. Check the lease file after the guest has been up for ~30 seconds:
```sh
cat /var/db/dhcpd_leases
```
An entry with a recent `lease` timestamp and the guest's MAC means networking is fine and the
problem is timing — raise `scheduler.bootTimeoutSeconds`.
2. No entry at all: pre-authorize the VM subnet, then **reboot** (these are read at boot):
```sh
sudo defaults write com.apple.network.local-network \
AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/24"
sudo defaults write com.apple.network.local-network \
AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/24"
```
Match the range to what your host's NAT actually hands out. `doctor` reports `local network
access` as a pass once it can see this. This is the deterministic fix for an unattended host —
see [Local Network: the app is not listed in System Settings](#local-network-the-app-is-not-listed-in-system-settings)
for why the interactive grant is not.
3. Or grant it interactively: run `gitea-macos-runner vm boot --image default` from a Terminal in
the GUI session and click **Allow**. The app is not listed under **System Settings → Privacy &
Security → Local Network** until it has made that first attempt.
---
## Local Network: the app is not listed in System Settings
**Symptom.** `doctor` prints the `local network access` note telling you to approve the app under
**System Settings → Privacy & Security → Local Network**, but the runner is nowhere in that list —
so there is nothing to switch on.
**Cause.** This is expected, not a bug. The Local Network list is populated lazily: an app appears
there only *after* it has actually attempted a local-network connection and been evaluated. It
cannot be pre-approved. A freshly installed runner that has not yet reached a guest has no entry.
There is also nothing to query, so `doctor` cannot tell you the grant's state. Unlike most macOS
privacy controls, Local Network privacy is not stored in TCC — per
[TN3179](https://developer.apple.com/documentation/technotes/tn3179-understanding-local-network-privacy)
the checks live "deep in the networking stack" as a Network Extension packet filter, so the
permission is absent from `TCC.db` and `tccutil reset` does not apply to it.
**Fix — interactive.** Make the app connect once, from a GUI session where a human can answer:
```sh
gitea-macos-runner vm boot --image default
```
Click **Allow**. The entry now exists and can be toggled later. Do not wait for the LaunchAgent to
trigger it; a background agent cannot answer the prompt, so it simply fails to reach the guest.
**Fix — deterministic, and what to use on a CI box.** Allowlist the VM subnet instead. It is keyed
on the network rather than on the app, so no prompt is involved and nothing needs redoing:
```sh
sudo defaults write com.apple.network.local-network \
AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/24"
sudo defaults write com.apple.network.local-network \
AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/24"
```
Reboot afterwards. `doctor` then reports `local network access` as a **pass**. Both keys are
Apple's, documented in TN3179; the same pair is what [Tart's FAQ](https://tart.run/faq/) recommends
for this exact problem on CI hosts.
> **Why the interactive grant does not stick here.** This project ships an **ad-hoc signed** bundle
> (`codesign --sign -`), and TN3179 notes that "local network privacy uses your main executable UUID
> as part of its implementation". The linker writes a new `LC_UUID` on essentially every rebuild, so
> a rebuilt-and-reinstalled runner can read as a *different* program and prompt again — while the
> old row lingers, since macOS provides no way to reset a Local Network decision to undetermined.
> Expect duplicate entries after a few upgrades. The subnet allowlist avoids all of this.
---
## SSH times out on a freshly built image
**Symptom.** The VM boots and gets an IP, but SSH never connects. Attaching a display to the guest
shows **Setup Assistant** — the region/Apple ID welcome flow — rather than a login window.
**Cause.** The image builder uses `VZMacGuestProvisioningOptions` to create the admin account, enable
Remote Login, and skip Setup Assistant. That API requires **macOS 27 or newer in the guest as well
as the host**. An older guest **silently ignores** the options: no error, no account, no SSH server
— it just sits at first-run setup forever.
**Fix.** Rebuild the base image from a macOS 27+ IPSW:
```sh
gitea-macos-runner image delete default
gitea-macos-runner image build --ipsw ~/Downloads/UniversalMac_27.0_XXXXX_Restore.ipsw
```
Verify the host is also 27+ (`sw_vers`). There is no way to make a pre-27 guest work unattended
with this builder.
---
## `SecKeyCreateRandomKey` / "Interaction is not allowed"
**Symptom.** The daemon starts but fails during VM setup with a Security-framework error mentioning
`SecKeyCreateRandomKey`, `errSecInteractionNotAllowed`, or "Interaction is not allowed". Often it
works when you run the daemon by hand in Terminal and fails under launchd.
**Cause.** The **`login.keychain` is locked.** macOS 15+ requires it unlocked for key operations the
VM lifecycle performs, and it is only unlocked inside a live, logged-in GUI session. A LaunchDaemon,
an SSH-only session, or a Mac sitting at the login window all fail this.
**Fix.**
1. Confirm the service is installed as a **LaunchAgent**, not a LaunchDaemon:
`gitea-macos-runner service install` does the right thing; a hand-written plist in
`/Library/LaunchDaemons` does not.
2. Ensure the runner user is actually logged in with the desktop loaded. Enable auto-login:
**System Settings → Users & Groups → Automatically log in as**.
3. Prevent the machine from returning to a locked state:
```sh
sudo pmset -a sleep 0 disablesleep 1
```
and disable "Require password after screen saver begins" for the runner user.
Connecting over Screen Sharing to a Mac at the login window does not unlock `login.keychain` for
launchd's session — auto-login is the reliable answer on a dedicated CI Mac.
---
## Job stays queued and no VM boots
**Symptom.** The workflow shows as queued in Gitea indefinitely. Nothing appears in the daemon logs
about it.
**Causes and fixes, in the order worth checking:**
1. **Label mismatch.** `runs-on` must use **bare label names** (`macos-arm64`), and every label
listed must appear in the host config's `runner.labels`. The `:host` suffix used at registration
is runner-side only and must never appear in workflow YAML. A single typo produces exactly this
symptom with no error anywhere.
2. **Daemon not running or not polling.**
```sh
gitea-macos-runner service status
log show --predicate 'process == "gitea-macos-runner"' --info --last 15m
```
You should see a poll every `scheduler.pollIntervalSeconds`.
3. **Gitea too old.** The queued-jobs API with the `labels` field requires **Gitea ≥ 1.25**. On an
older instance the daemon can never see jobs. `doctor` reports the server version.
4. **Admin PAT wrong or under-scoped.** The token must belong to a **site admin** and carry
`read:admin` + `write:admin`. Test it:
```sh
curl -H "Authorization: token $TOKEN" \
"https://gitea.example.com/api/v1/admin/actions/jobs?status=queued"
```
A 403 means scope or admin status; a 404 usually means the Gitea version predates the endpoint.
5. **A VM booted but its runner never came online.** Then the job is queued *and* you see VM
activity in the logs. Look at registration failures — most often an invalid or invalidated
registration token (see below).
6. **Job expired.** Gitea abandons a job after `ABANDONED_JOB_TIMEOUT` (default 24h). If the daemon
was down longer than that, the job is gone; re-run it.
### Registration fails with an invalid token
Registration tokens are reusable, but **creating a new token for a scope invalidates the previous
one**. If someone clicked "create new registration token" in the Gitea UI, the token in your
`registrationTokenFile` is now dead. Re-seed
`GITEA_RUNNER_REGISTRATION_TOKEN` on the server (and restart Gitea), or switch to
`fetchRegistrationTokenViaAPI: true`. See
[setup.md §1.3](setup.md#13-choose-a-registration-token-strategy).
---
## `actions/checkout` fails instantly
**Symptom.** The job starts, the runner connects, and the very first step fails immediately —
typically a spawn error naming `node`, or an unhelpful non-zero exit before any output.
**Cause.** Gitea Actions' JavaScript actions (`actions/checkout` and most of the ecosystem) run by
spawning `node` inside the guest. **Node.js is required in the image**, and if provisioning was
interrupted it may be absent.
**Fix.** Confirm, then reprovision:
```sh
gitea-macos-runner vm boot
ssh admin@<guest-ip> 'node --version && git --version'
gitea-macos-runner image provision default
```
If `git` is also missing, the provisioning step failed early — check the build log and rerun
provisioning.
---
## Runner rows piling up in the Gitea UI
**Symptom.** **Site Administration → Actions → Runners** accumulates offline `macos-vm-…` entries.
**Cause.** Gitea deletes an ephemeral registration when its job completes normally. A VM that is
killed uncleanly — daemon crash, host power loss, `jobTimeoutMinutes` kill — never reaches that
point, so the row is orphaned.
**Fix.** Usually nothing: the daemon's reconcile loop sweeps orphaned registrations every
`scheduler.reconcileIntervalSeconds` (default 300) via
`DELETE /api/v1/admin/actions/runners/{id}`, plus a daily midnight sweep. Orphans should clear
within a few minutes.
If they persist, the daemon's admin PAT probably lacks `write:admin` — check the logs for delete
failures. To clear them by hand, delete the rows in the Gitea UI; they are inert (offline
registrations that have already been spent cannot receive jobs).
---
## `doctor` warns "runner download url → 403"
**Symptom.** `doctor` reports the runner download URL as a 403, but pasting the same URL into a
browser downloads the binary fine:
```
! runner download url https://gitea.com/gitea/runner/releases/download/v3.0.2/… → 403
```
**Cause.** gitea.com does not serve release assets itself. It answers with a `303 See Other`
pointing at a presigned object-storage URL, and that signature covers the HTTP method of the
request that minted it. A `HEAD` probe therefore gets a HEAD-signed URL — and then the HTTP client,
following the 303, rewrites the method to `GET` (RFC 7231 §6.4.4) and replays the signed URL with
the one verb it was not signed for. The store answers `403 SignatureDoesNotMatch`. Nothing is
actually wrong with the asset.
**Fix.** Upgrade — the check now probes with a one-byte ranged `GET` (`Range: bytes=0-0`), the same
verb `image provision` uses for the real download, and falls back to a `HEAD` only if the range is
rejected. If you still see a 403 after upgrading, the asset really is gone: check `runner.version`
and `runner.runnerDownloadURL` for a `darwin-arm64` build.
---
## Disk filling up
**Symptom.** Free space falls steadily; the daemon starts refusing to launch VMs, citing
`storage.minFreeDiskGB`.
**Cause.** Each VM is an APFS copy-on-write clone of the base image. The clone is free at creation
but **grows as the job writes** — dependency caches, build outputs, Xcode's derived data. Clones
from uncleanly-killed VMs are not reclaimed automatically.
**Fix.**
```sh
# See what's there.
du -sh ~/Library/Application\ Support/gitea-macos-runner/*
ls -la ~/Library/Application\ Support/gitea-macos-runner/vms
# With the daemon stopped, remove stale clones.
gitea-macos-runner service uninstall # or stop the daemon
rm -rf ~/Library/Application\ Support/gitea-macos-runner/vms/<stale-clone>
gitea-macos-runner service install
```
Only delete entries under `vms/` — `images/` holds the base images you'd otherwise have to rebuild.
Longer term: raise `storage.minFreeDiskGB` so the guard trips earlier, delete unused base images
with `image delete`, and remember an Xcode image needs 140 GB+ of headroom, more with two
concurrent clones diverging.
---
## `image build` prints the banner and then nothing at all
**Symptom.** `image build` prints
```
building image 'default' (this takes a while; the IPSW alone is ~15 GB)
```
and then stops — no stage lines, no progress bar, no error. Activity Monitor shows the process
using no CPU and no VM services running. Ctrl-C does nothing.
**Cause.** A bug in releases before this fix. `image build` hosts an `NSApplication` run loop
(Virtualization.framework requires one), and the work was started with a `Task` that inherited the
main actor. Because `NSApplication.run()` is itself reached from Swift's async `main`, the main
dispatch queue already had a block in flight and would not re-enter — so the build task was queued
behind a run loop that never yields and never got a first tick. Nothing ran, including the code
that would have reported the error. `SIGINT`/`SIGTERM` handling was stuck the same way, which is
why Ctrl-C did not work either.
**Fix.** Upgrade. The build task is now detached and signals are handled off the main queue. You
should see stage lines within a second or two:
```
resolving restore image…
loading restore image metadata…
creating VM bundle (disk 64 GB)…
installing macOS [########----------------------] 27%
```
If a build still goes quiet, the line last printed tells you which stage owns the silence — see
below.
---
## `image build` sits at "loading restore image metadata…"
**Symptom.** The build reaches `loading restore image metadata…` and stays there. After a minute it
adds:
```
still loading — a truncated or partially downloaded .ipsw can block here; verify the download completed
```
**Cause.** `VZMacOSRestoreImage.image(from:)` reads the whole archive's metadata and reports no
progress while it does. On a healthy ~21 GB IPSW this takes seconds to a minute or two. On a
*partial* download it can block for a very long time instead of failing.
**Fix.** The obvious checks are now made before the framework is handed the file — it must exist, be
a regular file, be at least 1 GB, and start with the zip magic `PK` — so an incomplete download now
fails immediately with its actual size rather than hanging. If you are on an older build, check by
hand:
```sh
ls -la ~/Downloads/UniversalMac_*.ipsw # ~15-22 GB, no .download/.crdownload sibling
xxd -l 2 -p ~/Downloads/UniversalMac_*.ipsw # must print 504b
```
Anything smaller, or not starting `504b`, is an incomplete or wrong file: delete it and download it
again.
> A pattern that matches **more than one** IPSW is refused outright, listing the matches — for
> example `~/Downloads/UniversalMac_27.0_*.ipsw` when two 27.0 builds are sitting in `~/Downloads`.
> Name exactly one of them.
---
## `image build` hangs at install
**Symptom.** `image build` sits for a long time at the macOS install phase with little visible
progress.
**Cause.** Usually none — installing macOS from an IPSW genuinely takes a long time (tens of
minutes, longer on slower storage). The install phase is largely silent.
**Fix.** **Wait, and do not stop the VM mid-install.** Interrupting the installer leaves the disk
image in an undefined state; the resulting image may boot and then fail in confusing ways later.
There is no resume.
If you did interrupt it, or the build genuinely failed:
```sh
gitea-macos-runner image delete <name>
gitea-macos-runner image build --ipsw <path> --name <name>
```
Before rebuilding, verify the IPSW is complete and matches your host architecture (Apple Silicon)
and version requirement (macOS 27+ for unattended provisioning), and that you have enough free disk
for the IPSW plus the target disk size.