Files
gitea-macos-vm-orchestrator/docs/troubleshooting.md
T

28 KiB
Raw Blame History

Troubleshooting

Start with gitea-macos-runner doctor — it catches most misconfiguration before you go symptom-hunting. Then find your symptom below.

Useful log commands throughout:

# Daemon logs, live
log stream --predicate 'process == "gitea-macos-runner"' --info

# Daemon logs, last hour
log show --predicate 'process == "gitea-macos-runner"' --info --last 1h

gitea-macos-runner service status

Quick reference

Symptom Cause Fix
VM won't start; entitlement / com.apple.security.virtualization error Running an unsigned binary, or one outside the signed .app bundle make sign (or re-run make install); invoke the installed bundle, never .build/release/…
virtualMachineLimitExceeded at boot macOS allows at most 2 concurrent macOS VMs Set scheduler.maxConcurrentVMs ≤ 2; kill stray VMs from earlier runs
VM boots but never gets an IP DHCP lease not yet written, or Local Network privacy denial (macOS 15+) Check /var/db/dhcpd_leases; grant Local Network permission or pre-authorize the subnet
Runner not listed under Privacy & Security → Local Network Expected — the list is populated only after the app first attempts a local connection; it cannot be pre-approved Boot one VM by hand from a GUI Terminal to create the entry, or (better on CI) allowlist the subnet with defaults write com.apple.network.local-network
SSH times out on a freshly built image Guest macOS < 27, so provisioning options were ignored and Setup Assistant is waiting Rebuild the image from a macOS 27+ IPSW
ssh failed: cannot connect … No route to host) (errno: 65) part-way through provisioning macOS 15+ Local Network privacy blocking the app — the grant is keyed on the executable's UUID, so make install withdraws it Allowlist the subnet (192.168.64.0/18) and reboot; see SSH fails with "No route to host" mid-run
Allowlist is set but guests are still unreachable It names 192.168.64.0/24 while the NAT has moved to 192.168.65.x Widen it to 192.168.64.0/18 and reboot; doctor now warns about too-narrow allowlists
SecKeyCreateRandomKey / "Interaction is not allowed" login.keychain is locked — no GUI session Run as a LaunchAgent in an unlocked GUI session; enable auto-login
Job stays queued forever Label mismatch, or the daemon isn't running/reaching Gitea Use bare label names in runs-on; match runner.labels; check daemon logs
actions/checkout fails instantly Node.js missing from the guest image gitea-macos-runner image provision <name>
Runner rows piling up in the Gitea UI VMs killed uncleanly; registrations orphaned Reconcile loop cleans them; force it by restarting the daemon; delete manually if needed
doctor says runner download url → 403 but the URL works in a browser Old build: the check used HEAD, and the presigned redirect target is signed per method Upgrade — the check now uses a ranged GET. If it persists, the asset really is missing
Disk filling up Copy-on-write clones grow as jobs write Raise storage.minFreeDiskGB; delete stale clones in storeDir/vms
image build appears to hang during install Normal — macOS install is slow Wait. Do not stop the VM mid-install; if you did, delete the image and rebuild
image build prints its banner and then nothing, at 0% CPU Old build: the build task was queued behind the NSApplication run loop and never started Upgrade — the task is now detached. Stage lines should appear within seconds
image build stuck at loading restore image metadata… A truncated or partial .ipsw — the framework blocks rather than failing Upgrade (the file is now size- and magic-checked first); re-download the IPSW
provisioning failed: … Failed to lock auxiliary storage after install The installer's VM had not yet released nvram.bin when first boot started Upgrade — the install now drains its queue and first boot retries for 30 s. Re-run image build; it resumes
image build says image 'default' already exists after a failed first boot Old build: an installed-but-unprovisioned bundle was treated as a finished image Upgrade — image build now resumes it instead of refusing

VM won't start — entitlement error

Symptom. Any VM operation fails immediately with an error naming com.apple.security.virtualization, or a generic "operation not permitted" from Virtualization.framework.

Cause. Virtualization.framework checks the entitlement on the calling binary, and entitlements are only honoured on a signed binary inside a proper .app bundle. The raw product of swift build has neither.

Fix.

make install       # build → bundle → sign → install (make sign alone re-signs in place)
codesign -d --entitlements - ~/Applications/GiteaMacosRunner.app   # verify

The output must list com.apple.security.virtualization. Then confirm the command you're running resolves to the installed bundle's binary — which -a gitea-macos-runner should point at ~/Applications/GiteaMacosRunner.app/Contents/MacOS/gitea-macos-runner, not at .build/release/gitea-macos-runner. Ad-hoc signing is sufficient; you do not need a paid developer account.


virtualMachineLimitExceeded

Symptom. The first VM boots fine; a second or third fails with virtualMachineLimitExceeded.

Cause. macOS permits two concurrent macOS guests per host. This is an Apple kernel and licensing limit, not a resource constraint — more RAM will not raise it.

Fix. Set scheduler.maxConcurrentVMs to 2 or less. If you're already at 2 and still hitting the limit, a VM from a previous run is still alive — check for stray processes and for leftover directories under storeDir/vms, then restart the daemon so it starts from a clean state.

To handle more macOS jobs in parallel, add another Mac.


VM starts but never gets an IP

Symptom. The VM boots (you can see it progress if you use vm boot), but the daemon reports that it could not resolve the guest address, or gives up at scheduler.bootTimeoutSeconds.

Cause. The daemon resolves the guest's NAT address from the host's DHCP lease file, which is only written once the guest requests a lease — several seconds after the boot screen appears. If the address never appears at all, the usual culprit on macOS 15+ is the Local Network privacy prompt: a LaunchAgent that was never granted permission (or was denied) cannot talk to the guest.

Fix.

  1. Check the lease file after the guest has been up for ~30 seconds:

    cat /var/db/dhcpd_leases
    

    An entry with a recent lease timestamp and the guest's MAC means networking is fine and the problem is timing — raise scheduler.bootTimeoutSeconds.

  2. No entry at all: pre-authorize the VM subnet, then reboot (these are read at boot):

    sudo defaults write com.apple.network.local-network \
      AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
    sudo defaults write com.apple.network.local-network \
      AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"
    

    The /18 is deliberate: the NAT subnet is chosen at runtime and slides to the next free /24 (192.168.65.x, .66.x, …) when one is taken, so a pinned 192.168.64.0/24 breaks the day it moves. doctor reports local network access as a pass once it sees an allowlist covering that span. This is the deterministic fix for an unattended host — see Local Network: the app is not listed in System Settings for why the interactive grant is not.

  3. Or grant it interactively: run gitea-macos-runner vm boot --image default from a Terminal in the GUI session and click Allow. The app is not listed under System Settings → Privacy & Security → Local Network until it has made that first attempt.


Local Network: the app is not listed in System Settings

Symptom. doctor prints the local network access note telling you to approve the app under System Settings → Privacy & Security → Local Network, but the runner is nowhere in that list — so there is nothing to switch on.

Cause. This is expected, not a bug. The Local Network list is populated lazily: an app appears there only after it has actually attempted a local-network connection and been evaluated. It cannot be pre-approved. A freshly installed runner that has not yet reached a guest has no entry.

There is also nothing to query, so doctor cannot tell you the grant's state. Unlike most macOS privacy controls, Local Network privacy is not stored in TCC — per TN3179 the checks live "deep in the networking stack" as a Network Extension packet filter, so the permission is absent from TCC.db and tccutil reset does not apply to it.

Fix — interactive. Make the app connect once, from a GUI session where a human can answer:

gitea-macos-runner vm boot --image default

Click Allow. The entry now exists and can be toggled later. Do not wait for the LaunchAgent to trigger it; a background agent cannot answer the prompt, so it simply fails to reach the guest.

Fix — deterministic, and what to use on a CI box. Allowlist the VM subnet instead. It is keyed on the network rather than on the app, so no prompt is involved and nothing needs redoing:

sudo defaults write com.apple.network.local-network \
  AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
sudo defaults write com.apple.network.local-network \
  AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"

Reboot afterwards. doctor then reports local network access as a pass. Both keys are Apple's, documented in TN3179; the same pair is what Tart's FAQ recommends for this exact problem on CI hosts.

Why the interactive grant does not stick here. This project ships an ad-hoc signed bundle (codesign --sign -), and TN3179 notes that "local network privacy uses your main executable UUID as part of its implementation". The linker writes a new LC_UUID on essentially every rebuild, so a rebuilt-and-reinstalled runner can read as a different program and prompt again — while the old row lingers, since macOS provides no way to reset a Local Network decision to undetermined. Expect duplicate entries after a few upgrades. The subnet allowlist avoids all of this.


SSH times out on a freshly built image

Symptom. The VM boots and gets an IP, but SSH never connects. Attaching a display to the guest shows Setup Assistant — the region/Apple ID welcome flow — rather than a login window.

Cause. The image builder uses VZMacGuestProvisioningOptions to create the admin account, enable Remote Login, and skip Setup Assistant. That API requires macOS 27 or newer in the guest as well as the host. An older guest silently ignores the options: no error, no account, no SSH server — it just sits at first-run setup forever.

Fix. Rebuild the base image from a macOS 27+ IPSW:

gitea-macos-runner image delete default
gitea-macos-runner image build --ipsw ~/Downloads/UniversalMac_27.0_XXXXX_Restore.ipsw

Verify the host is also 27+ (sw_vers). There is no way to make a pre-27 guest work unattended with this builder.


SecKeyCreateRandomKey / "Interaction is not allowed"

Symptom. The daemon starts but fails during VM setup with a Security-framework error mentioning SecKeyCreateRandomKey, errSecInteractionNotAllowed, or "Interaction is not allowed". Often it works when you run the daemon by hand in Terminal and fails under launchd.

Cause. The login.keychain is locked. macOS 15+ requires it unlocked for key operations the VM lifecycle performs, and it is only unlocked inside a live, logged-in GUI session. A LaunchDaemon, an SSH-only session, or a Mac sitting at the login window all fail this.

Fix.

  1. Confirm the service is installed as a LaunchAgent, not a LaunchDaemon: gitea-macos-runner service install does the right thing; a hand-written plist in /Library/LaunchDaemons does not.

  2. Ensure the runner user is actually logged in with the desktop loaded. Enable auto-login: System Settings → Users & Groups → Automatically log in as.

  3. Prevent the machine from returning to a locked state:

    sudo pmset -a sleep 0 disablesleep 1
    

    and disable "Require password after screen saver begins" for the runner user.

Connecting over Screen Sharing to a Mac at the login window does not unlock login.keychain for launchd's session — auto-login is the reliable answer on a dedicated CI Mac.


Job stays queued and no VM boots

Symptom. The workflow shows as queued in Gitea indefinitely. Nothing appears in the daemon logs about it.

Causes and fixes, in the order worth checking:

  1. Label mismatch. runs-on must use bare label names (macos-arm64), and every label listed must appear in the host config's runner.labels. The :host suffix used at registration is runner-side only and must never appear in workflow YAML. A single typo produces exactly this symptom with no error anywhere.

  2. Daemon not running or not polling.

    gitea-macos-runner service status
    log show --predicate 'process == "gitea-macos-runner"' --info --last 15m
    

    You should see a poll every scheduler.pollIntervalSeconds.

  3. Gitea too old. The queued-jobs API with the labels field requires Gitea ≥ 1.25. On an older instance the daemon can never see jobs. doctor reports the server version.

  4. Admin PAT wrong or under-scoped. The token must belong to a site admin and carry read:admin + write:admin. Test it:

    curl -H "Authorization: token $TOKEN" \
      "https://gitea.example.com/api/v1/admin/actions/jobs?status=queued"
    

    A 403 means scope or admin status; a 404 usually means the Gitea version predates the endpoint.

  5. A VM booted but its runner never came online. Then the job is queued and you see VM activity in the logs. Look at registration failures — most often an invalid or invalidated registration token (see below).

  6. Job expired. Gitea abandons a job after ABANDONED_JOB_TIMEOUT (default 24h). If the daemon was down longer than that, the job is gone; re-run it.

Registration fails with an invalid token

Registration tokens are reusable, but creating a new token for a scope invalidates the previous one. If someone clicked "create new registration token" in the Gitea UI, the token in your registrationTokenFile is now dead. Re-seed GITEA_RUNNER_REGISTRATION_TOKEN on the server (and restart Gitea), or switch to fetchRegistrationTokenViaAPI: true. See setup.md §1.3.


actions/checkout fails instantly

Symptom. The job starts, the runner connects, and the very first step fails immediately — typically a spawn error naming node, or an unhelpful non-zero exit before any output.

Cause. Gitea Actions' JavaScript actions (actions/checkout and most of the ecosystem) run by spawning node inside the guest. Node.js is required in the image, and if provisioning was interrupted it may be absent.

Fix. Confirm, then reprovision:

gitea-macos-runner vm boot
ssh admin@<guest-ip> 'node --version && git --version'

gitea-macos-runner image provision default

If git is also missing, the provisioning step failed early — check the build log and rerun provisioning.


Runner rows piling up in the Gitea UI

Symptom. Site Administration → Actions → Runners accumulates offline macos-vm-… entries.

Cause. Gitea deletes an ephemeral registration when its job completes normally. A VM that is killed uncleanly — daemon crash, host power loss, jobTimeoutMinutes kill — never reaches that point, so the row is orphaned.

Fix. Usually nothing: the daemon's reconcile loop sweeps orphaned registrations every scheduler.reconcileIntervalSeconds (default 300) via DELETE /api/v1/admin/actions/runners/{id}, plus a daily midnight sweep. Orphans should clear within a few minutes.

If they persist, the daemon's admin PAT probably lacks write:admin — check the logs for delete failures. To clear them by hand, delete the rows in the Gitea UI; they are inert (offline registrations that have already been spent cannot receive jobs).


doctor warns "runner download url → 403"

Symptom. doctor reports the runner download URL as a 403, but pasting the same URL into a browser downloads the binary fine:

! runner download url    https://gitea.com/gitea/runner/releases/download/v3.0.2/… → 403

Cause. gitea.com does not serve release assets itself. It answers with a 303 See Other pointing at a presigned object-storage URL, and that signature covers the HTTP method of the request that minted it. A HEAD probe therefore gets a HEAD-signed URL — and then the HTTP client, following the 303, rewrites the method to GET (RFC 7231 §6.4.4) and replays the signed URL with the one verb it was not signed for. The store answers 403 SignatureDoesNotMatch. Nothing is actually wrong with the asset.

Fix. Upgrade — the check now probes with a one-byte ranged GET (Range: bytes=0-0), the same verb image provision uses for the real download, and falls back to a HEAD only if the range is rejected. If you still see a 403 after upgrading, the asset really is gone: check runner.version and runner.runnerDownloadURL for a darwin-arm64 build.


Disk filling up

Symptom. Free space falls steadily; the daemon starts refusing to launch VMs, citing storage.minFreeDiskGB.

Cause. Each VM is an APFS copy-on-write clone of the base image. The clone is free at creation but grows as the job writes — dependency caches, build outputs, Xcode's derived data. Clones from uncleanly-killed VMs are not reclaimed automatically.

Fix.

# See what's there.
du -sh ~/Library/Application\ Support/gitea-macos-runner/*
ls -la ~/Library/Application\ Support/gitea-macos-runner/vms

# With the daemon stopped, remove stale clones.
gitea-macos-runner service uninstall     # or stop the daemon
rm -rf ~/Library/Application\ Support/gitea-macos-runner/vms/<stale-clone>
gitea-macos-runner service install

Only delete entries under vms/ — images/ holds the base images you'd otherwise have to rebuild. Longer term: raise storage.minFreeDiskGB so the guard trips earlier, delete unused base images with image delete, and remember an Xcode image needs 140 GB+ of headroom, more with two concurrent clones diverging.


image build prints the banner and then nothing at all

Symptom. image build prints

building image 'default' (this takes a while; the IPSW alone is ~15 GB)

and then stops — no stage lines, no progress bar, no error. Activity Monitor shows the process using no CPU and no VM services running. Ctrl-C does nothing.

Cause. A bug in releases before this fix. image build hosts an NSApplication run loop (Virtualization.framework requires one), and the work was started with a Task that inherited the main actor. Because NSApplication.run() is itself reached from Swift's async main, the main dispatch queue already had a block in flight and would not re-enter — so the build task was queued behind a run loop that never yields and never got a first tick. Nothing ran, including the code that would have reported the error. SIGINT/SIGTERM handling was stuck the same way, which is why Ctrl-C did not work either.

Fix. Upgrade. The build task is now detached and signals are handled off the main queue. You should see stage lines within a second or two:

resolving restore image…
loading restore image metadata…
creating VM bundle (disk 64 GB)…
installing macOS  [########----------------------]  27%

If a build still goes quiet, the line last printed tells you which stage owns the silence — see below.


image build sits at "loading restore image metadata…"

Symptom. The build reaches loading restore image metadata… and stays there. After a minute it adds:

still loading — a truncated or partially downloaded .ipsw can block here; verify the download completed

Cause. VZMacOSRestoreImage.image(from:) reads the whole archive's metadata and reports no progress while it does. On a healthy ~21 GB IPSW this takes seconds to a minute or two. On a partial download it can block for a very long time instead of failing.

Fix. The obvious checks are now made before the framework is handed the file — it must exist, be a regular file, be at least 1 GB, and start with the zip magic PK — so an incomplete download now fails immediately with its actual size rather than hanging. If you are on an older build, check by hand:

ls -la ~/Downloads/UniversalMac_*.ipsw          # ~15-22 GB, no .download/.crdownload sibling
xxd -l 2 -p ~/Downloads/UniversalMac_*.ipsw     # must print 504b

Anything smaller, or not starting 504b, is an incomplete or wrong file: delete it and download it again.

A pattern that matches more than one IPSW is refused outright, listing the matches — for example ~/Downloads/UniversalMac_27.0_*.ipsw when two 27.0 builds are sitting in ~/Downloads. Name exactly one of them.


image build hangs at install

Symptom. image build sits for a long time at the macOS install phase with little visible progress.

Cause. Usually none — installing macOS from an IPSW genuinely takes a long time (tens of minutes, longer on slower storage). The install phase is largely silent.

Fix. Wait, and do not stop the VM mid-install. Interrupting the installer leaves the disk image in an undefined state; the resulting image may boot and then fail in confusing ways later. There is no resume from a partial install — only from a complete one (see the next section).

If you did interrupt it, or the build genuinely failed:

gitea-macos-runner image delete <name>
gitea-macos-runner image build --ipsw <path> --name <name>

Before rebuilding, verify the IPSW is complete and matches your host architecture (Apple Silicon) and version requirement (macOS 27+ for unattended provisioning), and that you have enough free disk for the IPSW plus the target disk size.


Failed to lock auxiliary storage right after the install finishes

Symptom. The macOS install runs to 100%, then:

first boot + guest provisioning…
error: provisioning failed: could not boot image 'default': provisioning failed: Invalid
virtual machine configuration. Failed to lock auxiliary storage.

Cause. A VZVirtualMachine holds an exclusive lock on its bundle's nvram.bin for its entire lifetime and releases it in dealloc. VZMacOSInstaller owns a VM of its own, and older builds resumed the caller from inside the installer's completion handler — before the framework's frame had unwound and dropped the last reference. First boot then constructed a second VM over the same bundle and lost the race.

Fix. Upgrade. Two changes address it:

  • The install now tears down its VM on its own serial queue and resumes the caller only from a later block on that queue, so the installer's VM is deallocated before install() returns.
  • First boot retries specifically on this failure for up to 30 s (2 s apart), printing waiting for installer to release the VM bundle…. Any other configuration error still fails immediately — an invalid configuration does not become valid by waiting.

Nothing is lost when it does happen: the install is complete, so re-running image build resumes.


image build resumes an installed-but-unprovisioned image

Symptom. A previous image build finished installing macOS and then failed at first boot or provisioning. Re-running it prints:

image default already installed — resuming first boot + provisioning

This is intended. An image build that fails after the install has left an hour of work on disk, and throwing it away to redo an identical install is not a reasonable default. When the bundle exists, is complete, and is not yet marked provisioned, the install phase is skipped and the build goes straight to first boot.

It is safe because provisioning is exactly what did not happen: the guest has never been booted, so its first boot is still ahead of it and VZMacGuestProvisioningOptions — which macOS evaluates only on the first boot after a restore — still applies.

The one ambiguous case. The bundle on disk cannot say whether an earlier run got far enough to boot the guest. If it did, that single chance at automated Setup Assistant is spent. Rather than guess, the resume boots and waits for SSH: if the guest answers, provisioning was never applied and the build continues normally. Only if SSH times out does it stop, and it then tells you so explicitly — that state is unrecoverable, and the fix is image delete followed by a fresh build.

An image that is already provisioned is untouched; image build still refuses with image '<name>' already exists. To re-run provisioning on a finished image, use gitea-macos-runner image provision <name> instead.


SSH fails with "No route to host" (errno 65) mid-run

Symptom. A command that had just talked to the guest successfully suddenly cannot reach it — most visibly image provision, which waits for SSH, reports the first provisioning step, and then dies on the upload:

first boot + guest provisioning…
provisioning: system configuration (provision.sh)
error: provisioning failed: could not upload provision.sh from …/provision.sh:
  ssh failed: cannot connect to 192.168.65.232:22: … No route to host) (errno: 65)

Cause. On macOS 15 and newer, an app that has not been granted Local Network access does not get a "permission denied": the packet filter answers EHOSTUNREACH — errno 65, "No route to host", which is indistinguishable from a guest that is genuinely off the network. Guests here live on a host-private NAT link that is reachable whenever the VM is up, so on this path errno 65 is far more often the privacy filter than a routing problem.

Two details make it look intermittent rather than like a permission problem:

  • The grant is keyed on the executable's UUID. Per TN3179, Local Network privacy "uses your main executable UUID as part of its implementation", and the linker mints a fresh LC_UUID on essentially every build. A make install after a code change therefore presents a program macOS has never seen, whose permission is undetermined again — even though the binary you ran ten minutes ago worked.
  • Processes started over SSH are exempt. Running the same command through ssh you@host … succeeds while running it from a GUI Terminal fails. A remote-shell success proves nothing about the interactive path.

Fix. Allowlist the subnet — it is keyed on the network, not on the app, so no rebuild can withdraw it and no prompt has to be answered:

sudo defaults write com.apple.network.local-network \
  AllowedEthernetLocalNetworkAddresses -array "192.168.64.0/18"
sudo defaults write com.apple.network.local-network \
  AllowedWiFiLocalNetworkAddresses -array "192.168.64.0/18"
sudo reboot

The values are only read at boot, so the reboot is not optional — until it happens, defaults read com.apple.network.local-network shows the new setting while the filter still behaves as before.

Use /18, not /24. Virtualization.framework's NAT starts at 192.168.64.0/24 but chooses the subnet at runtime and steps to the next free /24 when that one is in use, so hosts drift to 192.168.65.x and beyond. An allowlist naming a single /24 that the NAT has since moved off is the worst case: it reads as configured, doctor used to call it a pass, and every guest connection still fails with errno 65. doctor now warns instead when the allowlist does not cover 192.168.64.0–192.168.127.255.

Verifying. After the reboot, gitea-macos-runner doctor should show local network access as a pass naming the range. Re-run the command that failed; nothing else needs redoing, and image provision is idempotent, so a partially completed run is safe to repeat.