This commit is contained in:
@@ -26,6 +26,7 @@ gitea-macos-runner service status
|
||||
| VM boots but never gets an IP | DHCP lease not yet written, or Local Network privacy denial (macOS 15+) | Check `/var/db/dhcpd_leases`; grant Local Network permission or pre-authorize the subnet |
|
||||
| Runner not listed under Privacy & Security → Local Network | Expected — the list is populated only after the app first attempts a local connection; it cannot be pre-approved | Boot one VM by hand from a GUI Terminal to create the entry, or (better on CI) allowlist the subnet with `defaults write com.apple.network.local-network` |
|
||||
| SSH times out on a freshly built image | Guest macOS < 27, so provisioning options were ignored and Setup Assistant is waiting | Rebuild the image from a macOS **27+** IPSW |
|
||||
| VMs boot in a loop; every teardown says `reason=cancelled` and nothing is logged between the lease and the teardown | `scheduler.bootTimeoutSeconds` is below the guest's *worst-case* boot on a contended host, so each clone is killed while still starting — and each replacement makes the next one slower | Raise `scheduler.bootTimeoutSeconds` (default 900) and reduce the number of concurrent guests; see [The daemon boots VMs forever](#the-daemon-boots-vms-forever-and-every-teardown-says-reasoncancelled) |
|
||||
| `ssh failed: cannot connect … No route to host) (errno: 65)` part-way through provisioning | macOS 15+ Local Network privacy blocking the app — the grant is keyed on the executable's UUID, so `make install` withdraws it | Allowlist the subnet (`192.168.64.0/18`) and **reboot**; see [SSH fails with "No route to host" mid-run](#ssh-fails-with-no-route-to-host-errno-65-mid-run) |
|
||||
| Allowlist is set but guests are still unreachable | It names `192.168.64.0/24` while the NAT has moved to `192.168.65.x` | Widen it to `192.168.64.0/18` and reboot; `doctor` now warns about too-narrow allowlists |
|
||||
| `SecKeyCreateRandomKey` / "Interaction is not allowed" | `login.keychain` is locked — no GUI session | Run as a LaunchAgent in an unlocked GUI session; enable auto-login |
|
||||
@@ -195,6 +196,63 @@ with this builder.
|
||||
|
||||
---
|
||||
|
||||
## The daemon boots VMs forever and every teardown says `reason=cancelled`
|
||||
|
||||
**Symptom.** A job is queued, the daemon is running, and the log repeats the same three lines with a
|
||||
new runner name each time — but no runner ever appears in Gitea:
|
||||
|
||||
```
|
||||
info orchestrator: job=1 runner=macos-vm-1cd8e83f… slot=0 booting VM
|
||||
info orchestrator: ip=192.168.65.233 slot=0 guest leased address
|
||||
info orchestrator: reason=cancelled slot=0 tearing down slot
|
||||
```
|
||||
|
||||
Note what is missing: nothing between the lease and the teardown, and a teardown reason that names
|
||||
no cause.
|
||||
|
||||
**Cause.** The guest takes longer to reach `sshd` than `scheduler.bootTimeoutSeconds` allows, so the
|
||||
scheduler tears the slot down while it is still coming up — usually seconds before it would have
|
||||
succeeded. This is not a timeout that fires once; it is a **livelock**. The replacement clone starts
|
||||
from zero *and* adds load to an already contended host, so the next boot is slower still and the
|
||||
loop never converges.
|
||||
|
||||
Several Virtualization guests on one Mac is enough to cause it: a guest that reaches SSH in 40
|
||||
seconds on an idle host can take four or five minutes when it is sharing the machine, and each slot
|
||||
holds 4 vCPU and 8 GB for the whole attempt. Check with `uptime` inside a guest — a load average in
|
||||
the tens means the guest is starved, not broken.
|
||||
|
||||
**Fix.**
|
||||
|
||||
1. Raise `scheduler.bootTimeoutSeconds`. The default is 900; treat it as a ceiling on the guest's
|
||||
*worst* case, not its typical one. Timing out too early costs far more than noticing a genuinely
|
||||
wedged guest late.
|
||||
|
||||
2. Reduce contention. Count what is actually running:
|
||||
|
||||
```sh
|
||||
ps -Ao pid,rss,etime,comm | grep -i -e virtual -e vmnet
|
||||
```
|
||||
|
||||
Virtualization guests belonging to *other* tools compete for the same cores and the same two-VM
|
||||
macOS limit. Shut down what you are not using, or lower `scheduler.maxConcurrentVMs`.
|
||||
|
||||
3. Confirm the guest itself is fine, independently of the daemon, with `doctor` — its `guest ssh`
|
||||
check authenticates against whichever slot currently holds a lease:
|
||||
|
||||
```sh
|
||||
gitea-macos-runner doctor
|
||||
```
|
||||
|
||||
**If you are on an older build**, upgrade: the empty gap in that log was three bugs, all now fixed.
|
||||
`waitForSSH` was handed the full `bootTimeoutSeconds` even though the scheduler's clock had started
|
||||
before the clone — so the scheduler always fired first and `waitForSSH`'s own error was unreachable;
|
||||
its per-attempt failures were only ever reported in that unreachable error; and the teardown
|
||||
overwrote the planner's reason with `cancelled`. Current builds log `waiting for guest ssh` with the
|
||||
attempt count and the last error while it is happening, and report
|
||||
`reason="boot timeout: provisioning for 312s (limit 300s)"`.
|
||||
|
||||
---
|
||||
|
||||
## `SecKeyCreateRandomKey` / "Interaction is not allowed"
|
||||
|
||||
**Symptom.** The daemon starts but fails during VM setup with a Security-framework error mentioning
|
||||
|
||||
Reference in New Issue
Block a user