Files

17 KiB

M1 spikes (docs/WINDOWS_PORT.md §13)

Programs that answer questions the port is currently guessing at. They are checked in and runnable on purpose: a spike whose answer nobody can reproduce six months later is a rumour.

They are deliberately not in windows/Nucleic.sln. The solution's broker contract tests must keep running on any machine against FakeWslc; these need the Microsoft.WSL.Containers preview NuGet and a real WSL stack, so they are built by path.


WslcApiDump — what does the wslc API actually look like?

Status: it has served its original purpose (docs/WINDOWS_PORT.md §13.1). The real API is known, and it was obtained without a Windows machine at all — the package is public, so the .nupkg was downloaded, its projection assembly extracted, and its metadata read with MetadataLoadContext. What this tool is for has therefore changed: its assumption list now describes the surface actually observed in 2.9.3, so running it says what the next package version moved, not what we guessed wrong.

Why it is still reflection. Same reason as before: a typed check fails to compile on the first rename and reports one problem, where this reports all of them at once. That property is worth keeping for a preview API whose next version can break anything.

Two facts it discovered that anything referencing this package needs:

  • the assembly is wslcsdkcs.dll, not Microsoft.WSL.Containers.dll — loading it by package name fails;
  • the only published version is 2.9.3, targeting net8.0-windows10.0.19041.0. The 0.1.0-preview.1 pin this repo carried could never have restored.
cd windows/spikes/WslcApiDump
dotnet run                      # dump the API + check every facade assumption
dotnet run -- --probe           # + GetMissingComponents / GetVersion
dotnet run -- --session         # + create a session, a SECOND with the same name, identity-test, tear down
dotnet run -- --session --keep  # …and leave the sessions running afterwards
dotnet run -- --all-types       # include the ABI/marshalling plumbing in the dump
dotnet run -- --internal        # can the SERVICE-INTERNAL COM interface be reached? (QI only)
dotnet run -- --internal-call   # …and call through it (can crash — that is the finding)

If it reports missing components, the machine cannot run wslc yet, and the two components are NOT fixed the same way — the tool prints the specific remedy for each:

Missing Fix
VirtualMachinePlatform wsl --install, then reboot (an OS optional feature).
WslPackage wsl --update --pre-release, then wsl --shutdown. WSL is installed but older than the SDK, and 2.9.3 is pre-release-only — a plain wsl --update will not get there. Confirm with wsl --version.
SdkNeedsUpdate The NuGet pin is ahead of the installed service: update WSL further, or pin the package back.

The distinction is worth knowing because the failures look alike but mean opposite things: REGDB_E_CLASSNOTREG (0x80040154) is nothing installed, ERROR_NOT_SUPPORTED (0x80070032) is installed but too old. The tool decodes both, along with the WSLC_E_* range from wslc.idl, because these COM exceptions often carry an empty message and leave nothing but a hex code. --session is skipped while components are missing rather than failing the same way.

Note that the assumption check reads the Microsoft.WSL.Containers namespace only. That is correctness, not tidiness: a C#/WinRT projection also exports ABI.Microsoft.WSL.Containers.* marshalling types with the same short names, and matching on short name alone checks every member against the marshalling struct — which reported 42 false MISSINGs, with CreateMarshaler helpfully offered as the nearest name.

It writes the full public object model to wslc-api-dump.txt (--out to relocate) and prints an ok / MISSING line per assumption, each with why that member matters and a nearest-name hint. It exits 0 even when assumptions fail — a mismatch is the product, not an error. Only a genuinely broken run (the assembly won't load) exits non-zero.

--probe has been run clean on real hardware (Windows 11 amd64, WSL 2.9.3): 54/54 ok, no missing components. Every other path is exercised against a stand-in assembly carrying the observed 2.9.3 shape — the ABI. shadow types, a service reporting missing components, an 0x80070032 with an empty message, a ServiceVersion with no ToString() override, and a duplicate-name Session throwing 0x80040607 — so a failure on your machine is a finding about wslc, not about this tool.

Note the ToString() one, because it bit: a WinRT projection class does not override ToString(), so printing a returned object gives you its type name. Values are rendered by their properties instead — ServiceVersion { Major=2, Minor=9, Revision=3 }, not Microsoft.WSL.Containers.ServiceVersion.

What to send back: the console output, and wslc-api-dump.txt if anything is MISSING — which now means the package moved under us, not that we guessed wrong.

Why --session earns its risk

--session creates a real wslc session named nucleic-spike under %LOCALAPPDATA%\Nucleic\spike\wslc, starts it, and then creates a second one with the same name. That second construction is the whole point, and it is the last open question D13 turns on: the compat SDK exposes a Session constructor and no attach, so does constructing over an existing name re-adopt it, or refuse?

Answered on 2.9.4: it cannot. The constructor is lazy and always succeeds — judging by it is what made the first reading of this probe wrong. Start() is where the service is consulted, and a second Start() on a running name fails with ERROR_ALREADY_EXISTS (0x800700B7). Note it is not WSLC_E_SESSION_RESERVED, which exists in wslc.idl but evidently means something narrower — don't key on it.

So the probe judges on Start(). If a future version lets the second Start() through, it then runs an identity test: terminate the FIRST session and read from the SECOND. A read that worked before and fails after is one underlying session answering both handles; a read that keeps working means two independent VMs.

The identity test deliberately uses only the compat SDK. The obvious check would be wslc session ls — but wslc.exe is not on PATH by default, so a spike that depends on it answers nothing on a stock machine. (It ships beside wsl.exe; try C:\Program Files\WSL\wslc.exe. That it isn't on PATH is one more small argument for D13's no-CLI stance.)

Sessions are torn down at the end by default — an earlier version left a WSL VM running and told you to clean it up with a command that doesn't exist. Pass --keep to leave them, which is how you check the other half of the §2.3 question: whether session state outlives the process that created it. With --keep, wsl --shutdown clears everything.

The gateway address is not what this probe is for — §13.1 established that no API surfaces one, so it comes from GetAdaptersAddresses over vEthernet (WSL) instead, which WslcFacade now does.

--internal — is D13's internal arm even reachable?

The newest open question, and the one item 5 stopped at (docs/WINDOWS_PORT.md §13.2). D13 routes enumeration, reattach and the Terminal panel's pty to the service-internal IWSLCSessionManager (IID 82A7ABC8-…). But wslc.idl declares interfaces and no activatable class — there is no CLSID in it, so CoCreateInstance has nothing to name. The only registered coclasses in WSL's IDLs belong to the compat surface (WSLCCompatSessionManager and its factory) and to the WSL service proper (LxssUserSession, LxssUserSessionInBox).

Answered on hardware (WSL 2.9.4): yes. WSLCCompatSessionManager (a9b7a1b9-0671-405c-95f1-e0612cb4ce8f) — the same class the SDK itself activates — answers a QI for IWSLCSessionManager. One object, two faces. IWSLCVirtualMachine was refused by every class, confirming §13.1's retraction against the machine rather than against an IDL.

--internal-call then confirmed the vtable matches the IDL: GetVersion() (slot 3) returned 2.9.4, agreeing with the compat WslcService.GetVersion() in the same run — the same service answering through both faces — and ListSessions() (slot 6) returned S_OK. So hand-written ComImport is sufficient and the C++/WinRT shim §13.1 mused about is not needed.

Re-run as --session --keep --internal-call, ListSessions returned the live session by name (#6 "nucleic-spike"), so the entry-struct layout — inline wchar_t buffers, the same shape ListContainers uses — marshals as declared. Enumeration is proven.

That run also turned up a trap worth knowing before writing any of this into brokerd: OpenSessionByName failed 0x80070542 on the session ListSessions had just listed. That is ERROR_BAD_IMPERSONATION_LEVEL — the service impersonates the caller to resolve a per-user session, and .NET hands COM proxies RPC_C_IMP_LEVEL_IDENTIFY by default. A security error that reads exactly like "not found", and only on the methods reattach needs. The probe now raises the proxy blanket to RPC_C_IMP_LEVEL_IMPERSONATE; brokerd should call CoInitializeSecurity at startup instead, before its first COM call — including the compat SDK's, or it fails RPC_E_TOO_LATE. wsl --shutdown to clean up after --keep.

The hypothesis was that one of those objects also implements the internal interface — COM objects routinely expose several — so --internal activates each and QIs for the internal IIDs. It never calls a method, so a vtable mismatch cannot crash it, and QI alone answers the question. --internal-call then calls GetVersion (slot 3) and ListSessions to confirm the vtable really matches the IDL; that can take the process down, which is why it is opt-in.

It also QIs IWSLCVirtualMachine, which is expected to fail: per the IDL only IWSLCVirtualMachineFactory::CreateVirtualMachine produces one, and the SYSTEM service owns the factory. §13.1's hvsocket-primary decision was retracted on that reading, and a decision reversed by reading deserves confirming against the machine.

The answers that change the design

Most mismatches are a one-line edit in WslcFacade.cs — that is exactly what the IWslc seam is for, and nothing above it should move. These three are different:

Finding Consequence
No container enumeration, stats, pty or attach in the SDK D13's answer, now gated. They exist on wslc.idl, the service-internal interface wslc.exe calls — but that IDL declares no coclass, so reaching it at all is what --internal tests. Stats no longer wait on it: WslcFacade reads cgroup v2 in the guest instead.
Container has no Name, Session has no GetContainers() Sharper than "no enumeration": a container is reachable ONLY through the handle CreateContainer returned, so the broker keeps its own name→handle roster — which dies with the process. A restarted broker sees an empty sandbox (§13.2).
No create-or-attach on Session IWSLCSessionManager::OpenSessionByName / EnterSession — same gate as above. Until then session.ensure fails session_exists and names wsl --shutdown as the remedy.
No gateway address anywhere on the API Gateway TCP stays primary. The hvsocket alternative needed a VMID from IWSLCVirtualMachine::GetId, and that interface is unreachable from a client — retracted in §13.1. The address comes from GetAdaptersAddresses over vEthernet (WSL).
No uid on ProcessSettings Recoverable, and already planned for: exec wraps argv in setpriv/su agent -c. Interceptors and hydrashell don't care about the numeric uid (§3.2).


WslcSpike — does any of it actually work?

M1 spike (a2). The whole Windows container subsystem now compiles and none of it has ever run: WslcFacade (compat SDK), WslcInternal (D13 Tier 1 recovery), the §3.3 RPC surface. This drives all of it the way hostd does — spawn nucleic-brokerd.exe, speak NDJSON JSON-RPC over its stdio — so what passes here is the shipping code and not a rehearsal of it.

That is why it is a client of the broker rather than a second program calling wslc. Its csproj is plain net9.0 with no wslc reference at all, so it cannot drift into a parallel implementation, and it builds on any machine even though it only runs on Windows.

dotnet build windows\NucleicBroker\NucleicBroker.csproj -p:UseWslc=true   # the broker it drives
dotnet run --project windows\spikes\WslcSpike -- --repo C:\src\nucleic

Options: --image (default ghcr.io/abkslm/hydrangeaos-agent:26.07), --container, --session-name, --iterations (timing runs, default 5), --broker <path> if autodiscovery fails, --no-recovery to skip the crash test.

It answers four open questions:

  1. Does the happy path work? session → GHCR pull → container with an NTFS ContainerVolume → exec → stdio round-trip → signal → teardown. Also prints the §5 host gateway, which no wslc API surfaces and the control plane depends on entirely — an empty or loopback value there is M1 (b)'s answer arriving early.
  2. Is 9P fast enough for D8? §15 lists NTFS bind-mount performance as a top risk with no numbers behind it. The spike times git status in the mounted worktree and in a copy on the container's own ext4 — same container, same repo, so the mount is the only variable — then reports the ratio and what it implies. It deliberately does not "fix" a bad result by moving repos; §15 says surface that to the user, and a spike that silently relocated things would hide the very finding it exists to produce.
  3. Does D13 Tier 1 recovery work? It kills a broker with the session up — the exact §2.3 crash — starts a fresh one, and reports whether the orphan was cleared or whether it still fails session_exists. It distinguishes "Tier 1 never bound" from "Tier 1 bound and did not work", because those are different bugs.
  4. Can the guest reach the host at the gateway? (M1 b) The last thing M2 waits on. It binds a listener on the WSL-facing address only — never 0.0.0.0, which is the posture §5 requires — and has the container open a TCP round trip to it. Every agent session rides this: approvals, the git/gh interceptors, hydrashell's shell reports. On failure it prints the New-NetFirewallRule remediation §5 step 5 calls for, and notes that the AF_HYPERV fallback needs a VM GUID that §13.2 found no client route to — so gateway TCP is not merely preferred, it is the only path currently available.

Every step prints what happened rather than asserting: a spike's product is evidence. It exits non-zero only when a step that should have worked threw. The [brokerd] lines interleaved in the output are the facade's own stderr diagnostics — where WslcFacade/WslcInternal report what bound and what degraded — and are usually the fastest route to a cause.

--iterations on a large repo is the number that matters for D8; the default 5 on a small one is a smoke test, not a measurement.


Not yet written

  • The internal arm, Tier 1 — "recover" written: windows/NucleicBroker/Wslc/WslcInternal.cs. On broker restart it opens the orphaned session, lists its containers for the log, terminates it, and lets the facade start fresh — automatic clean recovery instead of a manual wsl --shutdown. Still needs one live test: kill a broker mid-session and confirm the next one recovers (see below).
  • Tier 2 — "adopt" (keep containers running across a broker restart) is blocked. The Session.FromAbi() handoff throws InvalidCastException even though the QI to IWSLCCompatSession succeeds: the WinRT layer appears to be a client-side wrapper in wslcsdk.dll, not the service object. --internal-call now compares COM identity to confirm. If confirmed, Tier 2 means driving exec/stdio through internal COM directly (IWSLCContainer::Exec, IWSLCProcess::GetStdHandle — handle-based, not event-based).
  • GatewaySpike — M1 (b), and the control-plane spike that actually matters now: can a Bridged wslc container reach the host at the vEthernet (WSL) address WslcFacade returns, under the default Windows firewall, and does control-bridge.js complete an MCP round trip over it? This is the committed path, not a fallback.
  • HvSocketSpike — downgraded back to upside (§13.2). It needs a VM GUID to bind AF_HYPERV to, and IWSLCVirtualMachine::GetId turned out to be unreachable from a client, so the VMID would have to come from outside wslc entirely — HCS enumeration, or the WSL VM's registry identity. Worth doing only if the gateway path shows friction in the wild; not on the M2 path.