Appendices · reference

Troubleshooting

18 min read · 2026-09-02 · vmlab 0.9

Each section below starts with the message vmlab prints, then says what causes it and what to do. Messages are quoted from the source; the parts in braces are filled in with a machine, lab or path name. Most of them reach you through the lab daemon's log as well as the terminal, so vmlab logs shows the same words when a command has already returned. Non-zero exit codes are the protocol error codes listed in Wire protocol and error codes.

KVM is unavailable and the guest runs under TCG

text
{machine}: KVM unavailable for {arch} — falling back to TCG (slow)

This is a warning, not an error. The machine still boots, with -accel tcg, and every instruction is emulated. Boot takes minutes rather than seconds and provisioning runs at the same pace.

vmlab uses KVM only when it can open /dev/kvm for reading and writing and the guest's architecture matches the host's. So there are two causes. Either the daemon user cannot open the device (the device is absent, or the user is not in the kvm group), or the lab asks for a foreign architecture, as examples/alpine-arm64 and examples/riscv64-ubuntu do on an x86 host. The second case is expected and has no fix.

For the first case, add your user to the kvm group and log in again, then restart the daemons with vmlab lab restart. On WSL 2, see Nested virtualisation is off on WSL 2 below.

A runtime binary is missing

text
missing required binaries on PATH: {names} — install the QEMU/swtpm packages (PRD §14 lists the runtime dependencies)

vmlab up and vmlab machine start check the binaries every targeted machine needs before anything boots. A VM needs qemu-img and the qemu-system-<arch> for its architecture, and swtpm when its resolved hardware has a TPM. A lab container needs qemu-img and the host architecture's qemu-system. vmlab validate does not run this check, so a lab that validates can still fail here.

Install the packages named in Install and run vmlab up again. Three more dependencies are not in this preflight and fail later, at the point of use, with their own words.

MessageCauseFix
{arch} UEFI firmware not found; tried: {paths}; vmlab's own firmware was not found either …The guest assets that carry vmlab's own OVMF and AAVMF are missing from every guest asset directory, and the host has no distribution firmware to fall back to. riscv64 has no bundled firmware, so only the host's counts.Reinstall the guest assets (install.sh, or just guest-install from a checkout), or point firmware_dir in the host config at a copy. For riscv64, install qemu-efi-riscv64.
{arch} UEFI VARS template not found; tried: {paths}Host fallback only: a distribution's firmware code was found but its variable-store template was not.Reinstall the guest assets, or install the same firmware package; the two ship together.
{arch} secure boot needs a UEFI VARS template with keys enrolled …secure_boot = true, the bundled firmware is missing, and the host has a secure-boot build but no variable store with Microsoft's keys enrolled beside it. Arch's edk2-ovmf is one such build. A blank store boots in setup mode and enforces nothing, so vmlab refuses it.Reinstall the guest assets: they carry OVMF_VARS_4M.ms.fd (x86_64) and AAVMF_VARS.ms.fd (aarch64), enrolled.
firmware_dir {dir} is the {arch} firmware this host config chose, and it lacks …The host config's firmware_dir has a directory for this architecture, which makes it the only place that architecture is looked up, and the pair the VM needs is not in it.Add the named file, or remove the architecture's directory to fall back to vmlab's own firmware.
this VM's UEFI VARS store is {n} bytes and the firmware vmlab would boot ({code}) takes {m} …The VM was created under a firmware of a different flash layout (a host's 2 MiB OVMF, say), and no build of that layout is installed any more. A VM keeps its VARS copy for life.Delete the VM's .vmlab/vms/<vm>/OVMF_VARS.fd to start it with a fresh store, losing its UEFI boot entries, or reinstall the firmware it was created with.
share "{share}" demands transport = "virtiofs" but no usable virtiofsd was found on this host …A share declares transport = "virtiofs" and no virtiofsd is on PATH or in the distribution's helper directories.Install virtiofsd, or point VMLAB_VIRTIOFSD at one.
… virtiofsd {version} ({path}) cannot serve a share: its --help lists no {flags} …The virtiofsd found lacks a flag every share needs. QEMU's retired C virtiofsd is one such binary.Install the Rust virtiofsd, 1.13.0 or later, or point VMLAB_VIRTIOFSD at one.
share "{share}" is read-only, and virtiofsd {version} ({path}) has no --readonly (added in virtiofsd 1.13.0) …A read-only share demands virtiofs and the virtiofsd predates 1.13.0.Install a newer virtiofsd, or set transport = "smb" or "auto".
{machine}: online snapshots of a machine with virtiofs devices need virtiofsd 1.11.0 or later …An online snapshot capture or restore on a machine whose shares or volumes ride a virtiofsd without --migration-mode. Ubuntu 24.04's 1.10.0 is one.Stop the machine and use an offline snapshot, or install virtiofsd 1.11.0 or later and restart the machine.
cannot spawn sqfstarA container image is being flattened and squashfs-tools is not installed.Install squashfs-tools.

A secure_boot = true machine whose firmware lookup fails is refused earlier, at hardware resolution, with `vm "{machine}": secure_boot = true (from {source}) but no firmware …`. Secure boot is enforced: a guest whose bootloader Microsoft's keys do not sign is refused by the firmware (Access Denied on its console). To see which firmware a running VM booted, read the -drive if=pflash arguments of its QEMU process; under a normal install they point into the guest asset directory's firmware/<arch>/.

No guest asset for a container micro-VM

text
no micro-VM guest asset for {arch} (need vmlinuz + initramfs.img); searched: {dirs}. Build one with `guest/build-asset.sh {arch}` and install it into one of those directories (or point VMLAB_GUEST_ASSET_DIR at guest/dist).

A lab container boots a micro-VM from a kernel and initramfs that ship with vmlab, not with the image. vmlab looks for them under $VMLAB_GUEST_ASSET_DIR/<arch>/, then /usr/share/vmlab/guest/<arch>/, then ~/.local/share/vmlab/guest/<arch>/. None of those held both files.

A packaged install puts the asset under /usr/share/vmlab/guest. From a source checkout, build it with guest/build-asset.sh <arch> and either copy guest/dist into one of the searched directories or export VMLAB_GUEST_ASSET_DIR=guest/dist for the daemons. The agent binary a template bake needs has the same search path and the same shape of message, naming `guest/build-agent.sh <os>-<arch>` instead.

The template is not in the store

text
template {arch}/{name}@{version} not found in the store

The lab names a template by store ref and no version of it is installed, or the pinned version is not. The reply carries the not_found error code, so the command exits 4.

vmlab template list shows what the store holds. If the template is published on a registry, point the template field at its registry ref and vmlab up pulls it, or run vmlab pull first to fetch every missing template without booting. If it is a local build, run vmlab template build in the template's directory, as each directory under examples/templates shows. A store ref that exists but has lost its disk is a different message, `template {ref} is corrupt: missing disk.qcow2; remove that version with vmlab template rm` and pull or build it again.

A container whose image has not been fetched yet says `{machine}: image not pulled yet — run vmlab pull or vmlab up`. Both verbs pull it.

The machine never becomes ready

text
{machine}: not ready after 600s
{machine} stopped while waiting for ready
no vmlab-agent answered on the agent channel
agent did not open the channel in time

A machine is *ready* when its agent answers on the vmlab.agent.0 virtio-serial port. VMs get ten minutes, containers five. vmlab up waits for readiness only when a later wave depends on the machine, or a provision or share must run on it, so the first symptom is often first-boot {name}: agent did not come up from a provision rather than the timeout itself.

The usual causes, in order of likelihood: the template was built without the agent, or with an agent from an older vmlab; the guest is still installing (a template build that has not finished); the guest is running under TCG and has not reached userspace yet; or the guest booted but its agent service failed. Open the display with vmlab console <machine> and look. vmlab machine capabilities <machine> prints agent - when the agent has not answered at all, and a feature list when it has.

If the guest is up and the agent is absent, rebuild the template so the bake installs the agent this vmlab ships, or push one into the running machine with vmlab machine repair-agent. A container's agent comes from the host's guest asset, so on a container the message ends with `there is nothing to rebuild or repair — an agent that is not answering here is a machine to restart, or a guest asset to reinstall (§19.4)`.

The machine is suspended

text
{machine} is suspended: the guest went to sleep (ACPI S3) and nothing in it answers until it is woken

QEMU's q35 machine advertises ACPI S3, and a guest that takes it, as a Windows client edition does after an idle timeout, leaves QEMU running with its vCPUs stopped. vmlab status shows the VM suspended, the event log has vm.suspended, and exec, shell, cp and every other agent command fail at once with this message and exit code 5 (conflict) rather than waiting out a handshake.

vmlab vm start <machine> wakes it where it left off. To stop it sleeping at all, set prevent_sleep = true on the vm block, or turn the guest's own sleep timer off. Template builds already run with prevent_sleep on.

The machine's agent is out of date

text
the guest agent has no `watch` support — rebuild the template to update it

The workspace syncer needs the agent's fileops and watch features, and an agent baked by an older vmlab may serve neither. vmlab machine capabilities <machine> lists what the agent negotiated. vmlab validate says nothing about it, deliberately, because it needs a running agent to know.

vmlab up normally fixes this by itself: when a VM's agent stamp differs from the agent this vmlab ships, it pushes the shipped agent before any provision runs and prints `agent: updated "<vm>" (<old> → <new>). If it printed warning: agent: could not update "<vm>"` instead, the reason follows on that line, and the old agent is still in place. A lab file with agent_update = false turns the refresh off.

Otherwise there are two remedies. Rebuilding the template is the durable one. vmlab machine repair-agent <machine> pushes the host's shipped agent into the running machine over its own channel and marks the machine *diverged* in vmlab status, so you remember the clone no longer matches its template. It refuses on three conditions, each in its own words.

RefusalMeaning
"{machine}" must be running with its agent answering before a new one can be pushed into it over that channelThere is no channel to push over. Start the machine, or fix the readiness problem first.
"{machine}"'s agent serves no fileops, so it cannot be handed a binary over its own channel — this one can only be replaced by rebuilding the template (§19.4)The old agent cannot receive a file. Rebuild the template.
this machine's agent lives in the initramfs guest asset this host installed, not in anything it boots — it already tracks the vmlab you are running and cannot go stale, so there is nothing to push into it. Refreshing it means reinstalling the guest asset (§19.4)The machine is a container. Reinstall the guest asset instead.

The workspace has stopped syncing

text
the workspace on "{machine}" has stopped, both directions, on {n} conflicting paths

The workspace syncer found a path changed on both sides since they last agreed, and halted the whole workspace on that machine, in both directions. It wrote nothing and deleted nothing: both copies are exactly where they were. The watch keeps running, so the halt lists every conflicting path in the batch rather than the first one. In the guest, a file named .vmlab-sync-halt at the workspace root carries the same list; it is the only signal the guest side gets, and it never syncs.

Resolution is host-side, from the lab directory, because vmlab opens every channel from the host and the guest only answers. `vmlab dev sync status` prints the halt with every path and the reason for each. Then, per path or for the batch: vmlab dev sync diff <path> shows the guest copy next to the host copy; vmlab dev sync resolve <path> --host keeps the canonical host copy and overwrites the guest's; --guest does the reverse; and --all takes every halted path with the flag you give. Making both sides identical by hand is a third route that needs no verb: the next pass adopts them as agreed. resolve with neither side flag refuses with say which side wins, and with no path and no --all it refuses with name the paths to resolve, or pass --all` to take the whole batch`.

The losing copy is gone

resolve overwrites the side that loses and vmlab keeps no copy of it. Run vmlab dev sync diff first when either side might hold something you want.

Two other messages come from the same loop and look like a halt but are not. The bulk-delete guard, `the guest deleted {n} of the {m} paths this workspace had agreed on, which is a rewrite of the canonical copy rather than an edit: nothing was removed on the host`, halts on a delete batch above a proportion with a floor; resolve it with --guest if the deletes were intended. The volume warning, `this pass is carrying {n} paths ({m} MiB) under {prefix} — syncing continues, and adding {prefix}/ to .vmlabignore makes that subtree guest-owned if it is build output`, never halts. A line under deferred while git holds a lock is timing, not a conflict, and clears itself when the lock goes.

A snapshot capture is refused

text
"{machine}"'s workspace is not in step with the canonical copy, so this snapshot would capture a tree the host has never agreed with. … There is no flag for this — a snapshot of a tree mid-transfer restores to somewhere meaningless. `vmlab dev sync status {machine}` says what is outstanding and `vmlab dev sync flush {machine}` waits for it; a halt has to be resolved first. Snapshots are not a workspace backup: a dev machine's source lives on the host, which is what survives `destroy` and what a restore re-converges the guest from.

Snapshot capture on a dev machine flushes the syncer first and refuses while the guest holds work the canonical copy has never seen: unsynced paths, a halt, an unfinished re-seed, or a workspace that has not completed a single pass. There is no escape flag, by design. Run vmlab dev sync flush <machine> and take the snapshot again; if the reason is a halt, resolve it as above. The middle of the message says which of these it is and names up to twenty of the paths still owed.

Restore has a matching refusal when the workspace is halted: `"{machine}"'s workspace is halted, and restoring would silently destroy the guest copy of every conflicting path.` Restore does have an escape flag, --discard-guest-changes, which throws away the guest copy of the whole workspace and re-converges it from the host after the rewind. Both refusals ride the ledger, so they apply to a stopped machine too.

Large reads fail on a Windows virtiofs share

text
WARNING: vm "{machine}": share "{share}" rides virtiofs into a Windows guest; virtio-win 0.1.302's VioFS driver fails large reads there at random ("Error performing inpage operation") and can leave a copy that reported success unreadable; use transport = "smb" (or "auto", which picks SMB on Windows) unless the template carries a VioFS without the fault, such as 0.1.285

vmlab up prints this for every share that rides virtiofs into a Windows guest. Listings and small files work, but reading a file of a few hundred megabytes or more fails in the guest with Error performing inpage operation. Some copy tools, robocopy /J among them, report success while leaving a copy that cannot be read. The fault is in virtio-win 0.1.302's VioFS driver, and no virtiofsd setting avoids it.

Set transport = "smb" on the share, or remove transport so that auto picks SMB. An SMB share needs the VM to have a nic {}. A template built with VioFS 0.1.285 reads large files correctly over virtiofs. The warning still appears for it, because vmlab cannot see the guest's driver version.

A port forward was skipped

text
{machine}: {forward}: host port {port} is already claimed by {machine}: {forward}

Every forward in the lab is planned before up installs any. When two claim the same host port, the first in plan order wins and the rest are dropped rather than left to a bind failure, because a bind failure names neither the winner nor the fact that there was a contest. The same plan skips a forward for three other reasons, each printed as {what}: {why}.

ReasonMeaning
no such vm or container in the labThe forward's to names a machine the lab does not declare.
needs a nic to reach it overThe target machine has no nic {}, so no segment reaches it.
no lease — is it running and ready?The target has a NIC but no DHCP lease yet. The forward installs once the machine is up; vmlab status shows the lease.

A host port held by an unrelated process is not detected by the plan. That forward fails at install time, which up treats as best effort: the rest of the lab still comes up, and vmlab logs carries a forward.skipped event in the same {what}: {why} shape, for example "vm1": forward: cannot listen on host 0.0.0.0:14330: Address already in use (os error 98). Change the host_port or free the port.

A forward that did install leaves a forward.installed event naming its host port and guest address. A forward with neither event was never planned: check that vmlab validate lists the segment and that the machine it names has started.

Nested virtualisation is off on WSL 2

vmlab prints nothing specific to WSL 2. When nested virtualisation is disabled in .wslconfig, /dev/kvm does not open and every machine gets the TCG warning at the top of this chapter. Enable it under [wsl2] in %USERPROFILE%\.wslconfig, run wsl --shutdown from Windows, and start again. Host configuration and WSL 2 has the full setup.

Some WSL 2 sessions have no $XDG_RUNTIME_DIR. vmlab does not fail on that: it falls back to /tmp/vmlab-<uid> for the supervisor socket and other runtime files, creates it with mode 0700, tightens an existing directory, and refuses one owned by another user. If you see the sockets under /tmp rather than /run/user/<uid>, that is why. Networking on WSL 2 needs no tap, bridge or macvlan, so the fast path degrades to userspace silently there; a Windows-side VNC viewer is reached with vmlab console --tcp, which bridges the display to a localhost port.

The supervisor did not come up

text
supervisor did not come up — check ~/.local/state/vmlab/vmlabd.log

Every vmlab verb that needs the supervisor connects to $XDG_RUNTIME_DIR/vmlab/vmlabd.sock, spawns vmlabd if nothing answers, and retries for about ten seconds. This message means the spawned process never answered a ping. The log it names holds vmlabd's stdout and stderr. The common causes are a stale socket from a previous user session, a runtime directory owned by another user, a vmlab-host.wcl that fails to parse, and a fast-path probe that crashed rather than degraded. The per-lab daemon's log is .vmlab/lab.log inside the lab directory, and vmlab logs reads it.

Related messages from the same code path are spawning vmlabd, when the binary itself cannot be executed, and opening {log_path}, when the state directory is not writable. Files and directories lists every location.

The lab name is already registered

text
lab `{name}` is already registered from {root} — run `vmlab down` there (or `vmlab lab stop {name}` from anywhere) to release it, or rename this lab

A lab's declared name is its host-global runtime identity, not its directory (ADR-0011). The supervisor compares the requested lab root with the registered one before it hands out a socket or starts a daemon, and a different root is a conflict in every registry state, so the command exits 5. This is what two clones or worktrees of the same repository hit when both declare lab "demo".

status, down, destroy, console and dev sync refuse the same way, so the second directory can neither read nor stop the first directory's lab.

Release the name with a full vmlab down (or vmlab destroy) in the directory the message names, or with vmlab lab stop <name> from any directory, which is also the way out when that directory no longer exists. Both stop the lab's machines, keep its clones and have the supervisor reap its daemon. Or change the lab name in one of the two files.

To carry the provisioned machines across instead of building them again, run vmlab lab move <name> in the directory that should own the lab. It refuses while a machine runs, releases the name itself, moves the working data and registers the lab from there (see vmlab lab move). Pass --from <dir> when the name has already been released.

The eBPF fast path is unavailable

text
network fast path: userspace (mode auto)
  afxdp unavailable: creating a tap needs CAP_NET_ADMIN. In a container, add `--cap-add NET_ADMIN` (the fast path also needs `--cap-add BPF`)
  sockmap unavailable: not used in auto mode: af_unix kernel splicing measures slower than the userspace fabric (psock backlog workqueue); force with `fastpath = "sockmap"` to evaluate it

vmlab fastpath asks the supervisor which tier the fabric selected and why each higher tier was skipped. Nothing inspects kernel versions or capability bits: each tier is proved end to end over throwaway sockets at daemon start, and a tier that fails its probe is logged as fast-path tier {tier} unavailable: {reason} and skipped. A forced tier that fails degrades to userspace rather than stopping the daemon. The lab runs unchanged on the userspace fabric in every case, so none of these reasons is an error.

ReasonFix
… — run the vmlab daemons with CAP_BPF + CAP_NET_ADMIN to enableGrant both capabilities to the daemon binaries, or run them where they hold them.
the tun driver is not loaded. Run modprobe tun on the host (not inside the container) and restart vmlabAs it says. If the kernel was upgraded in place, the longer variant of this message tells you to reboot first.
/dev/net/tun is missing. In a container, pass --device /dev/net/tun; …Expose the device to the container that runs the daemons.
ignoring VMLAB_FASTPATH={value} (want auto|off|sockmap|afxdp)The environment override is misspelled. The fastpath field of the host configuration takes the same four values.

The sockmap tier is never chosen in auto mode because it measures slower than the userspace fabric; forcing it is for evaluation only, and a runtime failure there falls back with `sockmap offload unavailable ({error}); using userspace switching`.