Lab Kiosk OS & Edge SaaS

Troubleshooting

Symptoms, root causes, and fixes, grouped by where the problem shows up. Nearly every entry here is a mistake that was actually made.


Control plane

SymptomRoot causeFix
SUPER_ADMIN_EMAIL and SUPER_ADMIN_PASSWORD must both be setThere is no default super admin once a D1 binding existswrangler secret put both
The D1 database is missing the current schemaMigrations not appliedwrangler d1 migrations apply labkiosk-db --remote (or --local)
No D1 database bound on startupenv.DB missingBind DB in wrangler.jsonc, or set ALLOW_LOCAL_DB=1 for local dev and tests only
D1_EXEC_ERROR: incomplete inputMiniflare/workerd parses multiline SQL poorly with CRLFSplit on ;, replace(/\r\n/g, "\n"), run each via db.prepare(stmt).run()
Columns of "x" differ between SCHEMA_SQL and migrations/The two homes of the schema driftedChange both: a new file in migrations/ and the mirror in SCHEMA_SQL
Env vars reset after deployA vars block in wrangler.jsonc overrides dashboard settingsRemove it; manage production values in the dashboard / wrangler secret
Custom domain 404sCloudflare is not routing that hostname to the workerProxied CNAME into your zone, or a Cloudflare for SaaS custom hostname

Admin console UI

SymptomRoot causeFix
A button does nothing; console says "Refused to execute inline event handler"A template gained an on*= attribute; the nonce CSP blocks itdata-action + a delegated listener, or addEventListener
A script block silently does not runThe <script> lacks nonce="${escapeAttr(nonce)}"Pass the response nonce to the renderer and stamp it
Resetting Broadcast lands on the SaaS landing page instead of the organization portalresetBroadcastToPortal() sent origin + "/" without tenant scopingPOST /api/command resolves portalUrlFor(tenant) authoritatively
Workstations disagree about the active broadcastBroadcast state was in isolate memory, which differs per coloIt lives in tenants.broadcast_url / broadcast_epoch in D1
A Single-Site Lockdown URL is rejectedNo scheme, e.g. canvas.example.comNone needed — safeHttpUrl() prepends https://. If it is still rejected the host itself is malformed
A broadcast to selected workstations reverts to the portal after 2–3 secondsOnly a broadcast to "all" was stored; the next heartbeat's targetUrl sent the screens backFixed by per-workstation broadcast state (migration 0010); apply it and deploy the Worker together
/admin/teachers or /super/schools bookmarksRenamed to /admin/staff and /super/organizationsBoth old paths redirect; update the bookmark

Workstation enrolment

SymptomRoot causeFix
Enrolment rejected with a valid-looking keyThe organization's key is empty, or the organization is pending/suspendedGenerate a key in Settings; have a super admin approve the organization
Enrolment rejected after several triesFailed attempts from that address are throttledWait, then retry with the correct key
Workstation never appears on the dashboardNot enrolled, or its token was revokedRe-run the wizard with the current key. Check /tmp/lab-agent.log for 401
Freshly enrolled kiosk shows "This page is blocked"Chromium reads managed policy only at startupThe agent sets pendingBrowserRestart and restarts the browser after the next sync — wait for it
Enrolment forgotten after a rebootoverlayroot="tmpfs" sends every write to RAM, including /etc/labkiosk/config.jsonThe installer creates LABKIOSK_DATA and mounts it at /etc/labkiosk. On an older image, enrol from the live session before installing
Wizard shows the installer tab on an installed machine/api/status returns "isLive": True as a literal instead of the computed valueCosmetic — POST /api/install still refuses with 400. → Client Agent
The school address is not a valid server URL (or organization address) for http://<private IP>:8787Plain http was accepted only for localhost and the container gatewaysA private IPv4 (10/8, 172.16/12, 192.168/16) is now accepted for a local test server; anything public still needs https

Client OS and browser

SymptomRoot causeFix
Nav bar and lock curtain vanish; Chromium logs "Loading of unpacked extensions is disabled by the administrator"A blanket ExtensionInstallBlocklist: ["*"] in the managed policyRemove it. An ExtensionInstallAllowlist entry does not override it
A site is blocked that should be allowedNot on the effective allowlist, or policy not reloaded yetAdd the domain in Settings; the browser restarts on the next sync
A Chromium policy rule is silently ignoredAn invalid pattern like http://localhost:* or 127.0.0.1:*Use the bare host: localhost, 127.0.0.1, host.containers.internal. Omitting the port matches all ports
A policy key you added has vanishedIt was added to only one of the two consumers of the policy baseEdit usr/share/labkiosk/chromium-policy-base.json, never the generated file; verify with generate-chromium-policy.py --check
Black screen on bootquiet loglevel=3 suppressed boot logs and PAM autologin was lockedconsoleblank=0 instead, passwd -d kiosk, pre-seed live-config markers
Boots to a GRUB password prompt every timeset superusers present with no --unrestricted entry--unrestricted must stay unconditional in 02-security.hook.chroot and on every entry of usr/share/labkiosk/boot/grub.cfg
mute does nothingalsa-utils missing from the imageIt is in kiosk.list.chroot; a drifted variant image caused this once
Remote shutdown accepted but nothing happensThe polkit power rule is missing; the agent runs as kiosk/etc/polkit-1/rules.d/50-labkiosk-power.rules grants exactly reboot and power-off

Disk installer

SymptomRoot causeFix
rsync: delete_file: rmdir(boot/efi) failed: Device or resource busy (16)An older installer mounted the ESP at /boot/efi before rsync --delete ranInstall from a current ISO: the installer no longer copies a root filesystem with rsync, it copies the system image into the image store and mounts ROOT, ESP and DATA at three separate directories
A disk is not offered by the installerIt is under 7 GiB, the minimum that holds two system images (a drive sold as 8 GB qualifies)Use a larger disk
The installer reports a missing system image or grub.cfgIt was not started from the Lab Kiosk live medium, so live/{vmlinuz,initrd.img,filesystem.squashfs}, /usr/share/labkiosk/version or the grub.cfg template is missingBoot the ISO and install from there; nothing is erased before these are found
No candidate internal drives detectedThe installer printed a log line to sys.stdout, corrupting the JSON the agent parsesAll logging goes to file=sys.stderr; stdout is JSON only
The installer offers the USB it booted from--list-disks recorded removable but never filtered on the live mediumlive_medium_disks() excludes it when listing and again before wipefs
parted aborts on the ROOT or DATA partitionNegative offsets (-513MiB) parsed as bundled optionsPass -- before them
Legacy BIOS will not boot the installed GPT diskNo BIOS Boot Partition for GRUB to embed core.imgPartition 1: bios_grub, 1–2 MiB, set 1 bios_grub on
UEFI boot entry missing after a rebootFirmware lost NVRAM boot variablesgrub-install --target=x86_64-efi --removable also runs, creating /boot/efi/EFI/BOOT/BOOTX64.EFI
Every installed machine shares a machine ID/etc/machine-id was truncated to a newline, not to emptyIt must be a genuinely empty file — that is the marker systemd replaces. The lockdown hook now ships it empty in the image itself, since an installed disk boots that very squashfs
An installed workstation is back on its previous version after an updateThe new image got its one try and failed it: it did not boot (a panic reboots after 10 s), or the agent and the browser did not stay up for a minute within the 10-minute health window, so labkiosk-boot-ok.service rebooted into the old imageSettings → Errors & Warnings lists the failed or rolled-back image. The workstation is running the last good image; fix the image before giving it another try
An installed workstation booted an image other than current, reported as a fallbackcurrent in boot/grub/grubenv names an image that is missing or incomplete, so GRUB booted previous or any complete imageListed as a warning in Errors & Warnings. Reinstall, or mount LABKIOSK_ROOT from another system and check images/ and grubenv
An installed workstation stops at "Please remove the live-medium … press ENTER" on every rebootIt is booting the USB stick rather than its disk (the firmware boot order still prefers USB), or the disk's grub.cfg is not the shipped template and lacks noeject / labkiosk.installed=1Remove the stick and make the internal drive the first boot target. The installed menu must be usr/share/labkiosk/boot/grub.cfg verbatim
Enrolment and Wi-Fi forgotten after a reboot, and the boot log says "no valid labkiosk.data=<uuid> on the kernel command line"boot/grub/labkiosk-data.cfg on LABKIOSK_ROOT is missing or damaged, or the data partition was reformatted and its UUID changed. /etc/labkiosk is mounted only by that UUID, never by label, so it stays unmounted and persistentStorage is falseReinstall from a current ISO, or mount LABKIOSK_ROOT from another system and write set data_uuid="<uuid>" (from blkid of the LABKIOSK_DATA partition) to boot/grub/labkiosk-data.cfg
Installing fails with 'en-US' is not a language taglabkiosk-localization anchored its language-tag check with \\Z in a raw string, which no tag can matchFixed; rebuild the ISO

Remote control

SymptomRoot causeFix
"This workstation is not connected to the console right now"POST /api/clients/remote-session answered 409: the workstation holds no control-channel WebSocket (offline, rebooting, or on the HTTP heartbeat)Confirm it is online and enrolled, then try again
The viewer waits for the workstation to answer, then the session endsIts agent is too old to join the console's relay; with no workstation side, the session ends after 60 secondsUpdate the workstation to a current image
The viewer says the workstation asked for a VNC password the console does not have yetNo vncPassword reported since boot, or /tmp/labkiosk/vnc.secret missingWait a few seconds after boot and try again; check vncPassword in GET /api/clients
"Enrolment failed unexpectedly" in the wizardThe agent hit an error that is not a network or input problem — most often it could not write /etc/labkiosk/config.jsonThe message now names the exception, and the wizard opens Agent Log & Diagnostics by itself. Read it before rebooting: the log is in RAM
Need the agent log on a real workstationIt has no terminal, and file:// is blockedSetup wizard → Agent Log & Diagnostics (administrator password once installed)
The wizard says There is a typing mistake somewhere in this keyThe key's last character does not match the rest: one character is wrong or two are swapped. A key copied before keys carried a check character also failsCopy the key again from Settings → Workstation Enrollment Key (old keys were replaced there)
The organization address and enrollment key are gone after a reboot, and the wizard shows no warning/etc/labkiosk was an overlay on RAM rather than the data partition: overlayroot was configured with an overlayroot_options= line it never reads, so it defaulted to recurse=1 and overlaid every fstab entry. It is mounted and writable, which is why nothing complainedReinstall from an ISO built with overlayroot="tmpfs:recurse=0". The agent now reports the filesystem type, so this state shows up as persistentStorage: false and an amber warning
The organization address and enrollment key are gone after a reboot/etc/labkiosk is a directory in the RAM overlay rather than the LABKIOSK_DATA partition, so the enrolment was never on disk. The wizard now says so in amber before you type anything, and persistentStorage in GET /api/status reports itReinstall from a current ISO. The boot-time repair mounts the partition when the boot has not, and refuses to fabricate a directory that would lose the next enrolment too
PermissionError: [Errno 13] ... /etc/labkiosk/config.json.tmp when enrollingThe data partition at /etc/labkiosk is owned by root, so the unprivileged agent cannot write there. Disks written by an older installer show thisReinstall from a current ISO: the installer verifies the kiosk user can write to the partition, and labkiosk-data-permissions.service corrects the ownership at every boot. The error text names the owner, mode and the agent's uid
Black or sluggish remote screenBandwidth or thin-client CPUCheck hardware acceleration in firmware; noVNC adapts to latency

ISO build

SymptomRoot causeFix
Build dies in the chroot stageA rootless engine forbids mknod even under --privilegedMake the engine rootful: podman machine set --rootful. Verify with the mknod one-liner in Building the ISO
Your changes have no effect on the ISOA stale or pulled builder image was used; the source is copied into itdocker build -t ghcr.io/akbhoi/labkiosk-iso-builder distro-builder first
Enormous build context, or a resurrected quiet loglevel=3.dockerignore missing, dragging in chroot/, cache/, old ISOs, and a stale config/binaryKeep .dockerignore intact
Build fails on a pinA pinned checksum or hash is wrongFix the value. Never invent one to make the build go green
A shell hook dies with $'\r': command not foundThe file was written or checked out with CRLF.gitattributes pins these to eol=lf. Never write them with a tool that translates newlines — Python's Path.write_text does, on Windows
__pycache__ directories appear inside includes.chrootpy_compile ran without PYTHONPYCACHEPREFIXAlways set PYTHONPYCACHEPREFIX=/tmp/labkiosk-pyc — live-build copies whatever is on disk into the ISO

Simulator

SymptomRoot causeFix
Running as root without --no-sandbox is not supportedThe container was started as root instead of as its kiosk userRun it the documented way (docker compose up); the entrypoint falls back to --no-sandbox only for uid 0, and never in the real image
Check failed: sys_chroot("/proc/self/fdinfo/"), black screen in the simulatorThe container lacks SYS_CHROOT, which Chromium's sandbox needsKeep cap_add: [SYS_CHROOT] from docker-compose.yml
chrome_crashpad_handler: --database is required, black screen$HOME is not writable (read-only root filesystem)The entrypoint moves the browser's home to /tmp; check that /tmp is a writable tmpfs
Your agent or extension edits do nothingA pulled image runs main's baked-in client sourcedocker compose up -d --build, or docker cp the files in
Extension changes do not appear after a restartMV3 extensions are parsed at browser launchRestart Chromium, not the agent
Cannot reach the wizard from the host browserThe agent's API is loopback-only, by designDrive it from the noVNC screen

Diagnostic commands

# Worker
pnpm --prefix cloudflare-control run typecheck
pnpm --prefix cloudflare-control test

# Client syntax
PYTHONPYCACHEPREFIX=/tmp/labkiosk-pyc python3 -m py_compile \
  distro-builder/config/includes.chroot/opt/labkiosk/agent/agent.py \
  distro-builder/config/includes.chroot/usr/local/bin/labkiosk-install \
  distro-builder/config/includes.chroot/usr/local/sbin/labkiosk-localization
PYTHONPYCACHEPREFIX=/tmp/labkiosk-pyc python3 -m unittest discover -s distro-builder/tests -t distro-builder/tests
node --check distro-builder/config/includes.chroot/opt/labkiosk/extension/content.js
node --check distro-builder/config/includes.chroot/opt/labkiosk/extension/background.js
python3 distro-builder/tools/generate-chromium-policy.py --check

# Simulator
docker exec labkiosk-client-01 tail -n 50 /tmp/lab-agent.log
docker exec -e DISPLAY=:0 labkiosk-client-01 scrot -o /tmp/screen.png
docker cp labkiosk-client-01:/tmp/screen.png .

# A real workstation has no shell (every getty is masked, no SSH): read the agent log in the
# setup wizard's Agent Log & Diagnostics. The same checks inside the simulator:
docker exec labkiosk-client-01 cat /etc/chromium/policies/managed/policies.json

Still stuck

  1. Read the agent log — it is the single most informative artefact on the client side.
  2. Take a screenshot. A log line proving the agent ran does not prove the user saw anything.
  3. Check the organization's audit log for what actually changed and who changed it, and Settings → Errors & Warnings for failed boots and rollbacks reported by installed workstations.
  4. Search existing issues, then open one with the ISO or worker version, the platform, and the exact error.

Do not open a public issue for a security vulnerability. → Security Model

This page is wiki/Troubleshooting.md in the repository.