Lab Kiosk OS & Edge SaaS

Security Model

What Lab Kiosk defends, how, and what it explicitly does not defend. A security page that claims completeness is worse than one that states its limits.


Threat model

AdversaryGoalPrimary defence
A curious user at the keyboardEscape the kiosk into a shell or an unapproved siteLayered lockdown → Kiosk Hardening
A hostile website the user visitsReach the local agent, remove the nav bar, or evade the lock curtainExtension architecture, closed Shadow DOM, loopback origin checks
Another organization's administratorRead or control this organization's fleetTenant scoping on every query and every route
An unauthenticated internet callerPost telemetry, drain a command queue, read an allowlistDevice bearer tokens, guards on every route
A cross-site attackerMake a signed-in operator's browser act on their behalfOrigin guard on cookie-authenticated mutations
Someone with physical access to a workstationBoot something else, or edit the kernel command lineBoot-menu password, firmware password. → limits

Control plane

Authentication

PropertyValue
AlgorithmPBKDF2-HMAC-SHA256 via crypto.subtle
Iterations100 000
Salt32 cryptographically random bytes per user
Derived key256 bits, hex-encoded
Session token32 random bytes, hex; only its SHA-256 is stored
CookieHttpOnly; Secure; SameSite=Lax, scoped to the parent domain
ComparisontimingSafeEqual()

There are zero runtime npm dependencies. No auth framework, no routing library, no ORM. Every primitive is a Web API, which removes the supply-chain surface entirely and keeps cold start under 10 ms.

Two-step sign-in. After the password, a super admin is always emailed a six-digit code (it works for that sign-in only, for ten minutes, five wrong tries). An organization account turns the same on under Settings → Two-factor sign-in; every account can add an authenticator app (TOTP) with ten recovery codes on top. A browser can be trusted for 30 days. Where email is not configured, an account with no app signs in with its password alone and the server logs why, rather than locking every account out. SUPER_ADMIN_EMAIL must therefore reach an inbox that can be read without signing in to the console.

Password changes go through one path — POST /api/auth/change-password — which verifies the current password and revokes the account's other sessions, so a stolen cookie does not outlive a password change.

Tenant isolation

Two rules, both asserted by the test suite:

  1. **Every query touching devices, commands, sessions, or portal apps filters by tenant_id.**
  2. Every route carries a guard. A route without one is a security defect, not an oversight.
GuardEnforces
resolveTenant()The organization comes from the Host header, authoritatively
requireTenantAdmin()The session owns this tenant — 401 anonymous, 403 wrong organization
requireSuperAdmin()Platform-level routes
requireDevice()A valid, unrevoked device token
rejectCrossSiteMutation()Cookie-authenticated POST/DELETE prove their origin

X-Forwarded-Host is never read. Only the Host header says where a request arrived; trusting a caller-supplied header to build a URL handed back to a workstation would let an attacker redirect a whole lab.

?tenant= and X-Tenant are honoured only on a dev host, for a super admin, for a session that already owns that tenant, or on an explicitly public route.

Device identity

A workstation's identity comes from its bearer token and nothing else. /api/telemetry **ignores any clientId or tenant the payload claims**.

Tokens are issued by exchanging the organization's enrollment key, stored only as SHA-256 hashes, and revoked instantly by decommissioning. Before device tokens existed (migration 0002), anyone who guessed a subdomain could post screenshots and drain that organization's command queue.

Output escaping

Tenant data is attacker-controlled from the platform's perspective: organization names, admin emails, portal card titles, and URLs all arrive through registration or the admin console.

ContextRequired
HTML textescapeHtml()
HTML attributeescapeAttr()
Inside a <script> blockescapeJson()
Navigable or href URLsafeHttpUrl()
Client-side DOMBuild nodes and assign textContent. Never concatenate into innerHTML.

escapeJson() also escapes U+2028 and U+2029, which are valid JSON but terminate a JavaScript line.

Content Security Policy

Every HTML response is built with buildHtmlHeaders(nonce, …):

Content-Security-Policy: … 'nonce-<random>' …
Strict-Transport-Security: …            (HTTPS only)
X-Frame-Options: DENY                   + frame-ancestors 'none'
Permissions-Policy: …
Cross-Origin-Opener-Policy: same-origin
X-Content-Type-Options: nosniff
Cache-Control: no-store                 (API responses)

No inline event handlers anywhere. No onclick=, onsubmit=, onmouseover=. Use a data-action attribute with a delegated listener. The CSP blocks inline handlers, so one silently breaks the feature — the test suite renders every page and fails if any script lacks a nonce or any on*= attribute appears.

Rate limiting

  • Sign-in: exponential back-off per identifier via login_attempts, answered 429.
  • Registration and failed enrolment: throttled per source address.
  • Reserved slugs: www, super, labkiosk, api, admin, portal, status, mail, app, kiosk, root — neither registerable nor resolvable.

Fail closed

ConditionBehaviour
No D1 bindinggetDatabase() throws, unless ALLOW_LOCAL_DB=1 (tests and local dev only)
Super-admin secrets unsetbootstrap() refuses to serve
Migrations unappliedassertSchemaCurrent() refuses to serve; tables are never created at runtime
A build pin unset or wrongWarns, or fails the build
Automatic bug reports without the AI binding, GITHUB_ISSUES_TOKEN and GITHUB_ISSUES_REPOUnavailable to every organization

Missing configuration is an error, never a reason to fall back to something weaker.

Automatic bug reports

Off for every organization until one of its administrators accepts the Automatic Bug Report Terms (/terms/bug-reports) and turns them on in Settings → Errors & Warnings; a new terms version pauses reports until it is accepted again. What leaves the platform is the kind of problem, the image version and the problem text with addresses, host names, e-mail addresses and identifiers masked — never the organization, the workstation or anything a person typed. At most five GitHub writes per hourly run. → docs/DEPLOYMENT.md


Client workstation

→ Kiosk Hardening for the full detail. In summary:

LayerDefends against
overlayroot="tmpfs"Persistence of anything — malware, cached credentials, saved files
Locked root, masked gettys, no SSHAny route to a shell
DontVTSwitch, DontZap, empty Openbox keybindingsEscaping the X session
sysctl hardening, no core dumpsKernel information disclosure
Chromium deny-all URLBlocklist + allowlistUnapproved browsing
AllowFileSelectionDialogs: falseA file-picker used as a file manager
Closed Shadow DOM + capture-phase event swallowingA page tampering with the kiosk UI or evading the curtain
GRUB --unrestricted + install-time passwordEditing the kernel command line
One-try boot of a new image, confirmed by labkiosk-boot-slots (root only, no sudo rule)A bad system image stranding a workstation: it falls back to the old one
Data partition mounted by UUID (labkiosk.data=), never by labelA USB stick labelled LABKIOSK_DATA standing in for the machine's own data

The loopback boundary

The agent binds only to 127.0.0.1:8888 and every request must satisfy both a loopback Host check and a loopback Origin check. The single exception is the kiosk extension's own origin (chrome-extension://hfjmbeplebjipenkfabncgkpadnjmmoe, pinned by the key in manifest.json): Chromium attaches it to the service worker's POST to /api/admin/verify. A web page cannot set Origin, and no other extension can hold that id.

The extension's service worker owns the host_permissions grant for that origin, so content.js never fetches the agent directly. This is why the agent can refuse cross-origin callers outright — it used to answer Access-Control-Allow-Origin: *, which meant any site a user visited could talk to it.

Mutating endpoints (/api/install, /api/reboot, /api/setup) re-validate their inputs inside the agent and inside the installer, because the sudoers rule lets kiosk invoke the installer directly. The agent is not a trust boundary.


Known limits

What is not defended

Physical disassembly. Anyone who can remove the drive can read it. Nothing on it is secret except an enrolment token, which is revocable from the dashboard in one click.

A compromised control plane. The client trusts its control plane by design. safe_navigable_url() restricts navigation to http(s), so a hostile response cannot inject javascript: or file: — but a compromised control plane can point a lab at anything the allowlist permits.

A stolen operator session. Remote Control is reached only through the console's relay, so what protects a live desktop is the console sign-in, the workstations permission and the one-time session token — not the 8-character RFB secret, which is about 32 bits, capped by the RFB protocol. Whoever holds an operator session with that permission can open a workstation's desktop. → Remote Control

Screen content in transit and at rest. Thumbnails are base64 JPEGs stored in D1 and served to authenticated operators over HTTPS. They are not end-to-end encrypted. Anyone with database access can see the latest frame from every workstation.

Boot reports are the workstation's word. POST /api/devices/boot-report takes what an enrolled device says about its last boot. The agent reads it from a root-owned file in /run that the kiosk user cannot write, and the Worker validates versions and timestamps and records a report only when it is at least a minute newer than the last one it holds — but a compromised workstation can still report a false outcome about itself.

A malicious operator. An organization admin can broadcast anything, read every screen, and take remote control. That is the product working as intended; the audit log records it, it does not prevent it.

Deliberate choices that look like gaps

ChoiceReason
kiosk has an empty, not locked, passwordA locked account deadlocked nodm's PAM stack into a black screen. No login path exists to use it — every getty is masked and no SSH server is installed.
No blanket ExtensionInstallBlocklistIt makes Chromium refuse --load-extension entirely, silently removing the kiosk's own nav bar and lock curtain.
Camera and microphone are not blockedLanguage labs and video pages need them.
grub.pin is empty in the repositoryA committed hash is one password shared by every customer, unrotatable and permanent in git history.
--unrestricted on every GRUB entry, unconditionallyOtherwise an installed disk gets set superusers with no unrestricted entry, and every workstation stops at a password prompt on every boot.

Reporting a vulnerability

Do not open a public GitHub issue. Use GitHub Security Advisories: the repository's Security tab → Report a vulnerability.

Particularly wanted:

  • Kiosk breakout — escaping the locked Chromium session into a shell or the Openbox desktop.
  • Tenant isolation bypass — reading or writing across organization boundaries in D1 or an organization's OrgHub.
  • Authentication bypass — flaws in the PBKDF2 implementation, session token generation, or cookie handling.
  • Remote code execution in agent.py or its loopback API.
  • Device impersonation — posting telemetry, draining a command queue, or reading an allowlist without a token issued through enrolment.
  • Supervision suppression — a visited page removing or disabling the nav bar or lock curtain, or keeping a locked workstation usable.

Response targets: acknowledgement within 48 hours, triage within 5 business days, critical patches within 14 days with an advisory.


Hardening recommendations for organizations

  1. Firmware password on every workstation, with USB and network booting disabled.
  2. Boot-menu password set at install time, unique per site.
  3. **Grant workstations sparingly** — it is the permission that opens Remote Control.
  4. User VLAN isolated from administrative networks.
  5. Rotate the enrollment key when it has been shared outside IT staff, and when a technician leaves.
  6. Review the audit log periodically — it records every command and settings change with the acting user.
  7. Decommission promptly. A retired workstation with a live token is a valid telemetry source until it is revoked.

→ Kiosk Hardening · Control Plane Internals · Remote Control

This page is wiki/Security-Model.md in the repository.