Files
agent-skills/skills/bastille-jail-provisioning/SKILL.md
T
MalinandClaude Sonnet 5 106e61cfbb docs: document NPM dedicated-port jail convention and zsh $path footgun
Two real gotchas from today's goddy->gringo Laravel migration: every jail
on an NPM-fronted host needs a dedicated listen port (8000+last-octet),
not just the shared port 80, or NPM cutover silently can't reach it even
though every direct/Host-header test passes. Separately, zsh ties a
variable literally named `path` to $PATH itself - overwriting it broke
every subsequent command in a config-editing script with no clear error.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-07 19:44:23 +02:00

9.9 KiB

name, description
name description
bastille-jail-provisioning Use when creating a new FreeBSD bastille jail (a new site/service on a shared-LAN jail host) via clone-from-known-good-base. Covers the IP-conflict gotcha that has caused real, time-costly incidents twice.

Bastille Jail Provisioning

Pattern for standing up a new jail on a FreeBSD host running Bastille, by cloning a known-working base jail (already has nginx+php-fpm+packages configured) rather than building from scratch.

CRITICAL FIRST STEP: arp-scan before picking an IP

Confirmed real incident, twice, on two different hosts on the same shared LAN: an IP address that looked free by every host-local check (bastille list all, ifconfig aliases, a manually-tracked "already used" list) turned out to already belong to a different physical device, or a jail on a different host sharing the same LAN segment. Symptom: the new jail's DNS resolution and outbound TCP connections silently fail or time out with no useful error — this looks exactly like a jail networking/pf config bug and can cost real time chasing the wrong theory before the actual cause (an IP conflict) is found.

Before assigning ANY candidate IP, always:

pkg install -y arp-scan   # if not already present
arp-scan --interface=vtnet0 <candidate-ip>

Zero responses = genuinely free. Any response (even from a device with no recognizable vendor match) means it's taken — pick a different IP and re-check. Do not rely on ping (misses devices with ICMP disabled/filtered — a real false-negative that happened once already) or on a per-host "what's already in use" list (each host only knows about its own jails, not other physical devices or other hosts' jails on the same shared subnet).

If you inherit a jail already provisioned on a conflicting IP: stop the jail, remove the bad alias, update jail.conf's ip4.addr, start the jail again (this re-adds the alias on the new IP), then propagate the IP change everywhere else it's referenced (DB user grants scoped by IP, host-level nginx proxy_pass target, etc.) — an IP conflict fixed in jail.conf alone but not in the DB grant/nginx config leaves the site broken even though the jail itself is now healthy.

Provisioning steps (once the IP is confirmed free)

  1. Clone from a known-good base jail rather than building from scratch — inherits working nginx/php-fpm/package configuration:
    bastille clone -a <known-good-base-jail> <new-jail-name> <new-ip>
    
  2. Wipe cloned content, do a fresh install if this is a brand-new site (not a migration) — the cloned base's site files are a template for the stack config, not content you want to keep.
  3. Create the database on the host's local MariaDB, with the grant scoped to the jail's own IP, not the host's IP:
    CREATE DATABASE `dbname` CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
    CREATE USER 'dbuser'@'<jail-ip>' IDENTIFIED BY '<random-password>';
    GRANT ALL PRIVILEGES ON `dbname`.* TO 'dbuser'@'<jail-ip>';
    
    The application's DB_HOST config value is the host's own LAN IP (where MariaDB actually listens), NOT the jail's IP — the jail IP is only used for the GRANT's scope. Confirmed gotcha across multiple real migrations: mixing these two up is an easy, non-obvious mistake.
  4. Wire the host-level nginx reverse-proxy vhost, proxy_passing to the jail's IP.
  5. Do NOT touch DNS or the reverse-proxy manager's public routing as part of automated provisioning — that's a deliberate, manual cutover step for a human to do once the site is verified working over a direct IP/Host-header test. Build and verify fully "dark" first.
  6. Smoke test both directly to the jail IP and via the host nginx proxy (with the right Host header and, if the site checks for HTTPS via X-Forwarded-Proto, that header too) before considering provisioning done.

Persist credentials immediately

Write the generated DB password (and any other generated secrets) to a durable location the moment they're generated — if a later provisioning step fails, an unrecorded generated password is otherwise lost with no way to recover it short of resetting it.

bastille create interface gotcha: don't assume bastille0 exists

Some hosts' bastille.conf still has the default bastille_network_loopback="bastille0" / bastille_network_vnet_type="if_bridge" values even though the host actually uses plain shared-IP aliasing on the physical interface (vtnet0) for every real jail on it — confirmed on the staging host (100.104.61.54) 2026-08-02. A bare bastille create <name> <release> <ip> fails with [ERROR]: bastille0 interface does not exist. Check an existing working jail's jail.conf (ip4.addr = vtnet0|<ip>;) to see which interface this host actually uses, then pass it explicitly:

bastille create <name> <release> <ip> vtnet0

After create/destroy churn, verify the IP alias actually attached

A bastille create immediately following a bastille destroy of a jail that held the same IP can silently fail to (re-)add the alias — the jail comes up "running" but has zero network connectivity (DNS and raw-IP fetch both time out, no error surfaced anywhere). Confirmed live 2026-08-02: pkg install inside the new jail failed with "Non-recoverable resolver failure" on the first attempt, which looks exactly like a DNS/resolv.conf problem but wasn't (resolv.conf was correct and identical to a working jail). Verify directly before assuming jail networking is up:

ifconfig <interface> | grep <jail-ip>   # must show the alias
bastille cmd <jail> fetch -o /dev/null http://1.1.1.1  # raw IP, bypasses DNS entirely

If the alias is missing, bastille restart <jail> (not just start) or a manual ifconfig <interface> inet <ip> netmask 255.255.255.255 alias reapplies it.

Testing a jail before its public vhost exists: hit the jail IP, not the host

If the site's domain isn't wired into the host's reverse-proxy nginx yet (no sites-enabled/<jail>.conf), curling the host's own 127.0.0.1 with a Host: header does NOT reach the new jail — it falls through to whichever server_name on the host's nginx matches first (often a different site's default_server), returning a normal-looking 200 OK for entirely the wrong site's content. This produced a long, wrong-track debugging session on 2026-08-02 (chasing a "cache never writes" theory against a site that was never actually being hit). Always curl the jail's own IP directly until the real vhost is wired:

curl -H 'Host: <domain>' http://<jail-ip>/

Only switch to testing via the host's 127.0.0.1 (or the public domain) once the sites-enabled vhost for this specific site actually exists.

Every jail needs a dedicated NPM-facing port, not just listen 80

Confirmed real gap, 2026-08-07: on jail hosts fronted by a shared Nginx Proxy Manager (NPM) instance, every existing jail's host-level nginx vhost has two listen directives — the shared port 80 (for same-host/Host-header testing) and a dedicated port unique to that jail, following the convention 8000 + <jail's last IP octet> (e.g. a jail at 10.20.0.43 gets listen 8043;, one at .50 gets listen 8050;). This dedicated port is what NPM's own proxy-host config actually targets when routing the public domain to this jail host — NPM does not rely on Host-header-based routing through the shared port 80 for this fleet, it connects to a specific port per site.

This was missed migrating 5 Laravel apps from a Linux host to a FreeBSD jail host (gringo) in one session: all 5 new vhosts were created with only listen 80;, matching the shared port but missing the dedicated one entirely. Every functional test the delegate ran (raw jail IP, and via the shared port 80 with the right Host: header) passed cleanly — the gap was invisible until specifically checked against sibling jails' configs, at which point every other jail on the host turned out to follow the two-port pattern without exception. Before considering any new jail's nginx vhost complete, diff its listen directives against at least one working sibling jail's config on the same host — don't just confirm the site loads via a manual test, confirm the routing surface matches the established convention, since NPM (external, not scriptable from this sandbox) is what will actually determine reachability once cut over, and a working manual test doesn't prove NPM's real path works.

# find the convention on any host that uses it
for f in /usr/local/etc/nginx/sites-available/*.conf; do
  echo "=== $(basename $f) ==="; grep -E '^\s*listen' "$f"
done

If a host doesn't show this pattern on ANY existing jail, it likely doesn't use per-jail dedicated ports — don't assume the convention applies fleet-wide without checking the specific host first.

Stock stack includes Valkey, not just nginx/php/MariaDB

The fleet's standard per-host stack (see docs/server-funky.md) is nginx + php-fpm + MariaDB (host-local) + Valkey (host-local, ACL-auth'd, aclfile-based, admin password at /root/.valkey_admin_pw). A host that's never hosted a Redis/object-cache-dependent site before may be missing Valkey entirely — install and configure it the same way (bind 127.0.0.1 <host-lan-ip>, aclfile, admin password file) rather than treating it as optional, so any plugin/site expecting a persistent object cache gets a real one to test against, not silently falls back to a weaker default. Inside a jail that needs to reach it, install the PHP Redis extension — package name is php85-pecl-redis (not php85-redis, which doesn't exist under that name in the FreeBSD ports tree) — and remember Valkey's ACL means any code connecting to it needs an explicit AUTH step, not just connect(); a connect()-only implementation will silently fail every subsequent command with NOAUTH against this fleet's standard Valkey setup.