From eeb08a0d77308c6a7f8096a682f361bef6de5d8a Mon Sep 17 00:00:00 2001 From: Malin Date: Sun, 2 Aug 2026 21:39:39 +0200 Subject: [PATCH] feat: capture staging-verification lessons from iWP Cache deployment New skill wordpress-plugin-staging-verification: use the persistent staging-1 jail (real-world plugin set left active) instead of a disposable clean-room jail, full verification checklist, and the cosmetic-proc_open-error-vs-real-fatal distinction. Extended bastille-jail-provisioning: bastille0-loopback-missing gotcha (pass the interface explicitly), IP-alias-can-silently-fail-after- recreate gotcha, host-reverse-proxy-vs-jail-IP curl mixup, Valkey as part of the stock stack (correct FreeBSD package name, ACL auth requirement). Extended wordpress-plugin-conventions: drop-in source files (object-cache.php/advanced-cache.php logic) must be excluded from any glob-based plugin autoloader, or they redeclare WP core's own wp_cache_*() functions and fatal -- real bug found and fixed live during iWP Cache's first staging activation attempt. Co-Authored-By: Claude Sonnet 5 --- README.md | 4 + skills/bastille-jail-provisioning/SKILL.md | 66 ++++++++++++++ skills/wordpress-plugin-conventions/SKILL.md | 30 ++++++ .../SKILL.md | 91 +++++++++++++++++++ 4 files changed, 191 insertions(+) create mode 100644 skills/wordpress-plugin-staging-verification/SKILL.md diff --git a/README.md b/README.md index cc9a276..c300ca7 100644 --- a/README.md +++ b/README.md @@ -54,6 +54,10 @@ than assuming the delegate can fetch it itself. - `parallel-delegate-shared-files` -- never have 2+ parallel delegates edit the same bootstrap/wiring file; use self-registration instead. Real incident from the iWP Cache build (lost updates, twice). +- `wordpress-plugin-staging-verification` -- use the persistent + `staging-1` jail (real-world plugin set left active) for verifying new + plugins, not a disposable clean-room jail; includes the full + cache-plugin verification checklist and cosmetic-vs-real error gotchas. ## Provenance diff --git a/skills/bastille-jail-provisioning/SKILL.md b/skills/bastille-jail-provisioning/SKILL.md index 629dd5f..a0b3c11 100644 --- a/skills/bastille-jail-provisioning/SKILL.md +++ b/skills/bastille-jail-provisioning/SKILL.md @@ -80,3 +80,69 @@ Write the generated DB password (and any other generated secrets) to a durable location the moment they're generated — if a later provisioning step fails, an unrecorded generated password is otherwise lost with no way to recover it short of resetting it. + +## `bastille create` interface gotcha: don't assume `bastille0` exists + +Some hosts' `bastille.conf` still has the default +`bastille_network_loopback="bastille0"` / `bastille_network_vnet_type="if_bridge"` +values even though the host actually uses plain shared-IP aliasing on the +physical interface (`vtnet0`) for every real jail on it — confirmed on the +`staging` host (100.104.61.54) 2026-08-02. A bare `bastille create + ` fails with `[ERROR]: bastille0 interface does not exist`. +**Check an existing working jail's `jail.conf`** (`ip4.addr = vtnet0|;`) +to see which interface this host actually uses, then pass it explicitly: +```bash +bastille create vtnet0 +``` + +## After create/destroy churn, verify the IP alias actually attached + +A `bastille create` immediately following a `bastille destroy` of a jail +that held the *same* IP can silently fail to (re-)add the alias — the jail +comes up "running" but has zero network connectivity (DNS and raw-IP fetch +both time out, no error surfaced anywhere). Confirmed live 2026-08-02: `pkg +install` inside the new jail failed with "Non-recoverable resolver failure" +on the first attempt, which looks exactly like a DNS/resolv.conf problem +but wasn't (resolv.conf was correct and identical to a working jail). +**Verify directly** before assuming jail networking is up: +```bash +ifconfig | grep # must show the alias +bastille cmd fetch -o /dev/null http://1.1.1.1 # raw IP, bypasses DNS entirely +``` +If the alias is missing, `bastille restart ` (not just `start`) or a +manual `ifconfig inet netmask 255.255.255.255 alias` +reapplies it. + +## Testing a jail before its public vhost exists: hit the jail IP, not the host + +If the site's domain isn't wired into the host's reverse-proxy nginx yet +(no `sites-enabled/.conf`), curling the **host's** own `127.0.0.1` +with a `Host:` header does NOT reach the new jail — it falls through to +whichever `server_name` on the host's nginx matches first (often a +different site's `default_server`), returning a normal-looking `200 OK` +for entirely the wrong site's content. This produced a long, wrong-track +debugging session on 2026-08-02 (chasing a "cache never writes" theory +against a site that was never actually being hit). **Always curl the +jail's own IP directly** until the real vhost is wired: +```bash +curl -H 'Host: ' http:/// +``` +Only switch to testing via the host's `127.0.0.1` (or the public domain) +once the `sites-enabled` vhost for this specific site actually exists. + +## Stock stack includes Valkey, not just nginx/php/MariaDB + +The fleet's standard per-host stack (see `docs/server-funky.md`) is +nginx + php-fpm + MariaDB (host-local) + **Valkey** (host-local, ACL-auth'd, +`aclfile`-based, admin password at `/root/.valkey_admin_pw`). A host that's +never hosted a Redis/object-cache-dependent site before may be missing +Valkey entirely — install and configure it the same way (`bind 127.0.0.1 +`, `aclfile`, admin password file) rather than treating it as +optional, so any plugin/site expecting a persistent object cache gets a +real one to test against, not silently falls back to a weaker default. +Inside a jail that needs to reach it, install the PHP Redis extension — +**package name is `php85-pecl-redis`** (not `php85-redis`, which doesn't +exist under that name in the FreeBSD ports tree) — and remember Valkey's +ACL means any code connecting to it needs an explicit `AUTH` step, not just +`connect()`; a `connect()`-only implementation will silently fail every +subsequent command with `NOAUTH` against this fleet's standard Valkey setup. diff --git a/skills/wordpress-plugin-conventions/SKILL.md b/skills/wordpress-plugin-conventions/SKILL.md index 7c174bf..338d4ae 100644 --- a/skills/wordpress-plugin-conventions/SKILL.md +++ b/skills/wordpress-plugin-conventions/SKILL.md @@ -112,6 +112,36 @@ needs `str_contains`/`str_starts_with`/`str_ends_with` (PHP 8.0+) on an environment that might run 7.4, guard with `function_exists()` polyfills rather than assuming they exist. +## Drop-in source files (`advanced-cache.php`, `object-cache.php`) must be excluded from any normal plugin autoload/glob + +If a plugin ships an `object-cache.php` or `advanced-cache.php` drop-in +(installed by copying a stub into `wp-content/` on activation), the PHP +file containing the *real logic* for that drop-in must **never** also be +loaded as a normal plugin include — not via an explicit `require`, and +not via a glob-based auto-loader that doesn't know to skip it. WordPress +core's own `wp-includes/cache.php` declares the same global `wp_cache_*()` +function names as the fallback used when no `object-cache.php` drop-in is +active; if the plugin's own copy of those functions loads a second time in +the same request (as a normal plugin file, in addition to — or instead +of — the standalone drop-in load), PHP fatals with "Cannot redeclare +function". Confirmed live during iWP Cache staging verification +(2026-08-02): a glob-based module auto-loader's exclusion list was written +before the object-cache class file existed, so it got swept in and +silently broke every activation. **If a plugin has any glob-based +autoload for its `includes/` directory, explicitly exclude every +drop-in-logic file by name** — don't rely on "it'll only load once" being +obviously true just because the code looks like a normal class file. + +## `object-cache.php`/`advanced-cache.php` code runs standalone, before ABSPATH-based guards mean what they normally mean + +The usual `if (!defined('ABSPATH')) exit;` guard doesn't prevent a +drop-in-logic file from loading twice in the same request, because by the +time it's `require`'d a second time (as a stray plugin include), ABSPATH +*is* already defined — WordPress has fully booted. A guard meant to stop +"direct access over HTTP" does nothing to stop "accidentally required +twice from two different code paths." Don't assume that guard is doing +more than it actually does. + ## Don't build what WordPress core already gives you Before writing custom code for: cron scheduling (`wp_schedule_event`), diff --git a/skills/wordpress-plugin-staging-verification/SKILL.md b/skills/wordpress-plugin-staging-verification/SKILL.md new file mode 100644 index 0000000..11048e8 --- /dev/null +++ b/skills/wordpress-plugin-staging-verification/SKILL.md @@ -0,0 +1,91 @@ +--- +name: wordpress-plugin-staging-verification +description: Use when verifying a new or updated first-party WordPress plugin before it ships. Persistent staging jail + real-world plugin set, not a disposable clean-room jail. +license: MIT +source: original, from the iWP Cache staging verification (2026-08-02) +--- + +# WordPress Plugin Staging Verification + +## Use the persistent staging jail, not a throwaway one + +`staging-1` on the `staging` host (Tailscale/SSH: `100.104.61.54`, jail IP +`192.168.0.187`, intended domain `staging.palmasolutions.net`) is a +standing verification environment — deploy new/updated plugins there and +**leave them installed and active** rather than spinning up a fresh +disposable jail per plugin and tearing it down after. + +**Why:** a clean-room jail with only the plugin under test installed +proves the plugin works in isolation, but says nothing about how it +behaves alongside the plugins real sites actually run. `staging-1` ships +with a representative baseline already active: WooCommerce, Classic +Editor, Contact Form 7, Yoast SEO (`wordpress-seo` is the correct wp.org +slug — `yoast-seo` doesn't exist), Wordfence, LiteSpeed Cache (a +*competing* cache plugin — deliberately, since drop-in conflicts are +exactly the kind of thing worth surfacing before a caching plugin ships). +Each new plugin verified here adds to that baseline instead of replacing +it, so the set of "things this plugin has been proven compatible with" +only grows over time. + +## Setup, once per jail (already done for `staging-1`) + +Standard stock stack per [[bastille_jail_provisioning]] (nginx, php-fpm, +MariaDB-client pointed at the host's local MariaDB, Valkey + +`php85-pecl-redis` for object-cache testing), plus WordPress installed +with **pretty permalinks set immediately** (`wp rewrite structure +"/%postname%/"` + `wp rewrite flush --hard`) — don't leave a fresh install +on Plain permalinks, see the informatiq.es incident in +`wordpress-cli-remote-execution`. + +## Verification checklist (adapt per plugin type; this is the cache-plugin version) + +1. Deploy the plugin, activate via `wp plugin activate `. If + activation fails, get the **real** error with `--debug`, not just the + truncated default output — the default output can look like a generic + sandbox/proc_open artifact (see next section) when it's actually a real + fatal. +2. Confirm any drop-in files (`advanced-cache.php`, `object-cache.php`) + actually landed in `wp-content/` with the plugin's own marker comment, + not silently skipped. +3. Cache hit/miss: purge, then two requests — first should be a MISS + (falls through to normal render), second a HIT (served from the + plugin's own header/marker). **Curl the jail's IP directly**, not the + host's own `127.0.0.1` — see the host-vs-jail routing gotcha below. +4. Logged-in-user bypass: send a request with a fake + `wordpress_logged_in_*` cookie, confirm it does NOT get a cache HIT. +5. Purge-on-publish: HIT the homepage, publish a new post via `wp post + create`, re-request the homepage, confirm it's no longer a HIT. +6. Object cache (if applicable): toggle it on, confirm the site stays + healthy (no fatals in `wp-content/debug.log` with `WP_DEBUG_LOG` + enabled), and confirm data is actually landing in the real backend — + for Redis/Valkey, `valkey-cli KEYS '*'` on the host and look for the + plugin's own key prefix, don't just trust "no fatal error" as proof + the persistent driver is actually being used instead of a silent + fallback. +7. Check `wp-content/debug.log` for anything beyond the plugin's own + expected output — a real PHP deprecation warning from the plugin's own + code (not a third-party plugin) is worth fixing even if it isn't fatal + yet, since deprecations become hard errors in later PHP versions. + +## Gotcha: cosmetic `proc_open`/`posix_spawn` errors vs. real fatals + +Some `wp-cli` write commands run inside this sandbox (`bastille cmd ... su +-m www -c '...'`) print a `proc_open(): posix_spawn() failed: Permission +denied` stack trace and report `[ERROR]`, **even when the underlying +database write actually succeeded**. Confirmed repeatedly (permalink +structure changes, plugin activation attempts). Don't treat this error +text alone as proof of failure — re-check the actual state directly (`wp +option get ...`, `wp plugin list`, look for the file/row that should have +been created) before concluding something didn't work. Conversely, don't +assume every failure is this cosmetic issue either — get the `--debug` +output and read past the proc_open noise for a *different*, real stack +trace underneath (this is exactly how the object-cache.php redeclare bug +in `wordpress-plugin-conventions` was actually found). + +## Gotcha: testing before the public vhost exists + +See [[bastille_jail_provisioning]]'s "curl the jail IP, not the host" +section — this bit hard during iWP Cache verification (a long stretch of +"why isn't the cache ever writing" chased against a completely different +site's WordPress install because `curl -H 'Host: staging....' 127.0.0.1` +on the host silently served Galex's site instead).