A release-readiness QA pass over the whole product. The commits split into
defects a user would hit and gates that were reporting green while measuring
nothing.
## Fixes that change behaviour
Rate limiting was bypassable on every install: TRUST_PROXY defaulted to true, so
request.ip came from a client-set header and a forged X-Forwarded-For got past
the login limiter. The default is now a private-network trust list.
A transient Postgres outage stranded in-flight jobs, leaving finished output on
disk with no row pointing at it. A reconciler now resolves those rows and adopts
the bytes rather than dropping the work.
A Redis connection that moved to a new address wedged every read-blocked
consumer, so completions stopped signalling while health still answered 200.
Socket timeouts plus subscriber pings recover it.
Installing more than one AI bundle left the shared venv multi-versioned and
silently broke three tools. The installer now reconciles distributions to one
version each.
Converting an image to JXL at quality 1 through 4 returned a 500, because
libjxl 0.7 rejects the distance those values compute. The quality is floored at
what the encoder honours. A missing ffmpeg was also reported to the user as a
corrupt upload; it now says the engine is unavailable.
RAW uploads reached an unpatched LibRaw on arm64, so it is built from source at
0.22.2, and the release scan was split so it can fail on an unfixed critical
instead of hiding it behind ignore-unfixed.
## Gates that could not fail
Two mutation lanes ran zero mutants because Stryker crawled the gitignored docs
build; coverage discarded its whole report on any failing test; the lint gate
skipped root tests, scripts, and two workspaces; and several generated matrices
counted a host missing ffmpeg as a passing tool. Each now measures what it
claims.
Full evidence and the outstanding release items are tracked locally and are not
part of this branch.
Fixes 15 defects found by a max-effort multi-agent review of the last 6
merged PRs (#388, #390, #391, #392, #393, #394), all adversarially
verified before fixing.
Install queue + dispatcher (the serious cluster):
- features.ts: finalize the installer child exactly once. A failed spawn
fires both "error" and "close", and the second event released the file
lock and active slot that pump() had just handed to the next queued
bundle, letting two pip processes write the same venv concurrently.
Outcome recording now happens before pump() so the next bundle's first
progress frame cannot race the previous install's bookkeeping.
- feature-status.ts: keep failed-install errors in a per-bundle map
instead of the single progress slot. With the queue auto-starting the
next install, the slot was overwritten within seconds and a failed
install vanished without ever surfacing to GET /features.
- bridge.ts: scope child lifecycle per process (stopped-children set +
request generation tags) instead of an instance-wide shuttingDown flag
that the next spawn reset. A stale SIGTERMed child's late close event
could record a phantom crash (5 of which permanently disable the
dispatcher), null out the freshly spawned child, and reject the new
child's pending requests. The request-timeout kill path still counts
as a real crash.
- install_feature.py: the pre-write disk re-check measured ai_dir's
filesystem even when budgeting the cross-filesystem copy that lands on
the venv's disk; now each budget is checked against the filesystem the
bytes actually land on, so ENOSPC cannot strike mid-write and leave
site-packages half overwritten.
Behavior regressions:
- embed-subtitles: preserve pre-existing subtitle tracks (0:s?) and MKV
attachments (0:t?) that the -map 0:v:0/0:a? rewrite silently dropped;
data streams stay unmapped on purpose (the actual MPEG remux fix). The
new subtitle maps first so the language tag hits the right stream.
- usage-survey-overlay: fail closed when the settings fetch fails; the
fail-open path rendered the blocking survey against an unhealthy API
and soft-locked admins, the lock-out class #392 fixed.
- features-store: queued bundles poll instead of each holding an SSE
connection (Install All could pin 7 EventSources and exhaust the
browser's 6-per-origin HTTP/1.1 limit, hanging the whole app);
listenToProgress closes any prior stream and stops any poll before
subscribing; installAll skips bundles already installing or queued.
Contracts, tests, i18n:
- openapi.yaml: add "queued" to the features status enum and document
downloadBytes/installedBytes (Schemathesis conformance).
- feature-lifecycle e2e: queue transcription (~0.5 GB) instead of ocr
(~6 GB) and give the test a budget that covers both install drains
(the stacked waits exceeded the old 900s timeout).
- docker-compose.qa.yml: parameterize the host port (QA_APP_PORT) so
QA_PROJECT_NAME concurrent stacks can actually bind.
- compare + watermark-image: restore per-input error attribution
("Invalid first/second image", "Invalid watermark image") lost in the
shared-handler migration.
- ai-features-section: the "{size} on disk" suffix now goes through
i18n; key added to all 21 locales.
- watermark-image + content-aware-resize: migrate to the shared
inputHandlerFor("image") chain like compare/vectorize/compose, fixing
drift in the inline copies (no SVG sanitize, no RAW extension hint,
no AVIF probe).
Verified: typecheck across 9 workspaces, Biome clean on all changed
files, 584 targeted unit tests and 249 integration tests green
(including real-ffmpeg embed-subtitles runs). One unit test updated to
the new poll-while-queued contract with a single-EventSource assertion.
Claude-Session: https://claude.ai/code/session_017mR1HiHaf3a1BmUtrHX4j3
Found and fixed during a full local Docker build validation (amd64/arm64, all
four fleet targets, AI bundle installs, QA harness) and the follow-up bug
sweep requested afterward. None of the affected scripts run in CI, so these
had been silently broken indefinitely.
- docker/feature-manifest.json: pythonVersion was a flat "3.11", but the
amd64 base (Ubuntu 24.04) ships Python 3.12 while arm64 (Debian bookworm)
ships 3.11. Changed to a per-arch object matching the file's existing
convention.
- tests/qa/api-sweep.mts and verify-ai.mts: bare "@snapotter/shared" import
can't resolve since tests/ is not a pnpm workspace member, making both
silently unrunnable via their own documented command on any fresh
checkout. Switched to a relative import.
- tests/qa/generate-ledger.mts: wrote to docs/qa/ without creating the
directory first; docs/ is gitignored except COMMUNITY_GUIDE.md, so a fresh
checkout threw ENOENT.
- Seven QA Playwright spec files (input-preview, settings,
settings-extended, multifile, output-preview, pipeline-ui, smoke) had
~115 fixture() calls using directory names that don't exist. Resolved
every call programmatically against the real fixture tree.
- packages/ai/src/bridge.ts: AI dispatcher restart (happens on every bundle
install) was falsely counted as a crash, risking permanent dispatcher
disable after enough legitimate restarts within the crash window. Added a
shuttingDown flag checked at all three recordCrash() call sites.
- packages/image-engine/src/operations/auto-enhance.ts: image-enhancement
hung 40+ seconds on large RAW photos (confirmed on a real 20.2MP file) in
Sharp's .clahe() step, whose cost scales with total pixel count regardless
of tile size. Added a 16-megapixel cap above which CLAHE is skipped;
verified against the real file (40+s -> 2.0s) with no regression to other
RAW formats or normal-sized images. Fixing this surfaced a second,
smaller bug where the saturation step's CLAHE compensation boost was
keyed off the raw toggle instead of whether CLAHE actually ran.
- Two QA-harness robustness gaps closed per "fix everything, even the small
bugs": the passport-photo/erase-object input-preview tests now skip
cleanly with a clear reason on a container without their AI bundle
installed, and docker-compose.qa.yml's hardcoded project/container name
(the actual root cause of a mid-validation container swap between two
concurrent sessions) is now parameterized via QA_PROJECT_NAME.
Full validation report is local-only per repo convention.
The binary search found the right quality but sharp(buffer).toBuffer()
re-encoded at default quality 80, inflating the output (e.g. 50KB target
producing 90KB). Replaced buffer-wrapping with proper Sharp pipelines
that include .toFormat() with the proven quality. Also added progressive
dimension reduction when quality alone cannot reach the target, and
tightened tolerance to only accept at-or-below-target results.
Add colorBlindness() operation with 8 simulation matrices (Vienot/Machado)
for protanopia, deuteranopia, tritanopia, protanomaly, deuteranomaly,
tritanomaly, achromatopsia, and blue cone monochromacy.
- Fix resize 20% failure rate: add Zod refine requiring at least one
dimension, enforce integer/max constraints, clamp percentage scaling
to minimum 1px, and guard against missing metadata in withoutEnlargement
- Fix PostHog init race condition: move consent check before async import
so frontend events (search, pageview) are no longer silently dropped
- Fix identify() passing nested $set/$set_once wrappers instead of flat
properties, so version person property now appears on PostHog profiles
- Add error_code and error_message to failed tool_used analytics events
for debugging tool failures from PostHog
- Unit: 1,353 tests (42 files) — +256 new tests covering AI bridge
modules, image-engine sharpen/optimize-for-web, Zustand stores, and
icon-map validation
- Integration: 1,640 tests (57 files) — +826 new tests across all
tool routes, pipeline/progress/batch infrastructure, user-files,
edit-metadata, and a 321-test cross-format matrix
- E2E-Docker: 389 passing (20 spec files) — 6 new spec files for
batch processing, format conversion, layout, optimization,
watermark/overlay, and pipeline chains. Tests verified against fresh
Docker container with all 6 AI bundles installed.
Bug fixes discovered during testing:
- fix(compress): SVG/BMP/exotic formats crashed Sharp encoder — added
format-safety fallback to PNG
- fix(rate-limit): increase default login attempt limit from 10 to 500
per minute — previous value caused false test failures and is too
restrictive for a self-hosted app
- fix(auth.setup): wait for consent button visibility before clicking
to prevent flaky E2E-Docker auth setup
- Add Cloudflare Pages deployment for landing page (snapotter.com) and
docs (docs.snapotter.com)
- Create deploy-landing.yml and update deploy-docs.yml workflows
- Update CI to ignore apps/landing/** paths
- Fix logo transparency (remove white background) across all apps
- Recreate social-preview.png with SnapOtter branding
- Update all docs URLs from GitHub Pages to docs.snapotter.com
- Update VitePress config: light theme default, fix llms.txt paths
- Add .vitepress/cache/ and .env.* to gitignore
Move sanitizeValue, parseExif, parseGps, parseXmp into the shared
image-engine package so both strip-metadata and edit-metadata can
reuse them. Includes 13 unit tests covering all four functions.