* fix(ai-bundles): lock the numpy-1.x ABI closure so the OCR bundle can't strand scipy
The OCR bundle installs paddleocr[doc-parser] 3.4, whose dependency closure drags
numpy 1.26.4 up to 2.5.1 and pulls scipy/scikit-learn/pandas wheels built against
the numpy 2.x ABI. build-bundle.sh re-pinned only numpy (basePackages), so those
numpy-2.x wheels stayed behind; the by-dir-name site-packages diff then shipped
them, and once merged onto the numpy==1.26.4 base they raise "numpy.dtype size
changed" on import.
Because the dispatcher pre-imports every ML library at startup and disables all AI
after 5 crashes in 60s, one stranded scipy takes down every AI tool, not just OCR
(observed on a CPU host: remove-background worked before the OCR bundle and broke
after). All-7 installs escaped it through last-writer-wins ordering; a subset
install did not, which is why it surfaced only intermittently.
Fix: add a manifest "constraints" list (numpy, scipy, scikit-learn, scikit-image,
pandas pinned to numpy-1.x-ABI versions) and apply it via PIP_CONSTRAINT to every
bundle pip install, so no bundle can pull a numpy-2.x wheel. paddleocr 3.4.1 still
resolves cleanly under the lock and the pinned stack imports without ABI error on
numpy 1.26.4 (validated on py3.12). Also import scipy/sklearn in the OCR path of
verify-bundle.sh so CI catches this class in isolation, and add a manifest
regression test.
Note: the published bundles must be rebuilt and republished (ai-bundles.yml) for
this to reach already-installed bases.
Claude-Session: https://claude.ai/code/session_01UvVCMNUBrgpghk8gye5gav
* chore(ai-bundles): sync OCR manifest sha256 to the rebuilt numpy-1.x bundles
Rebuilt the OCR bundle for both arches with the numpy-1.x-ABI constraints from
this PR and republished the tars to deepsafe/feature-bundles/v2.0.0, then updated
the baked manifest sha256 and sizes so installs verify against the fixed archives:
amd64-gpu 5.93 GB sha 2a00a3184f6a635f1fa9ae2a6517ad740a11f9e5ff58c098d2fd369a2bb1e16b
arm64-cpu 1.98 GB sha 6868c264069dcb74c6675c0b1f58dc1c9f60d9aa4459725e3dbde07a99a6a09a
Both tars ship scipy 1.12.0 / scikit-learn 1.4.2 / pandas 2.2.2 (numpy-1.x-ABI)
and zero numpy-2.x wheels, verified by listing the archive contents.
Stopgap note: these tars were built against the ghcr.io latest base (the 2.0.0
image is not published to GHCR), so they are not byte-identical to what the CI
build will produce. When ai-bundles.yml rebuilds at the 2.0.0 release, it will
mint fresh sha256 values and this manifest must be re-synced to them.
Claude-Session: https://claude.ai/code/session_01UvVCMNUBrgpghk8gye5gav
* fix: ship RealESRGAN_x2plus.pth in the upscale-enhance bundle for offline CodeFormer
codeformer-pip 0.0.4 downloads RealESRGAN_x2plus.pth at import of
codeformer.app, unconditionally, even though enhance_faces calls
inference_app with background_enhance=False and never uses the background
upsampler. The weight was not bundled, so explicit CodeFormer face-enhance
(enhance-faces model=codeformer) failed in strict offline mode
(SNAPOTTER_ALLOW_MODEL_DOWNLOAD=0) on a host that had never cached it -- the
guard raised before the import could complete.
Add RealESRGAN_x2plus.pth to the upscale-enhance bundle manifest (only that
bundle uses codeformer-pip; photo-restoration uses the CodeFormer ONNX path)
and link it in prepare_codeformer_weights alongside the other three weights,
replacing the download-or-error guard. Once the bundle ships it, the import
resolves offline and strict mode works.
Archive SHA256s updated in a follow-up once the bundle is rebuilt.
Claude-Session: https://claude.ai/code/session_01XGB4pGvTvb7sUX4JN745U7
* fix: require face-detection bundle for enhance-faces + point manifest at the x2plus archives
enhance-faces runs MediaPipe face detection (blaze_face_short_range.tflite)
before CodeFormer/GFPGAN. That model ships in the face-detection bundle, not
the tool's primary upscale-enhance bundle, so a standalone upscale-enhance
install failed face detection (offline: hard error; online: a surprise
download) before reaching the codeformer path. Declare the dependency in
TOOL_EXTRA_BUNDLES like passport-photo does.
Update the upscale-enhance archive SHA256/sizes to the rebuilt bundles that
include RealESRGAN_x2plus.pth (amd64-gpu + arm64-cpu), verified to install and
run enhance-faces model=codeformer in strict offline mode with zero downloads.
Claude-Session: https://claude.ai/code/session_01XGB4pGvTvb7sUX4JN745U7
Rebuilt the object-eraser-colorize, ocr, and transcription arm64-cpu
bundles with the protobuf<5 pin (PR #417) and republished them to
deepsafe/feature-bundles/v2.0.0. Update the manifest archive checksums,
compressed sizes, and (previously 0) extracted sizes to match the new
tarballs so install_feature.py's sha256 verification passes.
All three rebuilt bundles bake protobuf 4.25.9; verified gzip-clean and
that paddle 3.2.2 / onnxruntime coexist with protobuf 4.25.9 on aarch64.
Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
* fix(ai): pin protobuf<5 on arm64 so mediapipe face landmarks work
aarch64 has no mediapipe wheel above 0.10.18, and 0.10.18 calls
MessageFactory.GetPrototype (removed in protobuf 5+). With protobuf
unpinned, the paddle/onnxruntime deps pull protobuf 7.x into the shared
AI venv and break mediapipe FaceLandmarker, so red-eye-removal fails on
every input (blur-faces and smart-crop keep working via a prebuilt graph).
Split mediapipe by platform and pin protobuf>=4.25.3,<5 for aarch64 only.
x86_64 keeps mediapipe 0.10.35, which works with protobuf 7, so
requirements-gpu.txt (amd64 only) stays unpinned. Also fixes a latent
issue where mediapipe>=0.10.21 was unsatisfiable on aarch64.
Verified live on the arm64 container: red-eye-removal completes on real
jpg and heic faces; OCR (tesseract) and paddle import unaffected.
Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
* fix(ai): pin protobuf<5 in arm64 bundles lacking a mediapipe constraint
Bundles are built from docker/feature-manifest.json, not requirements.txt, so
this is the change that actually fixes the shipped arm64 bundles. On arm64,
object-eraser-colorize (onnxruntime), ocr (paddle) and transcription
(faster-whisper pulls onnxruntime) install a protobuf-dependent package with no
mediapipe to cap protobuf, so they bake protobuf 7.x. All bundles share one
/data/ai/venv at install time, so whichever of those installs last overwrites
protobuf to 7.x and breaks mediapipe FaceLandmarker (red-eye-removal). Pin
protobuf>=4.25.3,<5 in those three arm64 lists (appended last so it downgrades
after the puller installs). The four mediapipe bundles already resolve <5.
Dry-run on aarch64 confirmed paddle + protobuf 4.25.9 resolve with no conflict.
Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
* refactor(ai): keep protobuf fix in feature-manifest.json only
requirements.txt is not consumed by the Docker image build (the base
/opt/venv is installed from a hardcoded package list, and the ML libs
ship via bundles), so the requirements changes had no effect on shipped
artifacts and only tripped the dependency-review scanner on the protobuf
range. Revert them; the operative arm64 bundle fix lives entirely in
docker/feature-manifest.json.
Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
Found and fixed during a full local Docker build validation (amd64/arm64, all
four fleet targets, AI bundle installs, QA harness) and the follow-up bug
sweep requested afterward. None of the affected scripts run in CI, so these
had been silently broken indefinitely.
- docker/feature-manifest.json: pythonVersion was a flat "3.11", but the
amd64 base (Ubuntu 24.04) ships Python 3.12 while arm64 (Debian bookworm)
ships 3.11. Changed to a per-arch object matching the file's existing
convention.
- tests/qa/api-sweep.mts and verify-ai.mts: bare "@snapotter/shared" import
can't resolve since tests/ is not a pnpm workspace member, making both
silently unrunnable via their own documented command on any fresh
checkout. Switched to a relative import.
- tests/qa/generate-ledger.mts: wrote to docs/qa/ without creating the
directory first; docs/ is gitignored except COMMUNITY_GUIDE.md, so a fresh
checkout threw ENOENT.
- Seven QA Playwright spec files (input-preview, settings,
settings-extended, multifile, output-preview, pipeline-ui, smoke) had
~115 fixture() calls using directory names that don't exist. Resolved
every call programmatically against the real fixture tree.
- packages/ai/src/bridge.ts: AI dispatcher restart (happens on every bundle
install) was falsely counted as a crash, risking permanent dispatcher
disable after enough legitimate restarts within the crash window. Added a
shuttingDown flag checked at all three recordCrash() call sites.
- packages/image-engine/src/operations/auto-enhance.ts: image-enhancement
hung 40+ seconds on large RAW photos (confirmed on a real 20.2MP file) in
Sharp's .clahe() step, whose cost scales with total pixel count regardless
of tile size. Added a 16-megapixel cap above which CLAHE is skipped;
verified against the real file (40+s -> 2.0s) with no regression to other
RAW formats or normal-sized images. Fixing this surfaced a second,
smaller bug where the saturation step's CLAHE compensation boost was
keyed off the raw toggle instead of whether CLAHE actually ran.
- Two QA-harness robustness gaps closed per "fix everything, even the small
bugs": the passport-photo/erase-object input-preview tests now skip
cleanly with a clear reason on a container without their AI bundle
installed, and docker-compose.qa.yml's hardcoded project/container name
(the actual root cause of a mid-validation container swap between two
concurrent sessions) is now parameterized via QA_PROJECT_NAME.
Full validation report is local-only per repo convention.
rembg 2.0.75 pulls in numpy>=2.3.0 which conflicts with our pinned
numpy==1.26.4 and would break the entire AI dependency chain. The two
rembg CVEs (SSRF + path traversal) are in its server/CLI components
which we don't use; they're already in the pip-audit ignore list.
- fix(ocr): change paddlepaddle-gpu from --extra-index-url to --index-url
for the CUDA 12.6 package index. With --extra-index-url, pip could
resolve from PyPI (CUDA 11 build) instead of the cu126 index, causing
"libcusolver.so.11: undefined symbol" errors on CUDA 12 containers.
- revert version to 1.17.1 (v1.17.2 release was deleted)
Add pre-built release archives (Linux amd64/arm64) to the release
workflow, published as GitHub Release assets. Each archive is a
self-contained tar.gz (~240MB) with built frontend, API source,
and production node_modules. Users extract and run without needing
pnpm build.
Also includes AI install manifest fixes for Proxmox/bare-metal users:
- Pin setuptools<75 for Python 3.13 basicsr compatibility
- Pre-install basicsr with --no-build-isolation before realesrgan
- Loosen mediapipe pins from == to >= for Python 3.13 wheels
- Add retry logic to HuggingFace model downloads
The enhance-faces tool requires MediaPipe for face detection, but the
upscale-enhance feature bundle did not include mediapipe in its pip
packages. Users who installed only the upscale-enhance bundle got
"Face detection requires MediaPipe" errors. Added mediapipe to both
amd64 and arm64 package lists, matching the pattern used by the
face-detection and photo-restoration bundles.
Also added feature-manifest.test.ts with 25 tests validating bundle
dependency completeness to prevent similar missing-dependency bugs.
Closes#129
1. split batch 404: register split tool in batch registry via
registerToolProcessFn() so /api/v1/tools/split/batch works
2. CodeFormer crash: inference_app() expects a file path, not a numpy
array. Save to temp file before calling, read result back.
3. OCR fallback chain: fix case-sensitive "Segmentation fault" match
that prevented PaddleOCR crash from triggering Tesseract fallback.
Also add "process crashed" check. Upgrade ARM paddlepaddle to >=3.2.1.
4. blur-faces large images: downscale to 1920px max before MediaPipe
detection, scale coordinates back. Also add rotation retry for
portrait-oriented images where BlazeFace misses faces. Applied to
detect_faces.py, enhance_faces.py, and restore.py.
5. color-adjustments tool ID: fix mismatch in index.ts registration
array (was "color-adjustments", should be "adjust-colors").
- Add 8 new E2E specs for AI tools (upscale, enhance-faces, colorize,
restore-photo, erase-object, smart-crop, passport-photo, red-eye-removal)
closing all HIGH/MEDIUM coverage gaps from the test matrix audit
- Fix ensureAiDirs() crash on non-Docker environments by gating on
isDockerEnvironment() — prevents ENOENT when /data doesn't exist
- Bump torch 2.6.0→2.7.0 and torchvision 0.21.0→0.22.0 in feature
manifest for broader Python version compatibility
- Add Python 3.14 version guard warning in install_feature.py
- Remove duplicate torchvision shims from upscale.py and enhance_faces.py
(dispatcher.py already handles this at startup)
- Remove orphaned tools.batch i18n key and dead pipeline-builder filter
- Regenerate 4 visual regression baselines for current UI state
- Add data-testid to passport-photo generate button for E2E testability
- Pin torch==2.6.0+cu126 and torchvision==0.21.0+cu126 in feature
manifest to prevent NCCL symbol mismatch on CUDA 12.6 base images
- Move lpips after torch in install order to prevent wrong version
resolution from PyPI
- Add einops to upscale-enhance common deps (required by SCUNet)
- Update cpu_fallback_packages to handle multi-package CUDA torch
entries on amd64 without GPU
- Fix gpu.py ONNX CUDA detection: replace hardcoded .so path with
cross-platform session smoke-test
- Fix os.dup(1) crashes on Windows in upscale, enhance_faces, and
noise_removal by wrapping in try/except with sys.stderr fallback
- Guard top-level numpy/cv2 imports in colorize.py and restore.py
with helpful error messages
- Add weights_only=False fallback for torch.load in noise_removal
- Fix integration tests to accept 501 for uninstalled AI features
and 422 for missing system tools (exiftool, libheif)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Authoritative JSON manifest defining all 6 AI feature bundles with:
- Exact pip package versions and platform-specific variants (amd64/arm64)
- pip flags (--no-deps for codeformer, --extra-index-url for torch/paddle)
- postInstall re-pins (numpy==1.26.4 after codeformer)
- Model download entries (direct URL, rembg sessions, HuggingFace snapshots)
- Bundle-to-tool mapping matching shared/features.ts
Used by install_feature.py at runtime to install bundles on demand.