Commit Graph
21 Commits
Author SHA1 Message Date
SnapOtterandGitHub fd39f66f46 fix(ai-bundles): lock the numpy-1.x ABI closure so the OCR bundle can't strand scipy (#437)
* fix(ai-bundles): lock the numpy-1.x ABI closure so the OCR bundle can't strand scipy

The OCR bundle installs paddleocr[doc-parser] 3.4, whose dependency closure drags
numpy 1.26.4 up to 2.5.1 and pulls scipy/scikit-learn/pandas wheels built against
the numpy 2.x ABI. build-bundle.sh re-pinned only numpy (basePackages), so those
numpy-2.x wheels stayed behind; the by-dir-name site-packages diff then shipped
them, and once merged onto the numpy==1.26.4 base they raise "numpy.dtype size
changed" on import.

Because the dispatcher pre-imports every ML library at startup and disables all AI
after 5 crashes in 60s, one stranded scipy takes down every AI tool, not just OCR
(observed on a CPU host: remove-background worked before the OCR bundle and broke
after). All-7 installs escaped it through last-writer-wins ordering; a subset
install did not, which is why it surfaced only intermittently.

Fix: add a manifest "constraints" list (numpy, scipy, scikit-learn, scikit-image,
pandas pinned to numpy-1.x-ABI versions) and apply it via PIP_CONSTRAINT to every
bundle pip install, so no bundle can pull a numpy-2.x wheel. paddleocr 3.4.1 still
resolves cleanly under the lock and the pinned stack imports without ABI error on
numpy 1.26.4 (validated on py3.12). Also import scipy/sklearn in the OCR path of
verify-bundle.sh so CI catches this class in isolation, and add a manifest
regression test.

Note: the published bundles must be rebuilt and republished (ai-bundles.yml) for
this to reach already-installed bases.

Claude-Session: https://claude.ai/code/session_01UvVCMNUBrgpghk8gye5gav

* chore(ai-bundles): sync OCR manifest sha256 to the rebuilt numpy-1.x bundles

Rebuilt the OCR bundle for both arches with the numpy-1.x-ABI constraints from
this PR and republished the tars to deepsafe/feature-bundles/v2.0.0, then updated
the baked manifest sha256 and sizes so installs verify against the fixed archives:

  amd64-gpu  5.93 GB  sha 2a00a3184f6a635f1fa9ae2a6517ad740a11f9e5ff58c098d2fd369a2bb1e16b
  arm64-cpu  1.98 GB  sha 6868c264069dcb74c6675c0b1f58dc1c9f60d9aa4459725e3dbde07a99a6a09a

Both tars ship scipy 1.12.0 / scikit-learn 1.4.2 / pandas 2.2.2 (numpy-1.x-ABI)
and zero numpy-2.x wheels, verified by listing the archive contents.

Stopgap note: these tars were built against the ghcr.io latest base (the 2.0.0
image is not published to GHCR), so they are not byte-identical to what the CI
build will produce. When ai-bundles.yml rebuilds at the 2.0.0 release, it will
mint fresh sha256 values and this manifest must be re-synced to them.

Claude-Session: https://claude.ai/code/session_01UvVCMNUBrgpghk8gye5gav
2026-07-05 11:52:27 +00:00
SnapOtterandGitHub cf884b52cd fix: offline CodeFormer face-enhance (ship RealESRGAN_x2plus in upscale-enhance bundle) (#433)
* fix: ship RealESRGAN_x2plus.pth in the upscale-enhance bundle for offline CodeFormer

codeformer-pip 0.0.4 downloads RealESRGAN_x2plus.pth at import of
codeformer.app, unconditionally, even though enhance_faces calls
inference_app with background_enhance=False and never uses the background
upsampler. The weight was not bundled, so explicit CodeFormer face-enhance
(enhance-faces model=codeformer) failed in strict offline mode
(SNAPOTTER_ALLOW_MODEL_DOWNLOAD=0) on a host that had never cached it -- the
guard raised before the import could complete.

Add RealESRGAN_x2plus.pth to the upscale-enhance bundle manifest (only that
bundle uses codeformer-pip; photo-restoration uses the CodeFormer ONNX path)
and link it in prepare_codeformer_weights alongside the other three weights,
replacing the download-or-error guard. Once the bundle ships it, the import
resolves offline and strict mode works.

Archive SHA256s updated in a follow-up once the bundle is rebuilt.

Claude-Session: https://claude.ai/code/session_01XGB4pGvTvb7sUX4JN745U7

* fix: require face-detection bundle for enhance-faces + point manifest at the x2plus archives

enhance-faces runs MediaPipe face detection (blaze_face_short_range.tflite)
before CodeFormer/GFPGAN. That model ships in the face-detection bundle, not
the tool's primary upscale-enhance bundle, so a standalone upscale-enhance
install failed face detection (offline: hard error; online: a surprise
download) before reaching the codeformer path. Declare the dependency in
TOOL_EXTRA_BUNDLES like passport-photo does.

Update the upscale-enhance archive SHA256/sizes to the rebuilt bundles that
include RealESRGAN_x2plus.pth (amd64-gpu + arm64-cpu), verified to install and
run enhance-faces model=codeformer in strict offline mode with zero downloads.

Claude-Session: https://claude.ai/code/session_01XGB4pGvTvb7sUX4JN745U7
2026-07-04 16:30:08 +00:00
SnapOtterandGitHub 2a36b3dfe0 fix(ai): update arm64 bundle sha256/size after protobuf<5 rebuild (#421)
Rebuilt the object-eraser-colorize, ocr, and transcription arm64-cpu
bundles with the protobuf<5 pin (PR #417) and republished them to
deepsafe/feature-bundles/v2.0.0. Update the manifest archive checksums,
compressed sizes, and (previously 0) extracted sizes to match the new
tarballs so install_feature.py's sha256 verification passes.

All three rebuilt bundles bake protobuf 4.25.9; verified gzip-clean and
that paddle 3.2.2 / onnxruntime coexist with protobuf 4.25.9 on aarch64.

Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
2026-07-04 13:05:05 +08:00
SnapOtterandGitHub 2c2fb65fca fix(ai): pin protobuf<5 on arm64 so mediapipe face landmarks work (#417)
* fix(ai): pin protobuf<5 on arm64 so mediapipe face landmarks work

aarch64 has no mediapipe wheel above 0.10.18, and 0.10.18 calls
MessageFactory.GetPrototype (removed in protobuf 5+). With protobuf
unpinned, the paddle/onnxruntime deps pull protobuf 7.x into the shared
AI venv and break mediapipe FaceLandmarker, so red-eye-removal fails on
every input (blur-faces and smart-crop keep working via a prebuilt graph).

Split mediapipe by platform and pin protobuf>=4.25.3,<5 for aarch64 only.
x86_64 keeps mediapipe 0.10.35, which works with protobuf 7, so
requirements-gpu.txt (amd64 only) stays unpinned. Also fixes a latent
issue where mediapipe>=0.10.21 was unsatisfiable on aarch64.

Verified live on the arm64 container: red-eye-removal completes on real
jpg and heic faces; OCR (tesseract) and paddle import unaffected.

Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA

* fix(ai): pin protobuf<5 in arm64 bundles lacking a mediapipe constraint

Bundles are built from docker/feature-manifest.json, not requirements.txt, so
this is the change that actually fixes the shipped arm64 bundles. On arm64,
object-eraser-colorize (onnxruntime), ocr (paddle) and transcription
(faster-whisper pulls onnxruntime) install a protobuf-dependent package with no
mediapipe to cap protobuf, so they bake protobuf 7.x. All bundles share one
/data/ai/venv at install time, so whichever of those installs last overwrites
protobuf to 7.x and breaks mediapipe FaceLandmarker (red-eye-removal). Pin
protobuf>=4.25.3,<5 in those three arm64 lists (appended last so it downgrades
after the puller installs). The four mediapipe bundles already resolve <5.

Dry-run on aarch64 confirmed paddle + protobuf 4.25.9 resolve with no conflict.

Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA

* refactor(ai): keep protobuf fix in feature-manifest.json only

requirements.txt is not consumed by the Docker image build (the base
/opt/venv is installed from a hardcoded package list, and the ML libs
ship via bundles), so the requirements changes had no effect on shipped
artifacts and only tripped the dependency-review scanner on the protobuf
range. Revert them; the operative arm64 bundle fix lives entirely in
docker/feature-manifest.json.

Claude-Session: https://claude.ai/code/session_01VtvE6K8iEr5jGFJJpHEaPA
2026-07-04 12:34:14 +08:00
SnapOtterandGitHub bd1838e40b fix: repair docker validation QA tooling, dispatcher crash-accounting, and image-enhancement RAW hang (#391)
Found and fixed during a full local Docker build validation (amd64/arm64, all
four fleet targets, AI bundle installs, QA harness) and the follow-up bug
sweep requested afterward. None of the affected scripts run in CI, so these
had been silently broken indefinitely.

- docker/feature-manifest.json: pythonVersion was a flat "3.11", but the
  amd64 base (Ubuntu 24.04) ships Python 3.12 while arm64 (Debian bookworm)
  ships 3.11. Changed to a per-arch object matching the file's existing
  convention.
- tests/qa/api-sweep.mts and verify-ai.mts: bare "@snapotter/shared" import
  can't resolve since tests/ is not a pnpm workspace member, making both
  silently unrunnable via their own documented command on any fresh
  checkout. Switched to a relative import.
- tests/qa/generate-ledger.mts: wrote to docs/qa/ without creating the
  directory first; docs/ is gitignored except COMMUNITY_GUIDE.md, so a fresh
  checkout threw ENOENT.
- Seven QA Playwright spec files (input-preview, settings,
  settings-extended, multifile, output-preview, pipeline-ui, smoke) had
  ~115 fixture() calls using directory names that don't exist. Resolved
  every call programmatically against the real fixture tree.
- packages/ai/src/bridge.ts: AI dispatcher restart (happens on every bundle
  install) was falsely counted as a crash, risking permanent dispatcher
  disable after enough legitimate restarts within the crash window. Added a
  shuttingDown flag checked at all three recordCrash() call sites.
- packages/image-engine/src/operations/auto-enhance.ts: image-enhancement
  hung 40+ seconds on large RAW photos (confirmed on a real 20.2MP file) in
  Sharp's .clahe() step, whose cost scales with total pixel count regardless
  of tile size. Added a 16-megapixel cap above which CLAHE is skipped;
  verified against the real file (40+s -> 2.0s) with no regression to other
  RAW formats or normal-sized images. Fixing this surfaced a second,
  smaller bug where the saturation step's CLAHE compensation boost was
  keyed off the raw toggle instead of whether CLAHE actually ran.
- Two QA-harness robustness gaps closed per "fix everything, even the small
  bugs": the passport-photo/erase-object input-preview tests now skip
  cleanly with a clear reason on a container without their AI bundle
  installed, and docker-compose.qa.yml's hardcoded project/container name
  (the actual root cause of a mid-validation container swap between two
  concurrent sessions) is now parameterized via QA_PROJECT_NAME.

Full validation report is local-only per repo convention.
2026-07-02 14:21:13 +08:00
SnapOtterandGitHub 3b50bcdc5c fix(ai-bundles): repair bundle build + publish pipeline (deepsafe repo, CPU provider, manifest)
Bundle build/publish fixes: CPUExecutionProvider in rembg build, pip/import/arm64 deps, hf-CLI publish to deepsafe/feature-bundles, real manifest sha256+sizes, installer fallback repo.
2026-06-19 18:34:33 +08:00
SnapOtter a4ac7cf7d6 feat: bump feature manifest to v2 with archive metadata 2026-06-13 16:27:24 +08:00
SnapOtter 51666cdd5f feat(tools): 2.0 phase 5 wave 5b - ai pool: ocr-pdf, transcription, background composites (5 tools) (#226) 2026-06-13 10:19:47 +08:00
SnapOtter 3b8d529b44 fix(ci): revert rembg to 2.0.62 (2.0.75 requires numpy>=2.3)
rembg 2.0.75 pulls in numpy>=2.3.0 which conflicts with our pinned
numpy==1.26.4 and would break the entire AI dependency chain. The two
rembg CVEs (SSRF + path traversal) are in its server/CLI components
which we don't use; they're already in the pip-audit ignore list.
2026-06-10 21:21:39 +08:00
SnapOtter 8792080982 fix(deps): patch Dependabot security alerts
- Pillow 11.1.0 -> 12.2.0 (6 CVEs: OOB writes, decompression bomb, DoS)
- rembg 2.0.62 -> 2.0.75 (SSRF + path traversal in server component)
- @fastify/static ^8.1.0 -> ^9.1.3 (path traversal + route guard bypass)
- Remove redundant @fastify/static pnpm override
- Dismiss stale esbuild alert (already at 0.28.0)
- Dismiss file-type alert (16.5.4 is dev-only via @types/potrace)
2026-06-10 19:08:08 +08:00
SnapOtter 1d7bc00d2d fix(docker): use CUDA 12.6 index for PaddlePaddle GPU and revert version
- fix(ocr): change paddlepaddle-gpu from --extra-index-url to --index-url
  for the CUDA 12.6 package index. With --extra-index-url, pip could
  resolve from PyPI (CUDA 11 build) instead of the cu126 index, causing
  "libcusolver.so.11: undefined symbol" errors on CUDA 12 containers.

- revert version to 1.17.1 (v1.17.2 release was deleted)
2026-06-08 14:36:59 +08:00
SnapOtterandGitHub f1aae73397 feat: pre-built release archives + AI install fixes (#202)
Add pre-built release archives (Linux amd64/arm64) to the release
workflow, published as GitHub Release assets. Each archive is a
self-contained tar.gz (~240MB) with built frontend, API source,
and production node_modules. Users extract and run without needing
pnpm build.

Also includes AI install manifest fixes for Proxmox/bare-metal users:
- Pin setuptools<75 for Python 3.13 basicsr compatibility
- Pre-install basicsr with --no-build-isolation before realesrgan
- Loosen mediapipe pins from == to >= for Python 3.13 wheels
- Add retry logic to HuggingFace model downloads
2026-06-05 18:42:30 +08:00
SnapOtter c6a5d3f33e refactor: remove content-aware-crop tool
Remove the content-aware-crop tool entirely -- API route, frontend
settings component, e2e and integration tests, and all registry
entries.
2026-05-13 17:55:01 +08:00
SnapOtter 9763a10db4 fix: add tool registry entry and manifest for content-aware-crop 2026-05-11 21:23:47 +08:00
SnapOtter 6fc767523c fix: add mediapipe to upscale-enhance bundle for face enhancement
The enhance-faces tool requires MediaPipe for face detection, but the
upscale-enhance feature bundle did not include mediapipe in its pip
packages. Users who installed only the upscale-enhance bundle got
"Face detection requires MediaPipe" errors. Added mediapipe to both
amd64 and arm64 package lists, matching the pattern used by the
face-detection and photo-restoration bundles.

Also added feature-manifest.test.ts with 25 tests validating bundle
dependency completeness to prevent similar missing-dependency bugs.

Closes #129
2026-05-07 22:21:45 +08:00
SnapOtter 9db19a01a2 feat: add BiRefNet HR-matting model download and manifest entry 2026-05-05 22:56:31 +08:00
ashim-hq 77a60b24cc fix: resolve 5 bugs found during comprehensive tool testing
1. split batch 404: register split tool in batch registry via
   registerToolProcessFn() so /api/v1/tools/split/batch works

2. CodeFormer crash: inference_app() expects a file path, not a numpy
   array. Save to temp file before calling, read result back.

3. OCR fallback chain: fix case-sensitive "Segmentation fault" match
   that prevented PaddleOCR crash from triggering Tesseract fallback.
   Also add "process crashed" check. Upgrade ARM paddlepaddle to >=3.2.1.

4. blur-faces large images: downscale to 1920px max before MediaPipe
   detection, scale coordinates back. Also add rotation retry for
   portrait-oriented images where BlazeFace misses faces. Applied to
   detect_faces.py, enhance_faces.py, and restore.py.

5. color-adjustments tool ID: fix mismatch in index.ts registration
   array (was "color-adjustments", should be "adjust-colors").
2026-04-21 22:25:06 +08:00
ashim-hq f67a03bb36 fix: resolve all audit findings — e2e coverage, feature system hardening, visual baselines
- Add 8 new E2E specs for AI tools (upscale, enhance-faces, colorize,
  restore-photo, erase-object, smart-crop, passport-photo, red-eye-removal)
  closing all HIGH/MEDIUM coverage gaps from the test matrix audit
- Fix ensureAiDirs() crash on non-Docker environments by gating on
  isDockerEnvironment() — prevents ENOENT when /data doesn't exist
- Bump torch 2.6.0→2.7.0 and torchvision 0.21.0→0.22.0 in feature
  manifest for broader Python version compatibility
- Add Python 3.14 version guard warning in install_feature.py
- Remove duplicate torchvision shims from upscale.py and enhance_faces.py
  (dispatcher.py already handles this at startup)
- Remove orphaned tools.batch i18n key and dead pipeline-builder filter
- Regenerate 4 visual regression baselines for current UI state
- Add data-testid to passport-photo generate button for E2E testability
2026-04-20 18:47:59 +08:00
AshimandClaude Opus 4.6 01d30cfb61 fix: pin torch cu126 for GPU compatibility and fix cross-platform bugs
- Pin torch==2.6.0+cu126 and torchvision==0.21.0+cu126 in feature
  manifest to prevent NCCL symbol mismatch on CUDA 12.6 base images
- Move lpips after torch in install order to prevent wrong version
  resolution from PyPI
- Add einops to upscale-enhance common deps (required by SCUNet)
- Update cpu_fallback_packages to handle multi-package CUDA torch
  entries on amd64 without GPU
- Fix gpu.py ONNX CUDA detection: replace hardcoded .so path with
  cross-platform session smoke-test
- Fix os.dup(1) crashes on Windows in upscale, enhance_faces, and
  noise_removal by wrapping in try/except with sys.stderr fallback
- Guard top-level numpy/cv2 imports in colorize.py and restore.py
  with helpful error messages
- Add weights_only=False fallback for torch.load in noise_removal
- Fix integration tests to accept 501 for uninstalled AI features
  and 422 for missing system tools (exiftool, libheif)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 15:17:08 +08:00
ashim-hq 10bdc24a4a feat: on-demand AI feature install with progress indicators 2026-04-19 19:52:14 +08:00
ashim-hq 43a572adcf feat: add feature manifest with exact package versions and model URLs
Authoritative JSON manifest defining all 6 AI feature bundles with:
- Exact pip package versions and platform-specific variants (amd64/arm64)
- pip flags (--no-deps for codeformer, --extra-index-url for torch/paddle)
- postInstall re-pins (numpy==1.26.4 after codeformer)
- Model download entries (direct URL, rembg sessions, HuggingFace snapshots)
- Bundle-to-tool mapping matching shared/features.ts

Used by install_feature.py at runtime to install bundles on demand.
2026-04-18 02:26:54 +08:00