docs: deep-audit the API reference pages against code

rest.md: correct the tool response shape (JSON {jobId, downloadUrl, ...} for fast / 202 {jobId, async} for long, not the raw file), fix the SSE jobId source (body, with X-Job-Id only on batch/pipeline ZIP streams), complete the permissions list to all 17, and add ai-canvas-expand.

ai.md: reconcile ~14 AI tool parameter tables to the real Zod schemas (remove-background, upscale, ocr, enhance-faces, colorize, noise-removal, red-eye-removal, restore-photo, passport-photo, erase-object, smart-crop, content-aware-resize, transparency-fixer, image-enhancement), add the 6 missing AI tools (background-replace, blur-background, ocr-pdf, ai-canvas-expand, transcribe-audio, auto-subtitles), and fix the dispatcher script diagram and feature-bundle table.

image-engine.md: add 4 undocumented operations (sharpen, sharpen-advanced, color-blindness, edit-metadata) and fix flip/crop/rotate/convert/compress/strip-metadata/color-channels params, output-format count (13), and the info JSON example. supported-formats.md: output count 14 -> 13. Verified with a clean docs build.
This commit is contained in:
SnapOtter
2026-06-18 14:12:28 +08:00
parent 8532f3227b
commit 60e3008106
4 changed files with 377 additions and 161 deletions
+273 -139
View File
@@ -6,119 +6,173 @@ description: AI engine reference with all local ML tools. Background removal, up
The `@snapotter/ai` package bridges Node.js to a **persistent Python sidecar** for all ML operations. The dispatcher process stays alive between requests for fast warm-start performance. GPU is auto-detected at startup and used when available. The `@snapotter/ai` package bridges Node.js to a **persistent Python sidecar** for all ML operations. The dispatcher process stays alive between requests for fast warm-start performance. GPU is auto-detected at startup and used when available.
19 AI tool routes. All models run locally - no internet required after initial model download. 19 Python sidecar AI tools across four modalities (image, audio, video, document), plus 2 tools with optional AI capabilities. All models run locally -- no internet required after initial model download.
## Architecture ## Architecture
``` ```
Node.js Tool Route Node.js Tool Route
|
v
@snapotter/ai bridge.ts @snapotter/ai bridge.ts
(stdin/stdout JSON + stderr progress events) | (stdin/stdout JSON + stderr progress events)
v
Python dispatcher (persistent process) Python dispatcher (persistent process, "ai" profile)
|
├─ remove_bg.py (rembg / BiRefNet) |-- remove_bg.py (rembg / BiRefNet)
├─ upscale.py (RealESRGAN) |-- upscale.py (RealESRGAN)
├─ inpaint.py (LaMa ONNX) |-- inpaint.py (LaMa ONNX)
├─ ocr.py (PaddleOCR / Tesseract) |-- outpaint.py (LaMa canvas expansion)
├─ detect_faces.py (MediaPipe) |-- ocr.py (PaddleOCR / Tesseract)
├─ face_landmarks.py (MediaPipe landmarks) |-- ocr_pdf.py (page-by-page document OCR)
├─ enhance_faces.py (GFPGAN / CodeFormer) |-- ocr_preprocess.py (image enhancement for OCR)
├─ colorize.py (DDColor) |-- detect_faces.py (MediaPipe)
├─ noise_removal.py (tiered denoising) |-- face_landmarks.py (MediaPipe landmarks)
├─ red_eye_removal.py (landmark + color analysis) |-- enhance_faces.py (GFPGAN / CodeFormer)
├─ restore.py (scratch repair + enhancement + denoising) |-- colorize.py (DDColor)
├─ transparency_fix.py (BiRefNet HR-matting + defringe) |-- noise_removal.py (SCUNet / tiered denoising)
└─ seam_carving (Go caire binary - not Python) |-- red_eye_removal.py (landmark + color analysis)
|-- restore.py (scratch repair + enhancement + denoising)
|-- transcribe.py (faster-whisper speech-to-text)
+-- install_feature.py (on-demand bundle installer)
``` ```
A separate "docs" dispatcher profile replaces the AI allowlist with document-processing scripts (`doc_pagecount`, `doc_health`, `doc_flatten`, `doc_redact`, `doc_text`, `doc_to_word`, `doc_metadata`, `doc_html_pdf`) and skips heavy ML imports.
**Timeouts:** 300 s default; OCR and BiRefNet background removal get 600 s. **Timeouts:** 300 s default; OCR and BiRefNet background removal get 600 s.
## Feature Bundles
Each AI tool requires a model bundle to be installed before use. Bundles are installed on demand via the admin UI or `install_feature.py`.
| Bundle | Size | Tools |
|--------|------|-------|
| `background-removal` | 4-5 GB | remove-background, passport-photo, transparency-fixer, background-replace, blur-background |
| `face-detection` | 200-300 MB | blur-faces, red-eye-removal, smart-crop |
| `object-eraser-colorize` | 1-2 GB | erase-object, colorize, ai-canvas-expand |
| `upscale-enhance` | 4-5 GB | upscale, enhance-faces, noise-removal |
| `photo-restoration` | 800 MB - 1 GB | restore-photo |
| `ocr` | 3-4 GB | ocr, ocr-pdf |
| `transcription` | ~600 MB | transcribe-audio, auto-subtitles |
---
## Background Removal ## Background Removal
**Function:** `removeBackground`
**Tool route:** `remove-background` **Tool route:** `remove-background`
**Model:** rembg with BiRefNet (default) or U2-Net variants **Model:** rembg with BiRefNet (default) or U2-Net variants
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `model` | string | `birefnet-general` | Model variant - see table below | | `model` | string | - | Model variant (optional override) |
| `alphaMattingForeground` | number (1255) | 240 | Foreground threshold for alpha matting | | `backgroundType` | string | `"transparent"` | One of: `transparent`, `color`, `gradient`, `blur`, `image` |
| `alphaMattingBackground` | number (1255) | 10 | Background threshold for alpha matting | | `backgroundColor` | string | - | Hex color for solid background |
| `returnMask` | boolean | false | Return the mask instead of the cutout | | `gradientColor1` | string | - | First gradient color |
| `backgroundColor` | string | - | Fill removed area (hex color or "transparent") | | `gradientColor2` | string | - | Second gradient color |
| `gradientAngle` | number | - | Gradient angle in degrees |
| `blurEnabled` | boolean | - | Enable background blur effect |
| `blurIntensity` | number (0-100) | - | Blur intensity |
| `shadowEnabled` | boolean | - | Enable drop shadow on subject |
| `shadowOpacity` | number (0-100) | - | Shadow opacity |
| `outputFormat` | string | - | Output format: `png`, `webp`, or `avif` |
| `edgeRefine` | integer (0-3) | - | Edge refinement level |
| `decontaminate` | boolean | - | Remove color bleed from edges |
**Available models:** ## Background Replace
| Model ID | Best for | **Tool route:** `background-replace`
|----------|---------| **Model:** rembg / BiRefNet (shared with remove-background)
| `birefnet-general` | General purpose (default) |
| `birefnet-portrait` | People / portraits |
| `birefnet-dis` | Dichotomous Image Segmentation |
| `birefnet-hrsod` | High-resolution salient objects |
| `birefnet-cod` | Camouflaged objects |
| `u2net` | Fast general purpose |
| `u2net_human_seg` | Human segmentation |
| `isnet-general-use` | High quality general |
## Image Upscaling Removes the background and replaces it with a solid color or gradient.
**Function:** `upscale`
**Tool route:** `upscale`
**Model:** RealESRGAN (with Lanczos fallback on CPU-constrained systems)
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `scale` | 2 \| 4 | 4 | Upscale factor | | `backgroundType` | `"color"` \| `"gradient"` | `"color"` | Background mode |
| `model` | string | `realesrgan-x4plus` | Model variant | | `color` | string | `"#ffffff"` | Background hex color (when `backgroundType` is `color`) |
| `faceEnhance` | boolean | false | Apply GFPGAN face enhancement pass | | `gradientColor1` | string | - | First gradient hex color |
| `denoise` | number (01) | 0.5 | Denoising strength | | `gradientColor2` | string | - | Second gradient hex color |
| `format` | string | - | Output format override | | `gradientAngle` | integer (0-360) | `180` | Gradient angle in degrees |
| `quality` | number | 95 | Output quality (for JPEG/WebP) | | `feather` | integer (0-20) | `0` | Edge feathering radius |
| `format` | `"png"` \| `"webp"` | `"png"` | Output format |
## Blur Background
**Tool route:** `blur-background`
**Model:** rembg / BiRefNet (shared with remove-background)
Blurs the background while keeping the subject sharp.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `intensity` | integer (1-100) | `50` | Blur intensity |
| `feather` | integer (0-20) | `0` | Edge feathering radius |
| `format` | `"png"` \| `"webp"` | `"png"` | Output format |
## Image Upscaling
**Tool route:** `upscale`
**Model:** RealESRGAN (with Lanczos fallback when unavailable)
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `scale` | number | `2` | Upscale factor |
| `model` | string | `"auto"` | Model variant |
| `faceEnhance` | boolean | `false` | Apply GFPGAN face enhancement pass |
| `denoise` | number | `0` | Denoising strength |
| `format` | string | `"auto"` | Output format override |
| `quality` | number | `95` | Output quality (1-100) |
## OCR / Text Extraction ## OCR / Text Extraction
**Function:** `extractText`
**Tool route:** `ocr` **Tool route:** `ocr`
**Models:** Tesseract (fast), PaddleOCR PP-OCRv5 (balanced), PaddleOCR-VL 1.5 (best) **Models:** Tesseract (fast), PaddleOCR PP-OCRv5 (balanced), PaddleOCR-VL 1.5 (best)
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `quality` | `fast` \| `balanced` \| `best` | `balanced` | Processing tier | | `quality` | `"fast"` \| `"balanced"` \| `"best"` | `"balanced"` | Processing tier |
| `language` | string | `en` | Language code (ISO 639-1) | | `language` | string | `"auto"` | Language: `auto`, `en`, `de`, `fr`, `es`, `zh`, `ja`, `ko` |
| `enhance` | boolean | false | Pre-process image to improve OCR accuracy | | `enhance` | boolean | `true` | Pre-process image to improve OCR accuracy |
| `engine` | string | - | Deprecated. Maps `tesseract` to `fast`, `paddleocr` to `balanced` |
Returns structured results with bounding boxes, confidence scores, and extracted text blocks. Returns structured results with bounding boxes, confidence scores, and extracted text blocks.
## PDF OCR
**Tool route:** `ocr-pdf`
**Models:** Same tier system as image OCR
Extracts text from scanned PDF documents using AI-powered OCR, page by page.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `quality` | `"fast"` \| `"balanced"` \| `"best"` | `"balanced"` | Processing tier |
| `language` | string | `"auto"` | Language: `auto`, `en`, `de`, `fr`, `es`, `zh`, `ja`, `ko` |
| `pages` | string | `"all"` | Page selection: `"all"`, `"1-3"`, `"1,3,5"` |
## Face / PII Blur ## Face / PII Blur
**Function:** `blurFaces`
**Tool route:** `blur-faces` **Tool route:** `blur-faces`
**Model:** MediaPipe face detection **Model:** MediaPipe face detection
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `blurRadius` | number | 30 | Gaussian blur radius | | `blurRadius` | number (1-100) | `30` | Gaussian blur radius |
| `sensitivity` | number (01) | 0.5 | Detection confidence threshold | | `sensitivity` | number (0-1) | `0.5` | Detection confidence threshold |
## Face Enhancement ## Face Enhancement
**Function:** `enhanceFaces`
**Tool route:** `enhance-faces` **Tool route:** `enhance-faces`
**Models:** GFPGAN, CodeFormer **Models:** GFPGAN, CodeFormer
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `model` | `gfpgan` \| `codeformer` | `gfpgan` | Enhancement model | | `model` | `"auto"` \| `"gfpgan"` \| `"codeformer"` | `"auto"` | Enhancement model |
| `strength` | number (01) | 0.7 | Enhancement strength | | `strength` | number (0-1) | `0.8` | Enhancement strength |
| `sensitivity` | number (01) | 0.5 | Face detection threshold | | `sensitivity` | number (0-1) | `0.5` | Face detection threshold |
| `centerFace` | boolean | false | Focus enhancement on center face only | | `onlyCenterFace` | boolean | `false` | Enhance only the most central face |
## AI Colorization ## AI Colorization
**Function:** `colorize`
**Tool route:** `colorize` **Tool route:** `colorize`
**Model:** DDColor (with OpenCV DNN fallback) **Model:** DDColor (with OpenCV DNN fallback)
@@ -126,144 +180,182 @@ Converts black-and-white or grayscale photos to full color.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `intensity` | number (01) | 0.85 | Color saturation strength | | `intensity` | number (0-1) | `1.0` | Color saturation strength |
| `model` | string | `ddcolor` | Model variant | | `model` | `"auto"` \| `"ddcolor"` \| `"opencv"` | `"auto"` | Model variant |
## Noise Removal ## Noise Removal
**Function:** `noiseRemoval` **Tool route:** `noise-removal`
**Tool route:** `noise-removal` **Model:** SCUNet (tiered denoising pipeline)
Three-tier denoising pipeline (fast: OpenCV bilateral filter; balanced: frequency-domain; best: deep learning model).
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `quality` | `fast` \| `balanced` \| `best` | `balanced` | Processing tier | | `tier` | `"quick"` \| `"balanced"` \| `"quality"` \| `"maximum"` | `"balanced"` | Processing tier |
| `strength` | number (01) | 0.5 | Denoising strength | | `strength` | number (0-100) | `50` | Denoising strength |
| `preserveDetail` | boolean | true | Edge-preserving mode | | `detailPreservation` | number (0-100) | `50` | How much detail to preserve; higher keeps more texture |
| `colorNoise` | boolean | false | Target color noise specifically | | `colorNoise` | number (0-100) | `30` | Color noise reduction strength |
| `format` | string | `"original"` | Output format: `original`, `png`, `jpeg`, `webp`, `avif`, `jxl` |
| `quality` | number (1-100) | `90` | Output encoding quality |
## Red Eye Removal ## Red Eye Removal
**Function:** `removeRedEye`
**Tool route:** `red-eye-removal` **Tool route:** `red-eye-removal`
Detects face landmarks, locates eye regions, and corrects red-channel oversaturation. Detects face landmarks, locates eye regions, and corrects red-channel oversaturation.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `sensitivity` | number (01) | 0.5 | Red pixel detection threshold | | `sensitivity` | number (0-100) | `50` | Red pixel detection threshold |
| `strength` | number (01) | 0.9 | Correction strength | | `strength` | number (0-100) | `70` | Correction strength |
| `format` | string | - | Output format override (optional) |
| `quality` | number (1-100) | `90` | Output quality |
## Photo Restoration ## Photo Restoration
**Function:** `restorePhoto`
**Tool route:** `restore-photo` **Tool route:** `restore-photo`
Multi-step pipeline for old or damaged photos: scratch/tear detection and repair face enhancement denoising optional colorization. Multi-step pipeline for old or damaged photos: scratch/tear detection and repair, face enhancement, denoising, and optional colorization.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `mode` | `auto` \| `light` \| `heavy` | `auto` | Restoration intensity | | `scratchRemoval` | boolean | `true` | Detect and repair scratches, tears |
| `scratchRemoval` | boolean | true | Detect and repair scratches, tears | | `faceEnhancement` | boolean | `true` | Apply face enhancement pass |
| `faceEnhancement` | boolean | true | Apply face enhancement pass | | `fidelity` | number (0-1) | `0.7` | Face enhancement strength (higher = more conservative) |
| `fidelity` | number (01) | 0.7 | Face enhancement strength | | `denoise` | boolean | `true` | Apply denoising pass |
| `denoise` | boolean | true | Apply denoising pass | | `denoiseStrength` | number (0-100) | `25` | Denoising strength |
| `denoiseStrength` | number (0100) | 40 | Denoising strength | | `colorize` | boolean | `false` | Colorize after restoration |
| `colorize` | boolean | false | Colorize after restoration | | `colorizeStrength` | number (0-100) | `85` | Colorization intensity |
## Passport Photo ## Passport Photo
**Function:** Uses `detectFaceLandmarks` + `removeBackground`
**Tool route:** `passport-photo` **Tool route:** `passport-photo`
**Model:** MediaPipe face landmarks **Models:** MediaPipe face landmarks + BiRefNet background removal
Generates government-compliant ID photos. Supports **37 countries** across 6 regions (Americas, Europe, Asia, Africa, Oceania, Middle East). Each spec includes physical dimensions, DPI, head-height ratio, eye-line position, and background color requirements. Two-phase workflow: analyze (detect face + remove background) then generate (crop, resize, tile). Supports 37+ countries across 6 regions.
### Phase 1: Analyze
`POST /api/v1/tools/passport-photo/analyze`
Accepts an image file (multipart). Returns face landmark data, a base64 preview, and image dimensions.
### Phase 2: Generate
`POST /api/v1/tools/passport-photo/generate`
Accepts a JSON body with the Phase 1 results plus generation settings:
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `country` | string | `us` | ISO country code (see list in UI) | | `jobId` | string | (required) | Job ID from Phase 1 |
| `printLayout` | `4x6` \| `A4` \| `none` | `none` | Output as print sheet or standalone | | `filename` | string | (required) | Original filename from Phase 1 |
| `backgroundColor` | string | country default | Background fill color | | `countryCode` | string | (required) | ISO country code (e.g., `US`, `GB`, `IN`) |
| `documentType` | string | `"passport"` | Document type |
| `bgColor` | string | `"#FFFFFF"` | Background color hex |
| `printLayout` | string | `"none"` | Print layout: `none`, `4x6`, `a4`, `letter` |
| `maxFileSizeKb` | number | `0` | Max file size in KB (0 = no limit) |
| `dpi` | number (72-1200) | `300` | Output DPI |
| `customWidthMm` | number | - | Custom width in mm (overrides country spec) |
| `customHeightMm` | number | - | Custom height in mm (overrides country spec) |
| `zoom` | number (0.5-3) | `1` | Zoom factor |
| `adjustX` | number | `0` | Horizontal position adjustment |
| `adjustY` | number | `0` | Vertical position adjustment |
| `landmarks` | object | (required) | Landmarks from Phase 1 |
| `imageWidth` | number | (required) | Image width from Phase 1 |
| `imageHeight` | number | (required) | Image height from Phase 1 |
## Object Erasing (Inpainting) ## Object Erasing (Inpainting)
**Function:** `inpaint`
**Tool route:** `erase-object` **Tool route:** `erase-object`
**Model:** LaMa via ONNX Runtime **Model:** LaMa via ONNX Runtime
| Parameter | Type | Required | Description | The mask is sent as a **second file part** (fieldname `mask`), not as base64. White pixels in the mask indicate areas to erase. The `format` and `quality` settings are sent as top-level form fields.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `maskData` | string | Yes | Base64-encoded PNG mask (white = erase) | | `file` | file | (required) | Source image (multipart) |
| `maskThreshold` | number (0255) | No | Threshold for mask binarization | | `mask` | file | (required) | Mask image (multipart, fieldname `mask`, white = erase) |
| `format` | string | `"auto"` | Output format: `auto`, `png`, `jpg`, `jpeg`, `webp`, `tiff`, `gif`, `avif`, `heic`, `heif`, `jxl` |
| `quality` | integer (1-100) | `95` | Output quality |
GPU-accelerated when an NVIDIA GPU is available. GPU-accelerated when an NVIDIA GPU is available.
## Smart Crop ## AI Canvas Expand
**Function:** Uses MediaPipe + Sharp attention/entropy **Tool route:** `ai-canvas-expand`
**Tool route:** `smart-crop` **Model:** LaMa-based outpainting
**Model:** MediaPipe face detection
Expands the canvas of an image in any direction and fills new areas with AI-generated content that matches the existing image.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `mode` | `subject` \| `face` \| `trim` | `subject` | Crop strategy | | `extendTop` | integer | `0` | Pixels to extend at the top |
| `width` | number | - | Output width | | `extendRight` | integer | `0` | Pixels to extend at the right |
| `height` | number | - | Output height | | `extendBottom` | integer | `0` | Pixels to extend at the bottom |
| `facePreset` | string | - | Preset framing when `mode=face` | | `extendLeft` | integer | `0` | Pixels to extend at the left |
| `tier` | `"fast"` \| `"balanced"` \| `"high"` | `"balanced"` | Quality tier |
| `format` | string | `"auto"` | Output format: `auto`, `png`, `jpg`, `jpeg`, `webp`, `tiff`, `gif`, `avif`, `heic`, `heif`, `jxl` |
| `quality` | integer (1-100) | `95` | Output quality |
At least one extend direction must be greater than 0.
## Smart Crop
**Tool route:** `smart-crop`
**Model:** MediaPipe face detection (face mode only)
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `mode` | string | `"subject"` | Crop strategy: `subject`, `face`, `trim` |
| `strategy` | `"attention"` \| `"entropy"` | `"attention"` | Strategy for subject mode |
| `width` | integer | - | Output width |
| `height` | integer | - | Output height |
| `padding` | integer (0-50) | `0` | Padding percentage around subject |
| `facePreset` | string | `"head-shoulders"` | Preset framing when `mode=face` |
| `sensitivity` | number (0-1) | `0.5` | Face detection threshold |
| `threshold` | integer (0-255) | `30` | Background detection threshold (trim mode) |
| `padToSquare` | boolean | `false` | Pad trimmed result to a square |
| `padColor` | string | `"#ffffff"` | Background color for square padding |
| `targetSize` | integer | - | Target size for padded output (pixels) |
| `quality` | integer (1-100) | - | Output quality |
Legacy `mode` values `attention` and `content` are accepted and mapped to `subject` and `trim` respectively.
**Face presets:** **Face presets:**
| Preset | Head ratio | Best for | | Preset | Best for |
|--------|-----------|---------| |--------|---------|
| `close-up` | 1.8× face | Headshots | | `closeup` | Headshots |
| `head-and-shoulders` | 2.8× face | Profile photos | | `head-shoulders` | Profile photos |
| `upper-body` | 4.5× face | LinkedIn / formal | | `upper-body` | LinkedIn / formal |
| `half-body` | 7.0× face | Full upper body | | `half-body` | Full upper body |
## Image Enhancement ## Transcribe Audio
**Function:** `analyzeImage` + `applyCorrections` **Tool route:** `transcribe-audio`
**Tool route:** `image-enhancement` **Model:** faster-whisper
**Engine:** Analysis-based (Sharp histogram and statistics)
Analyzes the image and applies automatic corrections for exposure, contrast, white balance, saturation, sharpness, and noise. Supports scene-specific modes. Converts speech to text. Supports plain text, SRT, and VTT output formats.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `mode` | `auto` \| `portrait` \| `landscape` \| `low-light` \| `food` \| `document` | `auto` | Scene mode for tuning corrections | | `language` | string | `"auto"` | Language: `auto`, `en`, `de`, `fr`, `es`, `zh`, `ja`, `ko`, `id`, `th`, `vi` |
| `intensity` | number (0-100) | 50 | Overall correction strength | | `outputFormat` | `"txt"` \| `"srt"` \| `"vtt"` | `"txt"` | Output format |
| `corrections.exposure` | boolean | true | Apply exposure correction |
| `corrections.contrast` | boolean | true | Apply contrast correction |
| `corrections.whiteBalance` | boolean | true | Apply white balance correction |
| `corrections.saturation` | boolean | true | Apply saturation correction |
| `corrections.sharpness` | boolean | true | Apply sharpness correction |
| `corrections.denoise` | boolean | true | Apply denoising |
An additional analysis endpoint is available at `POST /api/v1/tools/image-enhancement/analyze` which returns the detected corrections without applying them. ## Auto Subtitles
## Content-Aware Resize (Seam Carving) **Tool route:** `auto-subtitles`
**Model:** faster-whisper (extracts audio from video, then transcribes)
**Function:** `seamCarve` Generates subtitle files from a video's audio track.
**Tool route:** `content-aware-resize`
**Engine:** Go `caire` binary (not Python - no GPU benefit)
Intelligently resizes images by removing or adding low-energy seams, preserving important content.
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `width` | number | - | Target width | | `language` | string | `"auto"` | Language: `auto`, `en`, `de`, `fr`, `es`, `zh`, `ja`, `ko`, `id`, `th`, `vi` |
| `height` | number | - | Target height | | `format` | `"srt"` \| `"vtt"` | `"srt"` | Output subtitle format |
| `protectFaces` | boolean | true | Protect detected face regions from seam removal |
| `blurRadius` | number | 0 | Pre-blur to reduce noise sensitivity |
| `sobelThreshold` | number | 10 | Edge sensitivity threshold |
| `square` | boolean | false | Force square output |
Max input edge before auto-downscaling: **1200 px**.
## PNG Transparency Fixer ## PNG Transparency Fixer
**Function:** `fixTransparency`
**Tool route:** `transparency-fixer` **Tool route:** `transparency-fixer`
**Model:** BiRefNet HR-matting (2048x2048 resolution) **Model:** BiRefNet HR-matting (2048x2048 resolution)
@@ -271,12 +363,11 @@ Fixes "fake transparent" PNGs where the background was removed but left behind f
**OOM fallback chain:** If BiRefNet HR-matting exceeds available memory, the tool automatically falls back to `birefnet-general`, then to `u2net`. **OOM fallback chain:** If BiRefNet HR-matting exceeds available memory, the tool automatically falls back to `birefnet-general`, then to `u2net`.
**Feature bundle:** Background Removal (shared with Remove Background and Passport Photo).
| Parameter | Type | Default | Description | | Parameter | Type | Default | Description |
|-----------|------|---------|-------------| |-----------|------|---------|-------------|
| `defringe` | number (0-100) | 30 | Edge defringe strength to remove color contamination | | `defringe` | number (0-100) | `30` | Edge defringe strength to remove color contamination |
| `outputFormat` | `"png"` \| `"webp"` | `"png"` | Output image format | | `outputFormat` | `"png"` \| `"webp"` | `"png"` | Output image format |
| `removeWatermark` | boolean | `false` | Apply watermark removal pre-processing (median filter) |
```bash ```bash
curl -X POST http://localhost:1349/api/v1/tools/transparency-fixer \ curl -X POST http://localhost:1349/api/v1/tools/transparency-fixer \
@@ -284,3 +375,46 @@ curl -X POST http://localhost:1349/api/v1/tools/transparency-fixer \
-F "file=@fake-transparent.png" \ -F "file=@fake-transparent.png" \
-F 'settings={"defringe":30,"outputFormat":"png"}' -F 'settings={"defringe":30,"outputFormat":"png"}'
``` ```
---
## Tools with Optional AI Capabilities
The following tools are not Python sidecar tools but use AI features when certain options are enabled.
### Image Enhancement
**Tool route:** `image-enhancement`
**Engine:** Analysis-based (Sharp histogram and statistics)
Analyzes the image and applies automatic corrections for exposure, contrast, white balance, saturation, sharpness, and noise. Supports scene-specific modes.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `mode` | `"auto"` \| `"portrait"` \| `"landscape"` \| `"low-light"` \| `"food"` \| `"document"` | `"auto"` | Scene mode for tuning corrections |
| `intensity` | number (0-100) | `50` | Overall correction strength |
| `corrections.exposure` | boolean | `true` | Apply exposure correction |
| `corrections.contrast` | boolean | `true` | Apply contrast correction |
| `corrections.whiteBalance` | boolean | `true` | Apply white balance correction |
| `corrections.saturation` | boolean | `true` | Apply saturation correction |
| `corrections.sharpness` | boolean | `true` | Apply sharpness correction |
| `corrections.denoise` | boolean | `true` | Apply denoising |
| `deepEnhance` | boolean | `false` | Enable AI noise removal via SCUNet (requires `upscale-enhance` bundle) |
An additional analysis endpoint is available at `POST /api/v1/tools/image-enhancement/analyze` which returns the detected corrections without applying them.
### Content-Aware Resize (Seam Carving)
**Tool route:** `content-aware-resize`
**Engine:** Go `caire` binary (not Python -- no GPU benefit)
Intelligently resizes images by removing low-energy seams, preserving important content.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `width` | number | - | Target width |
| `height` | number | - | Target height |
| `protectFaces` | boolean | `false` | Protect detected face regions (requires `face-detection` bundle) |
| `blurRadius` | number (0-20) | `4` | Pre-blur for energy calculation |
| `sobelThreshold` | number (1-20) | `2` | Edge sensitivity threshold |
| `square` | boolean | `false` | Force square output |
+96 -16
View File
@@ -32,6 +32,7 @@ Cut out a rectangular region from the image.
| `top` | number | Y offset from the top edge | | `top` | number | Y offset from the top edge |
| `width` | number | Width of the crop area | | `width` | number | Width of the crop area |
| `height` | number | Height of the crop area | | `height` | number | Height of the crop area |
| `unit` | string | `px` (default) or `percent` |
### rotate ### rotate
@@ -40,15 +41,16 @@ Rotate the image by a given angle.
| Parameter | Type | Description | | Parameter | Type | Description |
|---|---|---| |---|---|---|
| `angle` | number | Rotation angle in degrees (0-360) | | `angle` | number | Rotation angle in degrees (0-360) |
| `background` | string | Fill color for the exposed area (default: transparent or white) | | `background` | string | Fill color for exposed area (default: `#000000`). Only applies to non-90-degree angles. |
### flip ### flip
Mirror the image horizontally or vertically. Mirror the image horizontally, vertically, or both. At least one must be true.
| Parameter | Type | Description | | Parameter | Type | Description |
|---|---|---| |---|---|---|
| `direction` | string | `horizontal` or `vertical` | | `horizontal` | boolean | Mirror left to right |
| `vertical` | boolean | Mirror top to bottom |
### convert ### convert
@@ -56,9 +58,11 @@ Change the image format.
| Parameter | Type | Description | | Parameter | Type | Description |
|---|---|---| |---|---|---|
| `format` | string | Target format: `jpeg`, `png`, `webp`, `avif`, `tiff`, `gif`, `jxl`, `heic`, `heif`, `bmp`, `ico`, `jp2`, `qoi` | | `format` | string | Target format: `jpg`, `png`, `webp`, `avif`, `tiff`, `gif`, `jxl`, `heic`, `heif`, `bmp`, `ico`, `jp2`, `qoi` |
| `quality` | number | Compression quality (1-100, applies to lossy formats) | | `quality` | number | Compression quality (1-100, applies to lossy formats) |
The first seven formats (`jpg` through `jxl`) are encoded by Sharp in-process. The remaining formats use external encoders at the API layer: `heic`/`heif` via heif-enc, `bmp`/`ico` via ImageMagick, `jp2` via opj_compress, and `qoi` via an inline TypeScript codec.
### compress ### compress
Reduce file size while keeping the same format. Reduce file size while keeping the same format.
@@ -66,11 +70,20 @@ Reduce file size while keeping the same format.
| Parameter | Type | Description | | Parameter | Type | Description |
|---|---|---| |---|---|---|
| `quality` | number | Target quality (1-100) | | `quality` | number | Target quality (1-100) |
| `targetSizeBytes` | number | Optional target file size in bytes |
| `format` | string | Optional format override | | `format` | string | Optional format override |
### strip-metadata ### strip-metadata
Remove EXIF, IPTC, and XMP metadata from the image. Useful for privacy before sharing photos publicly. Takes no parameters. Remove EXIF, IPTC, XMP, and ICC metadata from the image. With no parameters (or `stripAll: true`), strips everything. Pass individual flags for selective stripping.
| Parameter | Type | Description |
|---|---|---|
| `stripAll` | boolean | Strip all metadata (default when no flags are set) |
| `stripExif` | boolean | Strip EXIF data (including GPS if `stripGps` is not separately set) |
| `stripGps` | boolean | Strip GPS location data |
| `stripIcc` | boolean | Strip ICC color profile |
| `stripXmp` | boolean | Strip XMP metadata |
### Color adjustments ### Color adjustments
@@ -94,35 +107,102 @@ These apply a fixed color transformation. They take no parameters.
### Color channels ### Color channels
Adjust individual RGB color channels. Adjust individual RGB color channels. Values are multipliers where 100 = no change.
| Parameter | Type | Description | | Parameter | Type | Description |
|---|---|---| |---|---|---|
| `red` | number | Red channel adjustment (-100 to 100) | | `red` | number | Red channel multiplier (0 to 200, 100 = unchanged) |
| `green` | number | Green channel adjustment (-100 to 100) | | `green` | number | Green channel multiplier (0 to 200, 100 = unchanged) |
| `blue` | number | Blue channel adjustment (-100 to 100) | | `blue` | number | Blue channel multiplier (0 to 200, 100 = unchanged) |
### sharpen
Simple sharpening controlled by a single value.
| Parameter | Type | Description |
|---|---|---|
| `value` | number | Sharpening intensity (0 to 100). Mapped to a Gaussian sigma of 0.5-10. |
### sharpen-advanced
Advanced sharpening with three selectable methods and an optional noise-reduction pre-pass.
| Parameter | Type | Description |
|---|---|---|
| `method` | string | `adaptive`, `unsharp-mask`, or `high-pass` |
| `sigma` | number | Gaussian blur radius, 0.5-10 (adaptive) |
| `m1` | number | Flat-area sharpening, 0-10 (adaptive) |
| `m2` | number | Textured-area sharpening, 0-20 (adaptive) |
| `x1` | number | Flat/jagged threshold, 0-10 (adaptive) |
| `y2` | number | Max brightening (halo clamp), 0-50 (adaptive) |
| `y3` | number | Max darkening (halo clamp), 0-50 (adaptive) |
| `amount` | number | Intensity percentage, 0-500 (unsharp-mask) |
| `radius` | number | Blur radius, 0.1-5.0 (unsharp-mask) |
| `threshold` | number | Minimum edge brightness, 0-255 (unsharp-mask) |
| `strength` | number | Blend strength, 0-100 (high-pass) |
| `kernelSize` | number | `3` or `5` for 3x3 / 5x5 kernel (high-pass) |
| `denoise` | string | Noise reduction pre-pass: `off`, `light`, `medium`, or `strong` |
Parameters are method-specific. Only supply the ones relevant to the chosen method.
### color-blindness
Simulate a color vision deficiency using a 3x3 color-recombination matrix.
| Parameter | Type | Description |
|---|---|---|
| `type` | string | One of: `protanopia`, `deuteranopia`, `tritanopia`, `protanomaly`, `deuteranomaly`, `tritanomaly`, `achromatopsia`, `blueConeMonochromacy` |
### edit-metadata
Write or remove individual EXIF/IPTC metadata fields without stripping the entire block.
| Parameter | Type | Description |
|---|---|---|
| `artist` | string | EXIF Artist tag |
| `copyright` | string | EXIF Copyright tag |
| `imageDescription` | string | EXIF ImageDescription tag |
| `software` | string | EXIF Software tag |
| `dateTime` | string | EXIF DateTime tag |
| `dateTimeOriginal` | string | EXIF DateTimeOriginal tag |
| `clearGps` | boolean | Remove all GPS tags |
| `fieldsToRemove` | string[] | List of EXIF field names to delete |
All parameters are optional. Fields listed in `fieldsToRemove` are deleted from the existing EXIF block. Fields set via the named parameters are written (or overwritten). Binary/unsafe keys like MakerNote are silently ignored.
## Format detection ## Format detection
The engine detects input formats automatically from file headers, not just file extensions. This means a `.jpg` file that is actually a PNG will be handled correctly. Detection uses a multi-layer approach: magic bytes first, then file extension as fallback. The engine detects input formats automatically from file headers, not just file extensions. This means a `.jpg` file that is actually a PNG will be handled correctly. Detection uses a multi-layer approach: magic bytes first, then file extension as fallback.
SnapOtter supports **55+ input formats** and **14 output formats**, including 23 camera RAW formats from 20+ brands, professional formats (PSD, EPS, OpenEXR, HDR), modern codecs (JPEG XL, AVIF, HEIC, QOI, JPEG 2000), and scientific/gaming formats (FITS, DDS). Decoding is handled by Sharp natively where possible, with automatic fallback to ImageMagick, LibRaw, and specialized CLI decoders. SnapOtter supports **55+ input formats** and **13 output formats**, including 23 camera RAW formats from 20+ brands, professional formats (PSD, EPS, OpenEXR, HDR), modern codecs (JPEG XL, AVIF, HEIC, QOI, JPEG 2000), and scientific/gaming formats (FITS, DDS). Decoding is handled by Sharp natively where possible, with automatic fallback to ImageMagick, LibRaw, and specialized CLI decoders.
See the [Supported Formats](/guide/supported-formats) page for the complete list. See the [Supported Formats](/guide/supported-formats) page for the complete list.
## Metadata extraction ## Metadata extraction
The `info` tool returns image metadata: The `info` tool returns image metadata. See [Image Info](/tools/image/info) for the full field reference.
```json ```json
{ {
"width": 1920, "filename": "photo.jpg",
"height": 1080, "fileSize": 2450000,
"width": 4032,
"height": 3024,
"format": "jpeg", "format": "jpeg",
"size": 245678,
"channels": 3, "channels": 3,
"hasAlpha": false, "hasAlpha": false,
"dpi": 72, "colorSpace": "srgb",
"exif": { ... } "density": 72,
"isProgressive": false,
"hasExif": true,
"hasIcc": true,
"hasXmp": false,
"bitDepth": "8",
"pages": 1,
"histogram": [
{ "channel": "red", "min": 0, "max": 255, "mean": 128.45, "stdev": 52.31 },
{ "channel": "green", "min": 2, "max": 253, "mean": 115.22, "stdev": 48.76 },
{ "channel": "blue", "min": 0, "max": 250, "mean": 102.89, "stdev": 55.14 }
]
} }
``` ```
+6 -4
View File
@@ -102,8 +102,9 @@ curl -X POST http://localhost:1349/api/v1/tools/<toolId>/batch \
- Upload is `multipart/form-data`. - Upload is `multipart/form-data`.
- `settings` is a JSON string with tool-specific options. - `settings` is a JSON string with tool-specific options.
- Response is the processed file directly (or a ZIP for batch). - **Fast tools** (200) return JSON: `{"jobId":"...","downloadUrl":"/api/v1/download/<jobId>/<filename>","originalSize":1234,"processedSize":567}`. Fetch the processed file from `downloadUrl`.
- Progress is tracked via SSE (see [Progress Tracking](#progress-tracking)). - **Long-running tools** (202) return JSON: `{"jobId":"...","async":true}`. Connect to SSE for progress, then download when complete (see [Progress Tracking](#progress-tracking)).
- **Batch** returns a ZIP archive streamed directly (with `X-Job-Id` header).
## Tools Reference ## Tools Reference
@@ -163,6 +164,7 @@ All AI tools run on your hardware (CPU or NVIDIA GPU). No internet required.
| `transparency-fixer` | PNG Transparency Fixer | BiRefNet HR-matting | `defringe` (0-100), `outputFormat` (png/webp) | | `transparency-fixer` | PNG Transparency Fixer | BiRefNet HR-matting | `defringe` (0-100), `outputFormat` (png/webp) |
| `background-replace` | Background Replace | rembg (BiRefNet) | `backgroundType` (color/gradient), `color` (hex), `gradientColor1`, `gradientColor2`, `gradientAngle`, `feather` (0-20), `format` (png/webp) | | `background-replace` | Background Replace | rembg (BiRefNet) | `backgroundType` (color/gradient), `color` (hex), `gradientColor1`, `gradientColor2`, `gradientAngle`, `feather` (0-20), `format` (png/webp) |
| `blur-background` | Blur Background | rembg (BiRefNet) | `intensity` (1-100), `feather` (0-20), `format` (png/webp) | | `blur-background` | Blur Background | rembg (BiRefNet) | `intensity` (1-100), `feather` (0-20), `format` (png/webp) |
| `ai-canvas-expand` | AI Canvas Expand | LaMa (outpainting) | `extendTop`, `extendRight`, `extendBottom`, `extendLeft` (px), `tier` (fast/balanced/high), `format`, `quality` |
### Watermark & Overlay ### Watermark & Overlay
@@ -435,7 +437,7 @@ Each step's output is the next step's input. Unlimited steps per pipeline by def
Long-running jobs (AI tools, batch, pipelines) emit real-time progress via Server-Sent Events: Long-running jobs (AI tools, batch, pipelines) emit real-time progress via Server-Sent Events:
```bash ```bash
# Connect to the SSE stream (jobId returned in X-Job-Id response header) # Connect to the SSE stream (jobId is in the JSON response body from the tool endpoint)
curl -N http://localhost:1349/api/v1/jobs/<jobId>/progress \ curl -N http://localhost:1349/api/v1/jobs/<jobId>/progress \
-H "Authorization: Bearer <token>" -H "Authorization: Bearer <token>"
``` ```
@@ -505,7 +507,7 @@ Custom role management with granular permissions.
| `PUT` | `/api/v1/roles/:id` | Admin (`users:manage`) | Update a custom role (cannot modify built-in roles) | | `PUT` | `/api/v1/roles/:id` | Admin (`users:manage`) | Update a custom role (cannot modify built-in roles) |
| `DELETE` | `/api/v1/roles/:id` | Admin (`users:manage`) | Delete a custom role (cannot delete built-in roles; affected users revert to `user` role) | | `DELETE` | `/api/v1/roles/:id` | Admin (`users:manage`) | Delete a custom role (cannot delete built-in roles; affected users revert to `user` role) |
Available permissions: `tools:use`, `files:own`, `files:all`, `apikeys:own`, `apikeys:all`, `pipelines:own`, `pipelines:all`, `settings:read`, `settings:write`, `users:manage`, `teams:manage`, `features:manage`, `system:health`, `audit:read`. Available permissions (17): `tools:use`, `files:own`, `files:all`, `apikeys:own`, `apikeys:all`, `pipelines:own`, `pipelines:all`, `settings:read`, `settings:write`, `users:manage`, `teams:manage`, `features:manage`, `system:health`, `audit:read`, `compliance:manage`, `webhooks:manage`, `security:manage`.
## Audit Log ## Audit Log
+2 -2
View File
@@ -8,7 +8,7 @@ SnapOtter processes files across five modalities: image, video, audio, document,
## Image Formats ## Image Formats
SnapOtter supports 55+ image formats for input and 14 formats for output. SnapOtter supports 55+ image formats for input and 13 formats for output.
## Input Formats ## Input Formats
@@ -101,7 +101,7 @@ SnapOtter supports 55+ image formats for input and 14 formats for output.
| PAM | .pam | Sharp (native) | Arbitrary map | | PAM | .pam | Sharp (native) | Arbitrary map |
| PFM | .pfm | Sharp (native) | Float map | | PFM | .pfm | Sharp (native) | Float map |
## Output Formats (14) ## Output Formats (13)
| Format | Encoder | Quality Control | Available In | | Format | Encoder | Quality Control | Available In |
|--------|---------|----------------|-------------| |--------|---------|----------------|-------------|