mirror of
https://github.com/guillaumemeyer/watermarks-remover.git
synced 2026-08-22 13:11:57 +02:00
* feat: split skill from service, add HTTP API and Docker distribution The agent skill (skills/remove-ai-marks/) is now a code-free remote client: all implementation moved to service/scripts/ and runs behind a stdlib HTTP service (server.py) with /health, /capabilities, /inspect, /clean and a dynamically generated OpenAPI 3.0.3 spec at /openapi.json. - Move scripts/ and the backend Dockerfiles under service/ - server.py: JSON/base64 HTTP entrypoint with size caps, binary guard, atomic writes, loopback default, optional bearer auth - Core Dockerfile (exiftool/qpdf/c2patool preinstalled) and a GHCR publish workflow for the core/markllm/markdiffusion images - compose.yaml (wr-* services, harness/heavy profiles) + compose-check.sh to validate the running stack (exit code only) - Fix markllm image build (tokenizers 0.22.2, CPU-only torch) and ctrlregen build (python:3.11 base for the 2023-era research pins) - Fix markllm/markdiffusion harness images missing common.py at runtime * docs: add .env.example and service configuration guide * fix: disable chain-of-thought for openai-compatible Layer B rewrites deepseek-v4-flash is a reasoning model: a one-line paraphrase burned 9,894 reasoning tokens (~100s) and hit the default timeout. Send reasoning_effort=none by default for the openai-compatible backend (--reasoning-effort / WATERMARKS_REWRITE_REASONING_EFFORT; 'off' omits the parameter), cutting the same rewrite to ~1s / 12 tokens. Tested end-to-end against api.deepseek.com. * fix: sanitize client-supplied filename in HTTP service CodeQL 'uncontrolled data in path expression' (server.py): a name like '../../x' flowed into Path(tmpdir) / name, letting an upload escape the request temp dir on write. Sanitize name to its basename in _decode_input (_safe_name) and refuse any joined path whose parent is not the tmpdir at the write sites (_tmp_path). Tests cover traversal names. * chore: gitignore .env (contains local rewrite credentials) * chore: deny-by-default gitignore and dockerignore; document compose env config .gitignore and service/.dockerignore now exclude everything by default and explicitly allow only what is publishable/needed: tracked source, docs, tests, .github, and (for images) the service/scripts/ tree that every Dockerfile COPYs. Root .dockerignore documents that all builds use service/ as context. README Configuration section now covers .env setup for docker compose, host-side export for CLI runs, and the full variable table.
36 lines
1.9 KiB
Bash
36 lines
1.9 KiB
Bash
# Copy to .env for `docker compose` (docker compose auto-loads .env from the
|
|
# repo root). Everything here is optional — the core service works with no
|
|
# configuration at all.
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Core HTTP service (used by wr-core)
|
|
# ---------------------------------------------------------------------------
|
|
# Optional bearer token for the HTTP API. When set, every request must send
|
|
# `Authorization: Bearer <key>`.
|
|
WATERMARKS_SERVER_API_KEY=
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Harness / heavy backends (only used by the harness/heavy profiles)
|
|
# ---------------------------------------------------------------------------
|
|
# Optional Hugging Face token for gated models (CtrlRegen, MarkLLM,
|
|
# MarkDiffusion score models). Env only — never on argv.
|
|
HF_TOKEN=
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Client-side (used by the skill or curl, NOT by compose)
|
|
# ---------------------------------------------------------------------------
|
|
# Where to reach the service. Defaults to http://127.0.0.1:8765.
|
|
# WATERMARKS_SERVICE_URL=http://127.0.0.1:8765
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Layer B statistical-watermark rewrite (only for the rewrite_text.py hook,
|
|
# which runs inside the core image or a local checkout; the agent skill does
|
|
# Layer B itself with its own model and does not need these)
|
|
# ---------------------------------------------------------------------------
|
|
# WATERMARKS_REWRITE_BACKEND=ollama # or: openai-compatible
|
|
# WATERMARKS_REWRITE_MODEL=llama3.2
|
|
# WATERMARKS_REWRITE_BASE_URL=http://127.0.0.1:11434
|
|
# WATERMARKS_REWRITE_API_KEY= # env only, never on argv
|
|
# WATERMARKS_REWRITE_ALLOW_REMOTE=1 # only for non-loopback endpoints
|
|
# WATERMARKS_REWRITE_REASONING_EFFORT=none # none/low/medium/high, or off to omit
|