mirror of
https://github.com/guillaumemeyer/watermarks-remover.git
synced 2026-08-22 13:11:57 +02:00
* feat: add reproducible SynthID-text removal benchmark bench_synthid_text.py orchestrates the existing Layer B machinery into a controlled, shareable experiment: generate watermarked + unwatermarked samples with the MarkLLM SynthID scheme, run removal variants (strength x candidates) plus controls (no-removal, Layer-A-only, optional re-stamp), and report clear rate, score suppression, quality, and cost (tokens, wall time, USD) with a clears-per-MTok efficiency ratio. Emits report.md / results.json / results.csv with the exact reproduction command and pinned commits; optional Gemini official-detector tier when WATERMARKS_GEMINI_API_KEY is set. Mock-based tests, no torch in CI. * docs: add README section on running the SynthID-text benchmark Explains what LLM performs the Layer B rewrite (an external model configured via WATERMARKS_REWRITE_* env vars or --rewrite-* flags; MarkLLM's opt-1.3b is only the watermark generator/detector) and how to run a benchmark with Ollama or an OpenAI-compatible endpoint, plus the non-origin-model re-stamp caveat. * fix: honor WATERMARKS_REWRITE_ALLOW_REMOTE in the SynthID-text benchmark The --rewrite-allow-remote flag now defaults from the env var (matching rewrite_text.py and the other WATERMARKS_REWRITE_* settings), so a non-loopback rewrite endpoint works after sourcing .env without an extra flag. * fix: MarkLLM sparse checkout and deps for the SynthID harness - setup_markllm.sh sparse-checkout omitted '/visualize/', which watermark/base.py imports at module load — every scheme (incl. SynthID) failed with 'No module named visualize' during generation/detection. - requirements-markllm.txt omitted scikit-learn, imported by the SynthID detector (watermark/synthid/detector_bayesian_torch.py). Both broke the MarkLLM harness at runtime; the benchmark's sanity gate then excluded every sample, producing empty per-variant results. * fix: drop 4 GiB RLIMIT_AS on benchmark subprocesses _run_cmd applied the common child RLIMIT_AS (default 4 GiB) via subprocess_preexec_fn to every MarkLLM/rewrite child. torch needs a much larger address space: CUDA init failed with 'out of memory' at cudaGetDeviceCount and the 5.2 GB fp32 opt-1.3b could not load, so every sample was excluded at generation. text_detectors.py already applies no address-space cap to MarkLLM by default; the benchmark now matches. * perf: keep MarkLLM resident via a serve worker (624 cold starts -> 1) The benchmark spawned a fresh torch + opt-1.3b process per operation (~60-90s each); a full run needs ~624 of them. detect_text_watermark.py gains a 'serve' mode (JSON-lines over stdin/stdout, ready handshake) that loads the model once; bench_synthid_text.py uses it via MarkLLMWorker with automatic fallback to one-shot subprocesses (--no-worker to force). Turns ~8h runs into ~40-60min. * perf: skip per-candidate Gemini detections in rewrite subprocess * feat: run the SynthID-text benchmark from the wr-markllm compose service - Dockerfile.markllm: add '/visualize/' to the sparse checkout (same fix as setup_markllm.sh) and COPY the benchmark + rewrite scripts (stdlib-only). - compose.yaml: wr-markllm gets the WATERMARKS_REWRITE_* and WATERMARKS_GEMINI_* env wiring, a bench-out volume for --out-dir, and a read-only mount of the bundled corpus (build context is service/, so the corpus cannot be COPY'd). - docs: docker compose run example. Note: the image ships CPU torch by design, so the container path is for portability/CI; GPU runs use the host setup_markllm.sh venv. * feat: per-sample progress logging in the benchmark The persistent worker returns samples in-memory, so nothing is written until the end of a run — runs looked stuck. eprint a [gen i/N] line per generated sample and a [removal] summary per sample. * chore: migrate Gemini config to gemini-3.6-flash; document SynthID-text retirement Google retired SynthID text watermarking on the Generative Language API (Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK is rejected on current 3.x models (confirmed by Google AI staff). Migrate the default detection model to gemini-3.6-flash, document the retirement in vendor-notes.md and the benchmark report caveat, and keep the detector seam fail-soft until a vendor endpoint (e.g. Vertex AI) returns. * feat: remove gemini-synthid-text detector (Google retired text watermarking) Google removed SynthID text watermarking from the Generative Language API (Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK is rejected on current 3.x models, so the vendor detector had nothing to detect. Remove GeminiSynthIDTextDetector and its wiring: - text_detectors.py: drop the Gemini class, HTTP helpers, and constants; keep MarkLLM + Claude seams (registry now markllm + claude-text). - server.py / rewrite_text.py: per-candidate detection now triggers on --markllm-scheme only. - bench_synthid_text.py: remove the Gemini tier (before/after, report table, --no-gemini flag); report caveat notes the retirement. - configs/docs: drop WATERMARKS_GEMINI_* from .env.example / compose / README / SKILL.md / vendor-notes.md; keep the retirement note. - tests: gemini tests removed or converted to MarkLLM (mocked subprocess). - Dockerfile.markllm: parameterize BASE_IMAGE + TORCH_INDEX_URL so a GPU/ arm64 image can be built (used for the --gpus all benchmark run). * fix: harden notes aggregation against non-string notes A run completed all samples but crashed at the final aggregate step with 'cannot use list as a set element' when a row's notes contained a non-string value. Filter notes to strings (aggregate + CSV) and add a regression test. * perf: let the rewrite subprocess reuse the resident MarkLLM worker The rewrite subprocess (rewrite_text.py) ran its own before/after MarkLLM detects, each a ~20s torch+model cold start (~12 per sample = ~5min of the ~6min/sample runtime). Now: - detect_text_watermark.py serve gains --port N: a loopback TCP JSON-lines listener (default -1 = off) sharing the resident model, with a lock so stdin and socket requests never run the model concurrently. - text_detectors.MarkLLMTextDetector checks WATERMARKS_MARKLLM_PORT and does a fast loopback detect when a worker is up, falling back to the one-shot subprocess otherwise. - The benchmark worker publishes its port via that env var, so the rewrite subprocess inherits it and its detects hit the resident model. Turns ~6 min/sample into ~1-2 min; a full run drops from ~2h to ~40-50min. Tests: loopback-client + fallback + env-publish coverage. * chore: add benchmark-smoke.sh / benchmark-full.sh wrappers Simple host wrappers: source .env, default MARKLLM_DIR to ~/MarkLLM, use a repo-local HF cache by default, and run bench_synthid_text.py with a quick (2 docs, 1 seed, paraphrase:1) or full (8 docs x 3 seeds, three variants, re-stamp control) configuration. OUT_DIR overrides the output location.
158 lines
5.4 KiB
YAML
158 lines
5.4 KiB
YAML
# Whole-infra bring-up for watermarks-remover.
|
|
#
|
|
# docker compose up --build -d # core HTTP service only
|
|
# docker compose --profile harness up --build -d # + markllm / markdiffusion harnesses
|
|
# docker compose --profile heavy up --build -d # + ctrlregen / synthid (local builds)
|
|
#
|
|
# The skill and any web app talk to the wr-core service at http://127.0.0.1:8765
|
|
# (loopback-only host mapping). The heavy/harness images are one-shot CLIs:
|
|
# `up` starts them with `--help` to confirm the image, and real jobs run via
|
|
# `docker compose run`, e.g.:
|
|
# docker compose run --rm wr-ctrlregen /data/shot.png -o /data/out.png
|
|
#
|
|
# ctrlregen and synthid bake in upstream code that is not publicly
|
|
# redistributable (all-rights-reserved / non-commercial Research License), so
|
|
# they build from source locally and are never pushed to GHCR.
|
|
|
|
name: watermarks-remover
|
|
|
|
services:
|
|
wr-core:
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile
|
|
image: ghcr.io/guillaumemeyer/watermarks-remover:latest
|
|
ports:
|
|
- "127.0.0.1:8765:8765"
|
|
environment:
|
|
# Empty by default (no auth). Set to require `Authorization: Bearer <key>`.
|
|
WATERMARKS_SERVER_API_KEY: ${WATERMARKS_SERVER_API_KEY:-}
|
|
# Optional SynthID image scorer sidecar (heavy profile).
|
|
WATERMARKS_SYNTHID_SCORER_URL: ${WATERMARKS_SYNTHID_SCORER_URL:-}
|
|
WATERMARKS_SYNTHID_SCORER_API_KEY: ${WATERMARKS_SYNTHID_SCORER_API_KEY:-}
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
restart: unless-stopped
|
|
|
|
wr-markllm:
|
|
profiles: [harness]
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile.markllm
|
|
image: ghcr.io/guillaumemeyer/watermarks-remover:markllm-latest
|
|
# One-shot CLI: `up` just confirms the image; run real jobs with
|
|
# `docker compose run --rm wr-markllm detect /data/wm.txt --scheme kgw`.
|
|
command: ["--help"]
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
environment:
|
|
HF_TOKEN: ${HF_TOKEN:-}
|
|
HF_HOME: /home/markllm/.cache/huggingface
|
|
# Layer B rewrite backend (openai-compatible / ollama).
|
|
WATERMARKS_REWRITE_BACKEND: ${WATERMARKS_REWRITE_BACKEND:-}
|
|
WATERMARKS_REWRITE_MODEL: ${WATERMARKS_REWRITE_MODEL:-}
|
|
WATERMARKS_REWRITE_BASE_URL: ${WATERMARKS_REWRITE_BASE_URL:-}
|
|
WATERMARKS_REWRITE_API_KEY: ${WATERMARKS_REWRITE_API_KEY:-}
|
|
WATERMARKS_REWRITE_ALLOW_REMOTE: ${WATERMARKS_REWRITE_ALLOW_REMOTE:-}
|
|
WATERMARKS_REWRITE_REASONING_EFFORT: ${WATERMARKS_REWRITE_REASONING_EFFORT:-}
|
|
volumes:
|
|
- markllm-cache:/home/markllm/.cache/huggingface
|
|
# Writable output for the benchmark + the bundled seed corpus (the
|
|
# build context is service/, so the corpus is mounted at runtime):
|
|
# docker compose run --rm wr-markllm \
|
|
# /app/bench_synthid_text.py --markllm-dir /opt/markllm \
|
|
# --corpus /bench-corpus --out-dir /data ...
|
|
- bench-out:/data
|
|
- ./benchmarks/corpus:/bench-corpus:ro
|
|
|
|
wr-markdiffusion:
|
|
profiles: [harness]
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile.markdiffusion
|
|
image: ghcr.io/guillaumemeyer/watermarks-remover:markdiffusion-latest
|
|
command: ["--help"]
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
environment:
|
|
HF_TOKEN: ${HF_TOKEN:-}
|
|
HF_HOME: /home/markdiffusion/.cache/huggingface
|
|
volumes:
|
|
- markdiffusion-cache:/home/markdiffusion/.cache/huggingface
|
|
|
|
wr-ctrlregen:
|
|
profiles: [heavy]
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile.ctrlregen
|
|
# Local-only tag: upstream `noai-watermark` ships no LICENSE (all-rights-
|
|
# reserved), so this image is never published.
|
|
image: watermarks-remover-ctrlregen:local
|
|
command: ["--help"]
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
environment:
|
|
HF_TOKEN: ${HF_TOKEN:-}
|
|
NOAI_WATERMARK_DIR: /opt/noai-watermark
|
|
volumes:
|
|
- ctrlregen-cache:/home/remover/.cache/huggingface
|
|
|
|
wr-synthid:
|
|
profiles: [heavy]
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile.synthid
|
|
# Local-only tag: upstream `reverse-SynthID` is under a non-commercial
|
|
# Research License, so this image is never published.
|
|
image: watermarks-remover-synthid-scorer:local
|
|
command: ["--help"]
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
volumes:
|
|
- synthid-cache:/home/scorer/.cache/huggingface
|
|
|
|
# HTTP SynthID scorer sidecar: lets wr-core score images before/after
|
|
# cleaning without bundling the non-commercial reverse-SynthID code in the
|
|
# published core image. Reach it via WATERMARKS_SYNTHID_SCORER_URL on
|
|
# wr-core (see .env.example) and require the same bearer key on both sides.
|
|
wr-synthid-score:
|
|
profiles: [heavy]
|
|
build:
|
|
context: service
|
|
dockerfile: Dockerfile.synthid
|
|
image: watermarks-remover-synthid-scorer:local
|
|
entrypoint: ["python3"]
|
|
command: ["/app/synthid_score_server.py", "--host", "0.0.0.0", "--port", "8766"]
|
|
read_only: true
|
|
tmpfs:
|
|
- /tmp
|
|
init: true
|
|
user: "10001:10001"
|
|
environment:
|
|
WATERMARKS_SYNTHID_SCORER_API_KEY: ${WATERMARKS_SYNTHID_SCORER_API_KEY:-}
|
|
REVERSE_SYNTHID_DIR: /opt/reverse-synthid
|
|
volumes:
|
|
- synthid-cache:/home/scorer/.cache/huggingface
|
|
|
|
volumes:
|
|
markllm-cache:
|
|
bench-out:
|
|
markdiffusion-cache:
|
|
ctrlregen-cache:
|
|
synthid-cache:
|