feat: add reproducible SynthID-text removal benchmark (#145)

* feat: add reproducible SynthID-text removal benchmark

bench_synthid_text.py orchestrates the existing Layer B machinery into a
controlled, shareable experiment: generate watermarked + unwatermarked
samples with the MarkLLM SynthID scheme, run removal variants (strength x
candidates) plus controls (no-removal, Layer-A-only, optional re-stamp),
and report clear rate, score suppression, quality, and cost (tokens,
wall time, USD) with a clears-per-MTok efficiency ratio.

Emits report.md / results.json / results.csv with the exact reproduction
command and pinned commits; optional Gemini official-detector tier when
WATERMARKS_GEMINI_API_KEY is set. Mock-based tests, no torch in CI.

* docs: add README section on running the SynthID-text benchmark

Explains what LLM performs the Layer B rewrite (an external model configured
via WATERMARKS_REWRITE_* env vars or --rewrite-* flags; MarkLLM's opt-1.3b is
only the watermark generator/detector) and how to run a benchmark with Ollama
or an OpenAI-compatible endpoint, plus the non-origin-model re-stamp caveat.

* fix: honor WATERMARKS_REWRITE_ALLOW_REMOTE in the SynthID-text benchmark

The --rewrite-allow-remote flag now defaults from the env var (matching
rewrite_text.py and the other WATERMARKS_REWRITE_* settings), so a
non-loopback rewrite endpoint works after sourcing .env without an extra
flag.

* fix: MarkLLM sparse checkout and deps for the SynthID harness

- setup_markllm.sh sparse-checkout omitted '/visualize/', which
  watermark/base.py imports at module load — every scheme (incl. SynthID)
  failed with 'No module named visualize' during generation/detection.
- requirements-markllm.txt omitted scikit-learn, imported by the SynthID
  detector (watermark/synthid/detector_bayesian_torch.py).

Both broke the MarkLLM harness at runtime; the benchmark's sanity gate then
excluded every sample, producing empty per-variant results.

* fix: drop 4 GiB RLIMIT_AS on benchmark subprocesses

_run_cmd applied the common child RLIMIT_AS (default 4 GiB) via
subprocess_preexec_fn to every MarkLLM/rewrite child. torch needs a much
larger address space: CUDA init failed with 'out of memory' at
cudaGetDeviceCount and the 5.2 GB fp32 opt-1.3b could not load, so every
sample was excluded at generation. text_detectors.py already applies no
address-space cap to MarkLLM by default; the benchmark now matches.

* perf: keep MarkLLM resident via a serve worker (624 cold starts -> 1)

The benchmark spawned a fresh torch + opt-1.3b process per operation
(~60-90s each); a full run needs ~624 of them. detect_text_watermark.py
gains a 'serve' mode (JSON-lines over stdin/stdout, ready handshake) that
loads the model once; bench_synthid_text.py uses it via MarkLLMWorker with
automatic fallback to one-shot subprocesses (--no-worker to force).
Turns ~8h runs into ~40-60min.

* perf: skip per-candidate Gemini detections in rewrite subprocess

* feat: run the SynthID-text benchmark from the wr-markllm compose service

- Dockerfile.markllm: add '/visualize/' to the sparse checkout (same fix as
  setup_markllm.sh) and COPY the benchmark + rewrite scripts (stdlib-only).
- compose.yaml: wr-markllm gets the WATERMARKS_REWRITE_* and
  WATERMARKS_GEMINI_* env wiring, a bench-out volume for --out-dir, and a
  read-only mount of the bundled corpus (build context is service/, so the
  corpus cannot be COPY'd).
- docs: docker compose run example.

Note: the image ships CPU torch by design, so the container path is for
portability/CI; GPU runs use the host setup_markllm.sh venv.

* feat: per-sample progress logging in the benchmark

The persistent worker returns samples in-memory, so nothing is written
until the end of a run — runs looked stuck. eprint a [gen i/N] line per
generated sample and a [removal] summary per sample.

* chore: migrate Gemini config to gemini-3.6-flash; document SynthID-text retirement

Google retired SynthID text watermarking on the Generative Language API
(Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK
is rejected on current 3.x models (confirmed by Google AI staff). Migrate
the default detection model to gemini-3.6-flash, document the retirement
in vendor-notes.md and the benchmark report caveat, and keep the detector
seam fail-soft until a vendor endpoint (e.g. Vertex AI) returns.

* feat: remove gemini-synthid-text detector (Google retired text watermarking)

Google removed SynthID text watermarking from the Generative Language API
(Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK
is rejected on current 3.x models, so the vendor detector had nothing to
detect. Remove GeminiSynthIDTextDetector and its wiring:

- text_detectors.py: drop the Gemini class, HTTP helpers, and constants;
  keep MarkLLM + Claude seams (registry now markllm + claude-text).
- server.py / rewrite_text.py: per-candidate detection now triggers on
  --markllm-scheme only.
- bench_synthid_text.py: remove the Gemini tier (before/after, report
  table, --no-gemini flag); report caveat notes the retirement.
- configs/docs: drop WATERMARKS_GEMINI_* from .env.example / compose /
  README / SKILL.md / vendor-notes.md; keep the retirement note.
- tests: gemini tests removed or converted to MarkLLM (mocked subprocess).
- Dockerfile.markllm: parameterize BASE_IMAGE + TORCH_INDEX_URL so a GPU/
  arm64 image can be built (used for the --gpus all benchmark run).

* fix: harden notes aggregation against non-string notes

A run completed all samples but crashed at the final aggregate step with
'cannot use list as a set element' when a row's notes contained a
non-string value. Filter notes to strings (aggregate + CSV) and add a
regression test.

* perf: let the rewrite subprocess reuse the resident MarkLLM worker

The rewrite subprocess (rewrite_text.py) ran its own before/after MarkLLM
detects, each a ~20s torch+model cold start (~12 per sample = ~5min of the
~6min/sample runtime). Now:

- detect_text_watermark.py serve gains --port N: a loopback TCP JSON-lines
  listener (default -1 = off) sharing the resident model, with a lock so
  stdin and socket requests never run the model concurrently.
- text_detectors.MarkLLMTextDetector checks WATERMARKS_MARKLLM_PORT and
  does a fast loopback detect when a worker is up, falling back to the
  one-shot subprocess otherwise.
- The benchmark worker publishes its port via that env var, so the rewrite
  subprocess inherits it and its detects hit the resident model.

Turns ~6 min/sample into ~1-2 min; a full run drops from ~2h to ~40-50min.
Tests: loopback-client + fallback + env-publish coverage.

* chore: add benchmark-smoke.sh / benchmark-full.sh wrappers

Simple host wrappers: source .env, default MARKLLM_DIR to ~/MarkLLM, use a
repo-local HF cache by default, and run bench_synthid_text.py with a quick
(2 docs, 1 seed, paraphrase:1) or full (8 docs x 3 seeds, three variants,
re-stamp control) configuration. OUT_DIR overrides the output location.
This commit is contained in:
Guillaume Meyer (The Opinionated Man)
2026-08-18 18:13:56 -07:00
committed by GitHub
parent 063119d7e5
commit d5f4f03f85
30 changed files with 2598 additions and 477 deletions
+118
View File
@@ -0,0 +1,118 @@
# SynthID-text removal benchmark
bench_synthid_text.py measures how well the Layer B rewrite
(rewrite_text.py) removes SynthID-text-class watermarks, and at what
cost. It generates a controlled corpus with the MarkLLM SynthID scheme, runs
removal variants, and emits a shareable report.
## What it measures
| Metric | Meaning |
| --- | --- |
| Clear rate | % of watermarked samples that flip to not-watermarked after removal (MarkLLM same-config detection) |
| Score suppression | mean/median drop in detector score (before - after) |
| Quality | lexical divergence (bigram Jaccard distance), length drift, number/URL survival |
| Cost | estimated tokens in/out, wall time per document, optional USD at your prices |
| Efficiency | clears per million output tokens - removal rate per unit of rewrite cost |
| Controls | Layer A only (expect ~0% - Unicode scrub must not clear a statistical mark), sanity-gate exclusions, optional re-stamp check |
## How to run
Prerequisites (all external, matching the repo's optional-harness model):
1. A MarkLLM checkout: run service/scripts/setup_markllm.sh (clones
THU-BPM/MarkLLM at a pinned commit and creates ~/MarkLLM/.venv).
2. A rewrite backend: Ollama (default, loopback) or any
OpenAI-compatible endpoint. The rewrite model must be a real model.
# minimal: 3 docs, 1 seed, paraphrase with 1 and 3 candidates (Ollama)
MARKLLM_DIR=~/MarkLLM \
python3 service/scripts/bench_synthid_text.py \
--markllm-dir ~/MarkLLM \
--rewrite-backend ollama --rewrite-model llama3.2 \
--out-dir out/bench-2026-06-01
# recommended full run: more docs/seeds, backtranslate variant, re-stamp control
python3 service/scripts/bench_synthid_text.py \
--markllm-dir ~/MarkLLM \
--docs 10 --seeds 3 \
--variants "paraphrase:1,paraphrase:3,backtranslate:1" \
--restamp-control \
--rewrite-backend openai-compatible \
--rewrite-model deepseek-v4-flash \
--rewrite-base-url https://api.deepseek.com \
--rewrite-allow-remote \
--out-dir out/bench-deepseek \
--tag deepseek-v4-flash
API keys are read from the environment only (WATERMARKS_REWRITE_API_KEY),
never argv. Non-loopback rewrite endpoints require --rewrite-allow-remote.
No vendor tier: Google retired SynthID text watermarking on its API in
Aug 2026 (DETECT_TEXT_WATERMARK is rejected on current models), so detection
here is MarkLLM same-config only. A vendor tier can be re-added if Google
exposes detection again (e.g. via Vertex AI).
Cost modeling: --cost-per-mtok-in 0.30 --cost-per-mtok-out 1.20 (example
prices) attaches an estimated USD figure per row; token counts are
chars / --chars-per-token estimates (default 4.0).
## Outputs (in --out-dir)
- report.md - self-contained Markdown you can paste anywhere: methodology,
config, results table, controls, caveats, exact reproduction command.
- results.json - full per-sample/per-row data + aggregates.
- results.csv - one row per (doc, seed, variant) for plotting.
- work/ - generated watermarked/unwatermarked samples (kept for inspection).
## Running from Docker (compose)
The `wr-markllm` service in compose.yaml can run the benchmark end-to-end
(image: pinned MarkLLM checkout at /opt/markllm + all scripts). The image
installs CPU torch by design, so use it for portability/CI, not for GPU
throughput on this machine — for GPU runs use the host `setup_markllm.sh`
venv instead (see README).
```bash
docker compose --profile harness build wr-markllm
docker compose run --rm wr-markllm \
/app/bench_synthid_text.py --markllm-dir /opt/markllm \
--corpus /bench-corpus --out-dir /data --tag docker-run \
--docs 10 --seeds 3 --variants "paraphrase:1,paraphrase:3,backtranslate:1" \
--restamp-control
```
Env (rewrite backend) is wired from your .env via compose interpolation;
results land in the `bench-out` volume (/data); the bundled
corpus is mounted read-only at /bench-corpus. The image runs the
persistent MarkLLM serve worker by default, so the ~2-4h one-shot runs
are not a constraint inside the container either.
## What it can and cannot claim
- Can claim: under the MarkLLM SynthID scheme config the benchmark
controls, at these seeds/docs, with this rewrite backend, this clear rate and
cost were observed. Same-config-only detection is deterministic and
reproducible (fixed seeds, pinned MarkLLM commit, recorded commands).
- Cannot claim: that Google's production SynthID-Text detector will fail.
MarkLLM's SynthID is a research reimplementation with a different keying,
and Google retired text watermark detection on its API (Aug 2026), so no
vendor tier exists to verify against. Rewriting with a watermarked model
can also re-stamp the text - run --restamp-control to check.
## Sharing a run
Share the --out-dir directory. report.md embeds the reproduction command,
the MarkLLM commit, the watermarks-remover commit, and the caveats, so a reader
can (a) trust what was measured and (b) rerun it. Keep work/ out of archives
unless you want the raw samples.
## Notes on statistical power
- A single document tells you nothing - the watermark is probabilistic. Use
several documents (--docs 10+) and several seeds per document
(--seeds 3+) so clear-rate differences are distinguishable.
- Longer text carries more watermark signal: default --max-new-tokens 300.
Very short samples are excluded by the sanity gate automatically.
- Compare variants (strength x candidates) within one run, not across runs
with different backends - the rewrite model dominates the outcome.