* feat: add reproducible SynthID-text removal benchmark bench_synthid_text.py orchestrates the existing Layer B machinery into a controlled, shareable experiment: generate watermarked + unwatermarked samples with the MarkLLM SynthID scheme, run removal variants (strength x candidates) plus controls (no-removal, Layer-A-only, optional re-stamp), and report clear rate, score suppression, quality, and cost (tokens, wall time, USD) with a clears-per-MTok efficiency ratio. Emits report.md / results.json / results.csv with the exact reproduction command and pinned commits; optional Gemini official-detector tier when WATERMARKS_GEMINI_API_KEY is set. Mock-based tests, no torch in CI. * docs: add README section on running the SynthID-text benchmark Explains what LLM performs the Layer B rewrite (an external model configured via WATERMARKS_REWRITE_* env vars or --rewrite-* flags; MarkLLM's opt-1.3b is only the watermark generator/detector) and how to run a benchmark with Ollama or an OpenAI-compatible endpoint, plus the non-origin-model re-stamp caveat. * fix: honor WATERMARKS_REWRITE_ALLOW_REMOTE in the SynthID-text benchmark The --rewrite-allow-remote flag now defaults from the env var (matching rewrite_text.py and the other WATERMARKS_REWRITE_* settings), so a non-loopback rewrite endpoint works after sourcing .env without an extra flag. * fix: MarkLLM sparse checkout and deps for the SynthID harness - setup_markllm.sh sparse-checkout omitted '/visualize/', which watermark/base.py imports at module load — every scheme (incl. SynthID) failed with 'No module named visualize' during generation/detection. - requirements-markllm.txt omitted scikit-learn, imported by the SynthID detector (watermark/synthid/detector_bayesian_torch.py). Both broke the MarkLLM harness at runtime; the benchmark's sanity gate then excluded every sample, producing empty per-variant results. * fix: drop 4 GiB RLIMIT_AS on benchmark subprocesses _run_cmd applied the common child RLIMIT_AS (default 4 GiB) via subprocess_preexec_fn to every MarkLLM/rewrite child. torch needs a much larger address space: CUDA init failed with 'out of memory' at cudaGetDeviceCount and the 5.2 GB fp32 opt-1.3b could not load, so every sample was excluded at generation. text_detectors.py already applies no address-space cap to MarkLLM by default; the benchmark now matches. * perf: keep MarkLLM resident via a serve worker (624 cold starts -> 1) The benchmark spawned a fresh torch + opt-1.3b process per operation (~60-90s each); a full run needs ~624 of them. detect_text_watermark.py gains a 'serve' mode (JSON-lines over stdin/stdout, ready handshake) that loads the model once; bench_synthid_text.py uses it via MarkLLMWorker with automatic fallback to one-shot subprocesses (--no-worker to force). Turns ~8h runs into ~40-60min. * perf: skip per-candidate Gemini detections in rewrite subprocess * feat: run the SynthID-text benchmark from the wr-markllm compose service - Dockerfile.markllm: add '/visualize/' to the sparse checkout (same fix as setup_markllm.sh) and COPY the benchmark + rewrite scripts (stdlib-only). - compose.yaml: wr-markllm gets the WATERMARKS_REWRITE_* and WATERMARKS_GEMINI_* env wiring, a bench-out volume for --out-dir, and a read-only mount of the bundled corpus (build context is service/, so the corpus cannot be COPY'd). - docs: docker compose run example. Note: the image ships CPU torch by design, so the container path is for portability/CI; GPU runs use the host setup_markllm.sh venv. * feat: per-sample progress logging in the benchmark The persistent worker returns samples in-memory, so nothing is written until the end of a run — runs looked stuck. eprint a [gen i/N] line per generated sample and a [removal] summary per sample. * chore: migrate Gemini config to gemini-3.6-flash; document SynthID-text retirement Google retired SynthID text watermarking on the Generative Language API (Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK is rejected on current 3.x models (confirmed by Google AI staff). Migrate the default detection model to gemini-3.6-flash, document the retirement in vendor-notes.md and the benchmark report caveat, and keep the detector seam fail-soft until a vendor endpoint (e.g. Vertex AI) returns. * feat: remove gemini-synthid-text detector (Google retired text watermarking) Google removed SynthID text watermarking from the Generative Language API (Aug 2026): text output is no longer watermarked and DETECT_TEXT_WATERMARK is rejected on current 3.x models, so the vendor detector had nothing to detect. Remove GeminiSynthIDTextDetector and its wiring: - text_detectors.py: drop the Gemini class, HTTP helpers, and constants; keep MarkLLM + Claude seams (registry now markllm + claude-text). - server.py / rewrite_text.py: per-candidate detection now triggers on --markllm-scheme only. - bench_synthid_text.py: remove the Gemini tier (before/after, report table, --no-gemini flag); report caveat notes the retirement. - configs/docs: drop WATERMARKS_GEMINI_* from .env.example / compose / README / SKILL.md / vendor-notes.md; keep the retirement note. - tests: gemini tests removed or converted to MarkLLM (mocked subprocess). - Dockerfile.markllm: parameterize BASE_IMAGE + TORCH_INDEX_URL so a GPU/ arm64 image can be built (used for the --gpus all benchmark run). * feat: iterative detection-guided Layer B rewriting (default 3 attempts) Layer B (rewrite_text.py) now rewrites iteratively and stops as soon as an attempt passes watermark evaluation: - --candidates defaults to 3 (WATERMARKS_REWRITE_CANDIDATES); each attempt is one rewrite + one evaluation, and the loop exits on the first attempt the evaluator reports as not watermarked. - Evaluator priority: MarkLLM same-config detection (when --markllm-scheme is passed) > bigram-Jaccard lexical divergence (fallback; no verdict, all attempts generated, most diverged selected). A vendor-detector seam is reserved ahead of MarkLLM for a future SynthID-text endpoint (Google retired text watermarking on its API in Aug 2026). - Best-effort fallback when the max is exhausted: the lowest-score attempt is returned with a note; detector errors are fail-soft and never fail the rewrite. - --json-stats now reports evaluator / attempts_made / passed and per-attempt candidate_scores records (passed, evaluation); markllm before/after/cleared is unchanged and the selected attempt's verdict is reused (no duplicate MarkLLM detection). Benchmark (bench_synthid_text.py): - --variants default becomes paraphrase:3 (candidates = max attempts). - Rows/report/CSV carry attempts per document (mean_attempts, att column; attempts / evaluator / passed columns). Tests, README, docs/synthid-text-benchmark.md and .env.example updated; 490 tests pass, ruff clean. * feat: split rewrite attempts into --candidates x --max-loops (defaults 1 x 1) Follow-up to the iterative Layer B rewrite: separate "variants per round" from "evaluation rounds", so the retry loop is explicit and defaults stay conservative. - rewrite_text.py: --candidates (WATERMARKS_REWRITE_CANDIDATES) is now the number of variants generated per loop iteration (default 1); new --max-loops (WATERMARKS_REWRITE_LOOPS) caps the evaluation rounds (default 1) -- each round generates --candidates variants and stops as soon as one passes, so raising --max-loops retries new variants until an evaluation passes. Stats now report max_loops and per-attempt records carry the loop index. - bench_synthid_text.py: new --rewrite-loops flag (default 1) passed through to --max-loops. - README / docs / .env.example updated; tests cover the 1x1 defaults, loop retry until pass, and cross-loop exhaustion. * feat: MarkLLM serve worker over loopback TCP (WATERMARKS_MARKLLM_PORT) detect_text_watermark.py serve can now also listen on a loopback TCP port, and MarkLLMTextDetector reuses a resident worker when WATERMARKS_MARKLLM_PORT is set (falls back to a one-shot subprocess when the worker is unreachable). This avoids a ~20s torch+model cold start per detect for callers that run a worker out-of-band. Tests: loopback worker protocol + detector worker-port routing (mock-based). * ci: add macOS runner to the test matrix
6.6 KiB
SynthID-text removal benchmark
bench_synthid_text.py measures how well the Layer B rewrite (rewrite_text.py) removes SynthID-text-class watermarks, and at what cost. It generates a controlled corpus with the MarkLLM SynthID scheme, runs removal variants, and emits a shareable report.
What it measures
| Metric | Meaning |
|---|---|
| Clear rate | % of watermarked samples that flip to not-watermarked after removal (MarkLLM same-config detection) |
| Score suppression | mean/median drop in detector score (before - after) |
| Quality | lexical divergence (bigram Jaccard distance), length drift, number/URL survival |
| Cost | estimated tokens in/out, wall time per document, optional USD at your prices |
| Efficiency | clears per million output tokens - removal rate per unit of rewrite cost |
| Attempts | mean rewrite attempts per document (the Layer B loop stops early on pass) |
| Controls | Layer A only (expect ~0% - Unicode scrub must not clear a statistical mark), sanity-gate exclusions, optional re-stamp check |
How to run
Prerequisites (all external, matching the repo's optional-harness model):
-
A MarkLLM checkout: run service/scripts/setup_markllm.sh (clones THU-BPM/MarkLLM at a pinned commit and creates ~/MarkLLM/.venv).
-
A rewrite backend: Ollama (default, loopback) or any OpenAI-compatible endpoint. The rewrite model must be a real model.
minimal: 3 docs, 1 seed, paraphrase with up to 3 attempts (default, Ollama)
MARKLLM_DIR=~/MarkLLM
python3 service/scripts/bench_synthid_text.py
--markllm-dir ~/MarkLLM
--rewrite-backend ollama --rewrite-model llama3.2
--out-dir out/bench-2026-06-01recommended full run: more docs/seeds, backtranslate variant, re-stamp control
python3 service/scripts/bench_synthid_text.py
--markllm-dir ~/MarkLLM
--docs 10 --seeds 3
--variants "paraphrase:3,backtranslate:3"
--restamp-control
--rewrite-backend openai-compatible
--rewrite-model deepseek-v4-flash
--rewrite-base-url https://api.deepseek.com
--rewrite-allow-remote
--out-dir out/bench-deepseek
--tag deepseek-v4-flash
API keys are read from the environment only (WATERMARKS_REWRITE_API_KEY), never argv. Non-loopback rewrite endpoints require --rewrite-allow-remote.
No vendor tier: Google retired SynthID text watermarking on its API in Aug 2026 (DETECT_TEXT_WATERMARK is rejected on current models), so detection here is MarkLLM same-config only. A vendor tier can be re-added if Google exposes detection again (e.g. via Vertex AI).
How variants map to rewrites: each : variant runs
the Layer B rewrite with candidates as the variants per evaluation round;
--rewrite-loops (default 1, mirrors --max-loops /
WATERMARKS_REWRITE_LOOPS) sets how many rounds run before the best-effort
variant is returned. The rewrite is iterative: it generates a variant, runs
MarkLLM detection (same-config) on it, and stops as soon as an attempt is not
watermarked — so a variant usually costs fewer rewrites than its candidate
count, and paraphrase:3 means "try up to 3 variants, stop on the first pass"
(raise --rewrite-loops to keep retrying new variants until one passes).
The report's att column (and mean_attempts in results.json / attempts in
results.csv) records the actual attempts per document.
Cost warning: with MarkLLM as the evaluator, each attempt also costs one MarkLLM detection — up to (candidates x loops) detections per input. The persistent serve worker (default) keeps the model loaded so detections are cheap; the --no-worker one-shot path re-loads the model per detection.
Cost modeling: --cost-per-mtok-in 0.30 --cost-per-mtok-out 1.20 (example prices) attaches an estimated USD figure per row; token counts are chars / --chars-per-token estimates (default 4.0).
Outputs (in --out-dir)
- report.md - self-contained Markdown you can paste anywhere: methodology, config, results table, controls, caveats, exact reproduction command.
- results.json - full per-sample/per-row data + aggregates.
- results.csv - one row per (doc, seed, variant) for plotting.
- work/ - generated watermarked/unwatermarked samples (kept for inspection).
Running from Docker (compose)
The wr-markllm service in compose.yaml can run the benchmark end-to-end
(image: pinned MarkLLM checkout at /opt/markllm + all scripts). The image
installs CPU torch by design, so use it for portability/CI, not for GPU
throughput on this machine — for GPU runs use the host setup_markllm.sh
venv instead (see README).
docker compose --profile harness build wr-markllm
docker compose run --rm wr-markllm \
/app/bench_synthid_text.py --markllm-dir /opt/markllm \
--corpus /bench-corpus --out-dir /data --tag docker-run \
--docs 10 --seeds 3 --variants "paraphrase:3,backtranslate:3" \
--restamp-control
Env (rewrite backend) is wired from your .env via compose interpolation;
results land in the bench-out volume (/data); the bundled
corpus is mounted read-only at /bench-corpus. The image runs the
persistent MarkLLM serve worker by default, so the ~2-4h one-shot runs
are not a constraint inside the container either.
What it can and cannot claim
- Can claim: under the MarkLLM SynthID scheme config the benchmark controls, at these seeds/docs, with this rewrite backend, this clear rate and cost were observed. Same-config-only detection is deterministic and reproducible (fixed seeds, pinned MarkLLM commit, recorded commands).
- Cannot claim: that Google's production SynthID-Text detector will fail. MarkLLM's SynthID is a research reimplementation with a different keying, and Google retired text watermark detection on its API (Aug 2026), so no vendor tier exists to verify against. Rewriting with a watermarked model can also re-stamp the text - run --restamp-control to check.
Sharing a run
Share the --out-dir directory. report.md embeds the reproduction command, the MarkLLM commit, the watermarks-remover commit, and the caveats, so a reader can (a) trust what was measured and (b) rerun it. Keep work/ out of archives unless you want the raw samples.
Notes on statistical power
- A single document tells you nothing - the watermark is probabilistic. Use several documents (--docs 10+) and several seeds per document (--seeds 3+) so clear-rate differences are distinguishable.
- Longer text carries more watermark signal: default --max-new-tokens 300. Very short samples are excluded by the sanity gate automatically.
- Compare variants (strength x candidates) within one run, not across runs with different backends - the rewrite model dominates the outcome.