Files
watermarks-remover/skills/remove-ai-marks/references/mark-classes.md
T
guillaume 759fd33acf Release v0.0.1: multi-vendor AI marks skill and cleaners
Rename remove-claude-marks to remove-ai-marks, add container metadata
support (SVG/PDF/DOCX/ODT/HTML/MD), Layer B rewrite hook, unified file
CLI, and multi-vendor documentation for the first public release.
2026-08-11 13:30:25 -07:00

1.8 KiB
Raw Blame History

Mark classes

1. Edit-based text (Unicode / rules)

Invisible or near-invisible characters, exotic spaces, bidi controls, tag characters, synonym tables.

Inspect kinds (Layer A) Examples
zwj_family ZWSP, ZWNJ, ZWJ, WJ, BOM
bidi LRE/RLO/LRI/…
tag_chars U+E0001U+E007F
variation_selector VS1VS256
space NBSP, em space, ideographic space
confusable Cyrillic/fullwidth Latin (aggressive)

Removal: clean_text.py / Layer A — deterministic, verifiable.

Maps to Nature paper “edit-based watermarking.”

2. Generative / statistical text (token sampling)

Bias next-token sampling toward a pseudo-random green list / score (Kirchenbauer, SynthID-Text / Tournament sampling, etc.). Signal lives in word choice, not metadata.

Removal: Layer B rewrite (paraphrase → back-translate → structural). Best-effort; no gold cert without vendor detector/key.

Maps to Nature paper primary method (SynthID-Text).

3. Data-driven / backdoor

Model trained or fine-tuned so trigger prompts produce marked or identifiable behavior.

Out of scope for this skill (model-side).

4. File provenance metadata (C2PA / EXIF / XMP / props)

Signed Content Credentials and AI generator tags in containers.

Format Support
PNG / JPEG Full strip (stdlib + optional exiftool)
SVG Drop metadata/XMP blocks
PDF Prefer exiftool; degraded stdlib XMP strip
DOCX / ODT Scrub zip XML props / customXml
HTML Meta generator / JSON-LD / data-ai*
Markdown YAML frontmatter AI keys

Removal: clean_file.py / clean_image.py — usually verifiable by re-inspect.

5. Pixel-domain image watermarks

Invisible image marks (e.g. SynthID for images). Out of scope.