mirror of
https://github.com/guillaumemeyer/watermarks-remover.git
synced 2026-08-22 13:11:57 +02:00
* fix: rewrite ODT/EPUB manifests and measure real zip bytes (#122) Two container correctness/security fixes from issue #122: - clean_odt dropped marker-bearing parts while leaving their entries in META-INF/manifest.xml, so readers flagged the package as damaged. It is now two-pass: compute the dropped set, then rewrite the manifest attribute-order-independently, and write each part exactly once. The same bug class in clean_epub (dropped parts left in the OPF manifest, plus dangling spine itemrefs) gets the same two-pass treatment. - The zip budget trusted ZipInfo.file_size from the archive's own central directory, so a crafted DOCX/ODT could declare a tiny size and still expand via zf.read. Budgets are now charged on actual decompressed bytes via _read_zip_member (streaming, cap enforced mid-read), with the declared size kept only as a fast-path pre-reject. * fix: classify unrecognized bytes as "unknown", not text (#122) Two classification defects from issue #122: - format_dispatch.classify_bytes fell back to "text" for any unrecognized file, so a binary with valid UTF-8 runs could be decoded and written back mangled (corrupted with --in-place) in clean_file auto mode. Unrecognized bytes now classify as "unknown"; clean_file refuses them in auto mode (exit 2, no write, router advice) and --as text / --force-text are the explicit opt-ins. inspect_file reports kind "unknown" (exit 0), audit_lib records a non-actionable item, and the HTTP server answers /inspect with kind "unknown" but rejects /clean of unknown formats (400). - classify(path) read the whole file to sniff a header, and only a full read could detect zip containers. It now routes known extensions without reading, sniffs a 4096-byte header once for images and prefix-based containers, and reads the whole file only when the header is a zip local header (PK), where the container signature lives in the central directory. * feat: distinct exit code for partial audits (#122) audit_dir and audit_website reported success (0) even when some files or URLs could not be scanned; the exit status was computed only over the items that succeeded. A scan that is missing items is not a clean scan. - common.EXIT_PARTIAL = 3, with precedence: partial (3) > actionable (1) > clean (0) — an incomplete audit is the more important CI signal. - audit_dir returns 3 when any file was skipped/failed; audit_website returns 3 when any URL failed to fetch or inspect. Both are independent of the output format (human/json/sarif already share one return). * fix: verify the pinned upstream ref on existing checkouts (#122) setup_ctrlregen.sh/setup_synthid.sh (and their .ps1 twins) only verified the pinned commit in the fresh-clone branch; an existing checkout at an unknown or drifted revision was silently reused, defeating the commit pin. All four scripts now check HEAD against the pinned ref in the existing-checkout branch too, and repair by fetch + detach checkout (re-applying the sparse-checkout set), failing hard if the ref cannot be reached or the re-pin does not land on it. * docs: unknown-format behavior, audit exit codes, backend isolation (#122) - README: clean_file no longer auto-cleans unrecognized formats (--as text / --force-text are the opt-ins), and the CtrlRegen bootstrap documents the isolation expectation for its research-era dependency pins plus the new re-pin check on existing checkouts. - SKILL.md: audit exit codes (0/1/2/3, partial=3) and a note that /clean requires a name with a known extension. - audit_website: document why stdlib ElementTree is used (stdlib-first) and that defusedxml is the fallback if that policy changes (DTD rejection stays). - requirements-ctrlregen.txt: advisory/isolation note for the pinned research dependencies. * test: ODT manifest and EPUB OPF dangling-ref regressions (#122) - clean_odt: dropped marker-bearing parts remove their META-INF/manifest.xml file-entry (attribute-order-independent), exactly one manifest entry, root and surviving entries kept, and the manifest is byte-identical when nothing is dropped. - clean_epub: dropped non-content parts lose their <item> entry in the OPF manifest, so the book no longer references removed members.
69 lines
2.2 KiB
Python
69 lines
2.2 KiB
Python
"""Tests for format routing: unknown kind + header-only classification."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import io
|
|
import sys
|
|
import zipfile
|
|
from pathlib import Path
|
|
|
|
ROOT = Path(__file__).resolve().parents[1]
|
|
SCRIPTS = ROOT / "service" / "scripts"
|
|
sys.path.insert(0, str(SCRIPTS))
|
|
|
|
from format_dispatch import classify, classify_bytes
|
|
|
|
|
|
def _boom(*_args, **_kwargs):
|
|
raise AssertionError("full read_bytes() was not expected here")
|
|
|
|
|
|
def test_classify_extension_wins_without_reading(monkeypatch, tmp_path):
|
|
f = tmp_path / "note.txt"
|
|
f.write_bytes(b"anything, the extension decides")
|
|
monkeypatch.setattr(Path, "read_bytes", _boom)
|
|
assert classify(f) == "text"
|
|
|
|
|
|
def test_classify_header_sniffs_image_without_full_read(monkeypatch, tmp_path):
|
|
f = tmp_path / "no_extension"
|
|
f.write_bytes(b"\x89PNG\r\n\x1a\n" + b"\x00" * 100)
|
|
monkeypatch.setattr(Path, "read_bytes", _boom)
|
|
assert classify(f) == "image"
|
|
|
|
|
|
def test_classify_header_sniffs_pdf_without_full_read(monkeypatch, tmp_path):
|
|
f = tmp_path / "no_extension"
|
|
f.write_bytes(b"%PDF-1.7\n" + b"0" * 100)
|
|
monkeypatch.setattr(Path, "read_bytes", _boom)
|
|
assert classify(f) == "container"
|
|
|
|
|
|
def test_classify_zip_needs_full_read_for_central_directory(tmp_path):
|
|
# Extension-less DOCX: the container signature lives in the central
|
|
# directory at the end of the archive, so the full file must be read.
|
|
buf = io.BytesIO()
|
|
with zipfile.ZipFile(buf, "w") as zf:
|
|
zf.writestr("word/document.xml", "<w:document/>")
|
|
f = tmp_path / "no_extension"
|
|
f.write_bytes(buf.getvalue())
|
|
assert classify(f) == "container"
|
|
|
|
|
|
def test_classify_truncated_zip_header_is_unknown(tmp_path):
|
|
f = tmp_path / "no_extension"
|
|
f.write_bytes(b"PK\x03\x04this is not a real zip")
|
|
assert classify(f) == "unknown"
|
|
|
|
|
|
def test_classify_unknown_for_unrecognized_bytes(tmp_path):
|
|
f = tmp_path / "no_extension"
|
|
f.write_bytes(b"just plain text with no magic and no extension")
|
|
assert classify(f) == "unknown"
|
|
|
|
|
|
def test_classify_bytes_unknown_fallback():
|
|
assert classify_bytes(b"random bytes \x00\xff", None) == "unknown"
|
|
assert classify_bytes(b"\x89PNG\r\n\x1a\nrest", None) == "image"
|
|
assert classify_bytes(b"PK\x03\x04rest", None) == "unknown" # not a full zip
|