* fix: rewrite ODT/EPUB manifests and measure real zip bytes (#122)
Two container correctness/security fixes from issue #122:
- clean_odt dropped marker-bearing parts while leaving their entries in
META-INF/manifest.xml, so readers flagged the package as damaged. It is
now two-pass: compute the dropped set, then rewrite the manifest
attribute-order-independently, and write each part exactly once. The same
bug class in clean_epub (dropped parts left in the OPF manifest, plus
dangling spine itemrefs) gets the same two-pass treatment.
- The zip budget trusted ZipInfo.file_size from the archive's own central
directory, so a crafted DOCX/ODT could declare a tiny size and still
expand via zf.read. Budgets are now charged on actual decompressed bytes
via _read_zip_member (streaming, cap enforced mid-read), with the declared
size kept only as a fast-path pre-reject.
* fix: classify unrecognized bytes as "unknown", not text (#122)
Two classification defects from issue #122:
- format_dispatch.classify_bytes fell back to "text" for any unrecognized
file, so a binary with valid UTF-8 runs could be decoded and written back
mangled (corrupted with --in-place) in clean_file auto mode. Unrecognized
bytes now classify as "unknown"; clean_file refuses them in auto mode
(exit 2, no write, router advice) and --as text / --force-text are the
explicit opt-ins. inspect_file reports kind "unknown" (exit 0), audit_lib
records a non-actionable item, and the HTTP server answers /inspect with
kind "unknown" but rejects /clean of unknown formats (400).
- classify(path) read the whole file to sniff a header, and only a full read
could detect zip containers. It now routes known extensions without
reading, sniffs a 4096-byte header once for images and prefix-based
containers, and reads the whole file only when the header is a zip local
header (PK), where the container signature lives in the central directory.
* feat: distinct exit code for partial audits (#122)
audit_dir and audit_website reported success (0) even when some files or
URLs could not be scanned; the exit status was computed only over the items
that succeeded. A scan that is missing items is not a clean scan.
- common.EXIT_PARTIAL = 3, with precedence: partial (3) > actionable (1)
> clean (0) — an incomplete audit is the more important CI signal.
- audit_dir returns 3 when any file was skipped/failed; audit_website
returns 3 when any URL failed to fetch or inspect. Both are independent
of the output format (human/json/sarif already share one return).
* fix: verify the pinned upstream ref on existing checkouts (#122)
setup_ctrlregen.sh/setup_synthid.sh (and their .ps1 twins) only verified
the pinned commit in the fresh-clone branch; an existing checkout at an
unknown or drifted revision was silently reused, defeating the commit pin.
All four scripts now check HEAD against the pinned ref in the
existing-checkout branch too, and repair by fetch + detach checkout
(re-applying the sparse-checkout set), failing hard if the ref cannot be
reached or the re-pin does not land on it.
* docs: unknown-format behavior, audit exit codes, backend isolation (#122)
- README: clean_file no longer auto-cleans unrecognized formats (--as text
/ --force-text are the opt-ins), and the CtrlRegen bootstrap documents the
isolation expectation for its research-era dependency pins plus the new
re-pin check on existing checkouts.
- SKILL.md: audit exit codes (0/1/2/3, partial=3) and a note that /clean
requires a name with a known extension.
- audit_website: document why stdlib ElementTree is used (stdlib-first) and
that defusedxml is the fallback if that policy changes (DTD rejection stays).
- requirements-ctrlregen.txt: advisory/isolation note for the pinned research
dependencies.
* test: ODT manifest and EPUB OPF dangling-ref regressions (#122)
- clean_odt: dropped marker-bearing parts remove their META-INF/manifest.xml
file-entry (attribute-order-independent), exactly one manifest entry, root
and surviving entries kept, and the manifest is byte-identical when nothing
is dropped.
- clean_epub: dropped non-content parts lose their <item> entry in the OPF
manifest, so the book no longer references removed members.