Files
SnapOtter/tests/fixtures/ocr-real/NOTICE.md
T
SnapOtterandGitHub 991c981529 fix: make OCR portable and reliable across AMD64 and ARM64 (#519)
* fix: make OCR portable and reliable

* fix: harden OCR installation portability

* fix: pin OCR partials across downloads

* fix: make OCR execution reliably asynchronous

* fix: harden OCR portability and docs routes

* fix: preserve decoder and docs safeguards
2026-07-15 03:34:24 +08:00

5.4 KiB
Raw Blame History

OCR fixture notices

These fixtures are redistributed only for local OCR regression testing. manifest.json pins each source revision, exact source Parquet shard or original-file identity, image/annotation hashes, privacy review, and all transformations. The image bytes are unmodified. Annotation records were extracted or manually transcribed into canonical JSON, and the UTF-8 text files contain annotation strings joined with LF as described in the manifest. No upstream author or organization endorses SnapOtter.

JaWildText

jawildtext-board-0001, jawildtext-board-0049, jawildtext-board-0127, and jawildtext-receipt-11120 come from JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding by Koki Maeda and Naoaki Okazaki (LLM-jp), pinned at Hugging Face revision 627ca7ea7c224ffe1accff8737991fc2240784fa.

The three board images form one frozen Fast release cohort. They come from distinct Parquet row groups and cover a dense colored product placard, a mildly skewed monochrome public notice, and a wide museum display with glare. Every board must independently meet the pre-existing Fast recall, precision, and F1 floors; aggregate performance cannot hide a failed image.

Visible business names, logos, and public-agency marks remain the property of their respective owners. Apache-2.0 does not grant trademark rights or imply endorsement.

Wikimedia Commons — Hagye station sign

commons-hagye-station-715 is the original Hagye01.jpg photograph by Wikimedia Commons user Marcopolis, file page ID 7592246, timestamped 2009-08-18T13:00:50Z. The file page is pinned at revision 1234274506; its original bytes are identified by SHA-1 fd1f0cc88f931af22576c7404270837916c60d8d and SHA-256 a9ae819505be17d87393695bdadd1aaff47b0a8b81faec98b38408397942b3dc.

The photograph forms the frozen Korean scene-text release cohort. Before any OCR output was observed, eligible original bitmaps in the Wikimedia Commons category for Seoul Subway Line 7 station signs were sorted by byte size. Public-domain landscape photographs at least 500×400 with no people or private data and manually legible Hangul, Latin, and Arabic digits were retained; Hagye01.jpg was the smallest eligible file.

CORD v2

cord-v2-test-0080 comes from CORD: A Consolidated Receipt Dataset for Post-OCR Parsing by Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee (NAVER Clova), test image ID 80, pinned at Hugging Face revision 7f0115a4b758a71d6473b8d085751692da2fef98.

The identifying header was already blurred in the upstream image; SnapOtter did not perform that alteration.

ClinOCR-Bench

clinocr-poor-t7-s2 comes from ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation by Enshuo Hsu, Jin Zhou, and Kirk Roberts, document poor_t7_s2, pinned at Hugging Face revision cb7c0c48a4f3d1c9054fb6548ccef84768983472.

ClinOCR-Bench describes its data as synthetic and protected-health-information-free. Names, addresses, identifiers, and clinical details visible in this sample are fictional.

Deliberately excluded

No TextOCR image is redistributed. TextOCR labels its dataset CC-BY-4.0, but its images come from Open Images. Open Images states that it makes no warranty about each image's license and requires users to verify the original image license themselves. We did not include a TextOCR sample without that image-level chain-of-title verification.