mirror of
https://github.com/snapotter-hq/SnapOtter.git
synced 2026-08-03 07:46:42 +02:00
The OCR bundle installs paddleocr[doc-parser] 3.4, whose dependency closure drags numpy 1.26.4 up to 2.5.1 and pulls scipy/scikit-learn/pandas wheels built against the numpy 2.x ABI. build-bundle.sh re-pinned only numpy (basePackages), so those numpy-2.x wheels stayed behind; the by-dir-name site-packages diff then shipped them, and once merged onto the numpy==1.26.4 base they raise "numpy.dtype size changed" on import. Because the dispatcher pre-imports every ML library at startup and disables all AI after 5 crashes in 60s, one stranded scipy takes down every AI tool, not just OCR (observed on a CPU host: remove-background worked before the OCR bundle and broke after). All-7 installs escaped it through last-writer-wins ordering; a subset install did not, which is why it surfaced only intermittently. Fix: add a manifest "constraints" list (numpy, scipy, scikit-learn, scikit-image, pandas pinned to numpy-1.x-ABI versions) and apply it via PIP_CONSTRAINT to every bundle pip install, so no bundle can pull a numpy-2.x wheel. paddleocr 3.4.1 still resolves cleanly under the lock and the pinned stack imports without ABI error on numpy 1.26.4 (validated on py3.12). Also import scipy/sklearn in the OCR path of verify-bundle.sh so CI catches this class in isolation, and add a manifest regression test. Note: the published bundles must be rebuilt and republished (ai-bundles.yml) for this to reach already-installed bases. Claude-Session: https://claude.ai/code/session_01UvVCMNUBrgpghk8gye5gav