mirror of
https://github.com/snapotter-hq/SnapOtter.git
synced 2026-08-03 07:46:42 +02:00
fix: repair docker validation QA tooling, dispatcher crash-accounting, and image-enhancement RAW hang (#391)
Found and fixed during a full local Docker build validation (amd64/arm64, all four fleet targets, AI bundle installs, QA harness) and the follow-up bug sweep requested afterward. None of the affected scripts run in CI, so these had been silently broken indefinitely. - docker/feature-manifest.json: pythonVersion was a flat "3.11", but the amd64 base (Ubuntu 24.04) ships Python 3.12 while arm64 (Debian bookworm) ships 3.11. Changed to a per-arch object matching the file's existing convention. - tests/qa/api-sweep.mts and verify-ai.mts: bare "@snapotter/shared" import can't resolve since tests/ is not a pnpm workspace member, making both silently unrunnable via their own documented command on any fresh checkout. Switched to a relative import. - tests/qa/generate-ledger.mts: wrote to docs/qa/ without creating the directory first; docs/ is gitignored except COMMUNITY_GUIDE.md, so a fresh checkout threw ENOENT. - Seven QA Playwright spec files (input-preview, settings, settings-extended, multifile, output-preview, pipeline-ui, smoke) had ~115 fixture() calls using directory names that don't exist. Resolved every call programmatically against the real fixture tree. - packages/ai/src/bridge.ts: AI dispatcher restart (happens on every bundle install) was falsely counted as a crash, risking permanent dispatcher disable after enough legitimate restarts within the crash window. Added a shuttingDown flag checked at all three recordCrash() call sites. - packages/image-engine/src/operations/auto-enhance.ts: image-enhancement hung 40+ seconds on large RAW photos (confirmed on a real 20.2MP file) in Sharp's .clahe() step, whose cost scales with total pixel count regardless of tile size. Added a 16-megapixel cap above which CLAHE is skipped; verified against the real file (40+s -> 2.0s) with no regression to other RAW formats or normal-sized images. Fixing this surfaced a second, smaller bug where the saturation step's CLAHE compensation boost was keyed off the raw toggle instead of whether CLAHE actually ran. - Two QA-harness robustness gaps closed per "fix everything, even the small bugs": the passport-photo/erase-object input-preview tests now skip cleanly with a clear reason on a container without their AI bundle installed, and docker-compose.qa.yml's hardcoded project/container name (the actual root cause of a mid-validation container swap between two concurrent sessions) is now parameterized via QA_PROJECT_NAME. Full validation report is local-only per repo convention.
This commit is contained in:
@@ -137,6 +137,7 @@ export class PythonDispatcher {
|
||||
private crashes = 0;
|
||||
private lastCrashTs = 0;
|
||||
private backoffEnd = 0;
|
||||
private shuttingDown = false;
|
||||
|
||||
constructor(opts: { profile: "ai" | "docs" }) {
|
||||
this.profile = opts.profile;
|
||||
@@ -174,6 +175,7 @@ export class PythonDispatcher {
|
||||
|
||||
private startChild(): ChildProcess | null {
|
||||
if (this.childFailed) return null;
|
||||
this.shuttingDown = false;
|
||||
|
||||
try {
|
||||
const proc = spawn(getPythonPath(), [resolve(PYTHON_DIR, "dispatcher.py")], {
|
||||
@@ -190,7 +192,13 @@ export class PythonDispatcher {
|
||||
req.reject(new Error("Python dispatcher stdin closed unexpectedly"));
|
||||
this.pending.delete(id);
|
||||
}
|
||||
this.recordCrash();
|
||||
// An intentional shutdown() ends stdin then SIGTERMs the child, which
|
||||
// can surface here as an EPIPE/ERR_STREAM_DESTROYED. That is not a
|
||||
// crash -- counting it would let repeated legitimate restarts (e.g.
|
||||
// shutdownDispatcher() on every AI bundle install) trip the crash
|
||||
// limit and permanently disable the dispatcher. Guard mirrors the
|
||||
// "close" handler below.
|
||||
if (!this.shuttingDown) this.recordCrash();
|
||||
this.child = null;
|
||||
this.childReady = false;
|
||||
}
|
||||
@@ -284,7 +292,9 @@ export class PythonDispatcher {
|
||||
console.error(`[bridge] Dispatcher error: ${err.message} (code: ${err.code})`);
|
||||
if (err.code === "ENOENT") {
|
||||
this.childFailed = true;
|
||||
} else {
|
||||
} else if (!this.shuttingDown) {
|
||||
// Skip crash accounting when we initiated the teardown (shutdown()
|
||||
// sets shuttingDown before killing the child); mirrors "close".
|
||||
this.recordCrash();
|
||||
}
|
||||
for (const [id, req] of this.pending.entries()) {
|
||||
@@ -300,7 +310,7 @@ export class PythonDispatcher {
|
||||
req.reject(new Error("Python dispatcher exited unexpectedly"));
|
||||
this.pending.delete(id);
|
||||
}
|
||||
if (code !== 0) {
|
||||
if (code !== 0 && !this.shuttingDown) {
|
||||
this.recordCrash();
|
||||
}
|
||||
this.child = null;
|
||||
@@ -531,6 +541,7 @@ export class PythonDispatcher {
|
||||
*/
|
||||
shutdown(): void {
|
||||
if (this.child && !this.child.killed) {
|
||||
this.shuttingDown = true;
|
||||
this.child.stdin?.end();
|
||||
this.child.kill("SIGTERM");
|
||||
this.child = null;
|
||||
|
||||
Reference in New Issue
Block a user