mirror of
https://github.com/snapotter-hq/SnapOtter.git
synced 2026-08-03 07:46:42 +02:00
feat(db)!: SnapOtter 2.0 phase 1 foundation: postgres, migrator, compose stack (#216)
* feat(infra): add dev compose stack with postgres and redis
* fix(infra): comment dev env defaults until wired; harden dev compose restart and start_period
* chore(deps): add pg driver and testcontainers for postgres migration
* feat(db): translate schema to drizzle pg-core (timestamptz, boolean, pgEnum, jsonb)
Schema translation (apps/api/src/db/schema.ts):
- sqlite-core -> pg-core, all 10 tables preserved 1:1
- integer(mode:'timestamp') -> timestamp({ withTimezone: true })
- integer(mode:'boolean') -> boolean
- jobs.status text enum -> pgEnum('job_status') with same 4 values
- 7 columns changed from text to jsonb: jobs.inputFiles, jobs.settings,
pipelines.steps, apiKeys.permissions, roles.permissions,
auditLog.details, userFiles.toolChain
- settings.value stays text, jobs.error stays text, jobs.progress stays real
jsonb call-site sweep (removed JSON.stringify on writes, JSON.parse on reads):
- apps/api/src/routes/roles.ts: permissions read/write (3 sites)
- apps/api/src/routes/api-keys.ts: permissions write + read (2 sites)
- apps/api/src/routes/audit-log.ts: details read (1 site)
- apps/api/src/routes/pipeline.ts: steps write + read (2 sites)
- apps/api/src/routes/progress.ts: inputFiles write (2 sites)
- apps/api/src/routes/tool-factory.ts: toolChain read + write (2 sites)
- apps/api/src/routes/user-files.ts: toolChain read + write (4 sites)
- apps/api/src/permissions.ts: roles.permissions read (1 site)
- apps/api/src/lib/audit.ts: details write (1 site)
- apps/api/src/plugins/auth.ts: apiKeys.permissions read (1 site)
* refactor(db): type jsonb columns via $type and note raw CTE conversion requirements
* feat(db): archive sqlite migrations and generate postgres baseline
* chore(db): dockerignore legacy migrations, add archive breadcrumb, fix trailing newline
* feat(db): pg pool connection, advisory-locked boot migrations, DATABASE_URL config
* fix(db): friendly fatal on unreachable postgres, idempotent closeDb, lock-key convention note
* refactor(db): async drizzle calls in plugins, lib, permissions
* fix(api): analytics never throws, typed permission guard, single-query session invalidation
* refactor(db): async drizzle calls across all routes and bootstrap
Convert every route file and index.ts from sync SQLite drizzle
patterns to async node-postgres drizzle:
- .all() removed (bare await on select)
- .get() converted to destructured [row] = await ...
- .run() removed (bare await on insert/update/delete)
- .changes replaced with .rowCount (null-guarded) in progress.ts
- sqlite import removed from user-files.ts; raw CTEs converted to
await db.execute(sql`...`) with postgres-dialect recursive CTEs
- ChainRow types updated: tool_chain is parsed jsonb (string[] | null),
created_at is Date (timestamptz) with no * 1000 conversion
- All requirePermission() guard calls awaited (security: unawaited
async guard returns truthy Promise, bypassing permission check)
- All hasEffectivePermission() and getPermissions() calls awaited
- All auditLog() calls awaited (preserves write-before-response order)
- trackEvent() and captureException() left un-awaited (fire-and-forget
by design, guaranteed never-throw)
- ensureAnonymousUser(), startCleanupCron(), recoverStaleJobs() awaited
in bootstrap sequence
- ensureInstanceId() and ensureDefaultSettings() made async
Files converted: 14 (index.ts + 12 route files + tools/index.ts)
* fix(db): await async checkStorageQuota in user-files upload/save routes
* fix(db): await checkStorageQuota in save-result route (missed second call site)
* feat(db): sqlite-to-postgres migrator with CLI and first-boot import
* fix(db): migrator error context, honest force semantics, boot-hook fatal, null-variance tests
* test: run suite against per-file postgres databases via testcontainers
- Add tests/global-setup.ts: spins up a Postgres testcontainer,
creates a migrated template database once per vitest run.
- Rewrite tests/setup/per-fork-env.ts: each test file (forks pool)
clones the template into its own database via CREATE DATABASE ...
TEMPLATE, preserving the same per-file isolation granularity.
- Update vitest.config.ts: add globalSetup, pg alias, update comment.
- Fix tests/integration/test-server.ts: remove DB_PATH mkdir, async
runMigrations, async db operations, remove SQLite WAL checkpoint.
- Fix 21 unit test db/index mocks: add pool and closeDb exports.
- Fix 8 unit test files: add async/await for now-async permission,
audit, and analytics functions.
- Fix 18 integration test files: convert sync .run()/.all()/.get()
to async drizzle patterns, add async to callbacks.
- Production change: apps/api/src/routes/teams.ts: cast COUNT(*)
to ::int so Postgres returns a number instead of bigint string.
* fix(db): seed built-in roles, reject NUL bytes, cast COUNT, serialize job persists
- Seed built-in roles (admin, editor, user) at boot via ensureBuiltinRoles()
with onConflictDoNothing, restoring data that legacy SQLite migration 0007
provided via INSERT statements (the pg baseline is DDL-only).
- Reject NUL bytes in login credentials with 401 (postgres rejects \x00 in
text columns; valid usernames never contain NUL, matching 1.x behavior).
- Cast COUNT(*)::int in user-files, audit-log, and roles listing queries so
postgres returns a JS number instead of bigint-as-string.
- Serialize fire-and-forget job progress DB writes per jobId so the final
"completed" status is never overwritten by a late-arriving "processing"
write (race condition exposed by async postgres round-trips).
* test: fix teams race, seed roles in test server, poll for job status
- Add missing await to resetTeams() in teams PUT beforeEach (the async
delete raced with the subsequent insert under postgres).
- Call ensureBuiltinRoles() in test server bootstrap so integration tests
have the same built-in roles as production.
- Replace fixed 100ms flushPersist delay with a polling helper that waits
for terminal job status, eliminating timing-dependent failures caused by
postgres network round-trip latency.
* test: make heic temp-file cleanup assertion resilient to concurrent workers
Use a set-based diff instead of raw file count when checking that
decodeHeic cleans up temp files. Other concurrent test workers can
create heic-in-*/heic-out-* files in the shared tmpdir, inflating the
"after" count and causing spurious failures under full-suite load.
* fix(db): align builtin-role seed to post-0010 legacy state; test polish
* feat(docker): three-container compose (app, postgres, redis) with boot wait and migrations
* fix(docker): set TEST_DATABASE_URL so containerized tests skip testcontainers
* chore(docker): test compose project name, clearer 1.x upgrade comment, unref probe timer
* feat(enterprise): enforce D15 license boundary; move s3 storage into packages/enterprise
* fix(enterprise): restore lazy aws-sdk loading; community installs load no s3 code at boot
* fix(enterprise): boundary check catches dynamic imports; document getS3 concurrency
* feat(db)!: SnapOtter 2.0 phase 1 foundation: postgres, migrator, compose stack
BREAKING CHANGE: SQLite is no longer the runtime database. Deployments now
require Postgres (and Redis, used from phase 2). Existing installs migrate
with SQLITE_MIGRATE_PATH or 'pnpm --filter @snapotter/api migrate:sqlite'.
* fix(ci): postgres service + fresh e2e database per run; ignore unfixable torch CVE-2025-3000
This commit is contained in:
@@ -47,18 +47,43 @@ const listeners = new Map<string, Set<(data: JobProgress | SingleFileProgress) =
|
||||
|
||||
// ── DB persistence helpers ──────────────────────────────────────────
|
||||
|
||||
function persistJobProgress(progress: JobProgress): void {
|
||||
/**
|
||||
* Per-job serialization queues. Fire-and-forget persist calls for the same
|
||||
* jobId must run sequentially so that the final "completed" write is never
|
||||
* overwritten by a late-arriving "processing" write. Without this, the
|
||||
* async Postgres round-trips can re-order concurrent writes.
|
||||
*/
|
||||
// TODO(phase-2): delete when progress persistence moves to BullMQ job events.
|
||||
const persistQueues = new Map<string, Promise<void>>();
|
||||
|
||||
/** Await any pending persist writes for a specific job (used by tests). */
|
||||
export async function drainPersistQueue(jobId: string): Promise<void> {
|
||||
const pending = persistQueues.get(jobId);
|
||||
if (pending) await pending;
|
||||
}
|
||||
|
||||
function enqueuePersist(jobId: string, fn: () => Promise<void>): void {
|
||||
const prev = persistQueues.get(jobId) ?? Promise.resolve();
|
||||
const next = prev.then(fn, fn); // run even if prior rejected
|
||||
persistQueues.set(jobId, next);
|
||||
// Clean up the map entry once the queue drains
|
||||
next.then(() => {
|
||||
if (persistQueues.get(jobId) === next) persistQueues.delete(jobId);
|
||||
});
|
||||
}
|
||||
|
||||
async function persistJobProgress(progress: JobProgress): Promise<void> {
|
||||
try {
|
||||
const completionRatio =
|
||||
progress.totalFiles > 0 ? progress.completedFiles / progress.totalFiles : 0;
|
||||
const existing = db
|
||||
const [existing] = await db
|
||||
.select({ id: schema.jobs.id })
|
||||
.from(schema.jobs)
|
||||
.where(eq(schema.jobs.id, progress.jobId))
|
||||
.get();
|
||||
.where(eq(schema.jobs.id, progress.jobId));
|
||||
|
||||
if (existing) {
|
||||
db.update(schema.jobs)
|
||||
await db
|
||||
.update(schema.jobs)
|
||||
.set({
|
||||
status: progress.status,
|
||||
progress: completionRatio,
|
||||
@@ -66,26 +91,25 @@ function persistJobProgress(progress: JobProgress): void {
|
||||
completedAt:
|
||||
progress.status === "completed" || progress.status === "failed" ? new Date() : null,
|
||||
})
|
||||
.where(eq(schema.jobs.id, progress.jobId))
|
||||
.run();
|
||||
.where(eq(schema.jobs.id, progress.jobId));
|
||||
} else {
|
||||
db.insert(schema.jobs)
|
||||
.values({
|
||||
id: progress.jobId,
|
||||
type: "batch",
|
||||
status: progress.status,
|
||||
progress: completionRatio,
|
||||
inputFiles: JSON.stringify({ totalFiles: progress.totalFiles }),
|
||||
error: progress.errors.length > 0 ? JSON.stringify(progress.errors) : null,
|
||||
})
|
||||
.run();
|
||||
await db.insert(schema.jobs).values({
|
||||
id: progress.jobId,
|
||||
type: "batch",
|
||||
status: progress.status,
|
||||
progress: completionRatio,
|
||||
inputFiles: { totalFiles: progress.totalFiles },
|
||||
error: progress.errors.length > 0 ? JSON.stringify(progress.errors) : null,
|
||||
});
|
||||
}
|
||||
} catch {
|
||||
// DB persistence is best-effort; don't break real-time SSE
|
||||
}
|
||||
}
|
||||
|
||||
function persistSingleFileProgress(progress: Omit<SingleFileProgress, "type">): void {
|
||||
async function persistSingleFileProgress(
|
||||
progress: Omit<SingleFileProgress, "type">,
|
||||
): Promise<void> {
|
||||
try {
|
||||
const status =
|
||||
progress.phase === "complete"
|
||||
@@ -93,33 +117,30 @@ function persistSingleFileProgress(progress: Omit<SingleFileProgress, "type">):
|
||||
: progress.phase === "failed"
|
||||
? "failed"
|
||||
: "processing";
|
||||
const existing = db
|
||||
const [existing] = await db
|
||||
.select({ id: schema.jobs.id })
|
||||
.from(schema.jobs)
|
||||
.where(eq(schema.jobs.id, progress.jobId))
|
||||
.get();
|
||||
.where(eq(schema.jobs.id, progress.jobId));
|
||||
|
||||
if (existing) {
|
||||
db.update(schema.jobs)
|
||||
await db
|
||||
.update(schema.jobs)
|
||||
.set({
|
||||
status,
|
||||
progress: progress.percent / 100,
|
||||
error: progress.error ?? null,
|
||||
completedAt: status === "completed" || status === "failed" ? new Date() : null,
|
||||
})
|
||||
.where(eq(schema.jobs.id, progress.jobId))
|
||||
.run();
|
||||
.where(eq(schema.jobs.id, progress.jobId));
|
||||
} else {
|
||||
db.insert(schema.jobs)
|
||||
.values({
|
||||
id: progress.jobId,
|
||||
type: "single",
|
||||
status,
|
||||
progress: progress.percent / 100,
|
||||
inputFiles: "[]",
|
||||
error: progress.error ?? null,
|
||||
})
|
||||
.run();
|
||||
await db.insert(schema.jobs).values({
|
||||
id: progress.jobId,
|
||||
type: "single",
|
||||
status,
|
||||
progress: progress.percent / 100,
|
||||
inputFiles: [],
|
||||
error: progress.error ?? null,
|
||||
});
|
||||
}
|
||||
} catch {
|
||||
// Best-effort
|
||||
@@ -130,27 +151,25 @@ function persistSingleFileProgress(progress: Omit<SingleFileProgress, "type">):
|
||||
* Mark any jobs left in "processing" or "queued" state as failed.
|
||||
* Called once at startup to recover from unclean shutdown.
|
||||
*/
|
||||
export function recoverStaleJobs(): void {
|
||||
export async function recoverStaleJobs(): Promise<void> {
|
||||
try {
|
||||
const result = db
|
||||
const result = await db
|
||||
.update(schema.jobs)
|
||||
.set({
|
||||
status: "failed",
|
||||
error: "Server restarted while job was in progress",
|
||||
completedAt: new Date(),
|
||||
})
|
||||
.where(eq(schema.jobs.status, "processing"))
|
||||
.run();
|
||||
const result2 = db
|
||||
.where(eq(schema.jobs.status, "processing"));
|
||||
const result2 = await db
|
||||
.update(schema.jobs)
|
||||
.set({
|
||||
status: "failed",
|
||||
error: "Server restarted while job was queued",
|
||||
completedAt: new Date(),
|
||||
})
|
||||
.where(eq(schema.jobs.status, "queued"))
|
||||
.run();
|
||||
const total = result.changes + result2.changes;
|
||||
.where(eq(schema.jobs.status, "queued"));
|
||||
const total = (result.rowCount ?? 0) + (result2.rowCount ?? 0);
|
||||
if (total > 0) {
|
||||
console.log(`Recovered ${total} stale jobs from previous run`);
|
||||
}
|
||||
@@ -166,7 +185,7 @@ export function recoverStaleJobs(): void {
|
||||
*/
|
||||
export function updateJobProgress(progress: JobProgress): void {
|
||||
jobProgressStore.set(progress.jobId, progress);
|
||||
persistJobProgress(progress);
|
||||
enqueuePersist(progress.jobId, () => persistJobProgress(progress));
|
||||
// Notify all SSE listeners (add type: "batch" so the frontend can distinguish
|
||||
// batch events from single-file events in the shared SSE stream)
|
||||
const subs = listeners.get(progress.jobId);
|
||||
@@ -187,7 +206,7 @@ export function updateJobProgress(progress: JobProgress): void {
|
||||
|
||||
export function updateSingleFileProgress(progress: Omit<SingleFileProgress, "type">): void {
|
||||
const event: SingleFileProgress = { ...progress, type: "single" };
|
||||
persistSingleFileProgress(progress);
|
||||
enqueuePersist(progress.jobId, () => persistSingleFileProgress(progress));
|
||||
|
||||
if (progress.phase === "complete" || progress.phase === "failed") {
|
||||
if (singleFileCompletions.size >= 10_000) {
|
||||
|
||||
Reference in New Issue
Block a user