6 Commits
Author SHA1 Message Date
Catubba f8ce39155b fix(power): lease a backup server per machine, not per configured device
Two datastores on one box are two devices by design, but one power switch. The
lease refcounted per device id, so a single sync route between them held two
independent leases on one machine: releasing the first shut the box down, and
releasing the second reached a machine already going down -- an SSL EOF from the
idle check, then a closed port 22 -- and recorded LEFT_ON. The run's
notification therefore warned "PBS left powered on" about a box that had gone to
sleep exactly as intended. Acquire was already correct by accident, because
_bring_up probes the host; only release, which never probes, was wrong.

lease_key() is the host, normalised the way discovery normalises it, falling
back to the device id when there is no host yet -- half-configured entries are
legal mid-wizard and would otherwise all collide on "". state() takes the device
rather than an id, and _pending_pbs_ids became _pending_pbs_keys: the queue's
answer and the lease's key must be the same space, or a queued run on a box's
other datastore stops holding it.

Re-keying alone was not enough. A same-machine sync then acquired one lease
twice and rendered "left on: still needed by another run" -- true only in the
sense that the same run holds it, so one wrong sentence for another. A run's
devices are now deduplicated by lease key: one machine, one wake, one power-off,
and the multi-device step labels fall away on their own.

Every device on a held machine now reports holders > 0, so the power button is
disabled on the sibling too. That is the point: an SSH poweroff takes down every
PBS instance on the box, including the one a run is using. Port is deliberately
not part of the key for the same reason.
2026-08-07 12:16:56 +02:00
Catubba f54c7f04e9 fix(jobs): stopping a run during a wake no longer leaves the server on
Two bugs on the same path, both found by stopping a real run on real
hardware.

The magic packet goes out while the lease still has no holder — `acquire`
registers one only once the box answers. Stop the run during that wait and
it raised before any lease existed, so `held` was empty and `_release_all`
had nothing to release: no power-off step in the timeline, no release, and a
machine that finished booting seconds later and stayed on until someone
noticed. The stop dialog's power-off toggle had nothing to act on.

The second only shows on a route with more than one server. The cancel probe
was passed to the *instant* reachability check as well, and that check runs
before the TCP connect — so once the flag was set the next device's probe
returned "unreachable" without touching the network, and the code sent a
magic packet to a box it had never even asked. A stopped run woke a machine
nobody wanted woken.

Both come down to one invariant, now held either side of the packet: don't
send one for a run that is already stopping, and once one is sent, see the
wake through so the caller gets a lease. Everything that puts a box back to
sleep — the power-off step, the queued-route check, the unmanaged check, the
"left powered on" warning, the stop dialog's toggle — already hangs off that
lease and now applies unchanged. The grace wait is bounded by what remains
of the interrupted attempt, so stopping a run never costs more than the wake
it interrupted.

A run stopped while its servers were coming up also no longer starts its
cycle: the first thing that cycle does is talk to a server that booted
seconds ago, and an error there filed the run as FAILURE and notified about
it — for a run the user stopped by hand.
2026-08-05 13:20:50 +02:00
Catubba 70c808edfe feat(i18n): localize run step details, and notify when a PBS never wakes
Step details were stored as English sentences, so the run timeline stayed
English whatever the UI language was. `run_steps` now carries a
`detail_key`/`detail_params` pair alongside the text, the same seam
`runs.error_key` already had: the English rendering stays on the row for
pre-1.0 rows and for foreign strings (a task UPID, another product's error
text), while the API rebuilds Joulenap's own copy in the user's language on
read. `set_detail()` writes both halves in one call so they cannot drift.

Fixes carried in the same files:

- a run whose PBS never came up recorded a history row and nothing else.
  The unreachable handler re-raised, unwinding past both the lease release
  and the notification, so the one failure this product exists to have an
  opinion about sent no Telegram, no ntfy, no email — a regression against
  0.9, which did notify. On a sync route it also left an already-woken box
  burning power with no POWEROFF step. It now falls through the same
  release-and-notify path a completed run uses.
- the notification's duration breakdown dropped the backup phase on every
  real backup route: the lookup matched bare step names, but a route labels
  its backup step per source (`backup:pve-alpha`). Matched with `_step_is`
  now, and summed per phase.
- a PBS sync remote outlived the route that created it, parking the peer's
  full API token in the executing box's `remote.cfg`. Remote and sync job
  are both removed once the sync is over.
- a step that set a detail and then failed rendered the detail instead of
  the error, because a resolvable key wins over the raw string.
2026-08-05 09:40:26 +02:00
Catubba 0ad459e14c fix(backend): resolve the nine findings of the backend-block review
The three silent ones first. A 0.9 config that fails to convert used to boot an
empty config that looks exactly like a fresh install: the reason now reaches the
UI (GET /api/status.config_error), the activity log and an ERROR line, and the
.bak parachute is written on the failure branch too, so a later save from the
Advanced tab cannot destroy the original.

"PBS left powered on" was wrong in both directions -- every successful run
against an always-on PBS warned, and a sync route that left its target awake did
not. lease.release() returned one False for four situations, only two of which
cost power; it now returns a ReleaseOutcome that names the reason, which becomes
both the POWEROFF step's detail and RunContext.left_on. The interrupted-run path
keeps a step-derived rule, now paired per device and filtered by managed_power.

422 bodies echoed the whole config, secrets included: a config-level validator
raises at loc=(), so pydantic attached every token, the secret key, the password
hash, the SMTP and bot tokens as the error's input. One helper with
include_input=False now serves all three config-shaped 422 sites.

Also: a redaction placeholder with nothing to resolve against is rejected instead
of silently clearing the credential (a renamed device id, or a create from a
copied body); the ad-hoc "Run verify" asks for outdated_after=0, since None meant
"only never-verified" and skipped exactly the snapshots the button exists for;
the manual power-off holds the single-run lock so it cannot cut a vzdump that
started in the check-then-act gap; _current_run_id is cleared when a run ends, so
a stop landing between two runs cannot hit the wrong one; and the pre-migration
.bak is chmod 0600 like every other secret-bearing file.

Tests: 617 passed, 2 skipped. Every finding was reproduced against the real code
before the fix, and each new test confirmed failing on the pre-fix code.
2026-08-03 00:56:40 +02:00
Catubba 5dc8c5719a feat(api): per-route scheduler, route/device REST API and notifications
Wire the route model to the outside world and retire the 0.9 single-PVE/PBS one.

Scheduler: one cron job per enabled route (route:<id>), built from schedule.time +
days or the schedule.cron escape hatch, gated by the new app.scheduler_enabled
kill-switch. Missed-run detection is per route.

API: new /api/routes and /api/devices (CRUD, connection test, power, ad-hoc GC and
verify), with a 409 removal guard that names the routes still using a device.
Status, dashboard and metrics report one entry per route and per PBS; metrics keep
their joulenap_ names and gain route=/pbs= labels. Cancel moved to
POST /api/runs/{id}/stop, /api/guests is scoped to one PVE, and the WoL smoke test
moved into the wizard router.

Notifications: the positional 5-tuple becomes a RunContext, filtered per route.
Cycles now return that context instead of sending it; the job service sends it after
releasing the power leases, so the message can report whether the box went back to
sleep, and the wake/power-off steps are recorded in the run's timeline again.

Removes pve:/pbs:/backup:/maintenance.gc/maintenance.verify from the schema (old
files still load: the keys are stripped after the 1.0 migration runs), the 0.9
cycles and job entry points, and the config-shaped connector factories.

Also fixes redacted secrets being matched by list position rather than by device id,
and git-ignores the migration's rollback copy of config.yaml, which holds the same
tokens as the original.
2026-08-02 20:33:41 +02:00
Catubba e5d0d5cf58 feat(jobs): FIFO run queue and per-PBS power lease
Per-route schedules mean two routes can target the same PBS minutes apart,
so no single cycle can decide when the box goes back to sleep.

Add a FIFO run queue to JobService (enqueue/pending/current/dequeue, backed
by a drain worker) and a PowerLease that refcounts each PBS: the first holder
wakes it or finds it awake, the last one powers it off, and only when the run
succeeded, the device manages its power, and no queued route still needs it.
Sync routes hold two leases, released independently. An unmanaged PBS is
probed, never woken, never powered off.

The queue sits beside the existing single-run lock rather than replacing it:
the worker takes the same lock, so queued runs and the 0.9 entry points still
serialise against each other while the cycle, scheduler and API are ported.
2026-08-02 01:18:26 +02:00