Files
roboco/tests/unit/runtime
2bd35e1c9e fix(run-hardening): stop three blocked-task respawn loops (#253)
* fix(run-hardening): stop three blocked-task respawn loops

Three independent fixes for blocked-task respawn loops observed in the live
run (the bleeders behind a wedged near-complete run):

- verb runner: re-check the working task after EACH composed atomic action,
  not just at entry. A concurrent transition between a verb's precondition
  gate and execution (e.g. a racing i_am_blocked moving a root from
  needs_revision to blocked) made claim() return None mid-sequence; the next
  composed step dereferenced None.id and crashed with the opaque
  "'NoneType' object has no attribute 'id'", looping the PM. Now fails fast
  with an actionable INVALID_STATE; the savepoint rolls the partial run back.

- blocker dispatch: never dispatch a Board role (product-owner / head-
  marketing) as a blocker resolver. Board roles have no unblock verb, so the
  dispatcher respawned one forever to "resolve" a blocker it could only
  notify/triage about — one incident burned ~6400 tool calls on a single
  mis-owned root. _blocker_resolver_slug now returns None for a Board
  assignee so the dispatch skips it.

- git push: recover a missing local task-branch ref from origin/<branch>
  before push-by-name. A re-provisioned shared clone can lack the branch
  locally though its commits are on origin, so push died on
  "src refspec <branch> does not match any" and the task wedged at i_am_done.
  Now materializes the ref (no-op push when already on origin) or fails loud
  with an unclaim+reclaim instruction when the work is on neither.

Adds regression tests for all three. Full no-DB gate green (ruff, reflow,
mypy, xenon); pytest+coverage validated by CI.

* fix(verb-runner): only raise on an INTERMEDIATE composed None, not the last

The mid-composition None-guard was too aggressive: it raised for a None
returned by the LAST composed action too (e.g. start()), preempting the
caller's existing `if task is None` handler that surfaces the verb-specific
message ("start failed for task ...", the board verb's decline envelope).
Three tests asserting those messages broke in CI.

Only an INTERMEDIATE None is fatal (the next action would deref None.id). A
None from the last action is the verb's own result and must flow out as the
runner's return value. Guard now fires only for position > 0, before the
next dispatch — still prevents the crash, preserves the last-action contract.

* fix(escalation): never hand a Main-PM coordination root to the Board

The upstream cause of the board catch-22 (which the orchestrator-side
blocker-dispatch guard only backstopped): the escalation chain points
main-pm -> product-owner, and i_am_blocked/escalate REASSIGNS the task to
that chain target. apply_escalation's board-advisory guard only refused
descendant cell tasks (both predicates require parent_task_id), so a
top-level Main-PM coordination root slipped through and the whole root was
reassigned to the Product Owner + marked blocked. The board has no unblock
verb, so it spam-notified the CEO and respawn-looped (~6400 tool calls on
one root).

Add _is_coordination_task (team == main_pm — covers a delivery root AND a
MegaTask root-subtask) and a shared _board_cannot_own predicate, applied at
all four board-refusal sites (escalation, reassign, reassign_active_claim,
dependency-revival). A main_pm coordination task escalated/reassigned onto a
board role is now diverted to the pool for a role-matched (Main-PM) reclaim.

Complements the blocker-dispatch backstop in the prior commits (defense in
depth). Tests: coordination-root predicate cases + apply_escalation divert;
existing teamless-root / board-root behavior unchanged.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-06-24 15:43:30 +02:00
..
2026-06-24 01:15:57 +02:00
2026-06-24 01:15:57 +02:00