Files
roboco/docs/backend/services/coordination-events.md
T
1114ee5ea0 [77719d3f] A2A team telemetry: coordination event notifications for 5 event types (#477)
* [13d03d5c] Add 5 coordination-event notification producers + wire at chokepoints (#472) (#474)

* [13d03d5c] Add 5 coordination-event notification producer methods

* [13d03d5c] Wire reassignment/collision/unblock/dependency-revival notifications

* [13d03d5c] Wire stale-claim-reaped notification into orchestrator reaper

* [13d03d5c] fix(runtime): guard reaper's UUID annotation + defensive attr access

The stale-claim-reaped notification hook added a runtime-unquoted
`UUID` type annotation (only imported under TYPE_CHECKING, so the
module raised NameError on import) and a direct `t.assigned_to`
attribute access that crashes against the minimal test doubles the
existing reaper test suite uses. Quote the annotation and switch to
getattr-defensive access, matching `_assignee_is_provider_parked`'s
existing convention in the same file.

* [13d03d5c] test(notification): unit coverage for 5 coordination-event producers

One test per new send_* method (reassignment, collision-sequencing,
unblock, dependency-revival, stale-claim-reaped) following the
existing _FakeDb/_patch_db_context pattern, asserting subject/body/
related_task_id/priority/recipient-count, plus a no-recipients no-op
case for reassignment.

* [13d03d5c] test(task): prove reassign + unblock don't double-fire notifications

Two chokepoint-level tests mocking NotificationService at its defining
module: a repeated reassign() to the same already-current target skips
the notification (guarded by comparing against the pre-mutation
assignee), and a repeated unblock() on the same task only notifies
once since the second call short-circuits on the status!=BLOCKED
guard.

* [13d03d5c] style(task): ruff format the collision-sequencing wiring block

No behavior change — reflows the newly-added _notify_collision_sequencing
call site to satisfy ruff format's line-length rules.

* [13d03d5c] docs(backend): add coordination-event notification producers guide

Documented the 5 new NotificationService producers (reassignment, collision-sequencing,
unblock, dependency-revival, stale-claim-reaped) with fire conditions, double-fire
prevention mechanisms, and implementation patterns. Updated backend README to link the
new services guide for developers integrating new coordination events.

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [3ee8150b] Frontend: render coordination-event notifications + e2e smoke coverage (#475)

* [69777c3a] test(e2e-smoke): add coverage for soft-block + unblock coordination notifications (#471)

Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>

* [8eb82639] Render 5 coordination-event notification types with task deep-links (#470)

* [8eb82639] feat(notifications): add APPROVAL type icon and deep-link component test

Add missing APPROVAL member to the frontend NotificationType enum to
match backend roboco/models/base.py, wire its icon into the existing
typeIcons Record in the notifications page, and add a component test
covering type rendering and the task deep-link.

* [8eb82639] docs(notifications): document 5 coordination-event types and APPROVAL enum addition

Added comprehensive reference guide explaining the 5 notification types
(TASK_ASSIGNMENT, BLOCKER_ESCALATION, REVIEW_REQUEST, DOCUMENTATION_REQUEST,
APPROVAL), their visual identities (icon + color), use cases, and
deep-linking behavior to related tasks. Updated panel README with quick
reference table. TypeScript Record pattern ensures exhaustive type coverage
at build time.

---------

Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>

---------

Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>

* [a27de2a8] fix(docs): reflow hard-wrapped notification-types.md to pass markdown gate (#479) (#481)

The Python quality gate on assembled PR #477 was red because the newly
added docs/frontend/components/notification-types.md (introduced by the
frontend coordination-event rendering commit) had manually wrapped prose
paragraphs, which scripts/reflow_md.py --check rejects as part of make
quality. Reflowed the file with scripts/reflow_md.py --apply (whitespace
only, no content change) so the check passes. ruff format/check, mypy,
xenon, vulture, bandit, and the full pytest suite (10284 passed) all
confirmed green on this commit; notification.py, task.py, and
orchestrator.py are untouched.

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>

* [705419d5] Remove duplicate unblock notification and fix its dependent tests (#485) (#488)

* [705419d5] fix(notifications): remove duplicate unblock notification, fix its tests

The /unblock route was still calling delivery.notify_assignee_of_unblock()
(TASK_ASSIGNMENT) after TaskService.unblock() already sent the
send_unblock_notification() ALERT wired in by an earlier task — a real
duplicate notification on every unblock. Delete the route-layer call and
the now-dead NotificationDeliveryService.notify_assignee_of_unblock
method, fix the integration test that mocked it, and fix/extend the e2e
notification-coordination-events test to assert the persisted ALERT rows
(exact subjects) for both the direct-unblock and dependency-revival
producers instead of the old TASK_ASSIGNMENT assertion.

* [705419d5] docs(backend): update coordination-events doc for unblock duplicate removal

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [6c142a73] docs(changelog): document restored coordination-event notification producers and add collision-sequencing double-fire test (#489) (#490)

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [77719d3f] Seed system agent in e2e harness to fix unblock/dependency-revival notifications

The e2e harness's seed_company omitted the system sentinel agent that
production seeds via initial_data.py. The unblock and dependency-revival
notification producers default to from_agent="system", which
_resolve_agent_uuid looks up by slug in the DB. With no system row the
resolver returns None and _create_notification silently skips the
notification, so the two ALERT assertions got 0 rows instead of 1.

The soft-block test passed because it uses NotificationDeliveryService
which creates the notification directly with a real agent UUID as
from_agent, bypassing the slug resolution path entirely.

* [77719d3f] Use foundation UUID for system agent to avoid slug collision

The first attempt seeded the system agent with a random UUID. Other
tests (_seed_system_and_secretary, _seed_video_agents) check by the
fixed foundation UUID via session.get(AgentTable, uuid); not finding
it they INSERT their own system row, hitting ix_agents_slug. Using the
foundation UUID makes their check find the seed_company row and skip.

* [77719d3f] Fix dependency-revival notification event loop mismatch

The dependency-revival test calls _unblock_dependents directly via
stack.run_db, which creates a new asyncio event loop. Inside,
_notify_dependency_revival -> NotificationService._create_notification
opened its own session via get_db_context(), which reuses the singleton
_DbHolder engine — bound to the FastAPI server's event loop. The
asyncpg connection raised 'Future attached to a different loop' and the
exception was silently caught + logged as a warning, so the notification
never persisted and the test saw 0 rows.

Fix: add an optional db_session parameter to _create_notification and
the two send methods. When provided, use the caller's session directly
and skip the internal commit (the caller owns the transaction). The
TaskService's _notify_unblock and _notify_dependency_revival now pass
self.session, keeping the notification in the same event loop + session
as the task transition.

* [77719d3f] Scope system-agent seeding to notification tests only

Seeding the system sentinel in seed_company (commits 3bba7b32/617b7890)
fixed the 0-notification bug but caused 3 i_documented gateway_timeout
failures: every e2e test now paid notification-creation latency for
system-origin notifications that were previously silently skipped,
pushing the already-slow i_documented verb past its 120s timeout.

Move system-agent seeding out of seed_company and into a scoped
_seed_system_agent helper called only by the two coordination-event
tests that exercise send_unblock_notification /
send_dependency_revival_notification (both resolve from_agent='system'
via DB lookup). dev_lifecycle and state_machine tests revert to the
pre-fix behavior (system-origin notifications silently skipped, no extra
latency).

The event-loop fix (commit 7b95d77d: pass db_session=self.session to
_create_notification) is unchanged — dependency_revival still needs it
because stack.run_db creates a new event loop while _DbHolder.engine is
bound to the FastAPI server loop.

* [77719d3f] Fix reassignment notification deadlock + suppressed-notification commit regression

Two fixes in notification.py / task.py:

1. Cross-session self-deadlock in send_reassignment_notification:
   TaskService.reassign() flushes an uncommitted row lock on the task,
   then calls _notify_reassignment -> send_reassignment_notification ->
   _create_notification(db_session=None) which opens a SEPARATE session
   via get_db_context() and INSERTs a notification with related_task_id
   FK -> tasks.id. The FK key-share lock blocks on the request session's
   uncommitted exclusive lock, but the request can't commit until the
   notify returns -> 120s verb hard-cut. Fix: pass db_session=self.session
   so the notification joins the verb's own transaction, same pattern as
   the unblock/dependency-revival fix in 7b95d77d.

2. Suppressed-notification commit regression: the 7b95d77d refactor moved
   await db.commit() out of _create_notification_with_session into
   _create_notification's db_session=None branch, where it ran
   unconditionally — even when _create_notification_with_session returned
   early (suppressed: unresolvable from_agent / no recipients /
   refire-guard / dedup-hit). Fix: _create_notification_with_session now
   returns bool (False at each early return, True after delivery);
   _create_notification commits only when created is True.

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-13 06:38:15 +02:00

9.0 KiB

Coordination-Event Notifications

The NotificationService provides five typed producers for coordination events—state transitions that affect multiple agents across a task lifecycle. Each producer fires at a specific chokepoint and is guarded against double-firing through idempotent upstream conditions.

Overview

Coordination events differ from generic task-state notifications: they signal changes that require coordination between agents (reassignments, dependency unblocks, sequencing blocks) or escalations (stale claims). Every producer follows the same pattern:

  1. Typed producer method in NotificationService (e.g., send_reassignment_notification)
  2. Best-effort wiring at the chokepoint via a helper method (e.g., _notify_reassignment in TaskService)
  3. Try/except guard — a notification failure never breaks the underlying state transition
  4. Double-fire prevention — idempotent upstream guards ensure no duplicate notifications

The Five Producers

1. Reassignment: send_reassignment_notification

Fired from: TaskService.reassign()_notify_reassignment()

Trigger: A task is reassigned to a different agent.

Recipients: Previous assignee + new assignee + CEO

Double-fire guard: _notify_reassignment checks new_assignee == previous_assignee and returns early if nothing changed. TaskService only captures the assignee once (pre-mutation) so if reassign is called twice in succession, the second call operates on an already-updated assignee and the guard catches it.

Subject & body:

Subject: Task {task_id} reassigned
Body: Task {task_id} was reassigned from {previous} to {new}.

Example scenario: A PM delegates a task from developer A to developer B. Both developers and the CEO receive notification of the change.


2. Collision Sequencing: send_collision_sequencing_notification

Fired from: TaskService.wire_sibling_collision_dag()_notify_collision_sequencing()

Trigger: A file/migration/shared-surface collision causes the collision-sequencing analyzer to add a dependency edge, holding back one task behind a blocking sibling.

Recipients: Held-back task's assignee + CEO

Double-fire guard: The wiring only fires the notification when add_dependency returns True (a freshly-inserted edge). Subsequent calls to wire_sibling_collision_dag over the same sibling pair contribute no new edge and therefore trigger no notification.

Subject & body:

Subject: Task {held_back_task_id} sequenced behind a sibling
Body: Task {held_back_task_id} was held back by the collision-sequencing 
       analyzer: it now depends on task {blocking_task_id}, which surfaced 
       an overlapping file/migration/shared-surface collision. It will 
       resume once that task reaches a terminal state.

Example scenario: Two backend dev tasks both touch roboco/models/task.py and one adds a migration. The analyzer detects a collision and adds a sequencing edge, notifying the developer of the held-back task.


3. Unblock: send_unblock_notification

Fired from: TaskService.unblock() / unblock_with_restore()_notify_unblock()

Trigger: A PM or resolver explicitly unblocks a task that was in BLOCKED status.

Recipients: Restored owner + CEO

Double-fire guard: Both unblock and unblock_with_restore are guarded by a status check (status != BLOCKED short-circuits early). A repeated call against an already-unblocked task is a no-op upstream and never reaches the notification handler.

Route-layer duplicate removed: The POST /api/tasks/{id}/unblock route previously called NotificationDeliveryService.notify_assignee_of_unblock() (a TASK_ASSIGNMENT notification) after TaskService.unblock() had already sent the ALERT above. That route-layer call and the now-dead delivery method have been removed, so the ALERT from the service chokepoint is the only notification fired for an unblock.

Subject & body:

Subject: Task {task_id} unblocked
Body: Task {task_id} has been unblocked and handed back to {owner}.
       It is ready to resume.

Example scenario: A task was blocked by an external dependency. The dependency resolves, the PM calls unblock(), and the task owner is notified it's ready to resume.


4. Dependency Revival: send_dependency_revival_notification

Fired from: TaskService._unblock_dependents()_notify_dependency_revival()

Trigger: A task's last outstanding dependency completes, automatically reviving the task.

Recipients: Revived task's assignee + CEO

Double-fire guard: _unblock_dependents prunes the dependency_ids list before firing the notification. A repeated call for the same completed dependency finds no matching dependent and never reaches the notification handler.

Distinct from send_unblock_notification: This fires when a dependency completes automatically (no resolver acted). send_unblock_notification fires when a PM explicitly calls unblock on a task blocked by escalation. The notification names which dependency unblocked it rather than who resolved it.

Subject & body:

Subject: Task {task_id} revived by dependency completion
Body: Task {task_id} was revived: its dependency {completed_dependency_id} 
       just completed and no other dependencies remain. It is ready to resume.

Example scenario: A task was blocked waiting on three dependencies. The first two complete and unblock nothing (others remain). The third completes, the last dependency clears, and the task is auto-revived with a notification.


5. Stale Claim Reaped: send_stale_claim_reaped_notification

Fired from: Orchestrator's _reap_with_service()_notify_stale_claim_reaped()

Trigger: The reaper detects a stale claim (no heartbeat updates) and releases it back to PENDING.

Recipients: Reaped agent + CEO

Priority: HIGH (higher than other coordination events; stale claims are operational issues)

Double-fire guard: A reaped task leaves list_in_progress_or_claimed once released to PENDING. A subsequent reaper tick never re-considers the same claim and cannot re-fire the notification.

Subject & body:

Subject: Task {task_id}: stale claim reaped
Body: Task {task_id}'s claim went stale (last heartbeat: {timestamp}) 
       and was reaped back to pending, releasing it from {reaped_agent}.

Example scenario: An agent crashed or became unresponsive while holding a task claim. The reaper detects the stale heartbeat, releases the claim, and notifies both the agent and CEO of the forced release.


Implementation Pattern

Every producer follows a consistent best-effort pattern at its call site:

async def _notify_<event>(self, ...) -> None:
    """Best-effort coordination notification for a <event>."""
    if <guard condition>:
        return
    try:
        from roboco.services.notification import NotificationService
        await NotificationService().send_<event>_notification(...)
    except Exception as e:
        self.log.warning(
            "<Event> notify failed", task_id=str(...), error=str(e)
        )

Why this pattern:

  • Localized guards in the helper prevent unnecessary notification attempts
  • Try/except ensures a notification failure never breaks the state transition
  • Logged & swallowed — operational visibility without crashing the flow
  • Lazy import avoids circular dependencies at the service layer

Adding a New Coordination Event

When a new coordination event arises:

  1. Add a new producer method to NotificationService following the existing signature (typed params, docstring describing fire condition + double-fire guard, calls _create_notification with related_task_id set)
  2. Add a helper in the originating service (TaskService, Orchestrator, etc.) following the best-effort pattern
  3. Wire at the chokepoint — the single place the state transition happens
  4. Document the guard in the helper's docstring so reviewers understand why no duplicate can fire
  5. Test the guard — add a chokepoint-level test proving a repeated call doesn't fire twice

See test_notification.py and test_task.py for worked examples of producer-level and chokepoint-level tests.

  • Implementation: roboco/services/notification.py (producers)
  • Wiring: roboco/services/task.py (TaskService helpers), roboco/runtime/orchestrator.py (reaper hook)
  • Route: roboco/api/routes/tasks.py (unblock endpoint; service-layer ALERT only, no route-level duplicate)
  • Unit tests: tests/unit/services/test_notification.py (producer unit tests), tests/unit/services/test_task.py (chokepoint double-fire proofs)
  • Route/chokepoint integration tests: tests/integration/test_tasks_routes.py (unblock returns 200 and leaves blocked with no duplicate notification), tests/e2e_smoke/test_notification_coordination_events.py (DB-truth checks for unblock and dependency-revival ALERT rows)
  • Data model: roboco/models/notification.py (CreateNotificationParams)