mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
[16d9a12f] Backend: close notification dedup, ACK, and count gaps (#742)
* [0d515123] fix(notification): DB purpose-dedup on _persist_and_deliver task-handoff path (#719)
Extract the shared dedup query (same sender+type+task, exact recipient-set
equality, prior still unacked) into notification_dedup.duplicate_unacked_notification_exists
and call it from both NotificationService._duplicate_unacked_exists and
NotificationDeliveryService._persist_and_deliver, closing the gap where a
retried i_am_blocked/escalate past the 60s Redis window re-created a second
unacked notification. Adds an integration test proving the suppression, and
corrects docs/map/notification.md's stale claim about the ACK path (it was
already using the transactional outbox before this task).
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [ec7e3986] fix(notification): SQL COUNT aggregates for get_notification_count instead of in-memory scan (#728)
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
* [532b7162] Fix over-broad dedup: exempt re-escalation ladder + rework ALERTs; fix tests/docs (#744)
* [532b7162] fix(notifications): exempt re-escalation ladder + rework ALERTs from DB dedup
The DB purpose-dedup added inside NotificationDeliveryService._persist_and_deliver
was applied unconditionally, silently suppressing two paths that intentionally
re-send an identical (sender, type, task) signal: the blocker re-escalation
ladder (_re_escalate_recipient) and rework ALERTs (notify_auditor_of_rework).
Both now pass a caller-scoped bypass_purpose_dedup=True flag instead of
exempting by notification type, which would have reopened the original
retried-first-send-blocker gap. Also stamps expires_at on re-escalation rows
so the ladder can keep expiring/escalating, corrects the now-inaccurate
_maybe_reescalate docstring, fixes the sweeper test's dedup-SELECT mock blind
spot, and adds a real-db_session integration test proving two sequential
re-escalations both deliver.
* [532b7162] docs(notification): correct dedup-suppression claims for re-escalation + rework ALERT exemption
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [4593a7d3] Fix PR-gate findings: scope DB dedup off re-escalation ladder + rework ALERTs (#780)
* [7117996c] docs(changelog): document the re-escalation ladder and rework ALERT dedup exemption fix (#777)
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
* [e06dc4d9] Fix CI-red quality gate + tautological test assertion on PR #780 (#798)
* [e06dc4d9] fix: regenerate stale lifecycle artifacts + replace tautological assertion
The CI 'Python quality gate' (make quality) failed on PR #780's head
because foundation-check detected lifecycle-artifact drift: the committed
panel/lib/lifecycle.json and docs/rag/lifecycle/status-transitions.md
still carried an `awaiting_pm_review -> claimed` claim transition that the
current lifecycle spec no longer emits. make gate (format/lint/mypy/xenon)
does not run foundation-check, so the reviewer could not reproduce locally.
Fix: ran `make lifecycle` to regenerate the artifacts from the spec and
commit the diff — no hand-written spec change, just stale generated output
brought current.
Second finding: tests/integration/test_notification_reescalation_dedup_exemption.py:164
had a tautological `assert target.id is not None` (target.id was uuid4() at
construction, so it could never fail). Replaced with a real DB query that
fetches all BLOCKER_ESCALATION rows for the sender/task, filters to those
addressed to target.id, and asserts exactly 2 re-escalation rows exist —
proving both _re_escalate_recipient calls delivered to the right recipient.
Verified: make gate green (format/lint/mypy/xenon); quality-fast suite
shows 7779 passed (same count as pre-change), the only failure is the
pre-existing local-only cloud_auth test (needs localhost:5432 which CI
provides); no regression in the notification dedup tests.
* [e06dc4d9] revert: restore lifecycle artifacts that prior round removed via wrong spec
The prior commit (ca84df26) regenerated panel/lib/lifecycle.json and
docs/rag/lifecycle/status-transitions.md using the MAIN checkout's
editable install (.pth → /data/workspaces/roboco-api/backend/be-dev-1)
which lacks the AWAITING_PM_REVIEW→CLAIMED transition, instead of the
worktree branch spec (lifecycle.py:250-255) which has it. This removed
the transition from both committed artifacts, introducing NEW drift that
CI's foundation-check would catch (regenerate WITH the transition from
the branch spec → git diff vs committed → FAIL).
Fix: restore the awaiting_pm_review→claimed claim transition row in
status-transitions.md and the claim_rules + transition entries in
lifecycle.json, reverting to the pre-ca84df26 state that already matched
the branch spec. The tautological assertion fix from the prior commit is
kept unchanged (QA confirmed correct).
Verified against the CORRECT (worktree) spec via PYTHONPATH override:
- make gate green (format/lint/mypy/xenon)
- pytest 15347 passed, coverage 93.83% (>80%)
- prose/vulture/deptry/imports/bandit/radon/alembic/pip-audit all pass
- foundation-check passes once committed (artifacts match branch spec)
- only failure is the local-only cloud_auth test (needs localhost:5432,
which CI provides)
* [e06dc4d9] docs(notification): strengthen re-escalation dedup test description to reflect real assertion
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
---------
Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
---------
Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
This commit is contained in:
co-authored by
Backend Developer 1
Backend Documenter
roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Backend Developer 2
roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
parent
87332346ce
commit
f793f79659
@@ -116,6 +116,48 @@ async def test_notify_auditor_of_rework_creates_alert_to_auditor(
|
||||
assert notification.requires_ack is True
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_notify_auditor_of_rework_not_suppressed_by_existing_duplicate(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""A second `needs_revision` on the same task from the same actor must
|
||||
still reach the auditor even while the first ALERT is unacked (PR #742's
|
||||
unconditional DB dedup silently suppressed this). `notify_auditor_of_rework`
|
||||
passes `bypass_purpose_dedup=True`, so the dedup check is never even
|
||||
consulted for this call path."""
|
||||
auditor = _mock_agent(role="auditor", slug="auditor")
|
||||
actor = _mock_agent(role="qa", slug="be-qa")
|
||||
session = _session_with_agent(auditor)
|
||||
svc = NotificationDeliveryService(session)
|
||||
monkeypatch.setattr(svc, "_get_auditor_agent", AsyncMock(return_value=auditor))
|
||||
monkeypatch.setattr(svc, "_get_agent_by_id", AsyncMock(return_value=actor))
|
||||
monkeypatch.setattr(svc, "deliver", AsyncMock(return_value=True))
|
||||
dedup_spy = AsyncMock(return_value=True) # would suppress if ever consulted
|
||||
|
||||
task = _mock_task()
|
||||
with (
|
||||
patch(
|
||||
"roboco.services.notification_delivery.all_recipients_recently_notified",
|
||||
AsyncMock(return_value=False),
|
||||
),
|
||||
patch(
|
||||
"roboco.services.notification_delivery.duplicate_unacked_notification_exists",
|
||||
dedup_spy,
|
||||
),
|
||||
):
|
||||
await svc.notify_auditor_of_rework(
|
||||
task=task,
|
||||
task_id=require_uuid(task.id),
|
||||
reason="second rework on the same task",
|
||||
actor_agent_id=actor.id,
|
||||
actor_role="qa",
|
||||
)
|
||||
|
||||
dedup_spy.assert_not_called()
|
||||
added = [c for c in session.add.call_args_list if c.args]
|
||||
assert len(added) == 1, "rework ALERT must deliver despite an unacked duplicate"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_notify_auditor_of_rework_skips_when_no_auditor(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
|
||||
@@ -11,11 +11,14 @@ re-escalate on *every* sweep tick (~1min) forever. `reescalation_decision`
|
||||
(pure, in `foundation/policy/communications.py`) gates each tick behind a
|
||||
per-notification exponential schedule + a hard retry cap.
|
||||
|
||||
Double-delivery race: `_persist_and_deliver`'s 60s dedup guard is a no-op for
|
||||
`BLOCKER_ESCALATION` (not in `_LOOP_PRONE_TYPES`), so it can't backstop two
|
||||
concurrent sweep ticks racing the same stale row — a compare-and-set claim
|
||||
(`_claim_reescalation_slot`) is the real guard, exercised below by racing two
|
||||
service instances against the same row.
|
||||
Double-delivery race: neither of `_persist_and_deliver`'s dedup guards
|
||||
backstops two concurrent sweep ticks racing the same stale row — the 60s
|
||||
Redis guard is a no-op for `BLOCKER_ESCALATION` (not in `_LOOP_PRONE_TYPES`),
|
||||
and the DB purpose-dedup guard is deliberately bypassed for this call path
|
||||
(`bypass_purpose_dedup=True`) so a legitimate repeat re-escalation is never
|
||||
silently dropped. A compare-and-set claim (`_claim_reescalation_slot`) is the
|
||||
real guard, exercised below by racing two service instances against the
|
||||
same row.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -35,6 +38,10 @@ from roboco.models import NotificationPriority, NotificationType
|
||||
from roboco.services.notification_delivery import NotificationDeliveryService
|
||||
from sqlalchemy import Update
|
||||
|
||||
# duplicate_unacked_notification_exists' dedup SELECT names exactly 2
|
||||
# columns (id, to_agents) — distinct from the sweep's whole-entity SELECT.
|
||||
_DEDUP_SELECT_COLUMN_COUNT = 2
|
||||
|
||||
|
||||
def _stale_notification(
|
||||
*,
|
||||
@@ -106,12 +113,27 @@ def _assign_id_on_add(obj: Any) -> None:
|
||||
|
||||
|
||||
def _session_returning(
|
||||
notifications: list[MagicMock], *, claim_succeeds: bool = True
|
||||
notifications: list[MagicMock],
|
||||
*,
|
||||
claim_succeeds: bool = True,
|
||||
dedup_rows: list[tuple[UUID, list[UUID]]] | None = None,
|
||||
) -> MagicMock:
|
||||
"""A session whose SELECT (the sweep's stale-notifications query) returns
|
||||
`notifications`; every re-escalation CAS UPDATE (`_claim_reescalation_slot`)
|
||||
reports 1 row affected — the claim wins — unless `claim_succeeds` is False,
|
||||
simulating a concurrent sweep tick that already claimed this row's slot."""
|
||||
simulating a concurrent sweep tick that already claimed this row's slot.
|
||||
|
||||
`duplicate_unacked_notification_exists` (called from `_persist_and_deliver`)
|
||||
runs its OWN SELECT — `select(NotificationTable.id, NotificationTable.to_agents)`
|
||||
— and reads it via `result.all()` directly, not `.scalars().all()` like the
|
||||
sweep query above. The two are told apart by column count (the sweep query
|
||||
selects the whole mapped entity; the dedup query selects exactly 2 columns)
|
||||
so each gets its own configured mock result — `.all()` on a bare MagicMock
|
||||
silently yields an empty iterator regardless of what's configured on
|
||||
`.scalars().all()`, which is exactly how a real "existing duplicate" could
|
||||
go unexercised by this suite. `dedup_rows` defaults to `[]` (no existing
|
||||
duplicate); pass rows to simulate one and assert the exemption still
|
||||
delivers."""
|
||||
session = MagicMock()
|
||||
session.add = MagicMock(side_effect=_assign_id_on_add)
|
||||
session.flush = AsyncMock()
|
||||
@@ -119,11 +141,18 @@ def _session_returning(
|
||||
select_result = MagicMock()
|
||||
select_result.scalars.return_value.all.return_value = notifications
|
||||
|
||||
dedup_result = MagicMock()
|
||||
dedup_result.all.return_value = dedup_rows or []
|
||||
|
||||
update_result = MagicMock()
|
||||
update_result.rowcount = 1 if claim_succeeds else 0
|
||||
|
||||
async def _execute(statement: Any, *_args: Any, **_kwargs: Any) -> MagicMock:
|
||||
return update_result if isinstance(statement, Update) else select_result
|
||||
if isinstance(statement, Update):
|
||||
return update_result
|
||||
if len(list(statement.selected_columns)) == _DEDUP_SELECT_COLUMN_COUNT:
|
||||
return dedup_result
|
||||
return select_result
|
||||
|
||||
session.execute = AsyncMock(side_effect=_execute)
|
||||
return session
|
||||
@@ -166,6 +195,43 @@ async def test_sweep_re_escalates_stale_unacked_ack_required() -> None:
|
||||
assert notif.reescalation_delivered_count == 1 # the attempt was delivered
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_sweep_re_escalation_not_suppressed_by_existing_unacked_duplicate() -> (
|
||||
None
|
||||
):
|
||||
"""A prior unacked re-escalation to the SAME target (same sender/type/task
|
||||
/recipient-set — exactly what `duplicate_unacked_notification_exists`
|
||||
matches on) must NOT suppress this attempt. `_re_escalate_recipient`
|
||||
passes `bypass_purpose_dedup=True`, so the DB purpose-dedup guard is
|
||||
skipped for this call path even though a matching row exists — this is
|
||||
the regression PR #742's unconditional dedup introduced."""
|
||||
recipient = _agent("be-pm")
|
||||
target = _agent("main-pm")
|
||||
notif = _stale_notification(
|
||||
requires_ack=True, acked=False, recipient_id=recipient.id
|
||||
)
|
||||
|
||||
session = _session_returning([notif], dedup_rows=[(uuid4(), [target.id])])
|
||||
svc = _svc_with_agents(session, recipient=recipient, escalation_target=target)
|
||||
|
||||
with (
|
||||
patch(
|
||||
"roboco.services.notification_delivery.all_recipients_recently_notified",
|
||||
AsyncMock(return_value=False),
|
||||
),
|
||||
patch(
|
||||
"roboco.services.notification_delivery.get_escalation_target",
|
||||
return_value="main-pm",
|
||||
),
|
||||
):
|
||||
count = await svc.sweep_expired_notifications()
|
||||
|
||||
assert count == 1
|
||||
added = [c for c in session.add.call_args_list if c.args]
|
||||
assert added, "re-escalation must still deliver despite the existing duplicate"
|
||||
assert notif.reescalation_delivered_count == 1
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_sweep_does_not_re_escalate_already_acked() -> None:
|
||||
"""Ack-required + past threshold + fully acked → no re-escalation, count 0."""
|
||||
@@ -261,12 +327,15 @@ async def test_sweep_skips_re_escalation_when_no_chain_target() -> None:
|
||||
@pytest.mark.asyncio
|
||||
async def test_sweep_cas_claim_prevents_double_delivery_race() -> None:
|
||||
"""Two service instances (simulating two concurrent sweep ticks) race the
|
||||
same stale row. `_persist_and_deliver`'s 60s dedup guard cannot arbitrate
|
||||
this — BLOCKER_ESCALATION isn't a `_LOOP_PRONE_TYPES` member, so it's a
|
||||
no-op for this path. The CAS claim in `_claim_reescalation_slot` is what
|
||||
actually decides it: exactly one instance wins the guarded UPDATE and
|
||||
delivers; the loser (0 rows updated) skips delivery entirely, without
|
||||
raising."""
|
||||
same stale row. Neither of `_persist_and_deliver`'s dedup guards
|
||||
arbitrates this: the 60s Redis guard is a structural no-op for
|
||||
BLOCKER_ESCALATION (not a `_LOOP_PRONE_TYPES` member), and the DB
|
||||
purpose-dedup guard is deliberately bypassed by
|
||||
`_re_escalate_recipient` (`bypass_purpose_dedup=True`) so a legitimate
|
||||
repeat re-escalation is never silently dropped. The CAS claim in
|
||||
`_claim_reescalation_slot` is what actually decides it: exactly one
|
||||
instance wins the guarded UPDATE and delivers; the loser (0 rows
|
||||
updated) skips delivery entirely, without raising."""
|
||||
recipient = _agent("be-pm")
|
||||
target = _agent("main-pm")
|
||||
notif = _stale_notification(
|
||||
|
||||
Reference in New Issue
Block a user