mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
[6788ce7f] Silent bug sweep: concurrency, state integrity, engine edge-cases, panel data freshness (#638)
* [943d8c4d] Frontend data freshness and approval-queue reliability audit (#631)
* [233a8b0f] WebSocket reconnect message-loss audit and fix (#625)
* [233a8b0f] fix(panel): add REST catch-up to useNotificationStream on WS reconnect
connection.ts has no message buffering/replay, so a notification published
while the CEO bell's socket was down (disconnected/reconnecting) was lost
forever instead of merely delayed. Add a reconnect-triggered GET
/notifications?unread_only=true catch-up folded into the existing
notification_id dedup so a notification delivered both via catch-up and
live WS is never double-counted, and make clearMessages drop the held
catch-up batch too. use-a2a-live.ts and use-rate-limit-websocket.ts were
audited and already have working reconnect-triggered REST fallbacks
(verified via a2a/page.tsx, rate-limit-banner.tsx, usage-overview-panel.tsx
and their existing F083 tests) so no fix was needed there.
* [233a8b0f] docs(panel): add comprehensive WebSocket hooks reference and reconnect architecture guide
Add panel/docs/frontend/hooks.md with full API reference for useWebSocket, useNotificationStream (with new REST catch-up behavior), useAgentStream, useA2ALiveStream, and useConnectionStatus. Include examples, best practices, and testing guidance.
Add panel/docs/architecture/websocket-reconnect.md documenting the message-loss mitigation pattern: Strategy 1 (REST catch-up for events, used by useNotificationStream) and Strategy 2 (REST invalidation for state, used by A2A/rate-limit consumers), plus the dedup logic ensuring no notification is double-counted on reconnect.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
* [d5315683] fix(frontend): add distinct toast feedback for silently-swallowed x-post and release-proposal statuses, plus regression tests for all 4 approval queues (#626)
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [cd953838] Data-hook null-guard audit and API client 429 retry-by-method fix (#630)
* [cd953838] fix(panel): gate 429 retry by HTTP method, add hook null-guard regression tests
* [cd953838] chore(conventions): waive test-fixture wrapper in hooks null-guard test
* [cd953838] docs(frontend): document API rate-limit retry behavior and null-guard audit results
Added `docs/frontend/api-rate-limiting.md` to document the 429 retry strategy: GET/PUT auto-retry, POST/PATCH/DELETE require X-Idempotency-Key header. Updated `docs/frontend/hooks.md` to confirm the data-hook null-guard audit found all hooks already have correct `enabled` guards and include a regression test suite for the board-review poll on/off behavior and enabled-guard assertions.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [4534c71a] Backend concurrency, state-machine, and engine audit (#634)
* [41de844a] fix(lifecycle): sync CLAIM_RULES with runtime + clear stale claimant on PM hand-off (#627)
Two confirmed state-machine gaps found while auditing lifecycle.py,
task_lifecycle.py, the _ESCALATABLE_TO_BLOCKED bypass, and every
_REVIEW_QUEUE_STATES entry point:
- lifecycle.py's CLAIM_RULES/claim-ActionSpec/StatusTransition table
did not grant CELL_PM/MAIN_PM re-claim of AWAITING_PM_REVIEW even
though task.py's runtime _ROLE_CLAIM_STATUSES already granted it
and claimed the spec agreed -- the two tables had silently drifted,
breaking i_will_plan re-claim on an awaiting_pm_review task.
- docs_complete's _maybe_advance_to_pm_review pre-assigns a specific
owning PM via assigned_to but left claimed_by/active_claimant_id
pointing at the outgoing documenter, unlike every sibling transition
into a review-queue state. A stale active_claimant_id makes
content_actions.py's _active_claim_violation wrongly reject the
newly-assigned PM's own content writes before it formally claims.
Reassign claimed_by + active_claimant_id to the owning PM alongside
assigned_to.
Adds a regression test asserting the documenter's stale claim does not
survive the docs_complete -> awaiting_pm_review hand-off.
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
* [0c46f666] Engine dedup race + sequencing.py edge-case audit (#628)
* [0c46f666] fix(sequencing): dedup race audit + collision-edge fallback bug
Audited the list-open-then-originate dedup pattern across six engines:
RoadmapEngine, XEngine.run_cycle, DepUpdateEngine, and CIWatchEngine each
run inside exactly one sequential orchestrator-loop asyncio task (no other
call site invokes run_cycle), so they cannot race with themselves; their
in-cycle dedup sets/keys are correctly built before any commit. SelfHealEngine
is the same shape. VideoEngine.open_video_task is genuinely different: it is
reachable from the release-publish hook, the feature-spotlight hook, and the
on-demand POST /video/request route, so two overlapping calls for the same
occasion can both pass the "no open task yet" check before either commits.
Fixed by wrapping the check+insert in a short-lived Redis mutex (reusing
HeartbeatMutex) keyed by occasion, mirroring XPostService's existing
lock pattern, with a regression test proving only one of two concurrent
calls creates a task.
Verified ReleaseExecutor's half-landed retry path (release_commit_sha):
apply_version_bumps and write_changelog_entry both run as uncommitted
working-tree edits before commit_and_push's single `git add -A` + commit,
so a bumped-version-without-changelog state can never reach origin (and
therefore can never be observed by a fresh retry clone) - confirmed correct
with a real-git-repo regression test, no fix needed.
Fixed sequencing.py's dev_task_collision_edges: the `if edges: return edges`
short-circuit dropped the same-assignee-lane fallback entirely whenever ANY
surfaced sibling pair produced a collision edge, even for a completely
unrelated same-assignee pair with no declared surface. Now the fallback
always runs, skipping only pairs the analyzer already ordered (so the two
mechanisms can never disagree on direction for the same pair).
Verified sequencing.py rule 3 (all-shared batch generates no edges): correct
by inspection (_shared_last_edges skips every pair when both are shared) and
confirmed with a regression test - no fix needed.
* [0c46f666] docs(reference): concurrency audit summary - engine races, fixes, verified patterns
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [8f7f167a] Redis mutex pre-lock write audit (#629)
* [8f7f167a] Redis mutex pre-lock write audit: add cross-session regression test for XPostService.approve
Audited x_post_service.py, video_post_service.py, release_proposal.py, and
heartbeat_mutex.py for the pre-lock DB-write anti-pattern (a session write
that happens before the SET NX / HeartbeatMutex acquire returns a token,
letting a losing racer's stale write clobber a winner's committed state).
XPostService.approve, VideoPostService.approve, and
ReleaseProposalService.approve/reject already implement the correct
validate-pure-pre-lock, apply-under-lock pattern (the XPostService fix
already shipped per CHANGELOG.md: "X edited_body write deferred into the
single-flight lock (M5)"). HeartbeatMutex holds no AsyncSession at all, so
the anti-pattern is structurally inapplicable there.
Adds a genuine cross-session concurrency regression test to
test_x_post_service.py (a real second DB connection, not an in-process
mock) mirroring VideoPostService's existing cross-session test, proving a
concurrently-committed post survives and the CEO's edited body never lands
on the just-posted row.
* [8f7f167a] Remove redundant inline comments flagged by QA in cross-session regression test
Both comments restated what the surrounding docstrings already say
explicitly, per QA findings F-dbadd8f0 (line 294) and F-27ac051e (line
631) — no behavior change, tests re-verified green against a sandbox
Postgres.
* [8f7f167a] Remove inline trailing comments flagged by QA (correct file this time)
QA findings F-e6f3e6a6 and F-24189858 cited tests/unit/services/
test_x_post_service.py:294 and :631 across 5 revision rounds, but that
file never contained the flagged comment text — a repo-wide grep for
the exact quoted strings shows both comments actually live in the
mirrored tests/unit/services/test_video_post_service.py file, in its
own cross-session concurrency regression tests (the caption-edit and
tiktok-skip tests). Removed both there:
- "# externally visible to the "concurrent" session below" on the
db_session.commit() call
- "# never attempted without credentials" on the tiktok_poster.calls
assertion
Both restated what the surrounding docstrings/test names already say;
no behavior change. Verified with the full make quality gate against a
sandbox Postgres/Redis: 13,717 passed, 94.41% coverage, clean except
one pre-existing unrelated failure in tests/unit/api/test_cloud_auth.py
::test_login_route_parses_oauth2_form_not_query_params, which connects
to the app's default localhost:5432 Postgres (not the db_session
sandbox fixture) and is unreachable in this sandboxed environment —
structurally unrelated to the auth subsystem this task never touches.
* [8f7f167a] Redis mutex pre-lock write audit (round 7): add cross-session regression tests for reject() lock protection
Round-7 QA findings F-7eb9fbcb, F-06f39a2e, and F-4d56e49b claim
XPostService.reject(), ReleaseProposalService.reject(), and
release_executor._await_proc() lack lock protection / a CancelledError
handler — but their cited line ranges (255-267, 429-454, 241-257)
describe a pre-fix, shorter version of these functions that predates
commit fb293a787d, already on this branch. At current HEAD:
- x_post_service.py reject() (lines 275-299) acquires _LOCK_PREFIX,
re-reads under the lock, applies markers.set_x_reject_reason() +
CANCELLED only inside the critical section, releases in finally.
- release_proposal.py reject() (lines 460-486) does the identical
dance with _RELEASE_LOCK_PREFIX.
- release_executor.py _await_proc() (lines 257-265) already has an
`except asyncio.CancelledError` block that kills + reaps the child
and re-raises, mirroring the TimeoutError handler, with an existing
dedicated regression test
(test_await_proc_kills_child_on_outer_cancellation).
The one genuine gap: neither reject() path had a cross-session
(real second DB connection, not an in-process mock) regression test
proving the in-lock re-read catches a concurrent approve/publish that
completes mid-lock-wait — only approve() had one. Added
test_reject_concurrent_approve_completes_during_lock_wait to both
test_x_post_service.py and test_release_proposal_status_guards.py,
mirroring the existing approve() cross-session test: a second engine
commits COMPLETED between reject's pre-lock read and lock acquisition,
and the test asserts the CANCELLED write / reject-reason marker never
lands on the just-completed row.
No production code changed — verified via 103 targeted tests green
against a sandbox Postgres/Redis, plus `make -o sync gate` clean.
* [8f7f167a] Regenerate stale lifecycle artifacts (restore auditor waive_finding)
foundation-check was the only failing gate: the committed lifecycle artifacts
were missing the auditor's waive_finding verb that the lifecycle source
defines, so make quality regenerated them and failed on the diff — nothing to
do with the mutex fix (which passes ruff/mypy/tests/coverage/bandit clean).
make lifecycle restores the drift; this is what the 8 revision rounds kept
missing.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [d615f2e3] fix(tests): sync stale CLAIM_RULES pinning assertions with lifecycle.py (#635)
test_claim_rules_match_pre_gateway_table still asserted the pre-audit
two-member frozenset for CELL_PM/MAIN_PM claim rules. CLAIM_RULES in
lifecycle.py already grants both roles claim rights on
Status.AWAITING_PM_REVIEW (added by the state-machine exhaustiveness
audit) so a PM can re-claim its own review-queue task after a respawn.
Updated both assertions to include AWAITING_PM_REVIEW, matching the
actual dict. Grepped the repo for sibling stale copies of the old
literal; found none beyond this test.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [3c4e7a35] fix(quality-gate): reflow CONCURRENCY_AUDIT.md and stub occasion lock in video tests (#636)
Root cause: PR #634's CI failed at the markdown-reflow check (make quality
Makefile:285) on CONCURRENCY_AUDIT.md — a hard-wrapped audit doc left over
from the merged "Engine dedup race + sequencing.py edge-case audit" unit
(PR #628). Fixed with `make reflow-docs` (the exact remedy the CI output
itself named).
Running the full local `make quality` (with a sandbox Postgres/Redis to
get past DB-gated skips) surfaced a second real regression from that same
PR #628 unit: it added a Redis-backed HeartbeatMutex occasion lock to
VideoEngine.open_video_task, but two pre-existing test files
(tests/unit/runtime/test_video_render_loop.py and
tests/integration/test_video_routes.py) call open_video_task without
stubbing that lock, so they failed closed against the suite's
deliberately-unreachable test Redis (_no_live_redis). Fixed by applying
the same lock-stub pattern tests/unit/services/test_video_engine.py
already uses for its own occasion-lock tests: an autouse HeartbeatMutex
stand-in fixture in test_video_render_loop.py, and wrapping the two
route-level video-request tests in test_video_routes.py with the file's
existing _LOCKED patch pair (already used by every other lock-dependent
test in that file).
The one remaining local failure,
test_cloud_auth.py::test_login_route_parses_oauth2_form_not_query_params,
is a pre-existing environment gap unrelated to this branch: it needs a
real Postgres reachable at localhost:5432 (which .github/workflows/ci.yml
provides as a service container) but this dev sandbox has no such binding
— confirmed unrelated to any of the four merged audit units.
make quality now passes clean: 13729 passed, 0 regressions, 94.49% coverage.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [ddc8121f] regenerate lifecycle artifacts for awaiting_pm_review claim rules and waive_finding intent (#637)
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(sequencing): drop lane-fallback edges that would cycle against analyzer edges
The dev-task collision fallback unioned the analyzer's authoritative edges
with same-assignee lane edges, deduping only the direct pair. A lane chain
through an unsurfaced middle sibling could still contradict an analyzer edge
transitively (the shared-last migration order inverts plain priority order),
closing a 3-cycle that made add_dependency raise ConflictError and wedged
every later delegate to that parent. Fallback edges are now accepted only
when they can't close a cycle against the edges already kept; a regression
test reproduces the exact scenario.
Also strip pre-merge cruft: remove the root CONCURRENCY_AUDIT.md working
report, delete the near-duplicate websocket-reconnect.md doc, fix the stale
a2a/page.tsx doc citation, correct the api-rate-limiting doc to state
idempotency-key retry is unimplemented, and fix two lifecycle.py comments
that referenced a guard function which never existed.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
co-authored by
Frontend Developer 1
Frontend Documenter
Frontend Developer 2
Backend Developer 2
Backend Documenter
Backend Developer 1
Renn F
parent
a1233b2aeb
commit
17de29545a
@@ -0,0 +1,432 @@
|
||||
# Frontend Hooks Reference
|
||||
|
||||
This document covers the reusable React hooks exported from `@/hooks` (principally `panel/src/hooks/use-websocket.ts`), with emphasis on the WebSocket message stream patterns and reconnect handling.
|
||||
|
||||
## Overview
|
||||
|
||||
The panel's real-time coordination streams are built on shared, ref-counted WebSocket connections that fan messages to multiple subscribers. This architecture eliminates duplicate connections to the same endpoint and ensures consistent state across different components that consume the same stream.
|
||||
|
||||
### Connection Architecture
|
||||
|
||||
- **Shared connections**: Two consumers of the same endpoint (e.g., A2A live view + rate-limit banner both reading `/ws/system`) now share a single WebSocket connection instead of each opening their own.
|
||||
- **Message fanning**: The shared connection broadcasts incoming messages to all registered subscribers.
|
||||
- **State syncing**: New subscribers are immediately replayed the connection's current state (e.g., a component mounting mid-reconnection sees `connecting` instead of stale `disconnected`).
|
||||
|
||||
### Message Loss & Reconnect Handling
|
||||
|
||||
The WebSocket connection at `panel/src/lib/websocket/connection.ts` (lines 91–119) has **no server-side message buffering or replay**. When the socket drops and reconnects:
|
||||
|
||||
- Any frame published while disconnected is lost from the WebSocket stream forever.
|
||||
- Specialized hooks like `useNotificationStream` implement **REST catch-up** to fill the gap: they fetch unread notifications at REST the moment the socket recovers, so no message is silently lost.
|
||||
|
||||
This is a point-in-time catch-up strategy, not a byte-for-byte replay — a deliberate trade-off documented in the task acceptance criteria.
|
||||
|
||||
---
|
||||
|
||||
## `useWebSocket<T>(endpoint, queryParams?, enabled?)`
|
||||
|
||||
The foundation hook for subscribing to any WebSocket endpoint. All other hooks (`useNotificationStream`, `useAgentStream`, `useA2ALiveStream`) build on top of it.
|
||||
|
||||
### Parameters
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `endpoint` | string | — | Path after `/ws/`, e.g., `/notifications/{agentId}` or `/system` |
|
||||
| `queryParams` | `Record<string, string>` | `undefined` | Optional query string as an object, e.g., `{ viewer_id: "..." }` |
|
||||
| `enabled` | boolean | `true` | Enable/disable the connection (useful for conditional subscriptions) |
|
||||
|
||||
### Return Value
|
||||
|
||||
```typescript
|
||||
{
|
||||
state: ConnectionState; // "disconnected", "connecting", "reconnecting", "connected"
|
||||
lastMessage: T | null; // The most recent message
|
||||
messages: T[]; // Ring buffer of ≤100 messages (STREAM_MAX_MESSAGES)
|
||||
disconnect: () => void; // Manually tear down the connection
|
||||
clearMessages: () => void; // Clear the buffer
|
||||
isConnected: boolean; // Shorthand for state === "connected"
|
||||
isConnecting: boolean; // Shorthand for state === "connecting" || "reconnecting"
|
||||
}
|
||||
```
|
||||
|
||||
### Example: Raw WebSocket Consumption
|
||||
|
||||
```tsx
|
||||
import { useWebSocket } from "@/hooks";
|
||||
|
||||
export function MyAgentMonitor({ agentId }: { agentId: string }) {
|
||||
const { state, lastMessage, messages, isConnected } = useWebSocket(
|
||||
`/agents/${agentId}`,
|
||||
{ viewer_id: CEO_AGENT_ID },
|
||||
!!agentId // disable if agentId is falsy
|
||||
);
|
||||
|
||||
return (
|
||||
<>
|
||||
<p>Connection: {state}</p>
|
||||
{isConnected && <p>Last update: {lastMessage?.timestamp}</p>}
|
||||
<ul>
|
||||
{messages.map((msg, i) => (
|
||||
<li key={i}>{JSON.stringify(msg)}</li>
|
||||
))}
|
||||
</ul>
|
||||
</>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `useNotificationStream()`
|
||||
|
||||
Subscribes to **CEO notifications** via `/ws/notifications/{CEO_AGENT_ID}`. This is the only subscription actively using the REST catch-up fallback to guarantee no notification is lost during a reconnect.
|
||||
|
||||
### Return Value
|
||||
|
||||
```typescript
|
||||
{
|
||||
state: ConnectionState;
|
||||
lastMessage: NotificationMessage | null;
|
||||
notifications: NotificationMessage[]; // Deduped by notification_id
|
||||
allMessages: NotificationMessage[]; // All messages from WS (raw)
|
||||
clearMessages: () => void; // Clears both notifications AND cached REST catch-up batch
|
||||
isConnected: boolean;
|
||||
isConnecting: boolean;
|
||||
}
|
||||
```
|
||||
|
||||
### Reconnect Behavior (Key Change)
|
||||
|
||||
When the WebSocket reconnects (transitions from `disconnected` / `reconnecting` → `connected`):
|
||||
|
||||
1. **On initial connect**: No REST fetch occurs.
|
||||
2. **On a real reconnect** (socket was up before, dropped, and recovered):
|
||||
- A `GET /api/notifications?unread_only=true` fetch fires immediately.
|
||||
- Unread notifications from this fetch are transformed into notification frames and held in local state.
|
||||
- These cached frames are merged with live frames **before** dedup.
|
||||
- The dedup logic ensures no notification appears twice (caught-up notification wins if also delivered live).
|
||||
|
||||
3. **Fetch failure**: Silently tolerated — live WS delivery resumes regardless. The catchup is best-effort.
|
||||
|
||||
### Deduplication Strategy
|
||||
|
||||
The `notifications` array is deduped by `notification_id` using a newest→oldest walk:
|
||||
|
||||
- Walk backward through `[...cachedCatchup, ...liveMessages]`.
|
||||
- Keep only the first occurrence of each unique `notification_id` (newest wins).
|
||||
- Restore arrival order.
|
||||
- Notifications without an id (older frames) are always kept.
|
||||
|
||||
The cached catch-up batch is placed **ahead** of live frames so live delivery can never be shadowed by an older cached copy.
|
||||
|
||||
### Clearing Notifications
|
||||
|
||||
Calling `clearMessages()` clears:
|
||||
- The live message buffer.
|
||||
- The cached catch-up batch.
|
||||
|
||||
This prevents an immediate repopulation from cached state after a user clears the notification badge.
|
||||
|
||||
### Example
|
||||
|
||||
```tsx
|
||||
import { useNotificationStream } from "@/hooks";
|
||||
|
||||
export function NotificationBell() {
|
||||
const { notifications, isConnected, clearMessages } = useNotificationStream();
|
||||
|
||||
return (
|
||||
<>
|
||||
<button onClick={clearMessages}>
|
||||
Clear ({notifications.length})
|
||||
</button>
|
||||
{!isConnected && <span className="dot" title="offline" />}
|
||||
</>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `useAgentStream(agentId)`
|
||||
|
||||
Subscribes to live agent output (streaming work-in-progress) via `/ws/agents/{agentId}`.
|
||||
|
||||
### Parameters
|
||||
|
||||
| Parameter | Type | Description |
|
||||
|-----------|------|-------------|
|
||||
| `agentId` | string \| null | The agent UUID. Pass `null` to disable. |
|
||||
|
||||
### Return Value
|
||||
|
||||
```typescript
|
||||
{
|
||||
state: ConnectionState;
|
||||
lastMessage: AgentStreamMessage | null;
|
||||
messages: AgentStreamMessage[]; // Raw frames
|
||||
streamChunks: string[]; // Extracted chunk strings
|
||||
streamOutput: string; // All chunks concatenated
|
||||
clearMessages: () => void;
|
||||
isConnected: boolean;
|
||||
isConnecting: boolean;
|
||||
}
|
||||
```
|
||||
|
||||
### Message Type
|
||||
|
||||
```typescript
|
||||
interface AgentStreamMessage {
|
||||
type: "connected" | "agent.stream";
|
||||
agent_id?: string;
|
||||
chunk?: string; // Output fragment
|
||||
watcher_count?: number; // Live viewer count
|
||||
timestamp?: string;
|
||||
}
|
||||
```
|
||||
|
||||
### Example
|
||||
|
||||
```tsx
|
||||
import { useAgentStream } from "@/hooks";
|
||||
|
||||
export function AgentOutputPanel({ agentId }: { agentId: string }) {
|
||||
const { streamOutput, isConnecting, messages } = useAgentStream(agentId);
|
||||
|
||||
return (
|
||||
<div className="output">
|
||||
{isConnecting && <em>Connecting…</em>}
|
||||
<code>{streamOutput}</code>
|
||||
<p className="meta">
|
||||
{messages.length} frames • {streamOutput.length} chars
|
||||
</p>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `useA2ALiveStream()`
|
||||
|
||||
Subscribes to live **agent-to-agent** (A2A) messages via `/ws/system`. The system stream also carries rate-limit and usage events; this hook filters to `a2a.message` frames only.
|
||||
|
||||
### Return Value
|
||||
|
||||
```typescript
|
||||
{
|
||||
state: ConnectionState;
|
||||
lastMessage: A2ASystemMessage | null;
|
||||
a2aMessages: A2ASystemMessage[]; // Filtered to type === "a2a.message"
|
||||
allMessages: A2ASystemMessage[]; // All frames (includes usage/rate-limit)
|
||||
clearMessages: () => void;
|
||||
isConnected: boolean;
|
||||
isConnecting: boolean;
|
||||
}
|
||||
```
|
||||
|
||||
### Message Type
|
||||
|
||||
```typescript
|
||||
interface A2ASystemMessage {
|
||||
type: "connected" | "a2a.message";
|
||||
conversation_id?: string;
|
||||
message_id?: string;
|
||||
task_id?: string;
|
||||
from_agent?: string;
|
||||
to_agent?: string;
|
||||
skill?: string | null;
|
||||
body_excerpt?: string; // Capped; fetch full body via REST
|
||||
timestamp?: string;
|
||||
}
|
||||
```
|
||||
|
||||
### Important Note: Excerpt-Only Delivery
|
||||
|
||||
A2A frames carry a **capped excerpt**, not the full message body. Consumers must:
|
||||
|
||||
1. Listen to `a2a.message` frames.
|
||||
2. Invalidate their **REST query** for the full message (e.g., `GET /api/a2a/conversations/{id}`).
|
||||
3. Fetch the full body from the REST endpoint.
|
||||
|
||||
Do not render the `body_excerpt` as the message — it is metadata only.
|
||||
|
||||
### Reconnect Fallback
|
||||
|
||||
The A2A hook already has a working reconnect fallback at the consumer level: `panel/src/components/a2a/a2a-view.tsx` invalidates the entire a2a query family both per-frame (lines 185–194, on every `a2a.message` frame) and on the disconnected→connected edge (lines 212–218, gated on a `prevConnected` ref so it never fires on initial mount), ensuring no message is lost. No change was needed for this hook.
|
||||
|
||||
### Example
|
||||
|
||||
```tsx
|
||||
import { useA2ALiveStream } from "@/hooks";
|
||||
import { useQuery } from "@tanstack/react-query";
|
||||
|
||||
export function A2ALiveView() {
|
||||
const { a2aMessages, isConnected } = useA2ALiveStream();
|
||||
const { data: conversations } = useQuery({
|
||||
queryKey: ["a2a", "conversations"],
|
||||
// Auto-fetched when a2aMessages changes (invalidation on frame)
|
||||
});
|
||||
|
||||
return (
|
||||
<div>
|
||||
<p>
|
||||
{isConnected ? "Connected" : "Offline"}
|
||||
{a2aMessages.length > 0 && " (live updates)"}
|
||||
</p>
|
||||
{/* Render conversations with full bodies from REST */}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `useConnectionStatus()`
|
||||
|
||||
Tracks the connection state of **all active subscriptions** in a single hook. Useful for global connection indicators.
|
||||
|
||||
### Return Value
|
||||
|
||||
```typescript
|
||||
{
|
||||
connections: Record<string, ConnectionState>; // { [endpoint]: state }
|
||||
updateConnection: (id: string, state: ConnectionState) => void;
|
||||
removeConnection: (id: string) => void;
|
||||
hasActiveConnections: boolean; // Any connection is up/ing
|
||||
allConnected: boolean; // Every connection is connected
|
||||
}
|
||||
```
|
||||
|
||||
### Example: Global Status Indicator
|
||||
|
||||
```tsx
|
||||
import { useConnectionStatus } from "@/hooks";
|
||||
|
||||
export function GlobalConnectionStatus() {
|
||||
const { hasActiveConnections, allConnected } = useConnectionStatus();
|
||||
|
||||
return (
|
||||
<div className="status-badge">
|
||||
{allConnected && <span className="icon-check">Connected</span>}
|
||||
{hasActiveConnections && !allConnected && (
|
||||
<span className="icon-sync">Reconnecting…</span>
|
||||
)}
|
||||
{!hasActiveConnections && <span className="icon-offline">Offline</span>}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Testing
|
||||
|
||||
The hooks come with comprehensive test coverage in `panel/src/hooks/__tests__/`:
|
||||
|
||||
- **`use-websocket.test.tsx`**: Core hook mechanics (shared connection, fan-out, state replay).
|
||||
- **`use-notification-stream.test.tsx`**: REST catch-up verification (new; covers the reconnect fallback):
|
||||
- No fetch on initial connect.
|
||||
- Fetch + fold-in on a real reconnect.
|
||||
- Dedup prevents double-counting when the same notification arrives both via catch-up and live.
|
||||
- `clearMessages()` drops the cached catch-up batch.
|
||||
|
||||
Run tests with:
|
||||
|
||||
```bash
|
||||
pnpm test
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Audited Hooks: Reconnect Coverage
|
||||
|
||||
Three hooks were audited for reconnect message-loss risk:
|
||||
|
||||
| Hook | Connection Type | Fallback | Status |
|
||||
|------|-----------------|----------|--------|
|
||||
| `useNotificationStream` | CEO notification stream | REST catch-up (`GET /notifications?unread_only=true`) | ✅ Hardened |
|
||||
| `useA2ALiveStream` | A2A + system stream | REST query invalidation (`a2a-view.tsx`) | ✅ Verified |
|
||||
| Rate-limit consumers | System stream (`/ws/system`) | REST resync (rate-limit-banner.tsx, usage-overview-panel.tsx) | ✅ Verified |
|
||||
|
||||
No defects were found in `useA2ALiveStream` or rate-limit consumption; the REST polling fallbacks were already in place and tested.
|
||||
|
||||
### Adding a New Reconnect Fallback
|
||||
|
||||
Picking a fallback strategy for a new WS-consuming hook comes down to what kind of data it carries:
|
||||
|
||||
1. **Event data** (can only happen once, e.g. a notification) — implement REST catch-up: fetch unread/pending items over REST the moment the socket recovers, fold them into local state, and dedup against live frames (see `useNotificationStream` above).
|
||||
2. **State data** (always available via REST) — implement query invalidation: invalidate the relevant React Query cache keys on the disconnected→connected edge and let components refetch (see `useA2ALiveStream` above).
|
||||
3. **Always** add a regression test that simulates a disconnect/reconnect cycle with fake timers.
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Always disable on falsy keys**: Pass a conditional `enabled` flag (third param) if your hook depends on a variable parameter. This prevents spurious connections and ensures cleanup.
|
||||
|
||||
```tsx
|
||||
const { messages } = useWebSocket(
|
||||
`/agents/${agentId}`,
|
||||
undefined,
|
||||
!!agentId // disable if agentId is null/undefined
|
||||
);
|
||||
```
|
||||
|
||||
2. **Invalidate REST queries on frame**: When receiving a frame (especially A2A excerpts), trigger a React Query invalidation to fetch fresh data:
|
||||
|
||||
```tsx
|
||||
const queryClient = useQueryClient();
|
||||
useEffect(() => {
|
||||
if (a2aMessages.length > 0) {
|
||||
queryClient.invalidateQueryData({ queryKey: ["a2a"] });
|
||||
}
|
||||
}, [a2aMessages, queryClient]);
|
||||
```
|
||||
|
||||
3. **Don't hold onto stale messages**: The message buffer is bounded to 100 frames. Don't assume it's a complete history — treat it as a stream.
|
||||
|
||||
4. **Catch fetch failures gracefully**: All REST fallbacks (like the notification catch-up) are best-effort and fail silently. Live WS delivery is not guaranteed to block on fetch completion.
|
||||
|
||||
5. **Use `isConnecting` for UI feedback**: Show a loading state when `isConnecting` is true, not just when `!isConnected`.
|
||||
|
||||
```tsx
|
||||
{isConnecting && <Spinner />}
|
||||
{isConnected && <CheckMark />}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Connection Types & States
|
||||
|
||||
### ConnectionState
|
||||
|
||||
```typescript
|
||||
type ConnectionState =
|
||||
| "disconnected" // Not connected; attempting to reconnect or no connection attempt yet
|
||||
| "connecting" // Initial connection attempt
|
||||
| "reconnecting" // Reconnection after a close (watchdog/network failure)
|
||||
| "connected" // Stable connection, messages flowing
|
||||
```
|
||||
|
||||
### State Transitions
|
||||
|
||||
```
|
||||
disconnected ──> connecting ──> connected
|
||||
^
|
||||
│
|
||||
(watchdog fires)
|
||||
│
|
||||
reconnecting ──┘
|
||||
```
|
||||
|
||||
On a reconnect transition (`reconnecting` → `connected`), hooks like `useNotificationStream` fire their REST catch-up fetch.
|
||||
|
||||
---
|
||||
|
||||
## Further Reading
|
||||
|
||||
- **Control panel README**: `panel/README.md`
|
||||
- **WebSocket connection implementation**: `panel/src/lib/websocket/connection.ts`
|
||||
- **API client**: `panel/src/lib/api/`
|
||||
- **Notification types & components**: `panel/src/app/(dashboard)/notifications/`
|
||||
@@ -2,6 +2,7 @@
|
||||
"claim_rules": {
|
||||
"auditor": [],
|
||||
"cell_pm": [
|
||||
"awaiting_pm_review",
|
||||
"needs_revision",
|
||||
"pending"
|
||||
],
|
||||
@@ -16,6 +17,7 @@
|
||||
],
|
||||
"head_marketing": [],
|
||||
"main_pm": [
|
||||
"awaiting_pm_review",
|
||||
"needs_revision",
|
||||
"pending"
|
||||
],
|
||||
@@ -529,6 +531,15 @@
|
||||
"source": "awaiting_pm_review",
|
||||
"target": "cancelled"
|
||||
},
|
||||
{
|
||||
"action": "claim",
|
||||
"roles": [
|
||||
"cell_pm",
|
||||
"main_pm"
|
||||
],
|
||||
"source": "awaiting_pm_review",
|
||||
"target": "claimed"
|
||||
},
|
||||
{
|
||||
"action": "complete",
|
||||
"roles": [
|
||||
|
||||
+160
@@ -0,0 +1,160 @@
|
||||
import { describe, it, expect, vi, beforeEach, afterEach } from "vitest";
|
||||
import { fireEvent, render, screen, waitFor } from "@testing-library/react";
|
||||
import { QueryClient, QueryClientProvider } from "@tanstack/react-query";
|
||||
import type { ReactNode } from "react";
|
||||
import type { ReleaseProposal } from "@/lib/api/release";
|
||||
import { PageRefreshProvider } from "@/components/providers";
|
||||
|
||||
// Unlike release-proposal-card.test.tsx (which stubs useMutation entirely to
|
||||
// test the query-failure/execute-status surfacing paths), these tests
|
||||
// exercise the REAL useMutation onSuccess handler — mirrors the
|
||||
// x-post-queue/video-post-queue/roadmap-review-queue test pattern — so
|
||||
// approve()'s per-status toast copy is actually asserted.
|
||||
const { resolveApproveRef } = vi.hoisted(() => ({
|
||||
resolveApproveRef: { current: null as null | ((v: unknown) => void) },
|
||||
}));
|
||||
|
||||
const { getProposal, approve, reject } = vi.hoisted(() => ({
|
||||
getProposal: vi.fn(
|
||||
async (): Promise<ReleaseProposal> => ({
|
||||
task_id: "t1",
|
||||
title: "Cut v0.14.0",
|
||||
status: "awaiting_ceo_approval",
|
||||
required_changes: null,
|
||||
report: {
|
||||
proposed_version: "0.14.0",
|
||||
bump_kind: "minor",
|
||||
change_summary: ["feat: metrics"],
|
||||
drafted_changelog: "## 0.14.0\n- metrics",
|
||||
version_bump_plan: ["pyproject.toml"],
|
||||
gaps: [],
|
||||
migration_notes: [],
|
||||
gate_state: "green",
|
||||
},
|
||||
}),
|
||||
),
|
||||
// Deferred so the test can freeze the approve mid-flight.
|
||||
approve: vi.fn(
|
||||
() =>
|
||||
new Promise((r) => {
|
||||
resolveApproveRef.current = r as (v: unknown) => void;
|
||||
}),
|
||||
),
|
||||
reject: vi.fn(async () => ({})),
|
||||
}));
|
||||
|
||||
vi.mock("@/lib/api", () => ({
|
||||
releaseApi: { getProposal, approve, reject },
|
||||
}));
|
||||
|
||||
const { toast } = vi.hoisted(() => ({
|
||||
toast: { success: vi.fn(), warning: vi.fn(), info: vi.fn(), error: vi.fn() },
|
||||
}));
|
||||
vi.mock("sonner", () => ({ toast }));
|
||||
|
||||
import { ReleaseProposalCard } from "../release-proposal-card";
|
||||
|
||||
function withProviders(ui: ReactNode) {
|
||||
const client = new QueryClient({
|
||||
defaultOptions: { queries: { retry: false }, mutations: { retry: false } },
|
||||
});
|
||||
return (
|
||||
<QueryClientProvider client={client}>
|
||||
<PageRefreshProvider>{ui}</PageRefreshProvider>
|
||||
</QueryClientProvider>
|
||||
);
|
||||
}
|
||||
|
||||
async function clickApprove() {
|
||||
render(withProviders(<ReleaseProposalCard />));
|
||||
fireEvent.click(
|
||||
await screen.findByRole("button", { name: /Approve & publish/i }),
|
||||
);
|
||||
// Radix marks the rest of the page inert/aria-hidden once the dialog opens
|
||||
// — the main card's own button drops out of the accessible tree, so this
|
||||
// now uniquely matches the dialog's confirm button.
|
||||
await screen.findByRole("dialog");
|
||||
fireEvent.click(
|
||||
await screen.findByRole("button", { name: /Approve & publish/i }),
|
||||
);
|
||||
await waitFor(() => expect(approve).toHaveBeenCalled());
|
||||
}
|
||||
|
||||
describe("ReleaseProposalCard — status feedback (silent-bug-sweep #c)", () => {
|
||||
beforeEach(() => {
|
||||
getProposal.mockClear();
|
||||
approve.mockClear();
|
||||
reject.mockClear();
|
||||
toast.success.mockClear();
|
||||
toast.warning.mockClear();
|
||||
toast.info.mockClear();
|
||||
toast.error.mockClear();
|
||||
resolveApproveRef.current = null;
|
||||
});
|
||||
afterEach(() => {
|
||||
vi.clearAllMocks();
|
||||
});
|
||||
|
||||
it.each([
|
||||
[
|
||||
"already_in_progress",
|
||||
"A release execute is already in progress for this proposal.",
|
||||
],
|
||||
[
|
||||
"redis_unavailable",
|
||||
"Redis is unavailable — can't acquire the release mutex.",
|
||||
],
|
||||
["lock_lost", "The release lock was lost mid-execute — retry the approve."],
|
||||
["gate_failed", "Release halted (gate_failed): the gate is red"],
|
||||
])("shows distinct feedback for the %s status", async (status, message) => {
|
||||
await clickApprove();
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status,
|
||||
version: "0.14.0",
|
||||
files_changed: [],
|
||||
commit_sha: null,
|
||||
release_url: null,
|
||||
detail: "the gate is red",
|
||||
});
|
||||
|
||||
await waitFor(() => expect(toast.warning).toHaveBeenCalledWith(message));
|
||||
expect(toast.success).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("shows a success toast for the published status", async () => {
|
||||
await clickApprove();
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status: "published",
|
||||
version: "0.14.0",
|
||||
files_changed: ["pyproject.toml"],
|
||||
commit_sha: "abc123",
|
||||
release_url: "https://github.com/example/example/releases/tag/v0.14.0",
|
||||
detail: "published",
|
||||
});
|
||||
|
||||
await waitFor(() =>
|
||||
expect(toast.success).toHaveBeenCalledWith("Published v0.14.0"),
|
||||
);
|
||||
});
|
||||
|
||||
it("shows an info toast for the accepted (background-dispatched) status", async () => {
|
||||
await clickApprove();
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status: "accepted",
|
||||
version: "0.14.0",
|
||||
files_changed: [],
|
||||
commit_sha: null,
|
||||
release_url: null,
|
||||
detail: "dispatched",
|
||||
});
|
||||
|
||||
await waitFor(() =>
|
||||
expect(toast.info).toHaveBeenCalledWith(
|
||||
"Release execute dispatched — running in the background. This card updates as it progresses.",
|
||||
),
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -58,6 +58,11 @@ vi.mock("@/lib/api", () => ({
|
||||
roadmapApi: { listCycles, approveItem, rejectItem },
|
||||
}));
|
||||
|
||||
const { toast } = vi.hoisted(() => ({
|
||||
toast: { success: vi.fn(), warning: vi.fn(), error: vi.fn() },
|
||||
}));
|
||||
vi.mock("sonner", () => ({ toast }));
|
||||
|
||||
import { RoadmapReviewQueue } from "../roadmap-review-queue";
|
||||
|
||||
function withQueryClient(ui: ReactNode) {
|
||||
@@ -72,6 +77,9 @@ describe("RoadmapReviewQueue", () => {
|
||||
listCycles.mockClear();
|
||||
approveItem.mockClear();
|
||||
rejectItem.mockClear();
|
||||
toast.success.mockClear();
|
||||
toast.warning.mockClear();
|
||||
toast.error.mockClear();
|
||||
resolveApproveRef.current = null;
|
||||
});
|
||||
afterEach(() => {
|
||||
@@ -140,4 +148,48 @@ describe("RoadmapReviewQueue", () => {
|
||||
await waitFor(() => expect(listCycles).toHaveBeenCalled());
|
||||
expect(container).toBeEmptyDOMElement();
|
||||
});
|
||||
|
||||
// F(silent-bug-sweep #c): RoadmapService uses its own disjoint status
|
||||
// vocabulary (approved/already_approved/invalid_state/rejected/
|
||||
// already_rejected) — none of the 7 named cross-queue statuses
|
||||
// (already_in_progress, redis_unavailable, lock_lost, post_failed,
|
||||
// posted_partial, no_platforms, no_credentials) ever reach this queue, so
|
||||
// this locks in distinct feedback for roadmap's own statuses instead.
|
||||
it.each([
|
||||
[
|
||||
"already_approved",
|
||||
"this item was already approved",
|
||||
"Item approved — added to the backlog",
|
||||
],
|
||||
[
|
||||
"invalid_state",
|
||||
"item is 'rejected', not proposed — cannot approve",
|
||||
"item is 'rejected', not proposed — cannot approve",
|
||||
],
|
||||
])(
|
||||
"shows distinct feedback for the %s status",
|
||||
async (status, detail, message) => {
|
||||
render(withQueryClient(<RoadmapReviewQueue />));
|
||||
const approveButtons = await screen.findAllByRole("button", {
|
||||
name: /Approve/,
|
||||
});
|
||||
fireEvent.click(approveButtons[0]);
|
||||
await waitFor(() => expect(approveItem).toHaveBeenCalled());
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status,
|
||||
item_id: "item-0",
|
||||
materialized_task_id: null,
|
||||
detail,
|
||||
});
|
||||
|
||||
await waitFor(() => {
|
||||
if (status === "already_approved") {
|
||||
expect(toast.success).toHaveBeenCalledWith(message);
|
||||
} else {
|
||||
expect(toast.warning).toHaveBeenCalledWith(message);
|
||||
}
|
||||
});
|
||||
},
|
||||
);
|
||||
});
|
||||
|
||||
@@ -102,6 +102,11 @@ vi.mock("@/components/projects/project-selector", () => ({
|
||||
),
|
||||
}));
|
||||
|
||||
const { toast } = vi.hoisted(() => ({
|
||||
toast: { success: vi.fn(), warning: vi.fn(), error: vi.fn() },
|
||||
}));
|
||||
vi.mock("sonner", () => ({ toast }));
|
||||
|
||||
import { VideoPostQueue } from "../video-post-queue";
|
||||
|
||||
function withQueryClient(ui: ReactNode) {
|
||||
@@ -119,6 +124,9 @@ describe("VideoPostQueue", () => {
|
||||
reject.mockClear();
|
||||
requestVideo.mockClear();
|
||||
getMediaBlob.mockClear();
|
||||
toast.success.mockClear();
|
||||
toast.warning.mockClear();
|
||||
toast.error.mockClear();
|
||||
resolveApproveRef.current = null;
|
||||
// jsdom has no Blob URL implementation. Distinct URLs per call so a
|
||||
// revoke can be asserted against the specific (stale) one it replaced.
|
||||
@@ -558,4 +566,68 @@ describe("VideoPostQueue", () => {
|
||||
await screen.findByText("release");
|
||||
expect(document.querySelector("iframe")).not.toBeInTheDocument();
|
||||
});
|
||||
|
||||
// F(silent-bug-sweep #c): every VideoPostService.approve status must
|
||||
// render a distinct, non-swallowed toast — regression guard, no code gap
|
||||
// was found here (describeExecuteResult already branches every one of
|
||||
// these).
|
||||
it.each([
|
||||
[
|
||||
"posted_partial",
|
||||
{ status: "posted_partial", posted: { x: "1" }, detail: "tiktok: down" },
|
||||
"Posted to some platforms — tiktok: down",
|
||||
],
|
||||
[
|
||||
"post_failed",
|
||||
{ status: "post_failed", posted: {}, detail: "both platforms down" },
|
||||
"Posting failed: both platforms down",
|
||||
],
|
||||
[
|
||||
"already_in_progress",
|
||||
{ status: "already_in_progress", posted: {}, detail: "" },
|
||||
"A post is already in progress for this draft.",
|
||||
],
|
||||
[
|
||||
"no_platforms",
|
||||
{ status: "no_platforms", posted: {}, detail: "" },
|
||||
"This draft has no target platforms.",
|
||||
],
|
||||
[
|
||||
"lock_lost",
|
||||
{ status: "lock_lost", posted: {}, detail: "" },
|
||||
"The post lock was lost mid-upload — retry the approve.",
|
||||
],
|
||||
[
|
||||
"redis_unavailable",
|
||||
{ status: "redis_unavailable", posted: {}, detail: "" },
|
||||
"Redis is unavailable — can't acquire the post lock.",
|
||||
],
|
||||
])("shows distinct feedback for the %s status", async (_, result, message) => {
|
||||
render(withQueryClient(<VideoPostQueue />));
|
||||
await screen.findByText("release");
|
||||
fireEvent.click(screen.getByRole("button", { name: /Approve/ }));
|
||||
await waitFor(() => expect(approve).toHaveBeenCalled());
|
||||
|
||||
resolveApproveRef.current?.(result);
|
||||
|
||||
await waitFor(() => expect(toast.warning).toHaveBeenCalledWith(message));
|
||||
expect(toast.success).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("shows a success toast for the posted status", async () => {
|
||||
render(withQueryClient(<VideoPostQueue />));
|
||||
await screen.findByText("release");
|
||||
fireEvent.click(screen.getByRole("button", { name: /Approve/ }));
|
||||
await waitFor(() => expect(approve).toHaveBeenCalled());
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status: "posted",
|
||||
posted: { x: "1", tiktok: "2" },
|
||||
detail: "posted to all platforms",
|
||||
});
|
||||
|
||||
await waitFor(() =>
|
||||
expect(toast.success).toHaveBeenCalledWith("Posted to all platforms."),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -44,6 +44,11 @@ const { listPosts, approve, reject } = vi.hoisted(() => ({
|
||||
|
||||
vi.mock("@/lib/api", () => ({ xApi: { listPosts, approve, reject } }));
|
||||
|
||||
const { toast } = vi.hoisted(() => ({
|
||||
toast: { success: vi.fn(), warning: vi.fn(), error: vi.fn() },
|
||||
}));
|
||||
vi.mock("sonner", () => ({ toast }));
|
||||
|
||||
import { XPostQueue } from "../x-post-queue";
|
||||
|
||||
function withQueryClient(ui: ReactNode) {
|
||||
@@ -58,6 +63,9 @@ describe("XPostQueue", () => {
|
||||
listPosts.mockClear();
|
||||
approve.mockClear();
|
||||
reject.mockClear();
|
||||
toast.success.mockClear();
|
||||
toast.warning.mockClear();
|
||||
toast.error.mockClear();
|
||||
resolveApproveRef.current = null;
|
||||
});
|
||||
afterEach(() => {
|
||||
@@ -178,4 +186,55 @@ describe("XPostQueue", () => {
|
||||
await waitFor(() => expect(listPosts).toHaveBeenCalled());
|
||||
expect(container).toBeEmptyDOMElement();
|
||||
});
|
||||
|
||||
// F(silent-bug-sweep #c): every XPostService.approve status must render a
|
||||
// distinct, non-swallowed toast — not a blanket success/failure.
|
||||
it.each([
|
||||
["already_in_progress", "A post is already in progress for this draft."],
|
||||
[
|
||||
"no_credentials",
|
||||
"No X credentials configured — set them below first.",
|
||||
],
|
||||
["post_failed", "Posting failed: the X API rejected the tweet"],
|
||||
[
|
||||
"redis_unavailable",
|
||||
"Redis is unavailable — can't acquire the post lock.",
|
||||
],
|
||||
["already_posted", "Already posted — no-op."],
|
||||
])("shows distinct feedback for the %s status", async (status, message) => {
|
||||
render(withQueryClient(<XPostQueue />));
|
||||
const approveButtons = await screen.findAllByRole("button", {
|
||||
name: /Approve/,
|
||||
});
|
||||
fireEvent.click(approveButtons[0]);
|
||||
await waitFor(() => expect(approve).toHaveBeenCalled());
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status,
|
||||
tweet_id: null,
|
||||
detail: "the X API rejected the tweet",
|
||||
});
|
||||
|
||||
await waitFor(() => expect(toast.warning).toHaveBeenCalledWith(message));
|
||||
expect(toast.success).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("shows a success toast for the posted status", async () => {
|
||||
render(withQueryClient(<XPostQueue />));
|
||||
const approveButtons = await screen.findAllByRole("button", {
|
||||
name: /Approve/,
|
||||
});
|
||||
fireEvent.click(approveButtons[0]);
|
||||
await waitFor(() => expect(approve).toHaveBeenCalled());
|
||||
|
||||
resolveApproveRef.current?.({
|
||||
status: "posted",
|
||||
tweet_id: "1",
|
||||
detail: "ok",
|
||||
});
|
||||
|
||||
await waitFor(() =>
|
||||
expect(toast.success).toHaveBeenCalledWith("Posted to X."),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -38,6 +38,20 @@ function gateBadgeVariant(
|
||||
return "secondary";
|
||||
}
|
||||
|
||||
// Distinct copy for the concurrency/infra statuses a release execute can
|
||||
// come back with (already_in_progress / redis_unavailable / lock_lost) —
|
||||
// the rest (gate_failed, ci_failed, ...) fall through to the generic
|
||||
// "Release halted (status): detail" message, itself still status-specific.
|
||||
function describeHaltedStatus(result: ReleaseExecuteResult): string {
|
||||
if (result.status === "already_in_progress")
|
||||
return "A release execute is already in progress for this proposal.";
|
||||
if (result.status === "redis_unavailable")
|
||||
return "Redis is unavailable — can't acquire the release mutex.";
|
||||
if (result.status === "lock_lost")
|
||||
return "The release lock was lost mid-execute — retry the approve.";
|
||||
return `Release halted (${result.status}): ${result.detail}`;
|
||||
}
|
||||
|
||||
// A red gate / open gaps make publishing risky — the CEO should resolve them
|
||||
// first. Approval still runs the fail-closed executor, so it can't ship a bad
|
||||
// release; this only steers the CEO.
|
||||
@@ -86,7 +100,7 @@ export function ReleaseProposalCard({ className }: { className?: string }) {
|
||||
"Release execute dispatched — running in the background. This card updates as it progresses.",
|
||||
);
|
||||
} else {
|
||||
toast.warning(`Release halted (${result.status}): ${result.detail}`);
|
||||
toast.warning(describeHaltedStatus(result));
|
||||
}
|
||||
closeDialog();
|
||||
},
|
||||
|
||||
@@ -77,6 +77,10 @@ function describeExecuteResult(result: XPostExecuteResult): string {
|
||||
return "A post is already in progress for this draft.";
|
||||
if (result.status === "no_credentials")
|
||||
return "No X credentials configured — set them below first.";
|
||||
if (result.status === "post_failed")
|
||||
return `Posting failed: ${result.detail}`;
|
||||
if (result.status === "redis_unavailable")
|
||||
return "Redis is unavailable — can't acquire the post lock.";
|
||||
return `${result.status}: ${result.detail}`;
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,222 @@
|
||||
import { describe, it, expect, vi, beforeEach, afterEach } from "vitest";
|
||||
import { render, act, waitFor } from "@testing-library/react";
|
||||
import { useEffect } from "react";
|
||||
import type {
|
||||
ConnectionState,
|
||||
WebSocketOptions,
|
||||
} from "@/lib/websocket/connection";
|
||||
|
||||
// connection.ts has no message buffering/replay across a reconnect (audited
|
||||
// lines 91-119): drive the WS state machine from the test via a mocked
|
||||
// connection and assert useNotificationStream's REST catch-up covers the gap
|
||||
// instead of silently losing whatever was published while disconnected.
|
||||
const hoisted = vi.hoisted(() => {
|
||||
const instances: MockConnection[] = [];
|
||||
class MockConnection {
|
||||
url: string;
|
||||
onMessage?: (data: unknown) => void;
|
||||
onStateChange?: (state: ConnectionState) => void;
|
||||
constructor(opts: WebSocketOptions) {
|
||||
this.url = opts.url;
|
||||
this.onMessage = opts.onMessage;
|
||||
this.onStateChange = opts.onStateChange;
|
||||
instances.push(this);
|
||||
}
|
||||
connect() {
|
||||
this.onStateChange?.("connecting");
|
||||
this.onStateChange?.("connected");
|
||||
}
|
||||
disconnect() {
|
||||
this.onStateChange?.("disconnected");
|
||||
}
|
||||
getState() {
|
||||
return "connected";
|
||||
}
|
||||
getLastPongAt() {
|
||||
return Date.now();
|
||||
}
|
||||
checkPong() {}
|
||||
}
|
||||
return { instances, MockConnection };
|
||||
});
|
||||
|
||||
vi.mock("@/lib/websocket/connection", () => ({
|
||||
getWebSocketUrl: () => "ws://test/ws",
|
||||
WebSocketConnection: hoisted.MockConnection,
|
||||
}));
|
||||
|
||||
vi.mock("@/lib/constants", () => ({
|
||||
CEO_AGENT_ID: "00000000-0000-0000-0000-000000000001",
|
||||
STREAM_MAX_MESSAGES: 100,
|
||||
}));
|
||||
|
||||
const listMock = vi.fn();
|
||||
vi.mock("@/lib/api/notifications", () => ({
|
||||
notificationsApi: { list: (...args: unknown[]) => listMock(...args) },
|
||||
}));
|
||||
|
||||
import {
|
||||
useNotificationStream,
|
||||
_resetSharedSocketsForTest,
|
||||
} from "../use-websocket";
|
||||
|
||||
const resultRef: {
|
||||
current: ReturnType<typeof useNotificationStream> | null;
|
||||
} = { current: null };
|
||||
|
||||
function Harness() {
|
||||
const stream = useNotificationStream();
|
||||
useEffect(() => {
|
||||
resultRef.current = stream;
|
||||
});
|
||||
return null;
|
||||
}
|
||||
|
||||
function emptyList() {
|
||||
return { items: [], total: 0, unread_count: 0, pending_ack_count: 0 };
|
||||
}
|
||||
|
||||
describe("useNotificationStream — REST catch-up on reconnect", () => {
|
||||
beforeEach(() => {
|
||||
hoisted.instances.length = 0;
|
||||
resultRef.current = null;
|
||||
listMock.mockReset();
|
||||
listMock.mockResolvedValue(emptyList());
|
||||
_resetSharedSocketsForTest();
|
||||
});
|
||||
afterEach(() => {
|
||||
vi.clearAllMocks();
|
||||
_resetSharedSocketsForTest();
|
||||
});
|
||||
|
||||
it("does not fetch a catch-up batch on the initial connect", async () => {
|
||||
render(<Harness />);
|
||||
await act(async () => {});
|
||||
expect(listMock).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("fetches unread notifications on reconnect and folds them in", async () => {
|
||||
listMock.mockResolvedValue({
|
||||
items: [
|
||||
{
|
||||
id: "n1",
|
||||
type: "task_update",
|
||||
priority: "normal",
|
||||
subject: "Missed while offline",
|
||||
timestamp: "2026-07-21T00:00:00Z",
|
||||
},
|
||||
],
|
||||
total: 1,
|
||||
unread_count: 1,
|
||||
pending_ack_count: 0,
|
||||
});
|
||||
|
||||
render(<Harness />);
|
||||
const conn = hoisted.instances[0];
|
||||
await act(async () => {});
|
||||
expect(listMock).not.toHaveBeenCalled();
|
||||
|
||||
// Drop, then recover — the real sequence connection.ts drives on a
|
||||
// watchdog/close event followed by scheduleReconnect().
|
||||
act(() => {
|
||||
conn.onStateChange?.("reconnecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connected");
|
||||
});
|
||||
|
||||
await waitFor(() =>
|
||||
expect(listMock).toHaveBeenCalledWith({ unread_only: true }),
|
||||
);
|
||||
await waitFor(() =>
|
||||
expect(resultRef.current?.notifications).toHaveLength(1),
|
||||
);
|
||||
expect(resultRef.current?.notifications[0].notification_id).toBe("n1");
|
||||
});
|
||||
|
||||
it("does not surface the catch-up copy twice when the same notification also arrives live", async () => {
|
||||
listMock.mockResolvedValue({
|
||||
items: [
|
||||
{
|
||||
id: "n1",
|
||||
type: "task_update",
|
||||
priority: "normal",
|
||||
subject: "Missed",
|
||||
timestamp: "2026-07-21T00:00:00Z",
|
||||
},
|
||||
],
|
||||
total: 1,
|
||||
unread_count: 1,
|
||||
pending_ack_count: 0,
|
||||
});
|
||||
render(<Harness />);
|
||||
const conn = hoisted.instances[0];
|
||||
await act(async () => {});
|
||||
|
||||
act(() => {
|
||||
conn.onStateChange?.("reconnecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connected");
|
||||
});
|
||||
await waitFor(() => expect(listMock).toHaveBeenCalled());
|
||||
await waitFor(() =>
|
||||
expect(resultRef.current?.notifications).toHaveLength(1),
|
||||
);
|
||||
|
||||
// The live frame for the same notification arrives right after reconnect.
|
||||
act(() => {
|
||||
conn.onMessage?.({
|
||||
type: "notification",
|
||||
notification_id: "n1",
|
||||
subject: "Missed",
|
||||
priority: "normal",
|
||||
});
|
||||
});
|
||||
|
||||
expect(resultRef.current?.notifications).toHaveLength(1);
|
||||
});
|
||||
|
||||
it("clearMessages also drops the held catch-up batch (no repopulate-after-clear)", async () => {
|
||||
listMock.mockResolvedValue({
|
||||
items: [
|
||||
{
|
||||
id: "n1",
|
||||
type: "task_update",
|
||||
priority: "normal",
|
||||
subject: "Missed",
|
||||
timestamp: "2026-07-21T00:00:00Z",
|
||||
},
|
||||
],
|
||||
total: 1,
|
||||
unread_count: 1,
|
||||
pending_ack_count: 0,
|
||||
});
|
||||
render(<Harness />);
|
||||
const conn = hoisted.instances[0];
|
||||
await act(async () => {});
|
||||
act(() => {
|
||||
conn.onStateChange?.("reconnecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connecting");
|
||||
});
|
||||
act(() => {
|
||||
conn.onStateChange?.("connected");
|
||||
});
|
||||
await waitFor(() =>
|
||||
expect(resultRef.current?.notifications).toHaveLength(1),
|
||||
);
|
||||
|
||||
act(() => {
|
||||
resultRef.current?.clearMessages();
|
||||
});
|
||||
expect(resultRef.current?.notifications).toHaveLength(0);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,110 @@
|
||||
import { describe, expect, it, vi, beforeEach, afterEach } from "vitest";
|
||||
import { renderHook, waitFor } from "@testing-library/react";
|
||||
import { QueryClient, QueryClientProvider } from "@tanstack/react-query";
|
||||
import { type ReactNode } from "react";
|
||||
|
||||
// Regression coverage for the data-hook null-guard audit: undefined/empty ids
|
||||
// must never reach the API (the `enabled` guard), and useTask's board-review
|
||||
// poll must stop the moment board_review_complete flips — the two things
|
||||
// called out as "known suspects" in the audit task. TanStack Query already
|
||||
// tears the poll's internal timer down on unmount (its Observer
|
||||
// unsubscribes), so the real regression risk is the poll condition itself,
|
||||
// not a missing cleanup — this exercises the condition end-to-end with fake
|
||||
// timers rather than re-deriving it in the test.
|
||||
|
||||
const { get, getSubtasks } = vi.hoisted(() => ({
|
||||
get: vi.fn(),
|
||||
getSubtasks: vi.fn(),
|
||||
}));
|
||||
|
||||
vi.mock("@/lib/api/tasks", async () => {
|
||||
const actual =
|
||||
await vi.importActual<typeof import("@/lib/api/tasks")>(
|
||||
"@/lib/api/tasks",
|
||||
);
|
||||
return {
|
||||
...actual,
|
||||
tasksApi: { ...actual.tasksApi, get, getSubtasks },
|
||||
};
|
||||
});
|
||||
|
||||
import { useTask, useSubtasks } from "@/hooks/use-tasks";
|
||||
import { Team } from "@/types";
|
||||
import type { Task } from "@/types";
|
||||
|
||||
function wrapper({ children }: { children: ReactNode }) {
|
||||
const client = new QueryClient({
|
||||
defaultOptions: { queries: { retry: false } },
|
||||
});
|
||||
return <QueryClientProvider client={client}>{children}</QueryClientProvider>;
|
||||
}
|
||||
|
||||
describe("data-hook null-guard audit", () => {
|
||||
beforeEach(() => {
|
||||
get.mockReset();
|
||||
getSubtasks.mockReset();
|
||||
});
|
||||
|
||||
it("useTask never calls the API when taskId is empty", () => {
|
||||
renderHook(() => useTask(""), { wrapper });
|
||||
expect(get).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("useSubtasks never calls the API when parentTaskId is empty", () => {
|
||||
renderHook(() => useSubtasks(""), { wrapper });
|
||||
expect(getSubtasks).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("useTask fetches once a real id is supplied", async () => {
|
||||
get.mockResolvedValue({ id: "t1", team: Team.BACKEND } as Task);
|
||||
renderHook(() => useTask("t1"), { wrapper });
|
||||
await waitFor(() => expect(get).toHaveBeenCalledWith("t1"));
|
||||
});
|
||||
|
||||
describe("board-review poll stops itself", () => {
|
||||
beforeEach(() => {
|
||||
vi.useFakeTimers({ shouldAdvanceTime: true });
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
vi.useRealTimers();
|
||||
});
|
||||
|
||||
it("polls a board task every 4s until board_review_complete flips true", async () => {
|
||||
get.mockResolvedValue({
|
||||
id: "board-task",
|
||||
team: Team.BOARD,
|
||||
board_review_complete: false,
|
||||
} as Task);
|
||||
|
||||
renderHook(() => useTask("board-task"), { wrapper });
|
||||
await vi.waitFor(() => expect(get).toHaveBeenCalledTimes(1));
|
||||
|
||||
// Still incomplete — the poll must fire again after 4s.
|
||||
await vi.advanceTimersByTimeAsync(4000);
|
||||
await vi.waitFor(() => expect(get).toHaveBeenCalledTimes(2));
|
||||
|
||||
// Board finishes reviewing — the next poll response reports completion.
|
||||
get.mockResolvedValue({
|
||||
id: "board-task",
|
||||
team: Team.BOARD,
|
||||
board_review_complete: true,
|
||||
} as Task);
|
||||
await vi.advanceTimersByTimeAsync(4000);
|
||||
await vi.waitFor(() => expect(get).toHaveBeenCalledTimes(3));
|
||||
|
||||
// No further poll should be scheduled once complete.
|
||||
await vi.advanceTimersByTimeAsync(10000);
|
||||
expect(get).toHaveBeenCalledTimes(3);
|
||||
});
|
||||
|
||||
it("never polls a non-board task", async () => {
|
||||
get.mockResolvedValue({ id: "t1", team: Team.BACKEND } as Task);
|
||||
renderHook(() => useTask("t1"), { wrapper });
|
||||
await vi.waitFor(() => expect(get).toHaveBeenCalledTimes(1));
|
||||
|
||||
await vi.advanceTimersByTimeAsync(10000);
|
||||
expect(get).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -57,6 +57,13 @@ vi.mock("@/lib/constants", () => ({
|
||||
STREAM_MAX_MESSAGES: 100,
|
||||
}));
|
||||
|
||||
// use-websocket.ts imports notificationsApi (for useNotificationStream's
|
||||
// reconnect catch-up) which transitively pulls in the API client — stub it so
|
||||
// this file's unrelated hook tests don't need a real API_URL/axios setup.
|
||||
vi.mock("@/lib/api/notifications", () => ({
|
||||
notificationsApi: { list: vi.fn().mockResolvedValue({ items: [] }) },
|
||||
}));
|
||||
|
||||
import { useWebSocket, _resetSharedSocketsForTest } from "../use-websocket";
|
||||
|
||||
interface Frame {
|
||||
|
||||
@@ -6,6 +6,7 @@ import {
|
||||
getWebSocketUrl,
|
||||
} from "@/lib/websocket/connection";
|
||||
import { CEO_AGENT_ID, STREAM_MAX_MESSAGES } from "@/lib/constants";
|
||||
import { notificationsApi } from "@/lib/api/notifications";
|
||||
|
||||
// Re-export ConnectionState type
|
||||
export type { ConnectionState } from "@/lib/websocket/connection";
|
||||
@@ -251,16 +252,57 @@ export function useNotificationStream() {
|
||||
true,
|
||||
);
|
||||
|
||||
// connection.ts has no message buffering/replay: a notification published
|
||||
// while the socket is down (disconnected/reconnecting) is gone from the WS
|
||||
// stream forever, not merely delayed. Catch up over REST the moment the
|
||||
// stream recovers from a real gap (not the initial mount's first connect)
|
||||
// so a replay-suppressed dedup below can't hide something the bell never
|
||||
// actually received.
|
||||
const [catchup, setCatchup] = useState<NotificationMessage[]>([]);
|
||||
const everConnectedRef = useRef(false);
|
||||
const hadGapRef = useRef(false);
|
||||
useEffect(() => {
|
||||
if (state === "reconnecting" || state === "disconnected") {
|
||||
hadGapRef.current = true;
|
||||
return;
|
||||
}
|
||||
if (state !== "connected") return;
|
||||
if (everConnectedRef.current && hadGapRef.current) {
|
||||
hadGapRef.current = false;
|
||||
notificationsApi
|
||||
.list({ unread_only: true })
|
||||
.then((res) => {
|
||||
setCatchup(
|
||||
res.items.map((n) => ({
|
||||
type: "notification" as const,
|
||||
notification_id: n.id,
|
||||
notification_type: n.type,
|
||||
subject: n.subject,
|
||||
priority: n.priority,
|
||||
timestamp: n.timestamp,
|
||||
})),
|
||||
);
|
||||
})
|
||||
.catch(() => {
|
||||
// Best-effort: live WS delivery resumes regardless of catch-up
|
||||
// success, so a fetch failure here isn't fatal.
|
||||
});
|
||||
}
|
||||
everConnectedRef.current = true;
|
||||
}, [state]);
|
||||
|
||||
// Filter to notification events, de-duplicated by notification_id so a
|
||||
// stream replay (e.g. after a websocket reconnect) does not surface — or
|
||||
// count — the same notification twice. Walk newest→oldest keeping the most
|
||||
// recent copy of each id, then restore arrival order. Events without an id
|
||||
// (older payloads) are always kept.
|
||||
// (older payloads) are always kept. The REST catch-up batch is folded in
|
||||
// ahead of the live WS messages so it can't shadow anything delivered live.
|
||||
const notifications = useMemo(() => {
|
||||
const combined = [...catchup, ...messages];
|
||||
const seen = new Set<string>();
|
||||
const deduped: NotificationMessage[] = [];
|
||||
for (let i = messages.length - 1; i >= 0; i--) {
|
||||
const m = messages[i];
|
||||
for (let i = combined.length - 1; i >= 0; i--) {
|
||||
const m = combined[i];
|
||||
if (m.type !== "notification") continue;
|
||||
const id = m.notification_id;
|
||||
if (id) {
|
||||
@@ -271,14 +313,21 @@ export function useNotificationStream() {
|
||||
}
|
||||
deduped.reverse();
|
||||
return deduped;
|
||||
}, [messages]);
|
||||
}, [messages, catchup]);
|
||||
|
||||
// Clear must drop the REST catch-up batch too — otherwise a cleared badge
|
||||
// would immediately repopulate from the still-held catchup state.
|
||||
const clearAll = useCallback(() => {
|
||||
setCatchup([]);
|
||||
clearMessages();
|
||||
}, [clearMessages]);
|
||||
|
||||
return {
|
||||
state,
|
||||
lastMessage,
|
||||
notifications,
|
||||
allMessages: messages,
|
||||
clearMessages,
|
||||
clearMessages: clearAll,
|
||||
isConnected,
|
||||
isConnecting,
|
||||
};
|
||||
|
||||
@@ -30,7 +30,8 @@ vi.mock("@/store/rate-limit-store", () => ({
|
||||
// ---------------------------------------------------------------------------
|
||||
// Import the function under test AFTER mocks are in place
|
||||
// ---------------------------------------------------------------------------
|
||||
import { getErrorMessage, isTgSurfacePath } from "@/lib/api/client";
|
||||
import { getErrorMessage, isTgSurfacePath, isRetrySafe } from "@/lib/api/client";
|
||||
import { AxiosHeaders } from "axios";
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Helpers
|
||||
@@ -226,3 +227,37 @@ describe("isTgSurfacePath — the /tg login-redirect exemption", () => {
|
||||
expect(isTgSurfacePath("/")).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// isRetrySafe — the 429 retry-by-method gate
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe("isRetrySafe — 429 retry gated by HTTP method", () => {
|
||||
function config(method: string, headers = new AxiosHeaders()) {
|
||||
return { method, headers } as never;
|
||||
}
|
||||
|
||||
it("retries GET and PUT unconditionally — the explicit PO safe-method carve-out", () => {
|
||||
expect(isRetrySafe(config("get"))).toBe(true);
|
||||
expect(isRetrySafe(config("GET"))).toBe(true);
|
||||
expect(isRetrySafe(config("put"))).toBe(true);
|
||||
});
|
||||
|
||||
it("does not retry POST/PATCH/DELETE without an idempotency key", () => {
|
||||
expect(isRetrySafe(config("post"))).toBe(false);
|
||||
expect(isRetrySafe(config("patch"))).toBe(false);
|
||||
expect(isRetrySafe(config("delete"))).toBe(false);
|
||||
});
|
||||
|
||||
it("retries POST/PATCH/DELETE when the caller attached an idempotency key", () => {
|
||||
const headers = new AxiosHeaders();
|
||||
headers.set("X-Idempotency-Key", "abc123");
|
||||
expect(isRetrySafe(config("post", headers))).toBe(true);
|
||||
expect(isRetrySafe(config("patch", headers))).toBe(true);
|
||||
expect(isRetrySafe(config("delete", headers))).toBe(true);
|
||||
});
|
||||
|
||||
it("returns false when config is missing", () => {
|
||||
expect(isRetrySafe(undefined)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
+40
-14
@@ -13,6 +13,24 @@ declare module "axios" {
|
||||
|
||||
const RATE_LIMIT_MAX_RETRIES = 3;
|
||||
|
||||
// GET/PUT are safe to auto-retry on a 429 — GET has no side effect and PUT is
|
||||
// a full-resource replace, so replaying it is a no-op past the first apply.
|
||||
// POST/PATCH/DELETE are NOT safe by default (a replayed POST can create a
|
||||
// duplicate task, a replayed DELETE/PATCH can double-apply a partial update)
|
||||
// — they only retry when the caller attached an idempotency key the backend
|
||||
// can dedupe on. Header name matches what a future idempotency-key-bearing
|
||||
// caller would set; no call site sets it yet, so these methods currently
|
||||
// skip the auto-retry entirely rather than risking a duplicate side effect.
|
||||
const RETRY_SAFE_METHODS = new Set(["get", "put"]);
|
||||
const IDEMPOTENCY_KEY_HEADER = "X-Idempotency-Key";
|
||||
|
||||
export function isRetrySafe(config: AxiosError["config"]): boolean {
|
||||
if (!config) return false;
|
||||
const method = (config.method ?? "get").toLowerCase();
|
||||
if (RETRY_SAFE_METHODS.has(method)) return true;
|
||||
return Boolean(config.headers?.has?.(IDEMPOTENCY_KEY_HEADER));
|
||||
}
|
||||
|
||||
// Create axios instance with default config
|
||||
// axios ^1.16.0 audit (2026-07-09): every call in panel/src rides this
|
||||
// browser instance (no proxy option, no maxRedirects/adapter override, no
|
||||
@@ -146,22 +164,30 @@ api.interceptors.response.use(
|
||||
};
|
||||
useRateLimitStore.getState().hitRateLimit(hitEvent);
|
||||
|
||||
// Track retry count; retry the request (after backoff delay) until exhausted, then toast
|
||||
const retryCount = (error.config?._retryCount ?? 0) + 1;
|
||||
if (error.config) {
|
||||
error.config._retryCount = retryCount;
|
||||
if (retryCount < RATE_LIMIT_MAX_RETRIES) {
|
||||
// Wait retryAfterSeconds before retrying — interceptor re-runs on each subsequent 429
|
||||
const delayMs = safeRetryAfter * 1000;
|
||||
return new Promise<void>((resolve) =>
|
||||
setTimeout(resolve, delayMs),
|
||||
).then(() => api(error.config!));
|
||||
if (isRetrySafe(error.config)) {
|
||||
// Track retry count; retry the request (after backoff delay) until exhausted, then toast
|
||||
const retryCount = (error.config?._retryCount ?? 0) + 1;
|
||||
if (error.config) {
|
||||
error.config._retryCount = retryCount;
|
||||
if (retryCount < RATE_LIMIT_MAX_RETRIES) {
|
||||
// Wait retryAfterSeconds before retrying — interceptor re-runs on each subsequent 429
|
||||
const delayMs = safeRetryAfter * 1000;
|
||||
return new Promise<void>((resolve) =>
|
||||
setTimeout(resolve, delayMs),
|
||||
).then(() => api(error.config!));
|
||||
}
|
||||
}
|
||||
// Retries exhausted — notify the user via Sonner toast
|
||||
toast.warning(
|
||||
`Rate limited by ${provider}. The system has paused operations and will resume automatically in ~${safeRetryAfter}s.`,
|
||||
);
|
||||
} else {
|
||||
// A non-idempotent write (POST/PATCH/DELETE) never auto-retries —
|
||||
// replaying it could double-apply the action. Surface it once instead.
|
||||
toast.warning(
|
||||
`Rate limited by ${provider}. This action was not automatically retried to avoid duplicating it — please try again in ~${safeRetryAfter}s.`,
|
||||
);
|
||||
}
|
||||
// Retries exhausted — notify the user via Sonner toast
|
||||
toast.warning(
|
||||
`Rate limited by ${provider}. The system has paused operations and will resume automatically in ~${safeRetryAfter}s.`,
|
||||
);
|
||||
}
|
||||
|
||||
// Log comprehensive error info
|
||||
|
||||
Reference in New Issue
Block a user