Files
roboco/tests/integration/test_metrics_service.py
T
7901ea419e Retire channels/sessions/messages; A2A becomes primary agent comms (#306)
* feat(a2a): deliver latest incoming message preview into the claim briefing

list_unread_a2a now carries last_message_preview (the latest message from the
OTHER agent, never the agent's own reply), fetched via a correlated subquery in
the same query — no N+1 on the per-verb briefing path.

* feat(a2a): read_a2a verb delivers unread message bodies to the agent

A2AService.get_unread_messages returns the caller's unread INCOMING messages
(never its own sends), marking exactly those rows read atomically so a message
arriving mid-call is preserved. Wired as the read_a2a content verb (route +
do_server tool + granted to every delivery role) — the content-bearing read the
A2A inbox lacked (read_messages only zeroed the counter).

* docs(rag): document read_a2a as the A2A content-read path

* fix(task): backlog activation no longer requires a discussion session

Removes the SessionTaskTable gate in activate() (and its dangling log field),
deletes _inherit_parent_session + its create() call, and drops the now-unused
SessionTaskTable import. Coordination rides task state; the session subsystem is
being retired. Tests updated to the new (no-session) behavior.

* fix(orchestrator): drop session sweep from _run_sweep

Removes the messaging import + sweep_timed_out_sessions call. That import sat
outside the try/except, so once messaging.py is deleted it would have killed the
entire sweep cascade (budget kill-switch, token rollups, retention, image prune,
superseded-PR reconcile). Notification sweep + all maintenance sweeps unchanged.

* release-manager --no-tags read-clone fix

* test: update evidence_repo unit test for a2a last_message_preview

* refactor(gateway): drop session propagation on delegate

Removes propagate_sessions_to_subtask from delegate(), the ChoreographerDeps
messaging field + property, and the ChoreographerDeps messaging arg in deps.py
(ContentActions messaging + import stay until the verbs are removed). Deletes the
propagation test; strips the now-invalid messaging kwarg from ChoreographerDeps
test builders.

* refactor(gateway): remove say/open_session/link_session/channels verbs

Removes the four channel/session verbs across content_actions (impls +
ContentActionsDeps.messaging), do_server (tools + registry), role_config (grants
+ _CHANNEL_DISCOVERY), do.py (routes), schemas/v1/do.py (request models), and
deps.py (MessagingService import + construction). Regenerates the prompt verb
tables. dm/notify/read_messages/read_a2a stay. Tests deleted/updated accordingly.

* uv.lock Upgrade

* refactor: remove conversation RAG indexing; Secretary announces via notification

Drops the CONVERSATIONS index (index_conversation, ConversationsIndexPlugin,
IndexType.CONVERSATIONS enum, IndexConversationParams, mentor.py type-label, the
messaging index hook) and its chunk-table manifest entries. The Secretary's
ANNOUNCE/RELAY_MESSAGE now fan out a BROADCAST notification to every agent's
inbox (NotificationService.broadcast) instead of posting to a dead channel.

* fix(panel): label RAG health error lines by subsystem

A red llm_error (e.g. the glm-5.2:cloud weekly-limit 429) rendered under
the 'Embedding: ok' header with no label, reading as an embedding failure.
Prefix each error line with LLM / Embedding / Vector store.

* refactor: remove channel/message reads from metrics, dashboard, git, events

MetricsService drops get_communication_volume + the MessageTable
message-count in get_agent_metrics (and the now-dead messages_sent_week
field). DashboardService drops get_channel_feeds/_compute_channel_status
and the message read in get_recent_activity (task activity kept);
get_auditor_metrics no longer reports communication_volume.
GitService's two primary-session-id helpers always return None now
(callers already treat None as "no primary session"). events/handlers.py
drops the SESSION_CLOSED/SESSION_TIMEOUT subscriptions + the
handle_session_boundary handler.

Forced follow-on: api/routes/dashboard.py + api/schemas/dashboard.py
dropped the now-dangling live_feeds/ChannelFeed surface and the
/metrics/communication route, which wrapped the removed service calls
directly (mypy would otherwise fail on the missing attributes).

* refactor: delete MessagingService + channel seeding

Edited db/__init__.py and services/__init__.py first (drop the unconditional
Channel/Group/Message/Session table + MessagingService re-exports), then
deleted services/messaging.py, then trimmed db/seed.py to only create_agents
(create_channels/create_channel_memberships/create_initial_messages gone).

Forced expansion: api/routes/{channels,groups,sessions,messages}.py import
roboco.services.messaging directly (not through the package __init__), as
does api/routes/tasks.py (the session-links embed on GET /tasks/{id} and the
GET /{id}/sessions route). Deleting messaging.py without addressing these
breaks `import roboco.api.app` immediately, since app.py eagerly imports all
route modules at startup. Since the 4 CRUD route files are 100%
MessagingService-backed with zero independent logic (and are wholesale
deletes in the plan's later API-routes task anyway), deleted them now +
unmounted from app.py/routes/__init__.py; tasks.py got the same surgical
trim its later task already specified (drop session-links embed +
TaskSessionLinkResponse/TaskResponse.sessions). This pulls a slice of that
later work forward — the routes/schemas for channels/groups/sessions/messages
still need their own pass, but their messaging-coupled parts are gone.

Verified with a full-suite collection sweep (12010 tests collected, zero
import errors) beyond the directly touched test dirs, given the expanded
blast radius.

* refactor: remove channel/session/message models, tables, and channel policy

Models: deleted channel.py/group.py/session.py/messaging.py wholesale
(zero external consumers besides the models/__init__.py re-export).
message.py surgically trimmed: removed MessageCreate (dead) and MessageEdit
(never instantiated; ExtractedMessage.edit_history retyped to
list[dict[str, Any]] to match how it's actually persisted — confirmed
ExtractedMessage was never written to any DB table, so MessageTable's
removal carries no functional risk to the kept extraction pipeline).
base.py: removed SessionStatus + ChannelType, kept MessageType. Also
removed the confirmed-dead channels_read/channels_write fields from
models/agent.py:AgentPermissions and models/dashboard.py:ChannelFeedData.

db/tables.py: deleted ChannelTable/GroupTable/SessionTable/SessionTaskTable/
MessageTable, TaskTable.session_links, and JournalEntryTable.session_id —
cascaded through models/journal.py, services/journal.py, and
api/schemas+routes/journals.py (22 plumbing sites).

foundation/policy/communications.py: removed the ChannelSpec/CHANNELS
catalog + TEAM_SCOPED_ROLES/_CELL_*/_AUDITOR_ONLY helpers, kept the
notification policy (Priority/parse_priority/NOTIFY_SENDER_ROLES/
ACK_REQUIRED_BY_TYPE). enforcement/channel_access.py deleted (confirmed
fully dead in production). agents_config.py: removed CHANNEL_ACCESS
(kept A2A_ALLOWED_PAIRS). seeds/initial_data.py: removed
DEFAULT_CHANNELS/CHANNEL_MEMBERSHIPS/AUDITOR_SILENT_ACCESS + the
never-consumed INITIAL_MESSAGES. config.py: removed
session_idle_timeout_seconds (zero consumers). exceptions.py: removed
dead ChannelError/ChannelAccessDeniedError/SessionClosedError.

Forced expansion beyond the original file list — ChannelType cascaded
into a live, mounted surface the plan didn't trace: agents_config.
CHANNEL_ACCESS -> services/permissions.py's channel-RBAC methods (not
models/permissions.py, which turned out to have no channel code at all)
-> two real endpoints in api/routes/stream.py (GET /permissions,
GET /permissions/channel/{name}) and two dependency factories in
api/deps.py. Removed the channel methods + fields, deleted the
channel-specific stream.py endpoint, deleted require_channel_read/write.
Also deleted api/schemas/{channels,sessions}.py (hard dependency on the
removed enums; already fully dead after the Task 10 route deletions) and
api/schemas/messages.py (a TYPE_CHECKING-only import of the deleted
MessageTable; likewise already fully dead) + its dedicated test file.

Test updates: test_permissions.py -14 channel tests (matches the planned
count exactly), test_communications.py / test_communications_consumers.py
split to keep only notification-policy coverage, test_exceptions.py -9,
test_deps.py -4, plus the journal/stream/foundation-smoke fallout. Also
fixed a pre-existing (Task 7) broken assertion in
test_foundation_phase3_smoke.py that inspected a `say()` method already
removed from ContentActions.

Verified: full-suite collection (11961 tests, zero import errors) and a
complete test run (11567 passed, 394 skipped, 0 failed) in addition to
the targeted suites.

* migration: drop channels/groups/sessions/session_tasks/messages + enum types

alembic/versions/060_drop_messaging.py: drop_column journal_entries.
session_id (sidesteps hardcoding the FK constraint name — verified
empirically against a live migrated DB that it's actually
fk_journal_entries_session_id_sessions, but drop_column doesn't care
either way); drop_table in FK order (messages -> session_tasks ->
sessions -> groups -> channels); DROP TABLE IF EXISTS chunks_conversations
(runtime-provisioned, not alembic-managed, would otherwise orphan); DROP
TYPE IF EXISTS for messagetype/sessionstatus/sessionscope/channeltype
(messagetype's Python enum stays for ExtractedMessage, but the DB type
had zero live columns left once MessageTable was dropped in the prior
commit). downgrade() raises NotImplementedError — one-way removal.

Pruned scripts/reset_runtime_state.sql + .sh: removed the DELETE/COUNT
lines for messages/session_tasks/sessions/groups/channels and the
groups.active_session_id reset block.

Verified end-to-end against a scratch Postgres DB: full migration chain
001->060 applies cleanly, alembic heads shows a single head, all 6 dropped
tables + 4 enum types + the journal_entries.session_id column are
confirmed gone, journal_entries keeps only its journal_id/task_id FKs,
downgrade correctly raises NotImplementedError without corrupting DB
state, and the pruned reset_runtime_state.sql runs clean (no errors)
against a fully-migrated DB.

* refactor(api): remove channel/session/message routes + WS streams

Most of this task's file list was already forced through in earlier
commits (routes/{channels,groups,sessions,messages}.py + app.py/__init__.py
unmounting in the MessagingService-deletion commit; tasks.py's
session-links embed + GET /{id}/sessions + schemas/tasks.py's
TaskResponse.sessions in that same commit; deps.py's require_channel_read/
write + schemas/{channels,sessions}.py in the models/tables commit). This
closes out what was left:

- api/websocket.py: deleted the channel_stream + session_stream routes,
  ConnectionManager's channel_connections/session_connections dicts,
  connect_channel/connect_session, broadcast_to_channel/broadcast_to_session,
  get_channel_subscriber_count, and their cleanup lines in disconnect().
  Agent streams, notification streams, and the operator system stream are
  untouched.
- api/websocket_bridge.py: deleted _handle_session_event +
  _handle_message_event and their SESSION_CREATED/SESSION_CLOSED/
  SESSION_TIMEOUT/MESSAGE_SENT subscriptions. The A2A live-view, rate-limit,
  usage, agent-lifecycle, and notification bridges are untouched.
- api/schemas/websocket.py: removed NewMessageBroadcast, WSMessageNew,
  WSMessageEdit, WSMessageDelete, WSSessionClosed — kept the WSMessage base
  class (still subclassed by the kept WSAgentStream/WSNotification) plus
  those two.
- api/schemas/groups.py: deleted (already fully orphaned since routes/
  groups.py was removed; its GroupResponse/GroupDetailResponse had zero
  consumers).

Updated the 5 websocket test files accordingly (removed the channel/
session-specific tests + fixed imports); test_websocket_bridge.py's
registration-coverage test dropped the SESSION_*/MESSAGE_SENT assertions.

Verified: full-suite collection (11943 tests, zero import errors) and a
complete test run (11549 passed, 394 skipped, 0 failed).

* docs: retire channels/sessions/messages from agent-facing docs + CLAUDE.md

Rewrites docs/rag (RAG-indexed) + docs/map + CLAUDE.md to reflect A2A (dm +
read_a2a) as primary agent comms; deletes the channel docs, splits messaging-tools
+ messaging-notification (renamed notification.md), swaps the WS worked example to
A2A_MESSAGE_SENT. _complete_map.md still needs regeneration (generated file).

* refactor(panel): remove Communications surface (channels/sessions)

Deletes the /communications routes, message components, task-detail Sessions tab,
use-channels + channel/session WS hooks, and the channels/sessions/messages/groups
api clients; prunes the Channel/Session/Message/Group types + mock data. (Auditor
live-feeds + dashboard.ts dead-route cleanup is a follow-up.)

* refactor(panel): drop auditor channel-feed + dead communication-metric route

* docs(map): regenerate _complete_map from updated slices

* fix(a2a): reduce get_unread_messages complexity below xenon C + stale comments

Extract the per-conversation unread-counter recompute into _reset_unread_counter
(the CI quality gate flagged get_unread_messages as rank C). Also drop the deleted
open_session from a content_actions comment and reword an evidence_repo docstring
that cited the removed messaging._notify_mentions.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-04 03:10:33 +02:00

349 lines
11 KiB
Python

"""MetricsService coverage — velocity, blockers, team health, agent metrics."""
from __future__ import annotations
from datetime import UTC, datetime, timedelta
from typing import TYPE_CHECKING
from uuid import uuid4
import pytest
import pytest_asyncio
from roboco.db.tables import AgentTable, AuditLogTable, ProjectTable, TaskTable
from roboco.models import AgentRole, AgentStatus, Team
from roboco.models.base import (
TaskNature,
TaskStatus,
TaskType,
)
from roboco.services.metrics import MetricsService
if TYPE_CHECKING:
from collections.abc import AsyncIterator
from sqlalchemy.ext.asyncio import AsyncSession
@pytest_asyncio.fixture
async def metrics_setup(
db_session: AsyncSession,
) -> AsyncIterator[dict]:
agent = AgentTable(
id=uuid4(),
name="Dev",
slug=f"be-dev-{uuid4().hex[:8]}",
role=AgentRole.DEVELOPER,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="dev",
capabilities=[],
permissions={},
metrics={},
)
db_session.add(agent)
await db_session.flush()
project = ProjectTable(
id=uuid4(),
name="M-Proj",
slug=f"m-proj-{uuid4().hex[:8]}",
git_url="https://example.com/r.git",
assigned_cell=Team.BACKEND,
created_by=agent.id,
)
db_session.add(project)
await db_session.flush()
yield {
"svc": MetricsService(db_session),
"agent_id": agent.id,
"project_id": project.id,
"db": db_session,
}
def _task(
setup: dict,
*,
status: TaskStatus,
team: Team = Team.BACKEND,
**extras: object,
) -> TaskTable:
"""Build a TaskTable; extras forwarded (completed_at, started_at, ...)."""
return TaskTable(
id=uuid4(),
title="t",
description="d",
acceptance_criteria=["ac"],
status=status,
priority=2,
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
project_id=setup["project_id"],
created_by=setup["agent_id"],
team=team,
**extras,
)
# ---------------------------------------------------------------------------
# Velocity
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_get_velocity_returns_metrics(metrics_setup: dict) -> None:
"""Velocity returns counts; numbers depend on test ordering."""
svc = metrics_setup["svc"]
velocity = await svc.get_velocity(days=7)
# Counts may be non-zero due to test pollution from session-scoped fixtures
# that commit. We only verify the shape, not the empty count.
assert isinstance(velocity.tasks_completed, int)
assert isinstance(velocity.tasks_created, int)
assert velocity.completion_rate >= 0
@pytest.mark.asyncio
async def test_get_velocity_with_completed_tasks(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
db = metrics_setup["db"]
now = datetime.now(UTC)
started = now - timedelta(hours=2)
db.add(
_task(
metrics_setup,
status=TaskStatus.COMPLETED,
started_at=started,
completed_at=now,
)
)
await db.flush()
velocity = await svc.get_velocity(days=7)
assert velocity.tasks_completed == 1
assert velocity.avg_completion_hours is not None
@pytest.mark.asyncio
async def test_get_velocity_filtered_by_team(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
db = metrics_setup["db"]
now = datetime.now(UTC)
db.add(
_task(
metrics_setup,
status=TaskStatus.COMPLETED,
team=Team.FRONTEND,
completed_at=now,
)
)
await db.flush()
backend_v = await svc.get_velocity(days=7, team=Team.BACKEND)
assert backend_v.tasks_completed == 0
frontend_v = await svc.get_velocity(days=7, team=Team.FRONTEND)
assert frontend_v.tasks_completed == 1
# ---------------------------------------------------------------------------
# Blocker metrics
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_get_blocker_metrics_empty(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
bm = await svc.get_blocker_metrics()
assert bm.active_blockers == 0
@pytest.mark.asyncio
async def test_get_blocker_metrics_with_blocked(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
db = metrics_setup["db"]
db.add(_task(metrics_setup, status=TaskStatus.BLOCKED))
db.add(_task(metrics_setup, status=TaskStatus.BLOCKED, team=Team.FRONTEND))
await db.flush()
bm = await svc.get_blocker_metrics()
_BLOCKED = 2
assert bm.active_blockers == _BLOCKED
assert "backend" in bm.blockers_by_team or "frontend" in bm.blockers_by_team
@pytest.mark.asyncio
async def test_blocked_hours_use_audit_blocked_at_not_updated_at(
metrics_setup: dict,
) -> None:
"""``blocked_since`` is the real ``task.blocked`` transition time from the
audit log, not ``updated_at`` (#67). The old heuristic (``updated_at or
created_at``) over-counted when a blocked task was later touched for a
non-blocking reason (a note, an assignee nudge) — its ``updated_at`` moved
forward and shrank the reported blockage. The audit row marks the actual
entry into BLOCKED; fall back to ``updated_at`` only when no audit row exists.
"""
svc = metrics_setup["svc"]
db = metrics_setup["db"]
now = datetime.now(UTC)
blocked_five_hours_ago = now - timedelta(hours=5)
touched_one_hour_ago = now - timedelta(hours=1)
task = _task(
metrics_setup,
status=TaskStatus.BLOCKED,
created_at=blocked_five_hours_ago,
updated_at=touched_one_hour_ago,
)
db.add(task)
db.add(
AuditLogTable(
event_type="task.blocked",
target_type="task",
target_id=task.id,
severity="info",
details={"from_status": "in_progress", "to_status": "blocked"},
timestamp=blocked_five_hours_ago,
)
)
await db.flush()
bm = await svc.get_blocker_metrics()
_ONE = 1
assert bm.active_blockers == _ONE
assert bm.longest_blocked_task_id == task.id
# ~5h since the audit event, NOT ~1h since the post-block update.
_FIVE_HOURS_MINUS_SKEW = 4.5 # blocked ~5h ago; allow timing skew
assert bm.longest_blocked_hours is not None
assert bm.longest_blocked_hours >= _FIVE_HOURS_MINUS_SKEW
assert bm.avg_blocked_hours is not None
assert bm.avg_blocked_hours >= _FIVE_HOURS_MINUS_SKEW
@pytest.mark.asyncio
async def test_blocked_hours_fall_back_to_updated_at_without_audit_row(
metrics_setup: dict,
) -> None:
"""No ``task.blocked`` audit row -> fall back to ``updated_at or created_at``
(the legacy heuristic), so a blocked task with no audit trail still reports a
value rather than None."""
svc = metrics_setup["svc"]
db = metrics_setup["db"]
now = datetime.now(UTC)
db.add(
_task(
metrics_setup,
status=TaskStatus.BLOCKED,
created_at=now - timedelta(hours=2),
updated_at=now - timedelta(hours=2),
)
)
await db.flush()
bm = await svc.get_blocker_metrics()
assert bm.active_blockers >= 1
assert bm.avg_blocked_hours is not None
assert bm.avg_blocked_hours >= 1.0
# ---------------------------------------------------------------------------
# Team metrics
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_get_team_metrics_empty(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
tm = await svc.get_team_metrics(Team.BACKEND)
assert tm.team == Team.BACKEND
assert tm.active_tasks == 0
assert tm.documentation_coverage == 0
@pytest.mark.asyncio
async def test_get_team_metrics_with_data(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
db = metrics_setup["db"]
now = datetime.now(UTC)
db.add(_task(metrics_setup, status=TaskStatus.IN_PROGRESS))
db.add(
_task(
metrics_setup,
status=TaskStatus.COMPLETED,
started_at=now - timedelta(hours=2),
completed_at=now,
dev_notes="some notes",
)
)
db.add(_task(metrics_setup, status=TaskStatus.BLOCKED))
await db.flush()
tm = await svc.get_team_metrics(Team.BACKEND)
assert tm.active_tasks == 1
assert tm.completed_tasks_week == 1
assert tm.blocked_tasks == 1
@pytest.mark.asyncio
async def test_get_all_team_metrics(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
_TEAMS = 3
rows = await svc.get_all_team_metrics()
assert len(rows) == _TEAMS # backend, frontend, ux_ui
assert {r.team for r in rows} == {Team.BACKEND, Team.FRONTEND, Team.UX_UI}
# ---------------------------------------------------------------------------
# Agent metrics
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_get_agent_metrics_for_known_agent(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
am = await svc.get_agent_metrics(metrics_setup["agent_id"])
assert am is not None
assert am.agent_id == metrics_setup["agent_id"]
@pytest.mark.asyncio
async def test_get_agent_metrics_returns_none_for_unknown(
metrics_setup: dict,
) -> None:
svc = metrics_setup["svc"]
assert await svc.get_agent_metrics(uuid4()) is None
# ---------------------------------------------------------------------------
# Health status
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_get_health_status_empty(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
h = await svc.get_health_status(team=Team.BACKEND)
assert h["status"] == "ok"
assert h["team"] == "backend"
@pytest.mark.asyncio
async def test_get_health_status_critical_when_majority_blocked(
metrics_setup: dict,
) -> None:
svc = metrics_setup["svc"]
db = metrics_setup["db"]
# 4 blocked + 1 in_progress → 80% blocked → critical
for _ in range(4):
db.add(_task(metrics_setup, status=TaskStatus.BLOCKED))
db.add(_task(metrics_setup, status=TaskStatus.IN_PROGRESS))
await db.flush()
h = await svc.get_health_status(team=Team.BACKEND)
assert h["status"] == "critical"
@pytest.mark.asyncio
async def test_get_health_status_org_wide(metrics_setup: dict) -> None:
svc = metrics_setup["svc"]
h = await svc.get_health_status(team=None)
assert h["team"] == "all"
def test_determine_health_status_directly() -> None:
"""Cover the threshold-decision helper without DB."""
svc = MetricsService.__new__(MetricsService) # No DB needed.
assert svc._determine_health_status(0.5, 10, 0) == "critical"
assert svc._determine_health_status(0.2, 10, 5) == "slow"
assert svc._determine_health_status(0.0, 10, 0) == "slow"
assert svc._determine_health_status(0.0, 1, 5) == "ok"