Files
roboco/tests/unit/services/test_rate_limit_tracker.py
T
303c2db289 Fix: rate limit real probe (#110)
* fix(rate-limit): real provider liveness probe instead of time-based stub

The rate-limit recovery sweeper cleared a provider and resumed parked agents
purely on elapsed time — _do_probe was a stub that always returned True once
the retry_after window passed, so it never confirmed the provider had actually
stopped rate-limiting us. Under a sustained limit that resumes agents straight
into another 429, re-parking them: avoidable churn.

Make the probe real. _do_probe now issues a free, unmetered liveness call —
Anthropic GET /v1/models or Ollama GET /api/tags — and treats any non-429
response as the limit having lifted. A 429 keeps the provider parked; a
network error keeps it parked too (retry next sweep). When the provider can't
be probed (no API key, or an unrecognized provider), it falls back to the
prior time-expiry optimism rather than stranding agents. _probe_target keeps
URL/header resolution separate and testable, and _do_probe stays a
monkeypatchable boundary so the existing sweep tests are unaffected.

Also drop two acceptance-criteria-number labels from comments in this file.

* chore(rate-limit): clear merged gate debt in rate-limit tests + deps lint

The rate-limit PR landed with ruff violations the full gate flags but the
authors' runs missed: test_rate_limit_sweep.py was unformatted, and
test_rate_limit_tracker.py had unsorted/unused imports and magic-value
comparisons. Format the sweep test, drop the dead imports, and bind the
magic comparison values to locals. Also strip acceptance-criteria-number
labels from comments/docstrings across the three rate-limit test files
(leaving genuine acceptance_criteria=[...] test data untouched), and add
api/deps.py to the PLC0415 per-file-ignore — it is the DI wiring hub and
defers a couple of service imports to call time to avoid import cycles,
the same rationale already applied to api/routes, runtime, and services.

* fix(rate-limit): resolve redis type errors in RateLimitStateTracker

A cold mypy run (the gate's true state — prior passes were warm-cache only)
flagged four redis-typing errors in rate_limit_tracker.py that the merge
missed: three unused type:ignore[type-arg] on redis.Redis, and an
aclose() the bundled redis type stub doesn't expose.

Drop the now-unused ignores, and close the scan client via
'async with redis.from_url(...) as r:' instead of a finally-block
aclose(). The context manager closes the client on exit using the modern
redis.asyncio API — no deprecated close(), no stub-missing aclose(), no
suppression. Extend the test's redis mock to model the async
context-manager protocol so it returns itself on enter.

* test(prompter): pass route='main_pm' in the product main-PM routing test

Pre-existing master failure, unrelated to the rate-limit work. The test is
named ...product_routes_to_main_pm and asserts team=MAIN_PM, but called
confirm_live_draft without a route, so it got the 'board' default — which
assigns the Product Owner and yields team=BOARD by design (the board-review
path keeps the root at team=board until the CEO approves). The Main-PM path
is selected with route='main_pm', exactly as the sibling
...main_pm_route_assigns_main_pm test does. Add the missing kwarg so the test
verifies the path it names; behaviour under test is unchanged.

* Updated uv.lock

* refactor(complexity): bring all rank-C blocks under the xenon B ceiling

The full quality gate's xenon step (--max-absolute B --max-modules A
--max-average A) failed on eight rank-C blocks plus the extraction module
average — debt the rate-limit and token-analytics merges deferred. Reduce
each by extracting cohesive helpers, behaviour unchanged:

- orchestrator._probe_one_provider: split into _too_early_to_probe,
  _on_probe_success, _on_probe_failure, _parked_agents_for.
- rate_limit_tracker.list_rate_limited_providers: extract _read_rate_limited_entry
  and a _decode helper.
- trigger_filter.decide_spawn: extract _stale_trigger_decision (drops the
  PLR0911 suppression too).
- ollama_embedder (embed_query, _embed_batch_sync, aembed_query,
  _embed_batch_async): share _rl_backoff / _map_embed_error / _log_429 /
  _sleep_connect_retry / _asleep_connect_retry; remove a dead post-loop guard
  in aembed_query.
- mentor._synthesize_answer: extract _select_system_prompt and
  _answer_from_response.
- indexes/base.ask: extract the 429-retried LLM call into _ask_llm.
- extraction.__init__: extract _compile_patterns so the module average
  lands at rank A.

xenon now exits 0; rate-limit, optimal_brain, extraction, and events suites
all green.

* chore(deps): drop obsolete types-redis stub; honor redis 8.0 inline types

types-redis 4.6 (typed for redis 4.x) shadowed redis 8.0's own inline types,
which both masked real annotation mismatches in stream_bus.py and forced
awkward workarounds elsewhere. The stale stub is why the mypy gate only ever
passed warm-cached: a cold run under the wrong stub disagreed with the code.

Remove types-redis (and its orphaned transitive stubs) so mypy uses redis's
shipped types. That surfaces that xreadgroup/xclaim return bytes-keyed records
while _handle_message is annotated str — the code already decodes bytes
defensively, so this is an annotation gap, not a runtime bug. Make the types
honest: cast each result to its concrete shape and decode the stream name and
message id to str at the dispatch boundary via a _to_str helper.

mypy roboco/ is now clean cold (247 files) against redis's real types; events
suite green.

* Updated uv.lock

* fix(workspace): install the dev extra so agents can run make quality

Agent workspaces were set up with plain `uv sync`, which installs only the
project's default dependency group (pytest) — not the `dev` *extra* where the
gate tools live (ruff, mypy, xenon, radon, vulture, bandit, deptry). So an
agent's .venv had pytest but no linters, and `make quality` died immediately
on `ruff: command not found`. Agents literally could not lint, type-check, or
complexity-check their own work, which is how format/mypy/xenon debt merged
unseen. Sync the `dev` extra (`uv sync --extra dev`) so the workspace gets the
full toolchain the setup's own docstring already promised.

* fix(panel): rate-limit endpoint shape + websocket path

Two panel-facing breakages from the rate-limit rework:

- GET /api/system/rate-limits returned a raw list, but the panel store reads
  response.entries — so `r.entries is not iterable` crashed the banner sync on
  page load. Return the panel's contract: a { entries: [...] } envelope whose
  items are camelCase {provider, affectedAgents, hitAt, resumeAt,
  retryAfterSeconds}, derived from the raw Redis state (resumeAt = hitAt +
  retryAfter).
- The rate-limit websocket hook passed "/ws/system" while getWebSocketUrl()
  already supplies the "/ws" base, producing the doubled "/ws/ws/system" URL.
  Pass "/system" to match the agents/channels/notifications hooks.

Note: the backend /ws/system endpoint itself does not yet exist (the rework
shipped the panel hook only); the REST fix keeps the banner correct on load
and reconnect until that endpoint is built.

* test(workspace): assert uv sync installs the dev extra

Follow the workspace setup change: the dependency-install command is now
`uv sync --extra dev` so the agent workspace gets the lint/type/complexity
toolchain. Update the three assertions that pinned the old `uv sync`.

* feat(ws): add /ws/system stream and bridge rate-limit events to the panel

The rate-limit rework shipped the panel's websocket hook but no backend: there
was no /ws/system endpoint and nothing forwarded RATE_LIMIT_HIT/LIFTED to a
socket, so the banner got no live updates.

Build the missing half:
- ConnectionManager grows a system-wide connection set with connect_system /
  broadcast_system, and disconnect() now clears it.
- A /ws/system websocket endpoint (operator stream, no per-agent keying) with
  the same connected + ping/pong lifecycle as the other streams.
- websocket_bridge subscribes RATE_LIMIT_HIT/LIFTED and forwards each to
  broadcast_system tagged with the type the panel switches on. Both events
  ride the same StreamEventBus singleton, and the subscriptions register
  before start_listening(), so the consumer reads their streams.

Pairs with the panel hook now passing '/system' (getWebSocketUrl supplies the
'/ws' base). Covered by handler, manager, and endpoint-lifecycle tests.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-06-11 18:16:20 +02:00

243 lines
9.4 KiB
Python

"""Unit tests for RateLimitStateTracker.
These tests use mock Redis clients (no real Redis server required) to
verify the state-management logic. The cross-reconnection persistence
test constructs *two* RateLimitStateTracker instances that share the
same mock Redis store, proving that state written by one instance is
visible to a fresh instance.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import AsyncMock
from roboco.services.gateway.rate_limit_tracker import RateLimitStateTracker
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _make_redis_mock(initial_store: dict[str, Any] | None = None) -> AsyncMock:
"""Build an async Redis mock backed by a plain dict.
The mock supports ``get``, ``set``, and ``delete`` with the same
semantics as the real redis.asyncio.Redis client.
"""
# Use the dict AS-IS (no copy) so that two mocks sharing the same
# dict object see each other's writes and deletes — this is what the
# cross-reconnection persistence tests rely on.
store: dict[str, Any] = initial_store if initial_store is not None else {}
async def _get(key: str) -> bytes | None:
val = store.get(key)
if val is None:
return None
if isinstance(val, bytes):
return val
return str(val).encode()
async def _set(key: str, value: Any) -> None:
store[key] = value
async def _delete(key: str) -> int:
return 1 if store.pop(key, None) is not None else 0
mock = AsyncMock()
mock.get = AsyncMock(side_effect=_get)
mock.set = AsyncMock(side_effect=_set)
mock.delete = AsyncMock(side_effect=_delete)
# Stash the backing store so tests can inspect raw state
mock._store = store
return mock
def _make_tracker(
provider: str = "anthropic",
redis_mock: AsyncMock | None = None,
) -> RateLimitStateTracker:
"""Build a tracker with an injected mock Redis client."""
tracker = RateLimitStateTracker(provider=provider, redis_url="redis://unused")
if redis_mock is not None:
tracker._redis = redis_mock # type: ignore[assignment]
return tracker
# ---------------------------------------------------------------------------
# Tests: basic operations
# ---------------------------------------------------------------------------
class TestActivateAndRead:
async def test_is_rate_limited_false_by_default(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
assert await tracker.is_rate_limited() is False
async def test_get_state_empty_by_default(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
assert await tracker.get_state() == {}
async def test_activate_sets_rate_limited_true(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
assert await tracker.is_rate_limited() is True
async def test_activate_stores_retry_after(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
retry_after = 30.0
await tracker.activate(retry_after=retry_after)
state = await tracker.get_state()
assert state["retry_after"] == retry_after
async def test_activate_stores_affected_agents(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate(affected_agents=["be-dev-1", "be-dev-2"])
state = await tracker.get_state()
assert state["affected_agents"] == ["be-dev-1", "be-dev-2"]
async def test_activate_initialises_probe_failures_zero(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
state = await tracker.get_state()
assert state["probe_failures"] == 0
async def test_clear_removes_state(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
await tracker.clear()
assert await tracker.is_rate_limited() is False
assert await tracker.get_state() == {}
class TestProbeFailures:
async def test_increment_starts_from_zero(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
count = await tracker.increment_probe_failures()
assert count == 1
async def test_increment_accumulates(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
increments = 3
for _ in range(increments):
count = await tracker.increment_probe_failures()
assert count == increments
async def test_reset_sets_zero(self) -> None:
mock = _make_redis_mock()
tracker = _make_tracker(redis_mock=mock)
await tracker.activate()
await tracker.increment_probe_failures()
await tracker.increment_probe_failures()
await tracker.reset_probe_failures()
state = await tracker.get_state()
assert state["probe_failures"] == 0
# ---------------------------------------------------------------------------
# Tests: cross-reconnection persistence
# ---------------------------------------------------------------------------
#
# State persists across client reconnection: a test writes state via
# activate(), creates a new RateLimitStateTracker instance pointing at the
# same Redis URL, calls is_rate_limited() and get_state() and gets back the
# same values — proving state survives a process restart.
#
# We simulate this by sharing the same backing dict between two mock Redis
# clients — one injected into the first tracker and one injected into the
# second. Both clients read from and write to the same dict, so the second
# tracker "sees" everything the first wrote.
# ---------------------------------------------------------------------------
class TestStatePersistsAcrossReconnection:
async def test_is_rate_limited_survives_reconnection(self) -> None:
shared_store: dict[str, Any] = {}
# First "connection": write rate-limit state
mock_a = _make_redis_mock(initial_store=shared_store)
tracker_a = _make_tracker(provider="anthropic", redis_mock=mock_a)
await tracker_a.activate(retry_after=60.0, affected_agents=["be-dev-1"])
# The mock writes into shared_store directly (our _set stores raw).
# We need to seed the second mock from the same backing store.
# Because mock_a._store IS shared_store (same dict object), we only
# need to give mock_b access to the same dict.
mock_b = _make_redis_mock(initial_store=mock_a._store)
tracker_b = RateLimitStateTracker(
provider="anthropic", redis_url="redis://unused"
)
tracker_b._redis = mock_b # type: ignore[assignment]
assert await tracker_b.is_rate_limited() is True
async def test_get_state_survives_reconnection(self) -> None:
shared_store: dict[str, Any] = {}
retry_after = 45.0
mock_a = _make_redis_mock(initial_store=shared_store)
tracker_a = _make_tracker(provider="anthropic", redis_mock=mock_a)
await tracker_a.activate(retry_after=retry_after, affected_agents=["be-dev-2"])
mock_b = _make_redis_mock(initial_store=mock_a._store)
tracker_b = RateLimitStateTracker(
provider="anthropic", redis_url="redis://unused"
)
tracker_b._redis = mock_b # type: ignore[assignment]
state = await tracker_b.get_state()
assert state["rate_limited"] is True
assert state["retry_after"] == retry_after
assert state["affected_agents"] == ["be-dev-2"]
async def test_clear_via_first_instance_visible_to_second(self) -> None:
shared_store: dict[str, Any] = {}
mock_a = _make_redis_mock(initial_store=shared_store)
tracker_a = _make_tracker(provider="anthropic", redis_mock=mock_a)
await tracker_a.activate()
# Second instance points at the same store
mock_b = _make_redis_mock(initial_store=mock_a._store)
tracker_b = RateLimitStateTracker(
provider="anthropic", redis_url="redis://unused"
)
tracker_b._redis = mock_b # type: ignore[assignment]
# Write clear via tracker_a
await tracker_a.clear()
# tracker_b observes the cleared state
assert await tracker_b.is_rate_limited() is False
assert await tracker_b.get_state() == {}
# ---------------------------------------------------------------------------
# Tests: different providers are isolated
# ---------------------------------------------------------------------------
class TestProviderIsolation:
async def test_activating_one_provider_does_not_affect_another(self) -> None:
store: dict[str, Any] = {}
mock_a = _make_redis_mock(initial_store=store)
mock_b = _make_redis_mock(initial_store=store)
tracker_anthropic = _make_tracker(provider="anthropic", redis_mock=mock_a)
tracker_ollama = _make_tracker(provider="ollama_cloud", redis_mock=mock_b)
await tracker_anthropic.activate()
assert await tracker_anthropic.is_rate_limited() is True
assert await tracker_ollama.is_rate_limited() is False