mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* refactor(analytics): reduce cyclomatic complexity in usage/pricing/rollup
Collapse the three near-identical get_by_* aggregation methods in
UsageService into a shared _aggregate_by helper parameterized by group
column and key name, and centralize token null-coalescing in a
_row_tokens helper. Extract the per-row upsert in _sweep_daily_rollup
into _upsert_rollup_row, and the pricing-table lookup into
_lookup_prices. All blocks now rank <= B and both modules rank A, so the
xenon gate passes; behavior is unchanged and existing tests stay green.
* feat(billing): make token pricing provider-aware
Distinguish three cases when a model has no per-token rate: a non-Anthropic
model (local Ollama, or an Ollama Cloud ":cloud" model billed by flat
subscription / GPU-time) legitimately has no per-token cost and returns 0.0
silently; an unpriced Anthropic ("claude"-named) model also returns 0.0 but
logs a warning, since that is real spend being undercounted and catches new or
renamed Claude models missing from the table. Folds the old ollama/ prefix
special-case into the general non-Anthropic path so there is one code path,
and replaces the blanket 'no pricing data' warning that fired even for
self-hosted models.
* fix(tasks): preserve ownership when force-unclaiming to pending
The stale-claim reaper and the dependency-blocked release both routed through
_force_unclaim_to_pending, which nulled assigned_to and left the task in a
pending state owned by nobody — no dispatcher re-spawns an ownerless pending
task, so it went dormant. The dispatcher-side claimed_by fallback only masked
half the cases.
Capture the owner before releasing the claim and keep both assigned_to and
claimed_by pointed at it (mirroring the unblock restore), releasing only the
live claim (active_claimant_id + heartbeat) and the WorkSession. The same agent
now resumes the task once it re-dispatches. Updates the reaper test that
asserted the old orphaning behavior and adds owner-preservation coverage for
both the reaper and dependency-release paths.
* fix(tasks): unblock restores the owner into both ownership fields
Audit follow-up to the force-unclaim ownership fix. unblock() only restored
assigned_to from blocker_raised_by, which block() stashes solely from
assigned_to. A task claimed via give_me_work (claimed_by set, assigned_to null)
therefore unblocked into a split-owner state — assigned_to null but claimed_by
set — that both the dev dispatcher and the PM pool-router race to pick up. It
also left claimed_by pointing at the resolver PM after an escalation.
Resolve the owner as blocker_raised_by or assigned_to or claimed_by and write
it to both fields, matching the force-unclaim and reassign convention so the
original worker resumes cleanly. Adds coverage for the give_me_work-claim case
and asserts owner restoration on the existing in_progress-resume test.
* test(orchestrator): cover dev owner resolution and the claimed_by fallback
_resolve_dev_owner_uuid had no coverage. Add the status-dependent precedence
(claimed/blocked prefer the live claimant; other statuses prefer the
PM-assigned owner) and the half-reap fallback where a pending task with
assigned_to nulled still resolves its owner from claimed_by instead of going
dormant.
* fix(tasks): wire the pre-block snapshot so unblock(restore=True) works
The restore=True path on a PM unblock was a no-op: pre_block_state /
pre_block_assignee (migration 006) were read by unblock_with_restore but never
written, so it always fell through to legacy unblock() and the restore flag did
nothing.
Snapshot the resting status + owner at every block entry (dependency block,
soft block, escalation) before mutating, capturing only the first block in a
chain so a re-block doesn't overwrite the original state. Escalation snapshots
the outgoing owner, not the escalation target, so restore returns the original
worker. The restore path applies the same branchless guard legacy unblock()
relies on — a snapshotted in_progress with no branch diverts to pending instead
of looping the dispatcher — and is extracted into _apply_pre_block_restore to
keep complexity under the gate. Adds coverage for snapshot capture, restore,
the branchless divert, and escalation owner restoration.
* test(tasks): update orphan-reconciler and dependency-release tests for owner preservation
Both the startup orphan reconciler and the dependency-blocked claim release
route through unclaim_for_reaper / _force_unclaim_to_pending, which now preserve
the owner instead of nulling assigned_to. Update the two tests that asserted the
old orphaning behavior to assert the owner is kept (so the same agent resumes)
while the live claim is released.
* chore(tests): scrub internal work-item labels from test names, docstrings, comments
Rename four test files that carried audit work-item IDs in their filenames
(test_p0_7_branch_atomicity, test_p2_8_orphan_reconciler,
test_p2_9_autogen_prompt_layer, test_p2_7_attempt_id) to describe what they
test, and strip the matching P-/D-/S- cluster labels from docstrings, comments,
and assertion messages across the test suite and two orchestrator comments.
These are internal references with no meaning in the codebase; behavior is
unchanged.
* style: reformat assertion line shortened by the internal-ref scrub
* build: waive unreachable torch CVE-2025-3000 in pip-audit gate
torch is a transitive CPU-pinned dep (piragi / sentence-transformers) never
loaded at runtime — the stack uses Ollama over HTTP for all embeddings/LLM, so
the vulnerable torch.jit.script path is unreachable. CVE-2025-3000 is MEDIUM,
local-only, with no published fix. Documented --ignore-vuln waiver; revisit when
a fixed torch ships.
* fix(orchestrator): route unplaceable pending tasks to main-pm instead of dropping them
_get_routing_target returned None when a 'dev'-classified task had no cell
agent (no team, or a non-cell team like fullstack/system) or when the routing
classification was unrecognized. _route_unassigned_pm_task logged 'no routing
target found' and returned, leaving the task ownerless and pending — and no
dispatcher re-spawns an unrouted pending task, so it went dormant for 10+ min
until the stuck-task detector caught it.
Fall back to main-pm (the same default cell_pm routing and escalation already
use) so the task is always owned and triaged, never stranded. Logs the fallback
so unplaceable tasks stay visible. Adds a test asserting no (routing, team)
combination ever resolves to None.
* fix(panel): make intake chat markdown inherit the bubble's text color
MarkdownBody is shared by the assistant (text-foreground) and user
(text-primary-foreground) bubbles. [&_*]:!text-inherit only colored the prose
div's descendants, so the prose div itself kept the prose typography body color
(gray) and children inherited that — unreadable on the muted assistant bubble.
Add !text-inherit on the prose div itself so it inherits the bubble's color
too; descendants then inherit the correct foreground. Fixes both bubbles without
hardcoding a color.
* fix(prompter): keep a board-reviewed product on the board team so Approve & Start shows
A product coordination root confirmed via 'Board review & Start' is assigned to
a board reviewer (product-owner) for review, but create_task_from_draft set
team=main_pm for every product unconditionally. The CEO's Approve & Start gate
keys on team=board, so the button never appeared — and because the owner stayed
a board agent while the team said main_pm, the dispatcher routed it to the board
path (nothing left to do after review) and the task stranded at pending, with
the board agent fruitlessly trying to escalate it up.
Route a product by its assignee: a board reviewer keeps it team=board (so the
gate appears and approve_and_start later hands it to Main PM), while a main-pm
assignee — the 'Approve & Start' straight-through path — is team=main_pm. Adds
_assignee_is_board mirroring TaskService's board-role check, and a test.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
332 lines
12 KiB
Python
332 lines
12 KiB
Python
"""
|
|
Unit tests for roboco.billing.pricing — calculate_cost().
|
|
|
|
Covers:
|
|
- Each model tier (opus, sonnet, haiku) with all 4 token types.
|
|
- Unknown model name returns 0.0 without raising.
|
|
- Empty model string returns 0.0 without raising.
|
|
- Substring match correctness: longer fragment wins
|
|
(e.g. 'claude-sonnet-4-6' matches 'claude-sonnet-4' not bare 'sonnet').
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import pytest
|
|
from roboco.billing.pricing import _is_anthropic_model, calculate_cost
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Named constants (ruff PLR2004: magic values in comparisons must be named).
|
|
# ---------------------------------------------------------------------------
|
|
|
|
# Token counts
|
|
_M = 1_000_000 # 1 million tokens
|
|
|
|
# Pricing — per-1M USD, matches the _PRICING table in pricing.py
|
|
_OPUS_INPUT = 5.00
|
|
_OPUS_OUTPUT = 25.00
|
|
_OPUS_CACHE_READ = 0.50
|
|
_OPUS_CACHE_WRITE = 6.25
|
|
|
|
_SONNET_INPUT = 3.00
|
|
_SONNET_OUTPUT = 15.00
|
|
_SONNET_CACHE_READ = 0.30
|
|
_SONNET_CACHE_WRITE = 0.75
|
|
|
|
_HAIKU_INPUT = 1.00
|
|
_HAIKU_OUTPUT = 5.00
|
|
_HAIKU_CACHE_READ = 0.10
|
|
_HAIKU_CACHE_WRITE = 1.25
|
|
|
|
_HAIKU3_INPUT = 0.25 # claude-haiku-3 is cheaper than haiku-3-5 / haiku-4
|
|
|
|
# Tolerance for floating-point comparisons
|
|
_TOL = 1e-4
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Opus tier
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestOpusTier:
|
|
"""claude-opus-4 family pricing."""
|
|
|
|
def test_input_only(self) -> None:
|
|
cost = calculate_cost("claude-opus-4-5", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _OPUS_INPUT) < _TOL
|
|
|
|
def test_output_only(self) -> None:
|
|
cost = calculate_cost("claude-opus-4-5", tokens_input=0, tokens_output=_M)
|
|
assert abs(cost - _OPUS_OUTPUT) < _TOL
|
|
|
|
def test_cache_read_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-opus-4-5",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_read=_M,
|
|
)
|
|
assert abs(cost - _OPUS_CACHE_READ) < _TOL
|
|
|
|
def test_cache_write_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-opus-4-5",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_write=_M,
|
|
)
|
|
assert abs(cost - _OPUS_CACHE_WRITE) < _TOL
|
|
|
|
def test_all_token_types(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-opus-4-5",
|
|
tokens_input=_M,
|
|
tokens_output=_M,
|
|
tokens_cache_read=_M,
|
|
tokens_cache_write=_M,
|
|
)
|
|
expected = _OPUS_INPUT + _OPUS_OUTPUT + _OPUS_CACHE_READ + _OPUS_CACHE_WRITE
|
|
assert abs(cost - expected) < _TOL
|
|
|
|
def test_short_alias(self) -> None:
|
|
"""Bare 'opus' alias resolves to the opus tier."""
|
|
cost = calculate_cost("opus", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _OPUS_INPUT) < _TOL
|
|
|
|
def test_returns_float(self) -> None:
|
|
cost = calculate_cost("claude-opus-4", tokens_input=100, tokens_output=50)
|
|
assert isinstance(cost, float)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Sonnet tier
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestSonnetTier:
|
|
"""claude-sonnet-4 family pricing."""
|
|
|
|
def test_input_only(self) -> None:
|
|
cost = calculate_cost("claude-sonnet-4-6", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _SONNET_INPUT) < _TOL
|
|
|
|
def test_output_only(self) -> None:
|
|
cost = calculate_cost("claude-sonnet-4-6", tokens_input=0, tokens_output=_M)
|
|
assert abs(cost - _SONNET_OUTPUT) < _TOL
|
|
|
|
def test_cache_read_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-sonnet-4-6",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_read=_M,
|
|
)
|
|
assert abs(cost - _SONNET_CACHE_READ) < _TOL
|
|
|
|
def test_cache_write_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-sonnet-4-6",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_write=_M,
|
|
)
|
|
assert abs(cost - _SONNET_CACHE_WRITE) < _TOL
|
|
|
|
def test_all_token_types(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-sonnet-4-6",
|
|
tokens_input=_M,
|
|
tokens_output=_M,
|
|
tokens_cache_read=_M,
|
|
tokens_cache_write=_M,
|
|
)
|
|
expected = (
|
|
_SONNET_INPUT + _SONNET_OUTPUT + _SONNET_CACHE_READ + _SONNET_CACHE_WRITE
|
|
)
|
|
assert abs(cost - expected) < _TOL
|
|
|
|
def test_short_alias(self) -> None:
|
|
"""Bare 'sonnet' alias resolves to the sonnet tier."""
|
|
cost = calculate_cost("sonnet", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _SONNET_INPUT) < _TOL
|
|
|
|
def test_35_variant(self) -> None:
|
|
"""claude-3-5-sonnet resolves to sonnet tier."""
|
|
cost = calculate_cost(
|
|
"claude-3-5-sonnet-20241022", tokens_input=_M, tokens_output=0
|
|
)
|
|
assert abs(cost - _SONNET_INPUT) < _TOL
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Haiku tier
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestHaikuTier:
|
|
"""claude-haiku family pricing."""
|
|
|
|
def test_input_only(self) -> None:
|
|
cost = calculate_cost("claude-haiku-4-5", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _HAIKU_INPUT) < _TOL
|
|
|
|
def test_output_only(self) -> None:
|
|
cost = calculate_cost("claude-haiku-4-5", tokens_input=0, tokens_output=_M)
|
|
assert abs(cost - _HAIKU_OUTPUT) < _TOL
|
|
|
|
def test_cache_read_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-haiku-4-5",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_read=_M,
|
|
)
|
|
assert abs(cost - _HAIKU_CACHE_READ) < _TOL
|
|
|
|
def test_cache_write_only(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-haiku-4-5",
|
|
tokens_input=0,
|
|
tokens_output=0,
|
|
tokens_cache_write=_M,
|
|
)
|
|
assert abs(cost - _HAIKU_CACHE_WRITE) < _TOL
|
|
|
|
def test_all_token_types(self) -> None:
|
|
cost = calculate_cost(
|
|
"claude-haiku-4-5",
|
|
tokens_input=_M,
|
|
tokens_output=_M,
|
|
tokens_cache_read=_M,
|
|
tokens_cache_write=_M,
|
|
)
|
|
expected = _HAIKU_INPUT + _HAIKU_OUTPUT + _HAIKU_CACHE_READ + _HAIKU_CACHE_WRITE
|
|
assert abs(cost - expected) < _TOL
|
|
|
|
def test_short_alias(self) -> None:
|
|
"""Bare 'haiku' alias resolves to the haiku tier."""
|
|
cost = calculate_cost("haiku", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _HAIKU_INPUT) < _TOL
|
|
|
|
def test_haiku3_variant(self) -> None:
|
|
"""claude-haiku-3 has lower pricing than haiku-3-5."""
|
|
cost = calculate_cost("claude-haiku-3", tokens_input=_M, tokens_output=0)
|
|
assert abs(cost - _HAIKU3_INPUT) < _TOL
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Unknown / edge cases — must return 0.0 without raising
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestUnknownModels:
|
|
def test_unknown_model_name_returns_zero(self) -> None:
|
|
cost = calculate_cost("gpt-4o", tokens_input=_M, tokens_output=_M)
|
|
assert cost == 0.0
|
|
|
|
def test_empty_string_returns_zero(self) -> None:
|
|
cost = calculate_cost("", tokens_input=_M, tokens_output=_M)
|
|
assert cost == 0.0
|
|
|
|
def test_gibberish_returns_zero(self) -> None:
|
|
cost = calculate_cost(
|
|
"totally-unknown-model-xyz", tokens_input=100, tokens_output=100
|
|
)
|
|
assert cost == 0.0
|
|
|
|
def test_zero_tokens_with_unknown_model_returns_zero(self) -> None:
|
|
cost = calculate_cost("unknown", tokens_input=0, tokens_output=0)
|
|
assert cost == 0.0
|
|
|
|
def test_does_not_raise_on_unknown_model(self) -> None:
|
|
"""Must not raise regardless of token counts."""
|
|
try:
|
|
calculate_cost(
|
|
"not-a-claude-model",
|
|
tokens_input=999_999,
|
|
tokens_output=999_999,
|
|
)
|
|
except Exception as exc:
|
|
pytest.fail(f"calculate_cost raised unexpectedly: {exc}")
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Substring match correctness
|
|
# ---------------------------------------------------------------------------
|
|
|
|
# Named constants for the comparison floor/ceiling used in these tests.
|
|
_ZERO_COST = 0.0
|
|
_SONNET_CHEAPER_THAN_OPUS = True # structural assertion in the test below
|
|
|
|
|
|
class TestSubstringMatchPriority:
|
|
def test_claude_sonnet_4_resolves_non_zero(self) -> None:
|
|
"""'claude-sonnet-4-6' must find a match (non-zero cost)."""
|
|
cost = calculate_cost("claude-sonnet-4-6", tokens_input=_M, tokens_output=0)
|
|
assert cost > _ZERO_COST
|
|
|
|
def test_haiku3_cheaper_than_haiku4(self) -> None:
|
|
"""claude-haiku-3 is cheaper than claude-haiku-4 — longest-match wins."""
|
|
haiku3_cost = calculate_cost("claude-haiku-3", tokens_input=_M, tokens_output=0)
|
|
haiku4_cost = calculate_cost("claude-haiku-4", tokens_input=_M, tokens_output=0)
|
|
# haiku-3 ($0.25/1M) < haiku-4 ($1.00/1M)
|
|
assert haiku3_cost < haiku4_cost
|
|
|
|
def test_non_claude_model_returns_zero(self) -> None:
|
|
"""A random non-Claude model must not match any Claude pricing entry."""
|
|
non_opus_cost = calculate_cost("llama-3-70b", tokens_input=_M, tokens_output=0)
|
|
assert non_opus_cost == _ZERO_COST
|
|
|
|
def test_opus_model_non_zero(self) -> None:
|
|
"""Claude opus model resolves to non-zero cost."""
|
|
opus_cost = calculate_cost("claude-opus-4", tokens_input=_M, tokens_output=0)
|
|
assert opus_cost > _ZERO_COST
|
|
|
|
def test_zero_tokens_returns_zero_for_known_model(self) -> None:
|
|
"""Known model with 0 tokens has 0 cost."""
|
|
cost = calculate_cost("claude-opus-4", tokens_input=0, tokens_output=0)
|
|
assert cost == _ZERO_COST
|
|
|
|
def test_case_insensitive_matching(self) -> None:
|
|
"""Model name matching is case-insensitive."""
|
|
lower_cost = calculate_cost(
|
|
"claude-sonnet-4-6", tokens_input=1000, tokens_output=1000
|
|
)
|
|
upper_cost = calculate_cost(
|
|
"CLAUDE-SONNET-4-6", tokens_input=1000, tokens_output=1000
|
|
)
|
|
assert lower_cost == upper_cost
|
|
assert lower_cost > _ZERO_COST
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Provider awareness — non-Anthropic models have no per-token cost
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestProviderAwareness:
|
|
"""Non-Anthropic models (local Ollama / Ollama Cloud) cost 0.0 per token."""
|
|
|
|
def test_ollama_prefixed_model_returns_zero(self) -> None:
|
|
"""Self-hosted Ollama models (``ollama/`` prefix) have no API cost."""
|
|
cost = calculate_cost("ollama/llama3", tokens_input=_M, tokens_output=_M)
|
|
assert cost == _ZERO_COST
|
|
|
|
def test_ollama_cloud_model_returns_zero(self) -> None:
|
|
"""Ollama Cloud (``:cloud`` tag) is subscription-billed, not per token."""
|
|
cost = calculate_cost("glm-5:cloud", tokens_input=_M, tokens_output=_M)
|
|
assert cost == _ZERO_COST
|
|
|
|
def test_bare_local_model_returns_zero(self) -> None:
|
|
"""A bare local embedding model has no per-token cost."""
|
|
cost = calculate_cost("qwen3-embedding:0.6b", tokens_input=_M, tokens_output=0)
|
|
assert cost == _ZERO_COST
|
|
|
|
def test_is_anthropic_model_true_for_claude_names(self) -> None:
|
|
for name in ("claude-opus-4-6", "claude-fable-5", "opus", "sonnet", "haiku"):
|
|
assert _is_anthropic_model(name) is True, name
|
|
|
|
def test_is_anthropic_model_false_for_non_claude_names(self) -> None:
|
|
for name in ("ollama/llama3", "glm-5:cloud", "qwen3-embedding", "gpt-4o"):
|
|
assert _is_anthropic_model(name) is False, name
|