Commit Graph
1117 Commits
Author SHA1 Message Date
roboco-app[bot]GitHubBackend Developer 2Backend Developer 1roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
db40ea82df [a3a68044] Revise factory-ai-rebrand-2026-06.md per PR-gate findings (#781)
* [bcae2134] docs(positioning): add response decision, customer list, and 4th verified fact to factory-ai-rebrand note (#779)

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [906fd2ba] docs(positioning): reword fact 4 to state self-heal loop ships dormant (#782)

Fact 4 stated the self-heal engine watches CI and opens fix tasks
without noting both gating flags default to False. PR #781 review
flagged this (F-8c80ee7f, major). Reworded to cite self_heal_enabled
and self_heal_originate_enabled defaulting to False in
roboco/config.py (lines 879, 907), so the loop does not run or
originate tasks until an operator enables both flags.

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [d1d36ba2] Format source URL as markdown link in factory-ai-rebrand positioning note (#789)

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
2026-08-01 21:34:46 +00:00
Main PM 3cd892f7dd Merge remote-tracking branch 'origin/slave' into feature/main_pm/984d8b7c 2026-07-31 23:40:43 +00:00
Renn F b19856dcbd fix(dispatch): route a code task that would land on main-pm to a cell PM instead
The routing layer's never-strand fallbacks — and the strategic
classifier's own cross-cell/high-complexity/teamless branches — could
hand a code-typed pending task to main-pm, but the dispatcher's
pre-claim route refuses that by org invariant (MAIN_PM_NO_CODE:
main_pm+code must not coexist; intake coerces such roots to
planning), so the task looped claim-reject every ~30s tick forever
instead of failing loud. Zero live instances in three days of logs —
preventive, found adversarially while fixing the creator-skip hold.

_get_routing_target now wraps the pure resolution: a code task whose
resolved target is main-pm redirects via _cell_pm_for_stranded_code_task
— nearest ancestor cell team wins (parent-chain walk mirroring
_resolve_pm_for_review), parentless/unresolvable defaults to backend.
Non-code routing is byte-for-byte unchanged; board precedence
untouched; the claim-route guard and the lifecycle spec's broader
i_will_plan allowance (cell PMs planning code parents, needs_revision
recovery) are both left exactly as designed.

Gate: full make quality green (suite passed through coverage; xenon/
bandit/pip-audit/deptry/alembic/import-linter/foundation-check exit 0).
2026-08-01 01:36:39 +02:00
roboco-app[bot]GitHubBackend Developer 1roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>Backend Developer 2Backend Documenter
cd011a1808 [a9a030d8] Add competitive-positioning note: RoboCo vs Factory.ai rebrand (#718)
* [4db76cc9] docs(positioning): add RoboCo vs Factory.ai rebrand positioning note (#716)

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [976a22aa] docs(positioning): add source citation line for Factory.ai rebrand note (F-1c3881c6) (#723)

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [261e7585] Diagnose and clear PR #718's red Python quality-gate check (#746)

* [261e7585] docs(positioning): reflow hard-wrapped prose to fix Python quality gate

Root-caused PR #718's red Python quality-gate CI check: it was NOT branch
staleness (sync_branch confirms this branch is already current with its
base a9a030d8) — docs/positioning/factory-ai-rebrand-2026-06.md has
hard-wrapped prose that fails make quality's "markdown prose (no
hard-wrapping)" step (scripts/reflow_md.py --check), a check distinct
from ruff/mypy/xenon/bandit/coverage that both prior PR-gate rounds
overlooked. Ran `make reflow-docs`, which only rewraps this one file to
one-line-per-paragraph; the tool's own token-invariant safety check plus
manual before/after comparison confirm zero content/meaning change.
make reflow-check and make gate both pass clean after the fix.

* [261e7585] docs(positioning): revert unauthorized reflow edit; report inherited PR #718 gate failure

Reverts commit 7745c980 which ran `make reflow-docs` on
docs/positioning/factory-ai-rebrand-2026-06.md, violating this task's
explicit "do not edit this file" instruction (finding F-ac93df6a). File
is restored byte-for-byte (1265 bytes, verified) to its pre-edit content
at parent commit 2e43db85, extracted directly from the git blob object
since evidence()/roboco_git_log/roboco_git_diff were all timing out this
session.

Diagnosis, confirmed locally: `make lint` (ruff format --check, ruff
check, mypy, vulture), `make xenon`, and `make bandit` are all green on
this reverted tree. `make reflow-check` (scripts/reflow_md.py --check,
wired into `make quality` as the "markdown prose (no hard-wrapping)"
step) is the sole red step — a check distinct from ruff/mypy/xenon/
bandit/coverage that prior rounds didn't isolate. sync_branch already
confirmed this branch is current with root (status: superseded), so this
is not staleness. The hard-wrap predates this task: it was merged in via
completed sibling tasks 4db76cc9 (added the file) and 976a22aa (added a
source line, still hard-wrapped). Fixing it requires editing the
protected file, which is out of this docs-only task's authorized scope
per its own branch-3 instruction — reporting instead so the PM can
authorize a proper follow-up (`make reflow-docs` merged through its own
task) to actually turn PR #718 green.

* [261e7585] docs(qa): document PR #718 reflow-check root cause as inherited, unfixable here

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [d4d255d1] Reflow docs/positioning/factory-ai-rebrand-2026-06.md to clear CI reflow-check (#770)

* [d4d255d1] docs(positioning): reflow factory-ai-rebrand note to clear markdown-prose check

* [d4d255d1] docs(qa): reflow ci-reflow-check-positioning-note.md to clear repo-wide markdown-prose check

scripts/reflow_md.py has no per-file scope — it globs the whole repo, so
PR #718's CI reflow-check step stayed red even after the target positioning
note was reflowed, because this sibling diagnostic note was also
hard-wrapped. Ran scripts/reflow_md.py --apply (token-invariant, no wording
change) per be-pm's scope authorization to clear finding F-d41b9bc7.

---------

Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>

* [aab89b51] docs(positioning): fix stray space and remove stale reflow-diagnosis doc (#773)

- docs/positioning/factory-ai-rebrand-2026-06.md: remove the stray space
  after "review/" (an artifact of an earlier reflow join) so the loop
  reads "build/test/review/secure/ship", matching the task description
  and the cited source. No other wording changed.
- docs/backend/qa/ci-reflow-check-positioning-note.md: deleted — it was
  stale against its own PR (claims an unresolved reflow failure that
  already landed) and added a second file beyond the root task's
  single-markdown-file acceptance criterion.

Verified scripts/reflow_md.py --check passes and make lint is fully
clean (ruff format/check + mypy + vulture, exit 0).

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
2026-07-31 23:00:19 +00:00
Renn F be9034363e fix(dispatch): age-gate the creator-skip so an unassigned PM-routed task can't be skipped forever
_route_unassigned_pm_task's creator-skip guard (don't auto-claim a
task back to the PM that just created it — it is about to assign it
one tool-call later) had no aging: when the creator-PM's session
ended without assigning, the task was skipped every ~30s dispatch
tick forever. Live: a be-pm-created fix task sat pending-unassigned
for 21 hours with 'Skipping auto-claim: routing target is the
creator' firing every tick.

The skip now applies only while the task is younger than
_CREATOR_ROUTE_GRACE_SECONDS (600) — past grace (or on an
unparseable/missing created_at, via the existing fail-open
_get_task_age) routing proceeds: the creator-PM claims its own task
and triages/delegates on spawn, which is legal and unguarded for a
cell PM. Extracted _creator_route_should_skip to hold the xenon B
ceiling.

Known-separate class, deliberately untouched here: a code-typed task
with an unmapped team (None/fullstack/system) routes to main-pm as a
deliberate never-strand fallback, but main-pm's claim is refused by
MAIN_PM_NO_CODE — a pre-existing every-tick claim-reject loop for
that shape regardless of creator, with no live instances today.

Gate: full make quality green (suite passed through coverage; xenon/
bandit/pip-audit/deptry/alembic/import-linter/foundation-check all
exit 0).
2026-08-01 00:28:28 +02:00
Renn F 3efc96e402 fix(lifecycle): pr_pass hands ownership to the owning PM — closes the passed-PR completion wedge
Removing the awaiting_pm_review re-claim edge (d87e2d9b, the #740
review-loop fix) exposed that the edge was load-bearing: pr_pass
cleared assigned_to/claimed_by to None and the re-claim was the only
way a PM ever re-acquired the task. Since then every assembled task
passing the PR gate wedged: the closure PM's complete rejected with
'not assigned to you', its fallback claim rejected (edge gone), and
its only exit was escalate_up — BLOCKING the task onto main-pm, who
is not the assignee either and burned spawns doing nothing. Live:
PR #741's task (7 ownership rejections, escalated 13:59Z), plus two
more tasks with 25 and 2 rejections in the same shape.

pr_pass now resolves the owning PM via _revision_pm_for_task and
assigns it, exactly as pr_fail always did — one chokepoint covering
cell (submit_up) and root (submit_root) tasks. mark_pr_created's
ready_for_pm branch (leaf docs-path, PR-arrives-second) had the same
clear-to-None wedge and now resolves via _resolve_pm_for_review like
its docs-first sibling. No claim edge is reintroduced; the #740
parity test is untouched.

Recovery for already-wedged tasks needs no DB surgery: new idempotent
TaskService.assign_review_pm + POST /tasks/{id}/assign-review-pm
(ASSIGN-gated, explicit commit) corrects an unassigned OR mis-assigned
awaiting_pm_review task to its owning PM; _dispatch_pm_review_work
ensures assignment before every spawn (cheap pre-check skips the
round-trip when the fetched assigned_to already matches), replacing
the dead _claim_task_for_agent call whose lifecycle claim the removed
edge now always rejects; _maybe_spawn_pm_closure routes through
_closure_review_pm, which keeps the team-resolved PM whenever the
assign route fails so a transient error can never spawn a stale
assignee.

Gate: 15506 passed, 459 skipped; xenon/ruff/mypy/vulture/bandit/
pip-audit/deptry/import-linter/foundation-check green.
2026-07-31 18:48:45 +02:00
Renn F 19d3c227c8 ROBOCO_VAULT_KB_INTERVAL_SECONDS:-3600 2026-07-31 17:05:42 +02:00
Renn F 984083c7e8 fix(gateway): bound the briefing's RAG memory leg; skip cold conventions re-resolution in claim_review
Two remaining 120s-wall consumers found tracing live 504s on the
redeployed stack (claim_review completing at exactly 120s with a
~74s silent tail after the bounded evidence legs gave up):

_institutional_memory's similar_memory call rode the Ollama
embedder's 120s read timeout unbounded inside every full briefing
(give_me_work / i_will_work_on / i_will_plan / resume pass task=t),
so a saturated embedder could eat a claim verb's whole budget. It
now runs under asyncio.wait_for with the new
institutional_memory_timeout_seconds (default 8s) and degrades to a
first-class {status: "timeout", lessons: []} block — the briefing
is enrichment, never the thing that kills a verb.

claim_review's advisory conventions leg trusted that the diff /
files_changed legs had warmed the workspace; when those legs timed
out, conventions_check_for_task's unwrapped setup phase re-hit the
same cold/contended git ops (30s-bounded subprocesses in sequence,
clone repair up to 300s) and ground on until the verb wall. The
evidence build now skips the conventions leg whenever the git legs
already recorded gaps, emitting the existing could_not_run shape
plus an evidence_gaps note instead of a 504. Fail-closed
enforcement (i_am_done / pr_pass) is untouched.

Gate: 15481 passed, 459 skipped; ruff/mypy/xenon/vulture/bandit/
pip-audit/deptry/import-linter/foundation-check green.
2026-07-31 16:52:54 +02:00
Renn F ce4d02b3c2 fix(gateway): bounded fail-open evidence assembly + PM decision transient-failure bypass
Evidence-assembly git legs (diff, changed-files, branch fetch, advisory
conventions run) ran unbounded inside claim_review / claim_doc_task /
claim_gate_review / evidence() / i_am_done's envelope build, so a slow
clone turned the whole verb into a silent 120s FlowVerbTimeout 504.
Each leg now runs through run_bounded_leg under a shared LegBudget
(evidence_assembly_timeout_seconds, 45s total): a timed-out leg — both
asyncio TimeoutError and git's own GitTimeoutError — degrades into an
evidence_gaps note on the envelope instead of hanging the verb, while
non-timeout git errors still propagate. The advisory conventions run
gets an inner-only timeout (conventions_validator_advisory_timeout_
seconds, 30s) threaded down to the subprocess so it is never orphaned
by an outer cancel; the fail-closed i_am_done/pr_pass conventions
gates keep their hardcoded 120s.

_ensure_pm_decision now reports a PmDecisionOutcome: a transient DB
failure recording the PM's decision journal (e.g. lock timeout under
load) no longer launders into a journal:decision gate rejection that
escalates and BLOCKS the task — the verb's own rationale satisfies the
gate with a structured warning, across all seven PM verbs.

Also: repo-wide ruff realignment to the lockfile-pinned ruff (8
format-only diffs, 14 UP038 isinstance conversions) that a transiently
newer venv ruff had masked.

Gate: 15474 passed, 459 skipped; ruff/mypy/xenon/vulture/bandit/
pip-audit/deptry/import-linter/foundation-check all green.
2026-07-31 13:56:19 +02:00
roboco-app[bot]andGitHub ef8849b0cd sync: master → slave 2026-07-31 01:26:15 +00:00
Renn F 07b5eecf27 fix(release): publish roboco-agent-kimi — v0.28.0 shipped without it
The kimi provider (#713) added docker/agent-kimi.Dockerfile and the
agent-kimi-image pull stanza in docker-compose.registry.yml, but never
the entry in release.yml's build/push map — the exact regression class
the map's own history records for the grok sub-images. v0.28.0
published every image except kimi and its pull-smoke job went red on
precisely that pull; the red run went unactioned. The image is
backfilled to both registries manually for :0.28.0/:latest; this entry
covers every future release.
2026-07-31 02:51:44 +02:00
Renzo FandGitHub a605607be4 Merge pull request #753 from rennf93/dependabot/github_actions/docker/login-action-4.5.2 2026-07-31 02:01:25 +02:00
Renzo FandGitHub 107f7af0e6 Merge pull request #752 from rennf93/dependabot/github_actions/actions/stale-11 2026-07-31 02:00:47 +02:00
dependabot[bot]andRenzo F 472d561ce2 chore(deps): bump github/codeql-action from 4 to 4.37.3
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4 to 4.37.3.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4...v4.37.3)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-31 02:00:14 +02:00
Renn F e6c9dde2a9 fix(runtime): stop the chown storm from starving claims, and savepoint the PM journal auto-record
Claim-shaped verbs were failing 7/7 (claim_review) and 6/6
(claim_doc_task) as silent 120s FlowVerbTimeout 504s on the NAS: the
per-claim ownership repair walked the whole clone issuing two stat
syscalls per entry (chown_ms 39502 vs git_ms 8 in the live log), several
passes stacked per claim, and the claim transaction held the task row
the whole time — so concurrent writers queued behind it into the 60s
lock_timeout. The walk now does one stat per entry shared by the
chown-skip and chmod-skip checks, and a .git/roboco-owned sentinel
(worktree-aware via _resolve_clone_root, written only after a
zero-failure pass) skips the walk entirely when the tree is already
agent-owned. Every root-side git write invalidates the sentinel BEFORE
its subprocess runs — GitService._run_git for scope != none, plus the
three raw-subprocess paths inside WorkspaceService the adversarial pass
proved bypass it deterministically on the common respawn shape
(_worktree_git for mutating verbs, _fetch_branch_ref,
_fetch_origin_best_effort) — so a live marker can never vouch for files
a root write is about to create.

One of those queued writers was the PM journal-decision auto-record:
its INSERT hit the lock timeout, _ensure_pm_decision's catch-all
swallowed it without rollback, and the poisoned session blew up
escalate_up with PendingRollbackError (live incident). The helper's try
body now runs in a savepoint — one fix covering all seven PM verbs that
route through it — verified empirically against real Postgres in both
directions: the failure path leaves the session healthy and the task
object readable, and create_entry's internal commit inside the savepoint
drains the transactional outbox exactly once.
2026-07-31 01:53:07 +02:00
dependabot[bot]andGitHub a1f7b2e339 chore(deps): bump docker/login-action from 4 to 4.5.2
Bumps [docker/login-action](https://github.com/docker/login-action) from 4 to 4.5.2.
- [Release notes](https://github.com/docker/login-action/releases)
- [Commits](https://github.com/docker/login-action/compare/v4...v4.5.2)

---
updated-dependencies:
- dependency-name: docker/login-action
  dependency-version: 4.5.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-30 23:34:55 +00:00
dependabot[bot]andGitHub fd11df2293 chore(deps): bump actions/stale from 10 to 11
Bumps [actions/stale](https://github.com/actions/stale) from 10 to 11.
- [Release notes](https://github.com/actions/stale/releases)
- [Changelog](https://github.com/actions/stale/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/stale/compare/v10...v11)

---
updated-dependencies:
- dependency-name: actions/stale
  dependency-version: '11'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-30 23:34:35 +00:00
Renn F d87e2d9b4e fix(lifecycle): stop PMs re-claiming tasks out of the closure queue — kills the i_will_plan review loop
A cell/main PM's i_will_plan could legally re-claim its own task from
awaiting_pm_review (a CLAIM_RULES edge added for post-respawn recovery),
resetting the task to in_progress and re-running submit_up -> pr_pass ->
awaiting_pm_review forever: one Sentinel conventions child looped eleven
full laps in four hours (14 reviews on one PR, 37 agent spawns) while
its root's closure check fired eighteen times and the Main PM could
never close anything.

The claim edge is gone from every table that carried it — CLAIM_RULES,
the claim ActionSpec's source statuses, the StatusTransition row, the
service-layer _ROLE_CLAIM_STATUSES twin, and the legacy enforcement
shim's operational-edge/role-gate entries (left divergent, it would be
the same silent two-table drift that produced this bug). The respawn
case the edge existed for is now served properly: _handle_pm_reentry
gained a third contract — a PM calling i_will_plan on its own
awaiting_pm_review task gets a steering envelope (no claim, no state
change) pointing at complete/request_changes, and give_me_work's next
hint for that status says the same instead of steering back into
i_will_plan. The no-transition review-claim path and the pm-review
dispatch prompt were already correct and are untouched, so closure
still converges through them.

A new bidirectional test asserts CLAIM_RULES and _ROLE_CLAIM_STATUSES
stay identical per PM role (the old comment claimed a sync test existed;
it checked one direction only). Lifecycle artifacts regenerated; the
parity suite's three unshaped session mocks fixed, zero AsyncMock
warnings remain.
2026-07-30 23:37:27 +02:00
Renn F 27e58f469f feat(runtime): stamp spawned containers with the stack's own compose-project labels
Agent spawns (all five provider paths), intake/secretary chats, and
sandbox sidecars now carry com.docker.compose.project/service/oneoff/
config-hash labels copied from the orchestrator's own compose project,
so a Docker UI (UGOS) groups them under the stack's project and they
die with the stack: bare compose stop/restart affects them, compose
down removes them (the config-hash label must be PRESENT for down to
even see the container — compose filters its API listing on that key
before the orphan predicate runs, verified live), and up -d deliberately
does not resurrect them since the orchestrator respawns its own agents.

Self-discovery reads the orchestrator's own container id from
/proc/self/mountinfo keyed on the root-independent /containers/<id>/
segment — the UGREEN NAS data-root is /volume1/@docker on btrfs, so its
mountinfo reads /@docker/containers/..., never the textbook
/var/lib/docker path (verified against the live NAS) — with a HOSTNAME
short-id fallback, then one docker inspect cached per process. Only
definitive outcomes cache; a transient inspect failure logs and retries
on the next spawn. Outside compose the helper yields nothing and every
spawn command is byte-for-byte unchanged.
2026-07-30 17:51:12 +02:00
93739a9dca fix(notifications): stop tick-wide row locks + poisoned-session swallows behind the PendingRollbackError 500s (#743)
* fix(notifications): release re-escalation row locks per row; stop swallowing DB errors in notify_get

The re-escalation sweep ran one tick-wide transaction, so each CAS
claim's row lock was held across every remaining delivery until the
single commit — a concurrent mark-read UPDATE on a claimed row starved
into the 60s lock_timeout. The sweep now commits per row (claim commit
releases the lock before delivery and makes the burned slot durable),
re-fetches each row by snapshotted id so one row's rollback can't
expire the rest of the tick, and savepoints each recipient's delivery.

notify_get's bare except swallowed the resulting LockNotAvailableError
into a false "notification not found" and returned a poisoned session
to the commit-at-send middleware, which blew up with
PendingRollbackError; it now catches only the two domain outcomes.

defer_after_commit's listeners fire on SAVEPOINT release too, which
would have drained deferred telegram/bus work before real durability —
they now skip savepoint boundaries via get_nested_transaction() (the
root get_transaction() is non-None inside the listener even at a real
commit). acknowledge_for_recipient's Redis dedup-clear moved before the
flush so the row lock never spans a Redis round-trip. The five
best-effort CEO-notify swallows that persist notification rows are
savepointed.

* fix(services): contain swallowed best-effort DB write failures instead of poisoning the session

Sweep of the same class as the notify_get incident: broad
except-Exception handlers that swallow a failure whose try-body writes
through the shared session leave the session rollback-pending, and the
verb/request then dies later with PendingRollbackError at
commit-at-send. Confirmed-dangerous sites now run the write inside a
savepoint (safe since defer_after_commit skips savepoint boundaries):
ceo_approve's verified-stamp, completion/pitch/postmortem-style CEO
notifies, _inherit_upstream_base, _link_commit_to_task (covers every
commit route), board-program LEARN records, the QA/PR-gate/PM-merge
verified-stamps, and the documenter->PM handoff.
_ack_pending_wake_notifications gets the same treatment so a wake-ack
failure can't fail the A2A read. telegram_inbound's per-update loop and
intake confirm roll back explicitly instead (their success paths commit
mid-flow, so a savepoint doesn't fit).

A swallowed savepoint rollback fully expires any ORM object mutated
inside the block, and the next attribute read raises MissingGreenlet —
strictly worse than the original bug. The two paths that keep using the
object after the swallow (doc handoff's envelope build, base
inheritance's claim continuation) refresh it in the except path;
regression tests run against a real session and were verified to fail
with the refresh reverted.

* test: shape mocked session.execute results so sync accessors stop leaking unawaited coroutines

An AsyncMock's auto-created children are themselves AsyncMock, so
production code that correctly awaits session.execute() and then calls
sync accessors (.scalars().all(), .scalar_one_or_none()) on the result
was silently collecting unawaited coroutines in 22 test files — 80
RuntimeWarnings per unit run, and in test_flow_soup_guard one mock
raised a real TypeError that a coincidentally-matching invalid_state
envelope masked. Each affected fixture now returns a plain MagicMock
shaped like a real Result. Zero AsyncMock warnings remain.

* docs: document per-row sweep commits and the savepoint/refresh containment pattern

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-30 16:35:42 +02:00
8b18dc3e95 fix(dispatch): stop burning breaker strikes on dependency-held blocked tasks; collision notify rides the delegate transaction (#726)
* fix(dispatch): stop burning breaker strikes on dependency-held blocked tasks; collision notify rides the delegate transaction

Two follow-ups from the 2026-07-29 blocked-task wave:

- _dispatch_blocker_work now skips a blocked task whose dependencies are
  non-terminal BEFORE the respawn gate. spawn_agent's readiness gate was
  refusing these anyway, but each refused attempt burned a respawn-
  breaker strike and, once tripped at 4, a duplicate CEO escalation
  every tick (seen live on the eslint-audit task held behind a paused
  backend dependency). The dispatcher re-checks each tick and proceeds
  the moment the dependency goes terminal — same resume path as before,
  minus the noise.

- send_collision_sequencing_notification gains the db_session threading
  its sibling senders already have, and TaskService passes its own
  session. delegate wires collision edges inside a transaction whose
  held-back child row is not yet committed; the notification INSERT ran
  on a separate auto-commit connection and died on the related_task_id
  FK (ForeignKeyViolationError seen live in cell_pm/delegate), silently
  losing the coordination alert. Riding the caller's transaction makes
  the row visible and commits both atomically.

* test(notification): cast the fake session for the gate's tests-scope mypy

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-29 23:23:29 +02:00
d7a1c2d203 fix(db): bound lock waits and idle transactions so a parked coroutine can't wedge the pool (#721)
* fix(db): bound lock waits and idle transactions so a parked coroutine can't wedge the pool

2026-07-29 production incident: verb/evidence handlers and background
dispatch coroutines held an open DB transaction across minutes of git
subprocess work and asyncio lock queues (per-workspace ensure locks).
Early writes in those transactions held tasks/agents row locks, every
other write convoyed behind them, and blocked statements camped on pool
connections until all 30 were waiters — 1000+ QueuePool timeouts per
hour, one transaction open 1h20m.

Two layers:

- get_engine now passes asyncpg server_settings:
  idle_in_transaction_session_timeout (default 120s) kills any session
  parked mid-transaction on non-DB work, releasing its locks and pool
  slot; lock_timeout (default 30s) makes a statement queued on someone
  else's row lock give up instead of holding a connection for the wait.
  Both env-tunable (ROBOCO_DATABASE_IDLE_IN_TRANSACTION_TIMEOUT_MS /
  ROBOCO_DATABASE_LOCK_TIMEOUT_MS), 0 disables. Alembic runs its own
  sync engine and is untouched; best-effort writers (proactive-context
  injection) already swallow errors and now fail in 30s instead of
  camping for an hour.

- ContentActions.evidence commits the request session before its
  fetch/diff git work, so a multi-minute evidence call no longer pins a
  pool connection for the duration (expire_on_commit=False keeps the
  loaded task usable; later reads reopen a transaction on demand).

The deeper restructuring — claim flows committing their transition
before briefing/workspace assembly — is scoped to the existing
evidence-assembly-timeout task and not attempted here.

* fix(db): bound lock waits and idle transactions so a parked coroutine can't wedge the pool

2026-07-29 production incident: verb/evidence handlers and background
dispatch coroutines held an open DB transaction across minutes of git
subprocess work and asyncio lock queues (per-workspace ensure locks).
Early writes in those transactions held tasks/agents row locks, every
other write convoyed behind them, and blocked statements camped on pool
connections until all 30 were waiters — 1000+ QueuePool timeouts per
hour, one transaction open 1h20m.

Two layers:

- get_engine now passes asyncpg server_settings:
  idle_in_transaction_session_timeout (default 20 min) kills any session
  parked mid-transaction on non-DB work, releasing its locks and pool
  slot; lock_timeout (default 60s) makes a statement queued on someone
  else's row lock give up with a clean retryable error instead of
  camping on a pool connection for the wait. Both env-tunable
  (ROBOCO_DATABASE_IDLE_IN_TRANSACTION_TIMEOUT_MS /
  ROBOCO_DATABASE_LOCK_TIMEOUT_MS), 0 disables. The idle default
  deliberately clears the longest LEGITIMATE in-transaction window — a
  cold-workspace claim holds its transaction across the clone (300s
  budget) + dep install (600s budget) under the 900s slow-verb wall —
  so routine claims never trip it while today's 80-minute parked
  transaction dies at 20 min. Alembic's env.py builds its own engine
  and never carries these; best-effort writers (proactive-context
  injection) already swallow errors and now fail in 60s instead of
  camping for an hour.

- ContentActions.evidence ends the request transaction (commit, or
  rollback on a poisoned session) before its fetch/diff git work, so a
  multi-minute evidence call no longer pins a pool connection for the
  duration (expire_on_commit=False keeps the loaded task usable; later
  reads reopen a transaction on demand).

The deeper restructuring — claim flows committing their transition
before briefing/workspace assembly — is scoped to the existing
evidence-assembly-timeout task and not attempted here.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-29 21:22:35 +02:00
roboco-app[bot]andGitHub e1f5e0950e sync: master → slave 2026-07-29 16:26:44 +00:00
dependabot[bot]andGitHub 6a5fe6958b chore(deps): bump tar from 7.5.19 to 7.5.21 in /video-renderer in the npm_and_yarn group across 1 directory (#715) 2026-07-29 18:12:32 +02:00
RoboCo Release Manager 06f50bd37f chore(release): 0.28.0 v0.28.0 2026-07-29 15:59:52 +00:00
Renzo FandGitHub 666f261a1a fix(kimi): spawn's own prepared instance no longer counts against the concurrency cap (#714) 2026-07-29 17:39:33 +02:00
Renn F f99215213a Claude Doctor Updates 2026-07-29 16:38:49 +02:00
Renn F 95e7d5df7c fix(kimi): cap concurrent Kimi agents to protect the shared auth chain
Every Kimi container redeems the same rotating refresh-token chain;
Moonshot rotates with a short reuse grace, so two containers refreshing
near-simultaneously fork the chain and a later stale redemption revokes
the whole family - fleet-wide re-login (observed twice in production,
each after paired spawns). With one consumer at a time refreshes are
strictly sequential and the chain stays coherent, so the spawn gate now
skips-and-retries Kimi spawns past ROBOCO_KIMI_MAX_CONCURRENT (default
1), sharing the provider-parked bail path. The compose files also gain
the four Kimi tunables their environment blocks silently dropped -
documented .env overrides never reached the orchestrator container.
2026-07-29 06:14:22 +02:00
Renn F 1c3c313c1e fix(coroner): playbook-kind postmortems stop rendering dead buttons
A kind=playbook process change drafts into the playbook queue at propose
time and both approve/reject refuse it - but legacy rows carry no marker
status, so the response defaulted to proposed and the panel rendered
approve/dismiss buttons that bounce forever. The list route now derives
not_applicable for playbook-kind changes regardless of stored status,
which the panel already renders as its Drafted-as-playbook badge.
2026-07-29 05:50:20 +02:00
Renn F 87332346ce docs(kimi): README and KB parity for the Kimi provider
README gains the Kimi setup block, tech-stack row, and checklist entry
(the provider prose also finally names Codex/Gemini, which it had
skipped). The agent-facing KB's provider enum sentence catches up too -
it still called OPENAI reserved and omitted GEMINI - and gains the Kimi
runtime detail plus a config-reference section for the four Kimi
settings.
2026-07-29 04:58:10 +02:00
6374bbbed0 feat(kimi): Kimi K3 provider on the official kimi-code CLI (#713)
* feat(kimi): Kimi K3 provider on the official kimi-code CLI (Wave 1)

ModelProvider.KIMI routes through KimiCliProvider driving Moonshot's kimi
CLI on a Kimi subscription (OAuth device-code, no metered key). One-shot
delivery roles only (V1), interactive ban wired in both guard lists.

Auth: one shared RW auth mount; containers symlink credentials/ and
oauth/ (the CLI's cross-process refresh-lock dir) into a container-local
KIMI_CODE_HOME so every container and the host redeem the SAME rotating
refresh chain - live-verified that per-copy chains cross-invalidate after
the reuse-grace window. No orchestrator refresh daemon; an expires_at
preflight exits 78.

Config renderer mirrors the login-managed provider/model blocks
field-for-field (live-captured; the model value is the CLI-side name,
never the raw API id), plus per-role deny rules and the bash-guard as a
PreToolUse hook via a wrapper script (an env key on a hooks entry makes
the CLI silently drop ALL hooks - live-verified). Usage capture sums
wire.jsonl usage.record 4-bucket events; sniff classifies rate-limit/auth
from structured error text only, mapped to the shared 75/78 park
contract. Image installs the CLI latest-at-build (no version pin, by
policy) with the resolved version stamped as provenance, binary split to
/usr/local away from mutable state.

Migrations 090 (enum) + 091 (provider seed); catalog, pricing, routing
mode, and orchestrator park/usage wiring mirror the codex integration.

* feat(kimi): surface sweep + fleet-wide pin drop (Wave 2)

Compose x3 gain the agent-kimi-image service and the orchestrator's
read-write ~/.kimi-code mount + kimi-usage dir; .env.example documents
the Kimi block. Panel mirrors ModelProvider.KIMI and adds the kimi
routing mode (catalog filter, mode button, mix-picker group, badge) with
tests; provider routes gain the kimi remediation entry. CLAUDE.md and
docs/map document the runtime. Per the no-pins policy, agent-grok/
gemini/codex Dockerfiles drop their version pins for latest-at-build
with resolved-version provenance stamps (grok resolves 0.2.112 vs the
old 0.2.56 pin - verified by real builds of all four images).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-29 01:48:55 +02:00
Renn F eb470dfb33 fix(x): Barfly drafts become standalone link-posts
X 403s programmatic replies into conversations that don't mention the
account (every non-Enterprise tier), which is Barfly's entire discovery
surface - so its drafts now carry commentary plus the conversation's
/i/web/status/ URL instead of a reply target. Handles are stripped
(mentions in plain posts are rejected too), the 280 budget accounts for
the URL, the exploration prompt asks for standalone commentary, and the
reject->redraft path re-appends a dropped link.
2026-07-28 23:52:04 +02:00
Renn F 98ee4d79e5 fix(x): thread mention replies and audit every post outcome
x_reply drafts (the one case X's 2026-02-23 reply policy allows - the
author summoned us) now post as real threaded replies via the carried
mention id; barfly drafts stop sending in_reply_to_tweet_id entirely.
Failed posts log a server-side warning and write an x_post.post_failed
audit row; successes write x_post.posted - at the _post chokepoint so
both approve routes are covered.
2026-07-28 23:52:04 +02:00
Renn F 709759fc9b chore(deps): bump prompt-toolkit to 3.0.53 2026-07-27 01:44:59 +02:00
Renn F 48feb033d6 fix(api): carry orchestration_markers in the task response
TaskResponse declares the field, but task_to_response never assigned it,
so every GET /tasks response served null. The orchestrator fetches its
work over that endpoint and builds prompts from the returned dict, so
all twelve marker reads in orchestrator.py resolved to `{}`.

Live symptom: a Barfly explorer was told "SCREENED CANDIDATES: (none)"
and correctly refused to invent a tweet — while BarflyEngine only opens
an exploration task when it HAS candidates, so the ones it gathered were
sitting on the row the whole time. The gateway path reads markers off the
ORM object directly, which is why propose_conversation_replies could see
what the prompt could not.

Same blindness applied to every other marker consumer on the dispatch
path: the spotlight's seen-features dedup, pest control's evidence
snapshot, coroner's autopsy subject, and the dev-redispatch's
original_developer lookup.
2026-07-27 01:35:18 +02:00
42f8a5d18e fix(board): nothing_to_propose exit for Board Program explorers + PR checks on slave-based PRs (#712)
* fix(board): give Board Program explorers a nothing_to_propose exit

Every propose_* verb requires at least one item, so an explorer that
legitimately found nothing — Barfly with no worthwhile X conversations,
Coroner with no autopsy subject — had no way to close its exploration
task. It declined, called i_am_idle(), and the task stayed PENDING
forever: the dispatcher re-matched it every tick and respawned the board
agent (~$0.61 a spawn, ~3 per 5-minute respawn-breaker cooldown window,
indefinitely), and BoardProgramEngine's one-open-cycle dedup wedged that
whole program shut, since the ledger row only closes once its exploration
task goes terminal.

nothing_to_propose(task_id, reason) is the explicit exit. task_id is
required rather than inferred: one explorer role owns several
independently-cadenced programs (head_marketing owns six) and each
assigns its exploration task to the same agent, so several are open at
once by design and guessing "the caller's oldest" completes the WRONG
cycle — stamping its reason onto an unrelated program's ledger while the
task actually being worked stays wedged. Resolution validates the named
task exists, carries a registered program source, is assigned to the
caller, and is non-terminal, then gates on the program's declared
explorer role from the registry, so a program registered later needs no
edit here.

The reason lands on board_program_cycles (migration 089) and renders into
the next cycle's LEARN context, replacing a bare "proposed 0, approved 0"
with why. That write runs in its own savepoint: it flushes on the same
session as the completion, and a bare try/except around a same-session
flush leaves the transaction pending-rollback, so a DB blip there would
discard the completion at the post-response commit while the verb
reported success.

All fourteen exploration prompts offer the exit, pinned by a
registry-parametrized test that fails when a future program is unwired.

* ci: fire PR checks on slave-based PRs, not master alone

All five gating workflows declared `pull_request: branches: [master]`,
but every fleet PR targets slave — cell->root, root->slave, and the CEO's
own. So `pull_request` never fired for any of them, and their only
coverage was the `push` trigger, which is gated on branch PREFIX
(feature/bug/chore/docs/hotfix). A branch named anything else got zero
checks — not a red run, an absent one — and a PR with no required check
present merges on a false green. PR #711 shipped that way on a `fix/`
branch.

Basing on the branch a PR merges INTO rather than what its head is named
makes coverage independent of branch naming, so a non-conforming prefix
can only ever cost the redundant push run, never the whole gate.

The same five also omitted slave from `push` (ci.yml aside, which added
it for the release gate's fail-closed CI read), so the panel suite, both
CodeQL analyses, and the e2e smoke never ran on the trunk master is cut
from.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-27 01:28:33 +02:00
a7b970a3b2 feat(board): materialize program items as Main-PM roots, make reports actionable (#711)
Two coupled gaps in the Board Program output path.

Approved items were created unowned and in BACKLOG. Nothing dispatches
BACKLOG, and once activated a cell PM claimed the parentless task as a root,
where _cell_pm_complete resolves its merge target through
resolve_parent_branch — which for a parentless task falls through to the
project head rung. The result was a cell branch merging straight into the
trunk, bypassing the Main-PM root, the root->master PR and the CEO gate
(live: PRs #703 and #704 both targeted slave directly).

All eight materializers now create a PENDING, main-pm-assigned root with
team=Team.MAIN_PM, matching what approve_and_start does for an intake draft.
The team is load-bearing, not cosmetic: _next_hint_pr_fail,
_deliver_pr_fail_to_owner, delegate's wave-chain dispatch and the PR layer
label all key on it, and a cell-teamed root drops the 'do NOT re-submit the
root' steer that exists because of PR #138's infinite pr_fail loop. The
item's own cell survives as a delegation hint in the description, which is
what the Main PM's briefing renders.

Periscope, Sentinel and Coroner produced artifacts with no way to act on
them — three panel surfaces carried explicit 'no approve/reject UI' comments
while each item already held a machine-readable suggested action. They now
have per-item approve and dismiss, modelled on the roadmap queue: idempotent
per item, CEO-gated, deep-copy-before-mutate so SQLAlchemy's dirty check
still fires, and every decision recorded through record_decision so it
reaches the next cycle's prompt. Approving materializes through the same
corrected Main-PM-owned path.

Target project resolves to each engine's own existing anchor — RoboCo's
project for Periscope and Sentinel, the incident's project for Coroner — and
fails with a clean invalid_state naming what is unresolvable rather than
guessing at a repo.

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-26 20:01:24 +02:00
roboco-app[bot]GitHubroboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>Backend Developer 1Backend Documenterroboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>Backend PM
66f0287d11 [3e7a5064] Eval harness: wire the real-spawn path so python -m roboco.eval run works end-to-end (#703)
* [431e73b7] Wire the real-spawn path: OrchestratorStageSpawner + disposable MCP config (#701)

* [431e73b7] Wire the eval harness real-spawn path: OrchestratorStageSpawner + disposable MCP config

_generate_mcp_config now prefers settings.api_url when set (both
PROJECT_HOST_PATH branches), so spawned MCP servers resolve to the
harness's disposable orchestrator URL instead of the real production
hostname or 127.0.0.1:port. OrchestratorStageSpawner.__init__ replaces
the NotImplementedError with a real AgentOrchestrator() constructed the
same way the production dispatcher builds it. The runner module
docstring + __main__.py docstring/run-subparser help drop the
NOT-YET-FUNCTIONAL wording. A new unit test pins the no-production-reach
guarantee: with settings.api_url patched, the MCP config's
ROBOCO_API_URL/ROBOCO_ORCHESTRATOR_URL point at the disposable URL (not
production), and the agent UUID is the real fixed UUID from
foundation.identity.AGENTS.

* [431e73b7] docs(eval): reflect the wired real-spawn path in tests map, CLAUDE.md, and CHANGELOG

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [5cc75f71] Fix disposable orchestrator container-reachability + document real-UUID isolation design (#705)

* [5cc75f71] Fix disposable orchestrator container-reachability + document real-UUID isolation

* [5cc75f71] docs(eval-harness): document container-reachability fix + real-UUID isolation in map docs

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [d96ec059] Rewrite stale eval-spawner pinning test to assert wired behavior (test-only, CI-green for PR #703) (#707) (#708)

* [d96ec059] test(eval): assert OrchestratorStageSpawner constructs a real AgentOrchestrator

Rewrite the stale pinning test that asserted the PRE-wiring
NotImplementedError (removed by commit 488e9e2f when the real-spawn
path was wired). The test now asserts the wired behavior:
OrchestratorStageSpawner() construction succeeds, _orchestrator is
an AgentOrchestrator instance, and _stage_timeout_seconds defaults
to 900.0. Renamed from test_orchestrator_stage_spawner_is_cut_and_
refuses_to_construct to reflect the new contract. Test-only — no
production code touched.

* [d96ec059] docs(map): note test_scoring spawner pinning test in eval-harness map entry

Add one clause to docs/map/tests.md's roboco/eval/ row naming
tests/unit/eval/test_scoring.py::test_orchestrator_stage_spawner_constructs_real_orchestrator
as the unit test that pins the OrchestratorStageSpawner wired-construction
contract (construction succeeds, _orchestrator is an AgentOrchestrator,
_stage_timeout_seconds defaults to 900.0). Consistent with the line's
existing pattern of citing test_eval_mcp_config_isolation.py and
test_eval_bench.py by name for the contracts they pin. The CHANGELOG #701
entry already covers the user-visible wiring; no CHANGELOG change needed.

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend PM <be-pm@roboco.tech>
2026-07-26 17:14:09 +00:00
roboco-app[bot]GitHubroboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>Frontend Developer 2Frontend DocumenterFrontend PM
b3dda00e41 [d700055f] Replace StubObjectivesSection with live charter objective cards (#702) (#704)
* [d700055f] feat(scorecard): replace StubObjectivesSection with live charter objective cards

Delete the StubObjectivesSection placeholder (fake 'Revenue growth'/'Customer
retention' labels) and render a real ObjectivesSection with three positional
charter objective cards, each showing its live metric against its target:
first_pass_yield (90%), median_lead_time_hours (<24h), escaped_defects (0).

- CockpitSummary: add optional first_pass_yield and escaped_defects fields
  (backend companion item not yet shipped; UI renders 'No data yet' until then)
- ObjectivesSection: positional mapping objectives[i].metric -> metric i,
  documented as a positional-by-convention assumption; canonical fallback
  labels when the charter objectives array is empty/shorter than 3; 'No data
  yet' italic muted fallback for null/undefined metrics (SpeedSection pattern);
  DeliveryMetric card styling; first_pass_yield formatted as pctOrDash does
- SpeedSection kept as-is; ObjectivesSection references the same
  median_lead_time_hours value as a peer card alongside the other two
- Tests: buildSummary carries the new fields; cover present/missing metrics
  and absence of the fake stub labels; existing lead-time/null tests adjusted
  for the now-shared value and multi-card 'No data yet'

* [d700055f] docs(scorecard): document live ObjectivesSection and new CockpitSummary fields

Add panel/docs/frontend/company-scorecard-card.md covering the new
ObjectivesSection: the three positional charter objective cards
(first_pass_yield 90%, median_lead_time_hours <24h, escaped_defects 0),
the positional-by-convention mapping, canonical fallback labels, the
'No data yet' fallback pattern, and the two new optional CockpitSummary
fields with the backend companion-item caveat. Add a Key Symbols row for
CompanyScorecardCard/ObjectivesSection in docs/map/panel.md.

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend PM <fe-pm@roboco.tech>
2026-07-26 17:13:22 +00:00
80ccf415cb feat(cockpit): expose first_pass_yield and a real escaped-defects metric (#709)
The Company Scorecard renders three charter objectives but the cockpit
summary only ever carried one of the metrics, so two cards read "No data
yet" permanently.

first_pass_yield is a pass-through — MetricsService.get_org_scorecard()
already computes it on the same 30d/org scope the rest of the delivery block
uses, and CockpitService.summary simply never forwarded it.

escaped_defects is new. The obvious definition — a blocker finding opened on
a task that already reached a terminal state — is unimplementable: every
producer of a task_review_findings row fires as part of a bounce whose
transition requires a non-terminal task, so it would read zero forever, and a
permanently-green card is the same fabrication the panel change removes.

What it counts instead: a blocker still at 'addressed', never 'verified', on
a task that has since completed. That is reachable because
stamp_addressed_verified only bulk-verifies rows matching its OWN origin, so
a blocker raised by one origin and never re-confirmed by that origin survives
to completion on the developer's word alone.

docs/map/metrics-observability.md documents what a zero actually means: the
one reachable trigger is a PM-origin blocker on a task escalated to the CEO
rather than completed by the PM, since escalate_to_ceo carries no
findings-resolved precondition and ceo_approve verifies only ceo-origin rows.
It also records that the count is per-finding over a rolling 30-day window,
which is not the same unit as the charter's "per release".

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-26 19:11:27 +02:00
0fe21b1f97 fix(ci): exempt roboco-app[bot] from the CLA check (#710)
The App the fleet pushes and opens PRs under authors the sync/merge commits
GitService creates when a task branch is brought up to date with its base, so
it is a committer on essentially every fleet PR. Unlike the agent identities
— whose roboco.tech emails map to no GitHub account, so CLA Assistant matches
them by the display names already allowlisted — the App maps to a real
account, and CLA Assistant demands a signature it cannot give: an App cannot
post the sign-off comment as itself.

The result is a permanently red `cla` check on fleet PRs (currently #703 and
#704) that no amount of rework can clear, which sends the PR reviewer round
another revision loop with nothing to fix.

Exempting it is also correct on the merits: the CLA exists to obtain
copyright assignment from human contributors, and the App commits on the
copyright holder's own behalf. Same form as the dependabot[bot] entry.

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-26 19:10:06 +02:00
a36a180b73 fix(orchestrator): let a failed Board Program exploration retry (#706)
`_board_dispatched` is an in-memory, never-expiring set of
(agent_slug, task_id). The solo exploration dispatchers consulted it, so an
exploration whose propose verb rejected was never retried for the life of the
process — Periscope, Sentinel, Scales and Barfly each spawned once on
2026-07-25, failed, and sat PENDING until the stack restarted.

The guard was written for the two-reviewer board REVIEW pass, where it is
correct: a reviewer has no verb to advance the task, so a respawn can only
loop. An explorer is the opposite — `propose_*` is exactly such a verb, so a
respawn can and should advance it.

Drop it from the 15 `_dispatch_*_exploration` functions. Bounding falls to
`_pm_respawn_should_gate`, which is what actually bounds a loop: DB-persisted,
reset by a status change, and cooled down so a deploy that fixes the cause
lets the work resume. `_dispatch_board_reviewer` keeps the set (its
`_board_review_complete` reads it as a has-run signal) as does vault
curation (same-process race guard behind its own durable marker).

The 14 `*_dispatch_is_one_shot` tests asserted the old contract on the false
rationale that board roles have no progression verb; they now assert a second
tick re-attempts.

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-26 15:52:33 +02:00
879afc14a4 Board Program LEARN context, ruff 0.16, and verb-rejection observability (#700)
* fix(board): LEARN decisions name the item, not its per-cycle index

A cycle's reject reasons are rendered into the NEXT cycle's exploration
prompt, but the ref recorded alongside each reason was the item's stored
id (item-0/item-1) — a per-cycle index that means something different
every cycle and appears nowhere the explorer can resolve. The reason
survived the loop; what it was about did not.

Record the item's title instead, via a shared learn_ref() helper (falls
back to the id when title-less, and reads target_task_title for Scales,
whose items name the live task they mutate).

* chore(lint): satisfy ruff 0.16 — keyword-only signatures and markdown formatting

The dev toolchain resolved ruff 0.16.0, which stabilises PLR0917 (too many
positional arguments) and formats python code blocks inside markdown. Both
fired repo-wide and neither had anything to do with the code they flagged.

- 36 signatures gain a `*` so their tail arguments are keyword-only, and
  the 104 call sites that passed them positionally are converted. mypy was
  the safety net for the static ones; the full suite caught nine more that
  only bind at runtime (the MCP tool functions, whose real callers already
  pass named JSON arguments).
- 28 markdown files reformatted by 0.16's code-block formatter.
- One RUF036 (`None` mid-union) autofixed in the GitLab provider.

* fix(gateway): log the reason when a verb rejects

A rejected envelope rides an HTTP 200, its body is never logged, and there
is no trace table — so in the access log a verb an agent could not satisfy
looks identical to one that worked. On 2026-07-25 four Board Programs
(Periscope, Sentinel, Scales, Barfly) each POSTed their propose verb three
or four times, persisted nothing, and left their exploration tasks PENDING;
the reason was unrecoverable afterwards, from the logs or from the agents'
own transcripts.

Log error/message/remediate/missing plus the calling agent at
envelope_to_response — the one chokepoint every v1 flow and do route
returns through. Success envelopes stay silent.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-26 15:07:28 +02:00
401f8a2cc9 feat(board): Board Programs — the complete twelve-program catalog (Phases 1-3) (#699)
* feat(board): Pest Control — the first project-scoped Board Program

The Product Owner hunts latent defects (what the org records but nobody
reads): a weekly cycle — accelerated off-schedule when the trailing-7-day
rework rate crosses pest_rework_threshold, with the cheap dedup/scope gates
evaluated before the metrics queries — opens one held exploration task
against the least-recently-explored opted-in project (deterministic
round-robin; opted_in_projects gains a stable ORDER BY), with server-
assembled evidence in the spawn prompt (rework hotspots, recurring-findings
and waived-minor ledger aggregates, all capped) plus prior-cycle LEARN
context. The PO calls the new PO-only propose_bug_hunt verb once: ≤5 items,
evidence required per item, targets validated against pest_control
participation. CEO decides per item — approve materializes a BACKLOG task
(source pest_control, never auto-starts), reject records the reason; both
feed the LEARN ledger by exploration task id; all-terminal completes the
cycle. Telegram queue pushes carry working Approve/Reject handlers
mirroring the roadmap kind. Doctrine: board.md Pest Control section +
product-owner verb entry + regenerated verb tables.

* feat(panel): Pest Control review queue

Command Center gains the pest review queue (per-item approve/reject with
reason, mirroring the roadmap queue); the Programs card and the project
settings participates-in checkboxes pick the new program up registry-driven
— the settings section renders for the first time now that a project-scoped
program exists.

* feat(board): Periscope — HoM market-research brief program

Weekly org-scoped cycle: a solo HoM spawn researches the market (web
research with mandatory source URLs — uncited findings are rejected) and
files one structured brief via the new HoM-only propose_market_brief verb:
headline, cited findings, threats/opportunities, positioning note, all
soup-checked and screened through the injection guard at persist time
(web-derived text later reaches prompts; flags recorded, content never
dropped). A brief is a report, not a proposal: the verb completes the
exploration in the same call (the x_feature asymmetry), the cycle ledger
auto-closes, and the CEO gets a best-effort notification with no
approve/reject surface (periscope deliberately never joins Telegram's
action kinds). The latest brief is injected into the roadmap exploration
prompt — Periscope feeds Printer, the first cross-role program input.

* feat(panel): Market Briefs tab (read-only)

Business page gains a Market Briefs tab listing Periscope briefs —
headline, cited findings, threats/opportunities — read-only by design; a
report has nothing to approve.

* feat(board): Coroner — event-triggered Auditor postmortems

The first EVENT program: no cron — three best-effort hooks open an autopsy
when a task bounces to its 3rd revision (the audit chokepoint), is
cancelled after work started, or is budget-blocked; all gated on arming +
one-open-autopsy dedup, none can fail the underlying transition. A solo
Auditor spawn reads the incident (server-assembled findings + transition
context) and files one propose_postmortem: incident summary, root cause,
failed stage (validated against the real status vocabulary), and ONE
process change — a playbook-kind change drafts via PlaybookService
directly into the normal pending-curation queue; the briefed draft_playbook
manifest grant was deliberately NOT added, preserving the existing
'auditor curates but never drafts' invariant test. Complete-at-propose
(report asymmetry), cycle ledger auto-closes, CEO notified link-only.
Integrated as a union with Periscope across the shared program surfaces.

* feat(panel): Coroner postmortems card

Read-only postmortems list under Business → Programs — incident, root
cause, failed stage, process change; nothing to approve, the process-change
artifact (a draft playbook) rides the existing curation queue.

* feat(board): Sentinel — Auditor drift-watch quality reports

Weekly org-scoped cycle: a solo Auditor spawn receives a server-assembled
drift context (waived-findings trend, open findings by severity,
conventions-violation hotspots, top spend — all capped, pure ORM) and files
one propose_quality_report: headline, 1-7 area-validated items with
evidence and suggested actions, overall assessment. Report semantics —
complete-at-propose, cycle auto-closes, CEO notified display-only (never on
Telegram's approve/reject surface); items are structured so a later
convert-to-task control is cheap. Integration adopts Sentinel's module-
level dict-dispatch for board-program routing (xenon-driven), folding all
prior programs in; app router mounting extracted to a helper for the same
budget.

* feat(panel): Quality Reports tab (read-only)

Business page gains the Sentinel quality-reports tab — headline, per-area
observations with evidence and suggested actions; read-only, a report has
nothing to approve.

* feat(board): Spackle — gap-fill audit program

Biweekly project-scoped PO cycle over the half-shipped surface area: API
routes without panel surfaces (and vice versa), armed flags without docs,
docs promises the code doesn't keep, dead-end tabs — the inventory diffing
is the PO's own read-tool work, ordered by the spawn prompt with file:line
citations required; the server injects only prior-cycle LEARN and the
rotation target. Rotation is now a shared module-level helper
(pick_rotation_target, parameterized by source) both project-scoped
engines use — pest_control delegates to it, behavior-identical, with a
cross-pollution test proving the two programs' rotations stay independent.
propose_gap_fill mirrors the bug-hunt verb (≤5 items, two-sided evidence
required, participation gate); per-item CEO decide materializes BACKLOG
source=spackle tasks; full Telegram kind incl. approve/reject handlers.
All seven program routers now mount from one helper.

* feat(panel): Spackle gap-fill review queue

Command Center gains the gap-fill queue mirroring the pest-control one —
per-item approve/reject with the two-sided gap evidence rendered.

* feat(board): Scales — monthly portfolio rebalance

Org-scoped PO cycle over the stale backlog: the spawn receives a capped
stale-task snapshot (BACKLOG/PENDING unclaimed >30 days) plus the charter
and prior-cycle LEARN, and files one propose_rebalance — 1-7 items, each a
resolvable task_ref with action reprioritize (validated new priority) or
cancel, rationale required. Per-item CEO decide: approve EXECUTES the
action (audited priority update, or the normal cancel path) — the first
program whose materializer mutates existing tasks instead of creating
them; reject records the reason; LEARN by exploration task id;
all-terminal completes the cycle. Full Telegram decide-kind wiring.
Integrated as the eight-program union (registry, dict dispatch, routers
helper, teardown enumerations).

* feat(panel): Scales rebalance review queue

Command Center gains the rebalance queue — per-item approve/reject with
the action, target task, and rationale rendered.

* feat(board): Mirror — quarterly positioning audit

Project-scoped HoM cycle over messaging surfaces: README claims vs shipped
reality, docs-site promises vs code, charter alignment — the audit is the
HoM's own read-tool work with citations required; the server injects the
charter, prior-cycle LEARN, and the shared rotation target. propose_
messaging_fixes mirrors the gap-fill verb (≤5 items, drift evidence naming
claim + contradicting reality, participation gate); per-item CEO decide
materializes BACKLOG source=mirror documentation tasks; full Telegram
decide-kind wiring. Nine-program union across the shared surfaces.

* feat(panel): Mirror messaging-fixes review queue

* feat(board): Megaphone — HoM standing editorial calendar

Cron cycle (3 days, org-scoped, gated on X credentials — drafting content
nobody can post is pointless): the HoM receives a shipped-this-week digest
plus Unreleased changelog bullets and files one propose_editorial_post
(angle-validated, ≤280, brand voice) that materializes a held x_editorial
draft through the SAME X-queue origination chokepoint release posts use —
zero new approval surface, notifications and CEO decide for free.
Complete-at-propose; cycle auto-closes. Ten-program union.

* feat(panel): x_editorial source labels in the X queue surfaces

* feat(board): Librarian — proactive playbook mining

Biweekly org-scoped Auditor cycle: mines recurring non-private learning
journals (≥2-count grouping with a recency fallback) against the existing
playbook-title inventory and files one propose_playbook_drafts — 1-3
drafts, each with the repeated-pattern evidence that justifies it,
duplicate titles rejected in-batch and against the live store. Drafts are
created via PlaybookService directly (the Coroner precedent — the
'auditor curates but never drafts' do-verb invariant stays intact and
tested) and land in the normal pending-curation queue the Auditor's own
triage already surfaces; no new panel surface. Complete-at-propose;
display-only CEO notification. Eleven-program union.

* feat(board): War Room — release campaign planning

EVENT program with a REAL originator (unlike coroner's stub): a release
publish hooks a campaign brief beside the release-post seam, and the CEO's
run-now originates on demand — the cron loop never fires it. The HoM
designs a 2-6 post arc (teaser → launch → follow-up → spotlight; 280-cap,
future strictly-ascending publish_after, stage vocabulary) and one
propose_campaign call materializes each post as a held x_campaign draft
through the X-queue chokepoint. V1 is manual-cadence by design: publish_
after renders as queue guidance and the CEO approves each post at its
moment — nothing auto-posts, ever; the auto-schedule upgrade is a
documented ceiling. Twelve-program union: full registry complete.

* feat(panel): x_campaign labels + publish-after guidance in the X queue

* feat(board): Barfly — adjacent-conversation replies

Cron cycle (2 days, org-scoped, X-credentials gated): the engine searches
X for conversations where RoboCo is relevant but unmentioned (new OAuth-
signed search_recent on the client; queries + candidate cap configurable),
screens every fetched tweet through the injection guard (stored unclamped
— a clamp was truncating the candidate under the envelope, caught by the
dev's own tests), dedupes via the existing x_seen_mentions ledger (no
migration; also prevents double-drafting against the mentions poll), and
opens one held HoM exploration carrying the screened candidates. propose_
conversation_replies enforces candidate-id-only replies (≤5, 280-cap);
each materializes a held x_barfly draft through the X-queue chokepoint,
threaded via a new in_reply_to seam on post_tweet that only x_barfly
drafts use. The X redraft machinery is now dict-dispatch over per-source
extractors with reply-ref carry for x_barfly. Thirteen-program registry.
War Room's test fakes gained the new abstract search_recent stub.

* feat(board): Dogfood — the PO walks the product

The fourteenth and final registry entry, completing the catalog. EVENT
program (release-publish hook beside the war-room hook + CEO run-now, both
through the same real originator; the cron loop never fires it), project-
scoped with shared rotation. The permission surface is the careful part:
the PO's dogfood spawn — and ONLY that spawn — gets the Playwright MCP
mounted, via a task-scoped fail-closed probe mirroring the video-authoring
precedent (a PO spawned for roadmap/pest/scales never sees browser tools;
tested both ways); the PM agent image bakes chromium unconditionally like
the ux image, the mount stays task-gated in code. The walk targets the
rotation target's live surfaces (panel_base_url only when the target is
the org's own project, honest degradation otherwise); propose_friction_
fixes files ≤5 walked-path-evidenced items; per-item CEO decide
materializes BACKLOG source=dogfood tasks; full Telegram decide kind.
Also: megaphone/librarian/war_room arming keys restored to the settings
validator — their panel toggles would have been rejected (dropped in
earlier unions; the same silent-arming class the drill killed once
already).

* feat(panel): Dogfood friction review queue

* chore(board): final whole-branch sweep fixes

The night's closing adversarial pass over the integrated fourteen-program
registry found ONE functional defect — the war-room test fakes' post_tweet
predated Barfly's in_reply_to_tweet_id kwarg (LSP violation, the only red
in an otherwise fully green gate) — plus doc/test drift, all fixed: the
source-parity test completes to fourteen (spackle/mirror were silently
absent while its neighboring comment claimed full coverage), the PO
identity doc gains its missing Dogfood verb, the auditor quick-list gains
propose_postmortem, three stale comments corrected (rotation docstring,
panel registry header, X source enumerations), the dogfood release-hook
gains the exception-swallow test its four sibling hooks already had, and
the CHANGELOG's Unreleased section documents the whole Board Programs
train. Full make quality: exit 0, all gates green.

* docs: full documentation sweep for the Board Programs train

CLAUDE.md's roadmap-engine entry superseded by the Board Program registry
entry (all fourteen programs, arming, scoping, LEARN, guardrails) with the
role verb tables and playwright row refreshed; docs/rag gains the agent-
facing architecture doc plus full propose_* call-shape sections in the
three board role docs, and corrects the strategy-engine section to shipped
reality (only idle→roadmap is wired); docs/map covers the registry + all
twelve engines with flags, gotchas, and drift notes. The 0.27.0 reference
inventory confirmed only the release-executor's canonical set carries the
version — left for the 0.28.0 cut.

* feat(board): human titles + descriptions on every program surface

Raw registry keys rendered as bare panel labels — an operator reading
x_feature had no idea what enabling or running it does. The registry
dataclass gains title/description (test-enforced non-empty for every
entry, unique titles), the API passes them through, and every surface
renders title-with-description-tooltip instead of the key: the Programs
card (label, toggle hint, run-now toast), and the project settings
participates-in/excluded-from checkboxes.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-25 17:13:32 +02:00
e77c3b7a63 feat(board): Board Program registry — Phase 1 (engine, LEARN ledger, per-project scoping, panel) (#689)
* feat(board): Board Program registry — generic trigger/dedup/originate/LEARN engine

One registry (foundation/policy/board_programs.py) + one BoardProgramEngine +
one orchestrator loop replace the bespoke roadmap/spotlight loops, behavior-
preserved: same sources, dispatch routing, one-open-cycle dedup (ledger rows
auto-close when their exploration task goes terminal, so x_feature's
complete-at-propose flow can't wedge), and live per-program interval
overrides with the tick capped at 1h.

program_armed() is the single arming chokepoint: the settings-store
board_program.<key>.enabled override when present, else the legacy flag —
routed through BoardProgramEngine, RoadmapEngine.run_cycle, and XEngine's
spotlight gate, so the panel toggle can never be a silent no-op against a
legacy boot flag.

LEARN: board_program_cycles (migration 087) accrues per-item CEO decisions
(exact attribution by exploration_task_id where the caller holds it) and
feeds the last closed cycles back into both exploration prompts. The
strategy engine's idle signal now opens a roadmap cycle (enabled+dedup
respected) instead of only nudging.

Per-project scoping (migration 088, projects.board_programs, dual polarity):
plain keys opt a project INTO project-scoped programs; "!key" opts it OUT
of an org-scoped program's outputs (default eligible — parity). Enforced at
propose_roadmap (names the excluded project) and defensively at materialize;
validation rejects unknown keys and meaningless polarity both directions.

API: GET /api/board-programs + POST /api/board-programs/{key}/run-now
(CEO-gated); settings keys for both migrated programs.

* feat(panel): Board Programs card + per-project program controls

Business page gains a Programs tab: per-program rows (role, trigger, scope,
open-cycle badge), enabled switch on the settings-store key, Run now
(disabled while a cycle is open). The edit-project dialog gains the
program controls next to the CI-watch/video toggles: participates-in
checkboxes for project-scoped programs, excluded-from checkboxes for
org-scoped outputs.

* test(board): full-gate hermeticity — mypy casts + shared-DB purge fixtures

make quality runs one pytest process over all suites against the shared
persistent DB: integration collects before unit, so the board-programs API
test's committed run-now state (settings-store overrides, an open cycle row,
its board_roadmap task) poisoned 13 downstream unit tests that pass in
isolation. The polluter now purges its own committed state in fixture
teardown, and the four consumer files get an autouse per-test purge
(board_program.% settings keys, ledger rows, open exploration tasks) so
they are hermetic regardless of collection order. Also the four
cast("UUID", ...) sites the tests-scope mypy run requires.

* feat(panel): re-home per-project program controls onto the settings page

Wave C deleted the edit-project dialog these controls originally landed in;
they now live on the project settings page's budget/ops card next to the
CI-watch/video toggles — participates-in switches for project-scoped
programs, excluded-from switches for org-scoped outputs, dual-polarity
tooltips, order-independent dirty tracking. Nine makeProject test fixtures
gain the required board_programs field the rebase left behind.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-25 05:39:44 +02:00
cbddfc7cb3 [c71959c3] Video: release 0.27.0 (#698)
* [c71959c3] feat(motion): add release-0.27.0 panel-demo clip

Extends the panel-demo kit register (kit/) with a v0.27.0 release
composition: identity cold open, four feature cards (WAF real-IP
resolution, Telegram Mini App V6, Codex/Gemini providers, cost-tiered
routing) flipping from in-progress to completed, a camera that pushes
per beat via pk-camera data-shots, a cursor that clicks and travels
via pk-cursor data-waypoints, a receipt stats overlay, a shipped
toast, and the roboco.tech outro. Ships vertical.html, square.html,
props.js, captions.json, and a vitest smoke test mirroring
release-0.26.0's invariants.

* [c71959c3] fix(motion): implement genuinely non-uniform card-beat gaps in release-0.27.0 clip

The four .pk-card animation-delay values were uniform 5.0s/10.0s/15.0s/20.0s
in both vertical.html and square.html -- byte-for-byte identical to
release-0.26.0's metronomic spacing the craft bar bans -- despite the prior
decision_log claiming 5.5s/6.4s/5.3s varied gaps had been added. Card delays
are now 5.0/10.5/16.9/22.2 (real 5.5s/6.4s/5.3s gaps). Camera data-shots,
cursor data-waypoints, and the stats/toast/outro tail in both files are
re-warped with a shared piecewise-linear time function anchored at each
card's new beat, so nothing desyncs and the tail still fits inside the
fixed 40s runtime. Added a regression test asserting the four card gaps
are not all identical.

* [c71959c3] docs(motion): add release-0.27.0 composition section to README

---------

Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
2026-07-25 03:39:07 +00:00
Renn F 069bb621f4 docs(changelog): the x-drafting fix bullet belongs under Unreleased, not 0.27.0 2026-07-25 03:05:13 +02:00
Renn F 037114338a fix(x): release-post drafts stop parroting changelog bullets
Highlights reorder to marketing order (Added/Changed before Fixed/Security
— the drafting model anchors on highlight #1 and the changelog opens with
Security), the prompt bans verbatim highlight copying and internal plumbing
jargon, and the deterministic fallback becomes a generic announcement that
can never quote a raw bullet.
2026-07-25 03:04:35 +02:00
roboco-app[bot]andGitHub 29243771e4 sync: master → slave 2026-07-25 01:02:49 +00:00
RoboCo Release Manager 18eaacb9ea chore(release): 0.27.0 v0.27.0 2026-07-25 00:38:57 +00:00