Board Program LEARN context, ruff 0.16, and verb-rejection observability (#700)

* fix(board): LEARN decisions name the item, not its per-cycle index

A cycle's reject reasons are rendered into the NEXT cycle's exploration
prompt, but the ref recorded alongside each reason was the item's stored
id (item-0/item-1) — a per-cycle index that means something different
every cycle and appears nowhere the explorer can resolve. The reason
survived the loop; what it was about did not.

Record the item's title instead, via a shared learn_ref() helper (falls
back to the id when title-less, and reads target_task_title for Scales,
whose items name the live task they mutate).

* chore(lint): satisfy ruff 0.16 — keyword-only signatures and markdown formatting

The dev toolchain resolved ruff 0.16.0, which stabilises PLR0917 (too many
positional arguments) and formats python code blocks inside markdown. Both
fired repo-wide and neither had anything to do with the code they flagged.

- 36 signatures gain a `*` so their tail arguments are keyword-only, and
  the 104 call sites that passed them positionally are converted. mypy was
  the safety net for the static ones; the full suite caught nine more that
  only bind at runtime (the MCP tool functions, whose real callers already
  pass named JSON arguments).
- 28 markdown files reformatted by 0.16's code-block formatter.
- One RUF036 (`None` mid-union) autofixed in the GitLab provider.

* fix(gateway): log the reason when a verb rejects

A rejected envelope rides an HTTP 200, its body is never logged, and there
is no trace table — so in the access log a verb an agent could not satisfy
looks identical to one that worked. On 2026-07-25 four Board Programs
(Periscope, Sentinel, Scales, Barfly) each POSTed their propose verb three
or four times, persisted nothing, and left their exploration tasks PENDING;
the reason was unrecoverable afterwards, from the logs or from the agents'
own transcripts.

Log error/message/remediate/missing plus the calling agent at
envelope_to_response — the one chokepoint every v1 flow and do route
returns through. Success envelopes stay silent.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-07-26 15:07:28 +02:00
committed by GitHub
co-authored by Renn F
parent 401f8a2cc9
commit 879afc14a4
84 changed files with 932 additions and 567 deletions
+7 -7
View File
@@ -6,10 +6,10 @@ A2A is direct peer-to-peer messaging between agents. There is **no** `roboco_age
```python
dm(
recipient="be-qa", # target agent slug
recipient="be-qa", # target agent slug
text="Please review my changes",
task_id="abc123...", # auto-filled from your active task if omitted
skill=None, # optional skill slug to scope the conversation
task_id="abc123...", # auto-filled from your active task if omitted
skill=None, # optional skill slug to scope the conversation
)
```
@@ -37,7 +37,7 @@ Same-cell peers (e.g. `be-dev-1` alongside `be-dev-2`/`be-qa`/`be-doc`/`be-pm`)
When another agent messages you, your claim briefing surfaces it under `unread_a2a` — each entry shows the sender and a preview of their latest incoming message. To read the full bodies (and clear them), call:
```python
read_a2a() # -> {"messages": [{from_agent, content, created_at}, ...]}
read_a2a() # -> {"messages": [{from_agent, content, created_at}, ...]}
```
`read_a2a()` returns only INCOMING messages (never your own sends) and marks them read. `read_messages()` is the lighter variant that only zeroes the unread counter without returning content — reach for `read_a2a()` when you actually need to see what was said. Either clears `i_am_idle()`'s unread-A2A soft-block.
@@ -45,9 +45,9 @@ read_a2a() # -> {"messages": [{from_agent, content, created_at}, ...]}
Formal, ack-required notifications are a separate inbox — see `docs/rag/tools/messaging-tools.md`:
```python
notify_list(unread_only=True) # list pending items
notify_get(notification_id) # read one (marks it read)
notify_ack(notification_id) # acknowledge after handling
notify_list(unread_only=True) # list pending items
notify_get(notification_id) # read one (marks it read)
notify_ack(notification_id) # acknowledge after handling
```
## When to use A2A
+22 -38
View File
@@ -16,31 +16,24 @@ roboco_kb_search(
query="rate limiting redis",
top_k=5,
project="roboco-api",
index_types=["code", "docs"]
index_types=["code", "docs"],
)
```
## AI-Generated Answers
```python
roboco_rag_query(
query="How does authentication work?",
top_k=5
)
roboco_rag_query(query="How does authentication work?", top_k=5)
```
## Mentor (Conversational)
```python
response = roboco_ask_mentor(
question="How do I handle auth?",
domain="coding"
)
response = roboco_ask_mentor(question="How do I handle auth?", domain="coding")
# Follow-up
roboco_ask_mentor(
question="What about refresh tokens?",
conversation_id=response["conversation_id"]
question="What about refresh tokens?", conversation_id=response["conversation_id"]
)
```
@@ -48,13 +41,15 @@ roboco_ask_mentor(
```python
# Write/update documentation (auto-dedup via RAG)
roboco_docs_write({
"task_id": "task-uuid",
"filename": "api-endpoints.md",
"doc_type": "api", # api, qa, guide, readme, changelog, architecture, design
"title": "API Endpoints",
"content": "# API Endpoints\n\n..."
})
roboco_docs_write(
{
"task_id": "task-uuid",
"filename": "api-endpoints.md",
"doc_type": "api", # api, qa, guide, readme, changelog, architecture, design
"title": "API Endpoints",
"content": "# API Endpoints\n\n...",
}
)
# List docs for a task
roboco_docs_list(task_id="task-uuid")
@@ -69,33 +64,24 @@ roboco_docs_read(path="backend/api/endpoints.md")
```python
# Index code (PM, Developer)
roboco_kb_index_code(
sources=["src/**/*.py"],
project="roboco-api"
)
roboco_kb_index_code(sources=["src/**/*.py"], project="roboco-api")
# Index docs (PM, Documenter) - for bulk/explicit indexing
# Note: roboco_docs_write() auto-indexes when writing
roboco_kb_index_docs(
sources=["docs/**/*.md"],
project="roboco-api"
)
roboco_kb_index_docs(sources=["docs/**/*.md"], project="roboco-api")
```
## Error Tracking
```python
# Search for similar errors
roboco_search_error(
error_message="Redis connection timed out",
context="startup"
)
roboco_search_error(error_message="Redis connection timed out", context="startup")
# Record solution
roboco_record_error_solution(
error_message="Redis connection timed out",
solution="Added retry with backoff",
worked=True
worked=True,
)
```
@@ -106,11 +92,9 @@ roboco_record_error_solution(
roboco_check_decision(topic="session storage")
# Record decision
roboco_record_decision(params={
topic: "Session storage",
decision: "Use Redis",
rationale: "Sub-ms reads"
})
roboco_record_decision(
params={topic: "Session storage", decision: "Use Redis", rationale: "Sub-ms reads"}
)
```
## Standards & Validation
@@ -135,7 +119,7 @@ def create_user(email, password):
user = User(email=email, password=password)
db.add(user)
return user
"""
""",
)
```
@@ -172,7 +156,7 @@ def create_user(email, password):
roboco_review_code(
code="def handle(...):",
file_path="src/api/auth.py",
change_type="modify" # add, modify, delete
change_type="modify", # add, modify, delete
)
```
+3 -3
View File
@@ -19,9 +19,9 @@ notify(target="be-dev-1", text="Task ready for you", priority="normal", task_id=
Every role with an inbox gets these (so `i_am_idle()` doesn't soft-block on unread items):
```python
notify_list(unread_only=True, limit=20) # your inbox
notify_get(notification_id) # read one (marks it read)
notify_ack(notification_id) # acknowledge after handling
notify_list(unread_only=True, limit=20) # your inbox
notify_get(notification_id) # read one (marks it read)
notify_ack(notification_id) # acknowledge after handling
```
When `i_am_idle()` reports unread A2A or @mentions, clear A2A with `read_a2a()` (see `a2a-tools.md`) and clear notifications with list -> get -> ack, then idle again. (The Auditor gets `notify_list`/`notify_get` for inbox visibility but does not ack.)
+5 -2
View File
@@ -31,8 +31,11 @@ There is **no** `roboco_git_commit / _push / _checkout / _create_pr / _merge_pr`
To learn how a project's codebase is laid out or how a subsystem works, query the knowledge base rather than a project tool:
```python
roboco_kb_search(query="rate limiting redis", project="roboco-api",
index_types=["code", "documentation"])
roboco_kb_search(
query="rate limiting redis",
project="roboco-api",
index_types=["code", "documentation"],
)
roboco_ask_mentor(question="How is auth wired up in this project?")
```
+67 -58
View File
@@ -7,20 +7,20 @@ The verbs below are grouped by who calls them.
## Developer flow
```python
give_me_work() # returns your most-actionable pending task
give_me_work() # returns your most-actionable pending task
i_will_work_on(task_id, plan="...")
# claims + sets plan + starts; auto-creates and
# checks out feature/{team}/{task-hierarchy}
commit(message, files=None) # content tool — repeat per change (auto-pushed)
open_pr(task_id) # pushes branch + opens the PR
# claims + sets plan + starts; auto-creates and
# checks out feature/{team}/{task-hierarchy}
commit(message, files=None) # content tool — repeat per change (auto-pushed)
open_pr(task_id) # pushes branch + opens the PR
i_am_done(task_id, notes="", resolved_findings=None)
# verifying -> awaiting_qa (PR must already be open);
# on a bounced task, name every open ledger finding
# via resolved_findings=[{finding_id, commit?, note?}]
i_am_blocked(task_id, reason) # external dependency; cell PM unblocks
unclaim(task_id) # release a claimed task back to the queue
resume(task_id) # recover a paused task after compact/restart
i_am_idle() # no work in your queue right now
# verifying -> awaiting_qa (PR must already be open);
# on a bounced task, name every open ledger finding
# via resolved_findings=[{finding_id, commit?, note?}]
i_am_blocked(task_id, reason) # external dependency; cell PM unblocks
unclaim(task_id) # release a claimed task back to the queue
resume(task_id) # recover a paused task after compact/restart
i_am_idle() # no work in your queue right now
```
There is no separate claim / start / pause verb — `i_will_work_on` composes claim + set-plan + start atomically, and `i_am_done` composes verify + submit-qa. Branches are auto-created on `i_will_work_on`; do not checkout by hand — every root task branches from the project's env-ladder **head rung**, not a hardcoded `default_branch`/`master` string (see `CLAUDE.md` "Env-branches ladder"; a project with no declared ladder resolves this identically to its `default_branch`, so nothing changes unless the project opted in).
@@ -48,11 +48,11 @@ The callable MCP tool names are `pass` / `fail` (`pass`/`fail` are reserved word
## Documenter flow
```python
give_me_work() # returns an awaiting_documentation task
claim_doc_task(task_id) # claim the doc phase
commit(message, files) # commit the doc files you write
give_me_work() # returns an awaiting_documentation task
claim_doc_task(task_id) # claim the doc phase
commit(message, files) # commit the doc files you write
i_documented(task_id, notes, files)
# awaiting_documentation -> awaiting_pm_review
# awaiting_documentation -> awaiting_pm_review
```
Documentation tasks are **not** delegated — the lifecycle auto-creates the doc phase after a code task passes QA.
@@ -60,33 +60,42 @@ Documentation tasks are **not** delegated — the lifecycle auto-creates the doc
## Cell PM flow
```python
triage() # list actionable tasks in your cell
triage() # list actionable tasks in your cell
i_will_plan(task_id, plan, approach)
# claim + plan + start a parent task
delegate(parent_task_id, title, description, assigned_to, team, task_type,
nature, estimated_complexity, acceptance_criteria,
covers_parent_criteria=[...])
# create a subtask; covers_parent_criteria maps
# it to the parent ACs it is responsible for —
# REQUIRED whenever the parent has any acceptance
# criteria (a ref that matches neither an AC id
# nor exact text is rejected, naming the valid
# criteria); omit only when the parent has none
# claim + plan + start a parent task
delegate(
parent_task_id,
title,
description,
assigned_to,
team,
task_type,
nature,
estimated_complexity,
acceptance_criteria,
covers_parent_criteria=[...],
)
# create a subtask; covers_parent_criteria maps
# it to the parent ACs it is responsible for —
# REQUIRED whenever the parent has any acceptance
# criteria (a ref that matches neither an AC id
# nor exact text is rejected, naming the valid
# criteria); omit only when the parent has none
reassign(task_id, assigned_to) # move a subtask to a different agent
unblock(task_id, reason) # blocked -> in_progress (PM only); reason is
# recorded as your journal:decision (no separate
# note needed)
unblock(task_id, reason) # blocked -> in_progress (PM only); reason is
# recorded as your journal:decision (no separate
# note needed)
submit_up(task_id, notes, resolved_findings=None)
# open cell->root PR; -> awaiting_pr_review
# (the cell PR reviewer gates it; after pr_pass
# the same Cell PM completes + merges); a re-submit
# after pr_fail must resolve every open finding first
complete(task_id, notes) # awaiting_pm_review -> completed (merges leaf PR)
# open cell->root PR; -> awaiting_pr_review
# (the cell PR reviewer gates it; after pr_pass
# the same Cell PM completes + merges); a re-submit
# after pr_fail must resolve every open finding first
complete(task_id, notes) # awaiting_pm_review -> completed (merges leaf PR)
request_changes(task_id, findings=[...])
# reject a subtask's merge review -> needs_revision,
# routed to whoever owns the revision; structured
# findings persist to the ledger + render into pm_notes
escalate_up(task_id, reason) # escalate to your escalation target
# reject a subtask's merge review -> needs_revision,
# routed to whoever owns the revision; structured
# findings persist to the ledger + render into pm_notes
escalate_up(task_id, reason) # escalate to your escalation target
```
After `i_will_plan` and each `delegate`, the envelope includes a coverage view of the parent — `parent_ac_coverage` (per-criterion `id` / `text` / `claimed` / `verified`) and `unclaimed_parent_acs` (criteria no subtask covers yet). A parent cannot idle with unclaimed criteria, nor `complete` / `submit_up` / `escalate_to_ceo` until every criterion traces to a child that passed QA. `delegate` refusing a child with no `covers_parent_criteria` (above) is what puts every parent with acceptance criteria under this coverage discipline from its first subtask on — a decomposition can no longer opt out by never declaring. A rejection now includes a copy-pasteable corrected `delegate(...)` skeleton with the parent's real criteria inlined (an id when the parent has one, its exact quoted text otherwise) — retry with that shape verbatim rather than re-deriving the field's syntax. `i_will_plan`'s planning briefing also carries `collision_context` (in `context_briefing`, not `evidence`) surfacing any same-parent siblings that already collide on file globs or migrations, so you can sequence your delegation before you commit to it. See `docs/rag/workflows/task-planning.md`.
@@ -98,15 +107,15 @@ After `i_will_plan` and each `delegate`, the envelope includes a coverage view o
The Main PM shares most Cell PM verbs (`i_will_plan`, `delegate`, `complete`, `request_changes`, `unblock`, `triage`, `escalate_up`), **adds** the verbs below, and — unlike a Cell PM — has **no** `submit_up` or `reassign`. Its bubble-up verb is `submit_root` (the root analogue of the Cell PM's `submit_up`):
```python
triage_all() # list actionable tasks across all teams
triage_all() # list actionable tasks across all teams
submit_root(task_id, notes, resolved_findings=None)
# open root->master PR; -> awaiting_pr_review
# (the main PR reviewer gates it; after pr_pass,
# complete escalates to the CEO); a re-submit
# after pr_fail must resolve every open finding first
# open root->master PR; -> awaiting_pr_review
# (the main PR reviewer gates it; after pr_pass,
# complete escalates to the CEO); a re-submit
# after pr_fail must resolve every open finding first
escalate_to_ceo(task_id, reason)
# awaiting_pm_review -> awaiting_ceo_approval
give_me_work() # Main PM may also pull work directly
# awaiting_pm_review -> awaiting_ceo_approval
give_me_work() # Main PM may also pull work directly
```
For a code root the Main PM **must** `submit_root` first — that opens the root→master PR and enters the in-path gate (`awaiting_pr_review`); only after the main reviewer `pr_pass`es it does `complete` escalate to the CEO. A branchless coordination root (product fan-out, no repo) skips the gate and is completed/escalated directly. The Main PM never merges to `master``complete` escalates and only the CEO merges the root→master PR.
@@ -114,7 +123,7 @@ For a code root the Main PM **must** `submit_root` first — that opens the root
## Board flow (Product Owner / Head of Marketing)
```python
triage() # list actionable tasks in scope
triage() # list actionable tasks in scope
escalate_to_ceo(task_id, reason)
i_am_idle()
```
@@ -126,7 +135,7 @@ The Product Owner additionally has `propose_roadmap(cycle_goal, items)` — a **
## Auditor flow
```python
triage() # read-only list of actionable tasks
triage() # read-only list of actionable tasks
i_am_idle()
```
@@ -135,10 +144,10 @@ The Auditor is a silent observer: read-only `triage`, no `notify`, no claim/comp
## PR Reviewer flow
```python
give_me_work() # returns an inbound-PR review task
claim_pr_review(task_id) # claim it (planless, branchless — read-only)
post_pr_review(task_id, ...) # posts one change-request on the PR; task -> completed
unclaim(task_id) # release a claimed inbound or gate review back to the pool
give_me_work() # returns an inbound-PR review task
claim_pr_review(task_id) # claim it (planless, branchless — read-only)
post_pr_review(task_id, ...) # posts one change-request on the PR; task -> completed
unclaim(task_id) # release a claimed inbound or gate review back to the pool
i_am_idle()
```
@@ -147,13 +156,13 @@ The PR Reviewer reviews inbound external/fork (and, behind a flag, internal) PRs
The same role also runs the **in-path PR-review gate** on the org's own assembled delivery PRs — the merge-level review before the PM merges:
```python
claim_gate_review(task_id) # claim an awaiting_pr_review task; returns the assembled
# diff + collision_context (colliding siblings, if any) +
# (on round >=2) prior_findings, the full ledger
pr_pass(task_id, notes) # assembled PR is correct -> awaiting_pm_review (the PM merges)
claim_gate_review(task_id) # claim an awaiting_pr_review task; returns the assembled
# diff + collision_context (colliding siblings, if any) +
# (on round >=2) prior_findings, the full ledger
pr_pass(task_id, notes) # assembled PR is correct -> awaiting_pm_review (the PM merges)
pr_fail(task_id, findings=[...])
# send it back -> needs_revision, like a QA fail;
# the deprecated issues=[str] shim still works this release
# send it back -> needs_revision, like a QA fail;
# the deprecated issues=[str] shim still works this release
```
Both verdicts are also posted on the assembled PR itself as a review (server-side, bot account) so the decision is visible on the PR the PM merges: `pr_pass` → APPROVE, `pr_fail` → REQUEST_CHANGES — except the root→master PR, which only ever gets a plain COMMENT (only the CEO acts on `master`). On a GitLab-backed project `pr_fail` posts as a plain MR note instead (GitLab has no request-changes review primitive) — the task still goes to `needs_revision` normally regardless of forge.