Merge 'master' into 'dogfood feature branch' (smoke test run) (#82)

* Fix: human surface lifecycle hardening (#81)

* fix: harden agent-idle, redis loop, git errors, escalation audit

- i_am_idle no longer 500s when auto-pausing a task whose commits are
  stored as dicts: tolerate dict-or-object commit refs and run the
  synthetic-checkpoint computation inside the swallowing try block.
- The stream event loop no longer logs an idle redis read-timeout as an
  ERROR every cycle; the blocking-read timeout is treated as a normal idle.
- Git command failures surface git's own (secret-scrubbed) stderr in the
  error message instead of a bare 'Command failed', so push/fetch
  rejections are diagnosable; the injected PAT is redacted.
- The escalate-to-pool redirect emits the task.pending audit event,
  closing a status mutation that previously skipped the audit log.

* fix: let privileged operators set task status via an audited override

The task update route silently dropped a 'status' field in the request body,
so a CEO/admin could not transition a task wedged in a state with no valid
in-band move (e.g. a blocked task whose work merged out-of-band) — the panel
returned 200 while nothing changed. Add 'status' to the update schema and
apply it through a new audited 'admin_set_status' that bypasses the strict
transition validator but always records the audit event. The override
requires elevated permissions; ordinary field updates are unchanged.

* fix: stop human chat sessions from expiring between messages

Messaging sessions fell back to a hardcoded 300s idle timeout, shorter than a
normal pause in a human conversation: the sweeper closed the session and the
next message opened a new one, so a person could not hold a continuous chat.
Make the idle timeout configurable (session_idle_timeout_seconds, default
3600) and resolve an unset timeout to it at every session-creation path
instead of the 300s column fallback.

* fix: resolve doubled doc paths and stop the indexer warning flood

The doc-path resolver returned absolute paths verbatim, so a documenter path
that doubled the base segment (/app/docs/docs/...) never resolved on disk and
the docs never indexed into RAG. Reduce an absolute path under the docs base
to a relative one before normalizing, leaving truly-external absolute paths
for the indexer to skip. The indexer now skips non-markdown source files and
logs a missing/non-doc source at debug instead of warning on every pass.

* fix: reject project repo URLs that point at a protected repository

Add a configurable denylist (protected_git_urls) enforced in the project
create and update paths, so a project cannot be registered against a
repository that must not receive agent commits or merges (e.g. the roboco
source repo during a smoke run). Empty by default (no behavior change);
operators set it to sandbox smoke-test projects.

* fix: let an agent release a blocked task back to the pool

A developer (or QA/doc) trapped on a 'blocked' task had no legal forward
move — every verb rejected from that state — so the dispatcher kept
respawning it with nothing to do. Allow 'unclaim' to release a blocked task
the agent owns back to pending (assignment cleared, work session abandoned,
audited), so the cell PM can re-delegate it instead of the agent churning.

* fix: keep blocked-dev churn out and cell tasks out of board hands

- The dispatcher no longer respawns the owner of a blocked task: from blocked
  the owner has no legal move, so respawning only churns; it is revived on
  unblock or released via unclaim.
- Escalation no longer hands a cell (backend/frontend/ux_ui) coordination task
  to a board/advisory role — such an escalation is diverted to the cell pool,
  matching the existing executable-task guard. main_pm targets are unaffected.

* chore: add an opt-in full clean-slate to the reset script

FULL_RESET=1 wipes everything under the roboco data root except the
persistent service stores (ollama/postgres/redis) and clears the persisted
agent Claude session dirs (ROBOCO_CLAUDE_STATE_DIRS), which otherwise replay
across runs. Default off — the existing DB/Redis wipe + workspace git-reset
is unchanged.

* ++

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>

* fix: read the real commit key (hash) and audit the restore-unblock path

- The auto-pause checkpoint and _extract_first_commit_sha read commit dicts by
  key 'sha', but persisted commits are keyed 'hash' (CommitRef.hash) — the prior
  change stopped the crash but silently dropped every ref. Read 'hash' (sha
  fallback) at both sites; the test now uses the production dict shape so the
  regression can't hide.
- unblock_with_restore set status directly and skipped the audit log; emit the
  status-transition audit there too, like the other direct-set paths.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-06-08 03:27:45 +02:00
committed by GitHub
co-authored by Renn F
parent 5a99140a55
commit 306de1e656
23 changed files with 534 additions and 33 deletions
+49
View File
@@ -22,6 +22,11 @@
# if none exist).
# SKIP_WORKSPACE_RESET=1 — skip the workspace cleanup step entirely
# (DB + Redis only).
# FULL_RESET=1 — opt-in aggressive clean-slate: after the DB + Redis wipe,
# delete everything under the roboco data root EXCEPT ollama/postgres/redis
# (overridable via ROBOCO_DATA_ROOT) and clear the persisted agent Claude
# session dirs listed in ROBOCO_CLAUDE_STATE_DIRS (space-separated). Skips
# the per-workspace git reset (the clones are removed and re-cloned).
set -euo pipefail
@@ -71,6 +76,50 @@ if $DOCKER ps --format '{{.Names}}' | grep -q '^roboco-redis$'; then
$DOCKER exec roboco-redis redis-cli FLUSHDB | sed 's/^/ /'
fi
# Optional full clean-slate (opt-in via FULL_RESET=1): wipe everything under the
# roboco data root EXCEPT the persistent service stores (ollama / postgres /
# redis), and clear any persisted agent Claude session state that would
# otherwise replay across runs. Default OFF — the workspace git-reset below is
# the usual path. The data root is resolved from the workspaces root's parent
# (or ROBOCO_DATA_ROOT); the Claude-state paths come from ROBOCO_CLAUDE_STATE_DIRS
# (space-separated) so nothing is guessed.
if [ "${FULL_RESET:-0}" = "1" ]; then
DATA_ROOT="${ROBOCO_DATA_ROOT:-}"
if [ -z "$DATA_ROOT" ]; then
for candidate in /volume1/roboco/data /data; do
if [ -d "$candidate/workspaces" ]; then
DATA_ROOT="$candidate"
break
fi
done
fi
if [ -n "$DATA_ROOT" ] && [ -d "$DATA_ROOT" ]; then
echo ">>> FULL_RESET: wiping $DATA_ROOT/* except ollama/postgres/redis ..."
for entry in "$DATA_ROOT"/*; do
[ -e "$entry" ] || continue
case "$(basename "$entry")" in
ollama | postgres | redis)
echo " keep $(basename "$entry")"
;;
*)
echo " wipe $(basename "$entry")"
rm -rf "$entry"
;;
esac
done
else
echo ">>> FULL_RESET: no data root resolved — skipping data wipe."
fi
for cdir in ${ROBOCO_CLAUDE_STATE_DIRS:-}; do
if [ -e "$cdir" ]; then
echo " clearing Claude session state $cdir"
rm -rf "$cdir"
fi
done
echo ">>> FULL_RESET done."
exit 0
fi
# Workspace reset — each agent has a private git clone at
# {root}/{project}/{team}/{agent}. Leftover staged/untracked edits and
# feature branches from a previous run will fail the claim→start