fix(agents): spawn create/start harness lazy — close the startup message-loss window

start_managed_agent_process was the last spawn_agent_child call site
passing lazy=false. Eager init makes buzz-acp serially initialize the
full worker pool BEFORE connecting to the relay (lib.rs:1317-1346), so
creating or starting a slow-starting harness (e.g. OpenClaw at 4-10s
per worker) took minutes — and any mention sent during that window
predated the startup watermark and could be missed permanently.

Lazy ordering (already used by manual start, restart, reconcile, and
restore per the F1 note in restore.rs) connects and subscribes first,
captures the watermark immediately, queues accepted work, and
initializes the pool in the cancellable wake task on first flushable
work (lib.rs:1714-1740, :2536-2563). Wake failures retry with capped
backoff; shutdown mid-wake drains and reaps partial pools
(lib.rs:2581-2602).

Regression: a source-contract test pins that no spawn_agent_child call
site passes a literal lazy=false (the function spawns a real OS process,
so the contract is pinned at source level, mirroring the reconnect
backoff source-parsing test pattern). Verified red under the exact
mutation (flipping the new site back to false) and green restored.

Demand-driven pool growth (seed 1, grow on contention) is the agreed
follow-up; this is the correctness half.

Team review in buzz-generic-acp-harnesses: Wren and Max both traced the
lifecycle independently and converged on shipping this alone — bundled
concurrent pool init was rejected (spawn_and_init lacks timeout/shutdown
plumbing; fleet-level thundering herd when restore starts many pairs
concurrently).

Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
Signed-off-by: Tyler Longwell <tlongwell@block.xyz>
This commit is contained in:
npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d
2026-07-26 17:32:14 -04:00
co-authored by Tyler Longwell
parent 1a56b7cc9e
commit d390b01039
2 changed files with 80 additions and 1 deletions
@@ -953,7 +953,16 @@ pub fn start_managed_agent_process(
// Scalar PIDs are migration-only and never establish pair liveness.
record.runtime_pid = None;
let mut process = spawn_agent_child(app, record, &key.relay_url, false, owner_hex)?;
// Lazy, matching manual start, restart, reconcile, and restore (see the
// F1 note in restore.rs). This was the last eager call site. Eager init
// ran the full serial pool spawn BEFORE the harness connected to the
// relay, so create/start of a slow-starting agent (e.g. a gateway-backed
// harness at several seconds per worker) took minutes — and a mention
// sent in that window predated the startup watermark and could be missed
// permanently. Lazy connects/subscribes first, queues accepted work, and
// initializes the pool in the cancellable wake task on first flushable
// work.
let mut process = spawn_agent_child(app, record, &key.relay_url, true, owner_hex)?;
let now = now_iso();
let receipt = super::ManagedAgentRuntimeReceipt {
key: key.clone(),
@@ -1052,3 +1052,73 @@ fn restart_eligible_false_when_orphan_has_no_drift() {
fn restart_eligible_false_when_non_orphan_has_no_drift() {
assert!(!super::restart_eligible(false, false, false));
}
// ── every spawn_agent_child call site is lazy ────────────────────────
//
// The eager path (`lazy=false` → BUZZ_ACP_LAZY_POOL=false) makes buzz-acp
// serially initialize the FULL worker pool before connecting to the relay.
// For a slow-starting harness that is minutes of startup, and any mention
// sent in that window predates the startup watermark and can be missed
// permanently. Every Desktop flow (create/start, manual start, restart,
// reconcile, restore) must therefore spawn lazy: relay connects first,
// accepted work queues, and the pool initializes in the cancellable wake
// task.
//
// `spawn_agent_child` spawns a real OS process, so this contract is pinned
// at the source level: no call site may pass a literal `false` for the
// `lazy` parameter. If a new call site legitimately needs eager init,
// it must carry a documented exemption here.
#[test]
fn no_spawn_agent_child_call_site_is_eager() {
let sources = [
(
"runtime.rs",
include_str!(concat!(
env!("CARGO_MANIFEST_DIR"),
"/src/managed_agents/runtime.rs"
)),
),
(
"runtime_commands.rs",
include_str!(concat!(
env!("CARGO_MANIFEST_DIR"),
"/src/managed_agents/runtime_commands.rs"
)),
),
(
"restore.rs",
include_str!(concat!(
env!("CARGO_MANIFEST_DIR"),
"/src/managed_agents/restore.rs"
)),
),
];
for (name, source) in sources {
for (idx, line) in source.lines().enumerate() {
let trimmed = line.trim();
// Skip comments; match only argument-position literal `false`
// in a spawn_agent_child call. The call spans multiple lines in
// restore.rs, so scan a small window after the call token.
if trimmed.starts_with("//") {
continue;
}
if trimmed.contains("spawn_agent_child(") {
let window: String = source
.lines()
.skip(idx)
.take(8)
.collect::<Vec<_>>()
.join("\n");
// The lazy flag is the 4th argument; a literal `false` in the
// window is the eager smell this test exists to catch.
assert!(
!window.contains("false"),
"{name}:{}: spawn_agent_child call site appears to pass \
lazy=false (eager pool init). All Desktop spawns must be \
lazy — see this test's doc comment. Call window:\n{window}",
idx + 1
);
}
}
}
}