mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
Windows installs of Goose and other harnesses failed at exactly five minutes with an empty error (#2401). The 300s ceiling was killing installs that were working, just slowly — the Goose step pulls a ~79MB release asset, and Windows Defender scans every file npm extracts. When the ceiling fired it discarded the output it had already read, so the user got a bare timeout string and no way to tell a hang from a large download. ## The ceiling `INSTALL_TIMEOUT` is 900s, and the error names the limit: `install command exceeded the 15m ceiling and was terminated`. It stays a pure wall-clock ceiling with no inactivity kill — nothing observable distinguishes a hung installer from one silently transferring a large artifact, so silence alone never kills an install. A ceiling kill remains non-retryable; re-running a command that already burned 15 minutes costs the user more time with no plausible path to success. The child's exit and both stream drains fold into one resumable settle governed by a single deadline. Waiting on the drains outside that deadline would let a descendant that outlived the install shell hold the output pipes — and the per-runtime install guard behind them — open with no bound, which is the failure the ceiling exists to prevent. So the deadline path terminates the process group on the normal-exit branch too: a leader that exited with a real status still gets its stragglers killed, and the guard cannot stick either way. Whether the leader had already exited only decides the verdict — its real status outranks a timeout. The install shell is a session leader and its descendants inherit the output pipes, so signalling only the leader left them running and the drains blocked on a pipe nobody would close. Escalation keys off the *group's* liveness rather than the leader's, since a descendant that ignores SIGTERM outlives the leader and would otherwise never receive the group SIGKILL. Reaping the killed child and finishing the drains share one bounded grace, so a termination that failed outright cannot extend the ceiling that just fired. ## Output capture Each stream drains into a bounded capture that is *shared* with the reader rather than returned by it, so whatever arrived before a stall is readable at the ceiling — exactly when the output matters most. Output of any size costs a fixed amount of memory. One capture holds two independently bounded views of the same bytes: | View | Head / tail | Cut marker | |------|-------------|------------| | UI (`InstallStepResult`) | 512 B / 1024 B | `... (N bytes omitted) ...` | | Log file | 128 KiB / 128 KiB | `... [N bytes omitted at cap] ...` | The UI budget is screen space; the log's is disk. Both markers are inline, so neither ever implies completeness it does not have. Both ends are cut at arbitrary byte offsets, so a partial character is trimmed and the partial token each cut left behind is dropped — the marker's byte count includes both trims. ## Install log `steps` carries only the last attempt of each step, truncated for display. Everything else — earlier retries, the prerequisite step that actually broke, the managed-Node bootstrap — used to be discarded. `InstallReporter` now appends one self-contained record per attempt of per step to `install-<runtime-id>.log` beside the agent logs, and `InstallRuntimeResult.log_path` carries the file to the UI, where a failure message ends with `Full log: <path>`. Each record is bounded independently by the log-scale capture that produced it, so a first attempt that printed megabytes cannot push out the later record explaining the failure; the run's total is bounded by steps × attempts × per-record cap. Every early return builds its result through one `InstallReporter::failed` helper, so no failure path can omit the log pointer, and synthesized steps go through `record_step` — a step that reaches the UI without passing it would be invisible in the file. Install output can echo a registry token or proxy credential from the environment it ran in, and the file is written unattended. Redaction keys off the *names* of the environment variables the install inherited, snapshotted once per run, rather than a list of known secret value prefixes: a credential with no recognisable shape is exactly the one a prefix match misses. Three name rules apply, because the variables need different treatment: | Rule | Variables | Redacted | |------|-----------|----------| | URL userinfo | `HTTP_PROXY`, `HTTPS_PROXY`, `ALL_PROXY`, `NPM_CONFIG_PROXY`, `NPM_CONFIG_HTTPS_PROXY`, `NPM_CONFIG_REGISTRY` | `user:password` only | | Exact name | `NPM_CONFIG_KEY`, `NPM_CONFIG__AUTH`, `NPM_CONFIG_OTP` | whole value | | Marker substring | `*TOKEN*`, `*SECRET*`, `*PASSWORD*`, `*_PAT`, … | whole value, 8-byte floor | A proxy or registry keeps its host and port, because an install that fails behind one is diagnosable only if the record still says which one it went through, and a bare `user@` with no password is not treated as a credential. npm's own settings are listed by exact name rather than matched on `KEY` or `AUTH` substrings — both occur throughout an ordinary environment on values that are paths and people's names — and they bypass the 8-byte floor, since a six-digit one-time password is a credential at that length. Matching is case-insensitive, which is what npm's lowercase `npm_config_*` spelling needs. `0o600` is set by the create rather than a later `chmod`, which would leave a window where the umask decides. A runtime id that cannot safely be a filename yields no log rather than a sanitized one — a rewritten id could collide with another runtime's log. The file holds exactly one run. A run opens its own session after the runtime id has been canonically resolved — the previous file rotates to `.1` and any older `.1` is removed before the rename, since a rename that will not replace its destination would otherwise wedge rotation permanently on Windows. The session writes a header naming the runtime, the app version (`app.package_info().version` on the Rust side — cannot be mocked or fail), the OS (`std::env::consts::OS`), and the start time: a Windows failure and a macOS one on the same runtime are different bugs, and a stale app version explains a failure that no longer reproduces. Each record carries its attempt's elapsed time. ## Live output line A 15-minute ceiling with nothing behind it but a spinner is indistinguishable from a hang. The same drain seam feeds an `acp-install-output` event carrying the newest complete line, and the three install entry points — Doctor harness rows, the harness catalog dialog, and onboarding runtime cards — render it under the spinner with `aria-live="polite"`. Ordering is keyed on a `seq` monotonic across the whole install, not on the attempt number, which restarts at 1 for every step: keyed on attempt, one step succeeding on attempt 2 would make the next step's attempt-1 output look stale and freeze the display for the rest of the install. Each executed attempt begins with an unthrottled `line: null` clear signal, so a stale failure line cannot sit under the spinner while the retry runs. Events are otherwise throttled to four per second, and the throttle *retains* the newest pending line and flushes it when the window reopens rather than dropping it — at an attempt boundary a drop would silently eat the new attempt's first line. The subscription is mounted for the runtime's whole lifetime rather than started when the install begins. The install command is invoked from the click handler, so the clear and a fast command's first lines can be emitted before React has committed the pending state, and nothing replays them — a subscription that waited for that state would lose the entire output of a short install. The run boundary resets the ordering key when the install settles, since `seq` restarts for the next run, and the line renders only while installing, so a straggler from a finishing drain cannot appear under a fresh Install button. The 15-minute ceiling deliberately stops waiting on stuck drain threads — a hung installer must not freeze the app. That means a drain thread can outlive its `InstallReporter`. Without a generation guard, a drain that calls `offer` after the run settles would publish an event with the run's high `seq`, poison the permanent listener's React state, and cause the next install's restarted `seq=0` events to be rejected. `Live` now carries a `lifecycle: Arc<RwLock<bool>>`; drain threads hold a **shared read guard** from the admission check through the `(self.emit)(...)` call, making the check-then-emit pair atomic with respect to shutdown. `InstallReporter::drop` takes the **exclusive write guard** and stores `false` — this blocks until every in-flight drain publication releases its read guard, then prevents any new admission. Deactivation is bounded: the write lock holds only for the flag store, so it can block at most for the duration of one emit call (microseconds to low milliseconds). Rust drops locals in reverse-declaration order, so `reporter` drops before `_guard`, ensuring the exclusive write completes before the per-runtime concurrency guard releases and a new install can start. ## Also Install result types move to `desktop/src/shared/api/installTypes.ts`, following the existing `searchTypes.ts` / `workflowTypes.ts` convention, and are re-exported from `tauri.ts` and `types.ts` — both already over the file-size cap, so neither can grow to carry them. Two comments described `AdapterOutdated` as applying only to the deprecated package; it also covers a version below the supported floor. Report: #2401 --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>