mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
The handoff gate fired at a hard-coded 90% of the context window, which is a reasonable default at 200k and a bad one at 1M: a single compaction near the top of a large window has to summarize ~900k tokens of history in one call. Make the fraction configurable and add an absolute ceiling (default 272k, OpenAI's long-context pricing boundary) so handoff fires at whichever binds first. The ceiling is inert at a 200k window, so existing deployments are unchanged. Rewrite the summary prompt around a fixed section contract so the next turn inherits the specifics it would otherwise have to rediscover, and raise the summary output budget from 8k to 32k tokens now that it has more to carry. Raise the handoff cap to 80, since it is a backstop against a compaction loop rather than a session length limit, and hitting it degrades the agent to dropping its oldest turns. Signed-off-by: Atish Patel <atish@squareup.com> Co-authored-by: Cursor <cursoragent@cursor.com>