Distinguish user-initiated stop (mic button) from send/navigation cleanup:
- stopRecording(): user stop — keeps the run valid during the 3s grace
window so the final commit's transcript is delivered to the composer.
- cancelRecording(): send/edit-save/navigation — immediately invalidates
the run so late transcripts cannot refill the cleared composer.
MessageComposer's stopDictationRef now uses cancelRecording (called on
send and edit-save). useComposerDictation's draftKey/disabled effects
also use cancelRecording. The user-facing toggleRecording/stopRecording
path preserves the final transcript for short recordings.
Address two review comments:
1. (useRealtimeDictation.ts) Increment activeRunIdRef immediately in the
manual-commit cleanup path so late transcript events from the 3s grace
window are rejected by handleRealtimeEvent. Previously, the run stayed
valid during the timeout, allowing transcripts to write back into the
composer after a send, edit-save, or navigation.
2. (transcribe.rs) Replace as_object_mut().unwrap() with a branch that
builds the correct JSON literal directly. Avoids introducing an unwrap
in a production path per AGENTS.md rules.
When BUZZ_TRANSCRIPTION_MODEL is a realtime-whisper variant, the relay
now omits server_vad from the session config. Without server VAD, OpenAI
buffers audio indefinitely until a manual input_audio_buffer.commit is
sent. This commit adds the client-side counterpart:
1. Export commitAudioBuffer() and requiresManualCommit() from
realtimeAudio.ts.
2. On session creation, track whether the model requires manual commit.
3. After flushing the pre-connection buffer, commit immediately and start
a 2s periodic commit interval so streaming transcripts flow during
recording.
4. On stop, send a final commit before teardown and keep the data channel
open briefly (3s) to receive the last transcript response.
5. Clear the commit interval on cleanup.
Without this, realtime-whisper sessions would stream/buffer audio but
never produce transcripts because no commit was ever sent.
Address two review comments:
1. /transcribe/status now returns configured: false for non-members on open
relays, preventing the mic button from appearing for users who would get
a 403 on session creation. Uses the same require_relay_member check.
2. The OpenAI Realtime session payload now omits turn_detection for
realtime-whisper models (which require manual audio commit per OpenAI
guidance). Other models continue to use server_vad.
Adds unit tests for the model-aware payload builder.
1. P1 — Force NIP-98 signed auth for /transcribe/* endpoints regardless of
BUZZ_REQUIRE_AUTH_TOKEN. The X-Pubkey dev fallback is spoofable, so a
billable endpoint must always require cryptographic proof of identity
before the membership check trusts the pubkey.
2. P2 — DictationButton now allows the stop action whenever isRecording is
true, even during startup (mic live but SDP exchange in progress) or
when the composer is disabled. Only blocks the button when idle and
disabled, or when startup hasn't captured the mic yet.
1. P1 — Require relay membership for billable /transcribe/session even on
open relays. Added require_relay_member() that always checks actual
membership (with NIP-OA fallback) regardless of the
BUZZ_REQUIRE_RELAY_MEMBERSHIP setting. Prevents arbitrary NIP-98 signers
from minting metered OpenAI sessions on the operator's bill.
2. P2 — Keep stop control usable while recording. DictationButton now
allows the stop action even when the composer is disabled — only
*starting* a new recording is blocked. Additionally, useComposerDictation
auto-stops the active session when the composer becomes disabled
mid-recording (e.g. channel becomes read-only).
3. P2 — .expect() already removed in prior commit (openai_client() returns
Result and propagates errors). No additional change needed.
4. P2 — Stop dictation before edit saves already addressed in prior commit
(stopDictationRef.current() at line 540). No additional change needed.
- Stop active dictation run before the edit branch clears/saves, matching
the normal-send path. Prevents late transcript events from writing back
into the restored or fresh draft after an edit save.
- Replace .expect() in openai_client() with proper error propagation via
Result, avoiding a panic if TLS backend initialization ever fails.
Addresses review feedback on PR #1511.
Addresses Wes's review feedback on PR #1511:
1. Per-pubkey rate limit on POST /transcribe/session (5/min, configurable
via BUZZ_TRANSCRIBE_SESSIONS_PER_MINUTE). Each session opens a metered
OpenAI Realtime connection on the operator's bill.
2. Auto-submit disabled by default (DEFAULT_AUTO_SUBMIT_PHRASE = '').
The infrastructure for configurable phrases remains in place and can
be wired to a user setting later.
Also:
- Move BUZZ_TRANSCRIPTION_MODEL into Config (consistent with other knobs)
- Use a shared reqwest::Client via OnceLock (connection pooling)
- Stop dictation on manual send (stopDictationRef wiring)
When the user manually sends (Enter/click) while recording, the submit
flow already calls stopRecording(). However, queued data channel messages
could still fire handleRealtimeEvent after the peer connection teardown
began, writing stale transcript text back into the now-empty composer.
Fix: pass the run ID captured at startRecording into the data channel
message handler closure. handleRealtimeEvent now checks activeRunIdRef
against the captured run ID and drops events from stale runs. This
guarantees that once stopRecording() increments the run counter, no
further transcript events from that session can mutate composer state.
Handle input_audio_buffer.committed events to register items in the
correct utterance order using previous_item_id before any transcript
events arrive. This ensures that when completions for different turns
arrive out of order (or when only completions are sent without deltas),
the composer reconstructs multi-utterance dictation in the correct
sequence rather than event-arrival order.
Added tests for committed-order preservation, out-of-order completions
with pre-registered order, and completion-only flows.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
The OpenAI client_secrets endpoint expects the body as
{ session: { type, audio: { input: { transcription, turn_detection } } } }
not as top-level fields. Also moves turn_detection under audio.input per
the Realtime transcription guide.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Relay: restructure OpenAI client-secrets payload to use the current
typed transcription schema (audio.input.transcription) instead of the
deprecated top-level input_audio_transcription field.
- realtimeAudio: insert space separators between transcript items when
neither the preceding nor following text has whitespace, preventing
multi-utterance runs from merging into unreadable text.
- useDictation: remove premature setText('') after auto-submit — the
send flow handles clearing on success, so dictated text survives if a
mention dialog opens or the send is blocked.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
When the composer's draftKey changes (channel or thread switch), stop any
active dictation session so transcript events from a stale WebRTC connection
don't leak into the wrong draft.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Use OpenAI typed transcription session format (type: "transcription")
instead of legacy realtime fields that would fail or produce no transcripts
- Sync editor content via syncContentRef before merging dictation text so
manually typed prefixes are preserved when dictation starts
- Read send-blocked state from refs at transcript time so uploads prevent
auto-submit from clearing the composer
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Switch relay from /v1/realtime/sessions to /v1/realtime/client_secrets
with the wrapped { session: { ... } } request shape per OpenAI's current
WebRTC guide. The old endpoint returns non-2xx, breaking dictation.
- Redesign TranscriptSegmentState to track per-item segments keyed by
item_id. Completed events for different turns can arrive out of order;
reconciling by item_id preserves utterance ordering and prevents text
reordering or partial-turn drops during fast consecutive speech.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Add nonce tag to NIP-98 auth events to prevent replay rejection when
multiple components call /transcribe/status in the same second.
- Wire dictation text into both the Tiptap editor and contentRef via
setComposerContent + setEditorContentRef, so dictated text actually
appears in the composer and is serialized on submit.
- Call submitMessageRef.current() synchronously in onSend instead of via
queueMicrotask, ensuring the editor content is consumed before the
subsequent setText('') clears it.
- Replace naive append-based transcript merging with segment-aware state
tracking (TranscriptSegmentState). Delta events accumulate into
pendingDelta; completed events replace accumulated deltas with the
finalized text, preventing duplication.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Both /transcribe/status and /transcribe/session now require NIP-98
authentication and relay membership (with NIP-OA fallback), matching
the security posture of /events, /query, and /count.
Promotes verify_bridge_auth, check_nip98_replay, and nip98_expected_url
to pub(crate) so the transcribe module can reuse them without duplication.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
New public API needs doc comments — clippy runs with -D missing-docs, so
TranscribeStatus and TranscribeSession were failing the Rust Lint gate.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Adds dictation support using OpenAI's Realtime API over WebRTC:
Relay:
- New /transcribe/status and /transcribe/session endpoints
- BUZZ_OPENAI_API_KEY env var gates the feature (hidden when absent)
- Proxies ephemeral client-secret minting from OpenAI
Desktop:
- New features/dictation module with:
- AudioWorklet for 24kHz PCM capture + buffering
- WebRTC peer connection to OpenAI Realtime API
- Real-time transcript merging into composer
- Auto-submit on trigger phrase ('submit')
- Mic button in composer toolbar (red pulse when recording)
- Integrated into MessageComposer via useComposerDictation hook
Signed-off-by: klopez4212 <klopez4212@gmail.com>