Addresses all 8 unresolved review threads on PR #1511:
1. **Accumulate local transcript segments** (already fixed) — each streaming
partial resets lastTranscriptRef so segments are appended, not replaced.
2. **Flush local speech before shutdown** (already fixed) — worker flushes
remaining speech_buf on loop exit.
3. **Scope dictation shortcut to one composer** — new activeComposer module
tracks which composer instance last received focus. The ⌘D handler only
dispatches to the active instance, preventing duplicate recordings when
both channel and thread composers are mounted.
4. **Keep listening through user-stop finalization** (already fixed) —
stopRecording no longer calls cleanup(); event listeners stay alive until
the native 'stopped' event arrives with the final transcript.
5. **Refresh dictation availability after model downloads** — availability
check now polls every 5s until the model is ready, then stops. Covers
the fresh-install case where the model downloads in the background.
6. **Stop native engine after startup failures** — catch block now calls
invoke('stop_dictation') before cleanup so the native SttEngine doesn't
linger when mic permission is denied or AudioWorklet setup fails.
7. **Prevent disabled composers from shortcut-starting** — keydown handler
checks disabledRef/isSendBlockedRef before starting dictation.
8. **Batch dictation audio before IPC** — worklet frames are accumulated in
a Float32Array batch and flushed every 100ms (~10 IPC calls/s instead of
~375). Also removed the erroneous worklet→destination connection that was
playing mic audio back through speakers.
Additionally:
- Removed the entire dead OpenAI relay proxy (transcribe.rs, routes, config
fields, .env.example entries, reqwest dep, bridge.rs visibility widening).
- Scoped the keyup event to only fire when ⌘D keydown was actually dispatched
(no more spurious events on normal 'd' typing).
1. Merge session + SDP into single POST /transcribe/connect
The two-step flow (POST /transcribe/session → POST /transcribe/sdp) stored
the OpenAI client secret in an in-memory DashMap, which breaks in HA
deployments where the two requests may land on different relay replicas.
Replaced with a single POST /transcribe/connect that accepts the SDP offer,
mints the OpenAI session, proxies the SDP exchange, and returns the SDP
answer + model — all in one request. No server-side session state is needed
between requests, so this works correctly across any number of replicas.
Removed the transcribe_sessions DashMap from AppState entirely.
2. Clear isTranscribing even when final text is unchanged
When a TRANSCRIPT_COMPLETED_EVENT produces no text change (final matches
accumulated deltas), the early-return in handleRealtimeEvent skipped the
setIsTranscribing(false) call. This left the 'Transcribing…' indicator
stuck indefinitely. Fixed by updating the transcribing flag before the
text-change early return.
1. Proxy SDP exchange through the relay (transcribe.rs)
The relay no longer returns the raw OpenAI client secret to the desktop
client. Instead, POST /transcribe/session returns an opaque session ID,
and the new POST /transcribe/sdp endpoint accepts the client's SDP offer,
looks up the cached secret, forwards it to OpenAI, and returns the SDP
answer. This prevents a compromised client from reusing the bearer token
to open non-transcription Realtime sessions under the operator account.
2. Allow send during the transcribing grace window (MessageComposer.tsx)
Previously, pressing Enter/Send while isTranscribing was true (during the
3s grace window after user-stop) blocked the send entirely. Now the send
proceeds with whatever content is already in the composer (transcript
deltas have been applied incrementally) and cancels the dictation run so
late events don't refill the cleared composer.
Two fixes:
1. (transcribe.rs) expires_after must be an object with anchor + seconds
fields per OpenAI's client-secret schema, not a bare integer. The bare
number would cause OpenAI to reject the request with a 400, making
dictation fail with a 502 from the relay.
2. (useRealtimeDictation.ts) The user-stop grace window (3s delayed
teardown with run kept valid) now applies to ALL models, not just
manual-commit ones. For server-VAD models, the final VAD completion
event may still be in-flight when the user clicks the mic to stop —
without the grace window, that event was dropped and the tail of the
dictated message disappeared.
Add expires_after: 60 (seconds) to the OpenAI Realtime client-secret
request payload. This limits the reuse window of minted secrets — a
client must establish its WebRTC connection within 60s, and cannot reuse
the secret to open additional sessions after that. Without this, the
default 600s TTL allowed a compromised or malicious client to bypass the
per-pubkey rate limiter by reusing a single minted secret for many
concurrent sessions.
Address two review comments:
1. (useRealtimeDictation.ts) Increment activeRunIdRef immediately in the
manual-commit cleanup path so late transcript events from the 3s grace
window are rejected by handleRealtimeEvent. Previously, the run stayed
valid during the timeout, allowing transcripts to write back into the
composer after a send, edit-save, or navigation.
2. (transcribe.rs) Replace as_object_mut().unwrap() with a branch that
builds the correct JSON literal directly. Avoids introducing an unwrap
in a production path per AGENTS.md rules.
Address two review comments:
1. /transcribe/status now returns configured: false for non-members on open
relays, preventing the mic button from appearing for users who would get
a 403 on session creation. Uses the same require_relay_member check.
2. The OpenAI Realtime session payload now omits turn_detection for
realtime-whisper models (which require manual audio commit per OpenAI
guidance). Other models continue to use server_vad.
Adds unit tests for the model-aware payload builder.
1. P1 — Force NIP-98 signed auth for /transcribe/* endpoints regardless of
BUZZ_REQUIRE_AUTH_TOKEN. The X-Pubkey dev fallback is spoofable, so a
billable endpoint must always require cryptographic proof of identity
before the membership check trusts the pubkey.
2. P2 — DictationButton now allows the stop action whenever isRecording is
true, even during startup (mic live but SDP exchange in progress) or
when the composer is disabled. Only blocks the button when idle and
disabled, or when startup hasn't captured the mic yet.
1. P1 — Require relay membership for billable /transcribe/session even on
open relays. Added require_relay_member() that always checks actual
membership (with NIP-OA fallback) regardless of the
BUZZ_REQUIRE_RELAY_MEMBERSHIP setting. Prevents arbitrary NIP-98 signers
from minting metered OpenAI sessions on the operator's bill.
2. P2 — Keep stop control usable while recording. DictationButton now
allows the stop action even when the composer is disabled — only
*starting* a new recording is blocked. Additionally, useComposerDictation
auto-stops the active session when the composer becomes disabled
mid-recording (e.g. channel becomes read-only).
3. P2 — .expect() already removed in prior commit (openai_client() returns
Result and propagates errors). No additional change needed.
4. P2 — Stop dictation before edit saves already addressed in prior commit
(stopDictationRef.current() at line 540). No additional change needed.
- Stop active dictation run before the edit branch clears/saves, matching
the normal-send path. Prevents late transcript events from writing back
into the restored or fresh draft after an edit save.
- Replace .expect() in openai_client() with proper error propagation via
Result, avoiding a panic if TLS backend initialization ever fails.
Addresses review feedback on PR #1511.
Addresses Wes's review feedback on PR #1511:
1. Per-pubkey rate limit on POST /transcribe/session (5/min, configurable
via BUZZ_TRANSCRIBE_SESSIONS_PER_MINUTE). Each session opens a metered
OpenAI Realtime connection on the operator's bill.
2. Auto-submit disabled by default (DEFAULT_AUTO_SUBMIT_PHRASE = '').
The infrastructure for configurable phrases remains in place and can
be wired to a user setting later.
Also:
- Move BUZZ_TRANSCRIPTION_MODEL into Config (consistent with other knobs)
- Use a shared reqwest::Client via OnceLock (connection pooling)
- Stop dictation on manual send (stopDictationRef wiring)
The OpenAI client_secrets endpoint expects the body as
{ session: { type, audio: { input: { transcription, turn_detection } } } }
not as top-level fields. Also moves turn_detection under audio.input per
the Realtime transcription guide.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Relay: restructure OpenAI client-secrets payload to use the current
typed transcription schema (audio.input.transcription) instead of the
deprecated top-level input_audio_transcription field.
- realtimeAudio: insert space separators between transcript items when
neither the preceding nor following text has whitespace, preventing
multi-utterance runs from merging into unreadable text.
- useDictation: remove premature setText('') after auto-submit — the
send flow handles clearing on success, so dictated text survives if a
mention dialog opens or the send is blocked.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Use OpenAI typed transcription session format (type: "transcription")
instead of legacy realtime fields that would fail or produce no transcripts
- Sync editor content via syncContentRef before merging dictation text so
manually typed prefixes are preserved when dictation starts
- Read send-blocked state from refs at transcript time so uploads prevent
auto-submit from clearing the composer
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Switch relay from /v1/realtime/sessions to /v1/realtime/client_secrets
with the wrapped { session: { ... } } request shape per OpenAI's current
WebRTC guide. The old endpoint returns non-2xx, breaking dictation.
- Redesign TranscriptSegmentState to track per-item segments keyed by
item_id. Completed events for different turns can arrive out of order;
reconciling by item_id preserves utterance ordering and prevents text
reordering or partial-turn drops during fast consecutive speech.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Both /transcribe/status and /transcribe/session now require NIP-98
authentication and relay membership (with NIP-OA fallback), matching
the security posture of /events, /query, and /count.
Promotes verify_bridge_auth, check_nip98_replay, and nip98_expected_url
to pub(crate) so the transcribe module can reuse them without duplication.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
New public API needs doc comments — clippy runs with -D missing-docs, so
TranscribeStatus and TranscribeSession were failing the Rust Lint gate.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Adds dictation support using OpenAI's Realtime API over WebRTC:
Relay:
- New /transcribe/status and /transcribe/session endpoints
- BUZZ_OPENAI_API_KEY env var gates the feature (hidden when absent)
- Proxies ephemeral client-secret minting from OpenAI
Desktop:
- New features/dictation module with:
- AudioWorklet for 24kHz PCM capture + buffering
- WebRTC peer connection to OpenAI Realtime API
- Real-time transcript merging into composer
- Auto-submit on trigger phrase ('submit')
- Mic button in composer toolbar (red pulse when recording)
- Integrated into MessageComposer via useComposerDictation hook
Signed-off-by: klopez4212 <klopez4212@gmail.com>