Handle input_audio_buffer.committed events to register items in the
correct utterance order using previous_item_id before any transcript
events arrive. This ensures that when completions for different turns
arrive out of order (or when only completions are sent without deltas),
the composer reconstructs multi-utterance dictation in the correct
sequence rather than event-arrival order.
Added tests for committed-order preservation, out-of-order completions
with pre-registered order, and completion-only flows.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
The OpenAI client_secrets endpoint expects the body as
{ session: { type, audio: { input: { transcription, turn_detection } } } }
not as top-level fields. Also moves turn_detection under audio.input per
the Realtime transcription guide.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Relay: restructure OpenAI client-secrets payload to use the current
typed transcription schema (audio.input.transcription) instead of the
deprecated top-level input_audio_transcription field.
- realtimeAudio: insert space separators between transcript items when
neither the preceding nor following text has whitespace, preventing
multi-utterance runs from merging into unreadable text.
- useDictation: remove premature setText('') after auto-submit — the
send flow handles clearing on success, so dictated text survives if a
mention dialog opens or the send is blocked.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
When the composer's draftKey changes (channel or thread switch), stop any
active dictation session so transcript events from a stale WebRTC connection
don't leak into the wrong draft.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Use OpenAI typed transcription session format (type: "transcription")
instead of legacy realtime fields that would fail or produce no transcripts
- Sync editor content via syncContentRef before merging dictation text so
manually typed prefixes are preserved when dictation starts
- Read send-blocked state from refs at transcript time so uploads prevent
auto-submit from clearing the composer
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Switch relay from /v1/realtime/sessions to /v1/realtime/client_secrets
with the wrapped { session: { ... } } request shape per OpenAI's current
WebRTC guide. The old endpoint returns non-2xx, breaking dictation.
- Redesign TranscriptSegmentState to track per-item segments keyed by
item_id. Completed events for different turns can arrive out of order;
reconciling by item_id preserves utterance ordering and prevents text
reordering or partial-turn drops during fast consecutive speech.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
- Add nonce tag to NIP-98 auth events to prevent replay rejection when
multiple components call /transcribe/status in the same second.
- Wire dictation text into both the Tiptap editor and contentRef via
setComposerContent + setEditorContentRef, so dictated text actually
appears in the composer and is serialized on submit.
- Call submitMessageRef.current() synchronously in onSend instead of via
queueMicrotask, ensuring the editor content is consumed before the
subsequent setText('') clears it.
- Replace naive append-based transcript merging with segment-aware state
tracking (TranscriptSegmentState). Delta events accumulate into
pendingDelta; completed events replace accumulated deltas with the
finalized text, preventing duplication.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Both /transcribe/status and /transcribe/session now require NIP-98
authentication and relay membership (with NIP-OA fallback), matching
the security posture of /events, /query, and /count.
Promotes verify_bridge_auth, check_nip98_replay, and nip98_expected_url
to pub(crate) so the transcribe module can reuse them without duplication.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
New public API needs doc comments — clippy runs with -D missing-docs, so
TranscribeStatus and TranscribeSession were failing the Rust Lint gate.
Signed-off-by: klopez4212 <klopez4212@gmail.com>
Adds dictation support using OpenAI's Realtime API over WebRTC:
Relay:
- New /transcribe/status and /transcribe/session endpoints
- BUZZ_OPENAI_API_KEY env var gates the feature (hidden when absent)
- Proxies ephemeral client-secret minting from OpenAI
Desktop:
- New features/dictation module with:
- AudioWorklet for 24kHz PCM capture + buffering
- WebRTC peer connection to OpenAI Realtime API
- Real-time transcript merging into composer
- Auto-submit on trigger phrase ('submit')
- Mic button in composer toolbar (red pulse when recording)
- Integrated into MessageComposer via useComposerDictation hook
Signed-off-by: klopez4212 <klopez4212@gmail.com>