mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
Phase 1: Extract reusable SttEngine from huddle/stt.rs into stt_engine.rs. The core STT logic (resample 48→16 kHz, earshot VAD, sherpa-onnx Parakeet inference) is now a standalone component configurable via SttEngineConfig. huddle/stt.rs becomes a thin wrapper that passes huddle-specific flags (TTS barge-in, PTT gating). drain_until_shutdown moves to stt_engine and is re-exported by huddle/mod.rs for backward compat. Phase 2: Add Tauri dictation commands (start_dictation, stop_dictation, push_dictation_audio, get_dictation_status) that create a standalone SttEngine instance with dictation-tuned settings (longer silence threshold, no TTS/PTT flags). Transcribed text is emitted to the frontend via 'dictation-transcript' Tauri events. Phase 3: Add useLocalDictation hook that captures mic audio via AudioWorklet and sends raw PCM to the native STT engine. useDictation now routes to local STT when available (offline, no API key), falling back to cloud (OpenAI Realtime via relay) when the model isn't downloaded. Key wins: - Works fully offline — no BUZZ_OPENAI_API_KEY needed - Self-hosters get dictation for free - Lower latency (no network round-trip) - No relay billing concern - Cloud fallback preserved for higher accuracy