Nadia
Warm narrator · en-GB
Engine v4 · now streaming at 180 ms
Timbre turns one short recording into a studio-grade synthetic voice that stays yours. Direct it with emotion tags, correct it phoneme by phoneme, then ship it to every channel you run through a single API.
The workflow
Read three prompt sentences into any laptop mic. Timbre strips room tone, tames plosives and normalises level before a single weight is trained.
The v4 engine fits prosody, breath pattern and accent on our GPUs, then hands back a private checkpoint. Nothing enters a shared model.
Steer delivery with inline tags, nudge any syllable in the phoneme editor, and stream the result to your app, your IVR or your render farm.
Voice library
Every voice below is licensed from a paid performer with recorded consent. Scroll the rack.
Warm narrator · en-GB
Broadcast · en-NG
Character · es-ES
Explainer · de-DE
Upload a take and it joins the rack in about four minutes.
Start a clonePlatform
Bracket tags sit inline with your script. Whisper a line, land a laugh, hold a beat, drop the pitch for the disclaimer. No re-record, no re-train.
// script.txt
[warm] Welcome back to the show.
[pause:400ms] [whisper] Now, the part they cut.
[laugh:soft] You are going to want to sit down.
Websocket streaming fast enough for live agents and two-way calls.
Click any word to open its IPA string. Fix a surname, stress a brand, lock the pronunciation across every future render.
/ˈtɪm.bər/ → /ˈtæ̃:bʁ/
Accent and identity carry across the whole set. Record in English, publish in Japanese without booking a second session.
# one endpoint, any voice
curl https://api.timbre.audio/v4/speak \
-H "Authorization: Bearer $TIMBRE_KEY" \
-d '{"voice":"nadia","stream":true,
"text":"Your car is two minutes away."}'
Consent and control
Voice cloning earns trust or it earns regulation. We built the consent layer first and the model second, so the person in the recording keeps the final say.
Read the policyEvery clone starts with a spoken consent phrase matched against the training audio. No match, no model.
All output carries a durable watermark that survives compression, re-recording and clipping. Free public detector.
Blocklists for public figures, plus per-voice rules that reject scripts touching payments, elections or medical advice.
Owners can kill a voice from any device. Checkpoints and cached audio are purged within 24 hours, receipts included.
Pricing
For one creator and one voice.
$19/month
For teams shipping on a schedule.
$89/month
For regulated and high volume work.
Custom
Questions
Still unsure? Write to studio@timbre.audio and a human answers within a day.
Thirty seconds of clean speech gets you a usable clone. Three minutes gets you one that holds up across long-form narration and unusual words. More than ten minutes stops helping.
You do. The checkpoint is yours, it is never folded into a shared base model, and you can export or delete it at any time. If you cancel, we hold it for 30 days and then purge it.
Only with their recorded consent through our capture flow, or a signed licence on Enterprise. Public-figure voices are blocked outright, and attempts are logged.
Yes. Identity and accent are modelled separately from language, so a Glaswegian speaker still sounds Glaswegian reading Spanish. You can also dial the accent toward native.
Rendering keeps working and overage bills at $0.11 per minute. No hard cut-off in the middle of a production, and you can cap overage per project.
Free trial, no card, and the clone is deleted the moment you ask.
Clone your voice