Engine v4 · now streaming at 180 ms

Clone a voice in
thirty seconds.
Keep it for life.

Timbre turns one short recording into a studio-grade synthetic voice that stays yours. Direct it with emotion tags, correct it phoneme by phoneme, then ship it to every channel you run through a single API.

  • No card to start
  • 48 kHz / 24-bit output
  • SOC 2 Type II
timbre / studio REC
founder_take_03.wav 00:31
Training 82%
Similarity98.4
Latency180ms
Languages32
[warm] [whisper] [read: news] +14
Audible NineNorthwind GamesRadio Vela Lumen HealthKestrel StudiosDuolang Postmark MediaHavn Airlines

The workflow

Three sentences in.
A working voice out.

  1. 01

    Record thirty seconds

    Read three prompt sentences into any laptop mic. Timbre strips room tone, tames plosives and normalises level before a single weight is trained.

  2. 02

    Train in four minutes

    The v4 engine fits prosody, breath pattern and accent on our GPUs, then hands back a private checkpoint. Nothing enters a shared model.

  3. 03

    Direct and ship

    Steer delivery with inline tags, nudge any syllable in the phoneme editor, and stream the result to your app, your IVR or your render farm.

Voice library

Cloned, cleared and ready to cast

Every voice below is licensed from a paid performer with recorded consent. Scroll the rack.

Nadia

Warm narrator · en-GB

  • audiobook
  • calm
  • 142 wpm

Kwesi

Broadcast · en-NG

  • news
  • authoritative
  • 165 wpm

Mireia

Character · es-ES

  • games
  • expressive
  • 151 wpm

Jonas

Explainer · de-DE

  • tutorial
  • neutral
  • 138 wpm

Your voice
belongs here

Upload a take and it joins the rack in about four minutes.

Start a clone

Platform

Everything a real production needs

Direct it like a performer

Bracket tags sit inline with your script. Whisper a line, land a laugh, hold a beat, drop the pitch for the disclaimer. No re-record, no re-train.

// script.txt
[warm] Welcome back to the show.
[pause:400ms] [whisper] Now, the part they cut.
[laugh:soft] You are going to want to sit down.
180ms

First audio byte

Websocket streaming fast enough for live agents and two-way calls.

Phoneme editor

Click any word to open its IPA string. Fix a surname, stress a brand, lock the pronunciation across every future render.

/ˈtɪm.bər/ → /ˈtæ̃:bʁ/

One clone, 32 languages

Accent and identity carry across the whole set. Record in English, publish in Japanese without booking a second session.

  • EN
  • ES
  • FR
  • DE
  • PT
  • JA
  • HI
  • +25

Ship it anywhere

# one endpoint, any voice
curl https://api.timbre.audio/v4/speak \
  -H "Authorization: Bearer $TIMBRE_KEY" \
  -d '{"voice":"nadia","stream":true,
       "text":"Your car is two minutes away."}'
  • REST
  • Websocket
  • Node
  • Python
  • Unity

Consent and control

A cloned voice should only ever be the owner's to lend.

Voice cloning earns trust or it earns regulation. We built the consent layer first and the model second, so the person in the recording keeps the final say.

Read the policy
  • 01

    Verified consent capture

    Every clone starts with a spoken consent phrase matched against the training audio. No match, no model.

  • 02

    Inaudible watermark

    All output carries a durable watermark that survives compression, re-recording and clipping. Free public detector.

  • 03

    Likeness lock

    Blocklists for public figures, plus per-voice rules that reject scripts touching payments, elections or medical advice.

  • 04

    Revoke in one click

    Owners can kill a voice from any device. Checkpoints and cached audio are purged within 24 hours, receipts included.

4minmedian training time
98.4%speaker similarity
32languages per clone
99.98%API uptime, 12 months

Pricing

Pay for minutes, not seats

Solo

For one creator and one voice.

$19/month

  • 2 cloned voices
  • 120 minutes of audio
  • 48 kHz WAV and MP3
  • Emotion tags and phoneme editor
  • Commercial use included
Start free trial
Most studios pick this

Studio

For teams shipping on a schedule.

$89/month

  • 15 cloned voices
  • 900 minutes, pooled
  • Streaming API at 180 ms
  • Shared pronunciation dictionary
  • Seat roles and audit log
  • Priority render queue
Start free trial

Enterprise

For regulated and high volume work.

Custom

  • Unlimited voices and minutes
  • Private or on-prem inference
  • Custom watermark keys
  • SSO, SCIM, DPA, HIPAA
  • Named solutions engineer
Talk to sales

Questions

Before you record

Still unsure? Write to studio@timbre.audio and a human answers within a day.

How much audio do I actually need?

Thirty seconds of clean speech gets you a usable clone. Three minutes gets you one that holds up across long-form narration and unusual words. More than ten minutes stops helping.

Who owns the resulting voice?

You do. The checkpoint is yours, it is never folded into a shared base model, and you can export or delete it at any time. If you cancel, we hold it for 30 days and then purge it.

Can I clone someone else's voice?

Only with their recorded consent through our capture flow, or a signed licence on Enterprise. Public-figure voices are blocked outright, and attempts are logged.

Does the accent survive translation?

Yes. Identity and accent are modelled separately from language, so a Glaswegian speaker still sounds Glaswegian reading Spanish. You can also dial the accent toward native.

What happens when I hit my minutes?

Rendering keeps working and overage bills at $0.11 per minute. No hard cut-off in the middle of a production, and you can cap overage per project.

Say one sentence.
Get every sentence.

Free trial, no card, and the clone is deleted the moment you ask.

Clone your voice