A voice is a set of decisions.
Sibilant builds a speaking model from a consented reference session, then keeps the decisions the speaker already made: where the tongue lands, how the release breaks, what the breath does in between.
Every session is recorded with the speaker in the room and a signed grant on file. No grant, no build.
What a voice is made of
Speech is a small set of gestures, repeated the same way every time. Place of articulation runs across, manner runs down. A build learns how one speaker performs this grid, not an average of everyone who ever has.
| Manner of articulation | Bilabial | Labio |
Alveolar | Post |
Velar | Glottal |
|---|---|---|---|---|---|---|
| Plosive | p b | No symbol | t d | No symbol | k ɡ | ʔ |
| Nasal | m | ɱ | n | No symbol | ŋ | Judged impossible |
| Fricative | ɸ β | f v | s z (the sound this studio is named for) | ʃ ʒ | x ɣ | h ɦ |
| Approximant | No symbol | ʋ | ɹ | No symbol | ɰ | Judged impossible |
Voiceless left, voiced right Hatched judged impossible Marked alveolar fricative, which is where the name came from
How a build runs
Three appointments. The speaker is present for the first and the last.
-
Reference session
The speaker reads a script that covers the grid above: every place, every manner, both voicings, then running prose so the model hears the joins rather than a bag of isolated sounds.
-
Build
We fit a model to that one speaker and nobody else. No pooled corpus, no timbre borrowed from a neighbouring voice to fill a gap the session did not cover.
-
Review
The speaker hears the output before any client does, on the same monitors the session was cut on, and signs off on it or does not.
Bring us one speaker who said yes.
Tell us the work and the language, and we will tell you what a reference session would need to cover.