Three sections are titled text to speech, voice conversion, and voice cloning. In text to speech, text input is processed by grapheme to phoneme or linguistic frontend, then by acoustic model, then by vocoder, and produces waveform output. In voice conversion, source speech is processed by F 0, pitch extractor, content encoder A S R bottleneck, and target speaker embedding. These components feed into conversion model, which is processed by vocoder and produces converted speech target voice. In voice cloning, text prompt and speaker encoder are processed by acoustic model pretrained T T S. The output is processed by vocoder and produces cloned voice speech. Arrows indicate left to right sequential processing in each section.An illustration of three standard audio synthesis pipelines. Top (text-to-speech): this pipeline generates speech from text. It begins with text input, which is converted into a phonetic representation by a linguistic frontend. An acoustic model generates acoustic features from the phonemes, and a vocoder synthesizes the final waveform output. Middle (voice conversion): this pipeline transforms the voice of a source speaker into that of a target speaker. Source speech is processed to extract its linguistic content (content encoder) and pitch (F0 extractor). A conversion model combines this information with a target speaker embedding to create new acoustic features. A vocoder then synthesizes these features into the converted speech target voice. Bottom (voice cloning): this pipeline synthesizes a specific person’s voice from a text prompt. A text prompt and a speaker representation from a speaker encoder are fed into a pre-trained acoustic model. A vocoder then generates the final cloned voice speech
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.