End-to-End Adversarial Text-to-Speech (Paper Explained)

Text-to-speech engines are usually multi-stage pipelines that transform the signal into many intermediate representations and require supervision at each step. When trying to train TTS end-to-end, the alignment problem arises: Which text corresponds to which piece of sound? This paper uses an alignment module to tackle this problem and produces astonishingly good sound.

OUTLINE:
0:00 – Intro & Overview
1:55 – Problems with Text-to-Speech
3:55 – Adversarial Training
5:20 – End-to-End Training
7:20 – Discriminator Architecture
10:40 – Generator Architecture
12:20 – The Alignment Problem
14:40 – Aligner Architecture
24:00 – Spectrogram Prediction Loss
32:30 – Dynamic Time Warping
38:30 – Conclusion

Paper:
Website:

Abstract:
Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from normalised text or phonemes in an end-to-end manner, resulting in models which operate directly on character or phoneme input sequences and produce raw speech audio outputs. Our proposed generator is feed-forward and thus efficient for both training and inference, using a differentiable monotonic interpolation scheme to predict the duration of each input token. It learns to produce high fidelity audio through a combination of adversarial feedback and prediction losses constraining the generated audio to roughly match the ground truth in terms of its total duration and mel-spectrogram. To allow the model to capture temporal variation in the generated audio, we employ soft dynamic time warping in the spectrogram-based prediction loss. The resulting model achieves a mean opinion score exceeding 4 on a 5 point scale, which is comparable to the state-of-the-art models relying on multi-stage training and additional supervision.

Authors: Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, Karen Simonyan

Links:
YouTube:
Twitter:
Discord:
BitChute:
Minds:

source

What's Hot

OpenAI’s $10 Million+ AI Consulting Business: Deployment Takes Center Stage

Why Cartken pivoted its focus from last-mile delivery to industrial robots

Tesla ‘Model Q’ gets bold prediction from Deutsche Bank that investors will love

End-to-End Adversarial Text-to-Speech (Paper Explained)

Energy-Based Transformers are Scalable Learners and Thinkers (Paper Review)

Yannic Kilcher Live Stream

Imagination-Augmented Agents for Deep Reinforcement Learning

Sam Gilliam Foundation, David Kordansky Sued Over ‘Disavowed’ Painting

Donors Reportedly Pulling Support from Florida University Museum after its Controversial Transfer

What will come of the Guggenheim Asher legal battle?

Painter Says DHS Stole His Work for Post About ‘Homeland’s Heritage’

OpenAI’s $10 Million+ AI Consulting Business: Deployment Takes Center Stage

Why Cartken pivoted its focus from last-mile delivery to industrial robots

Tesla ‘Model Q’ gets bold prediction from Deutsche Bank that investors will love

What's Hot

End-to-End Adversarial Text-to-Speech (Paper Explained)

Related Posts

Subscribe to Updates