In Text-to-Speech, Your Ears Will Always Matter More Than Metrics
I’ve been doing projects with various ML topics for multiple years before coming to Text-to-Speech (Classical ML, NLP, CV). I expected that the general pipeline would be pretty similar: dataset, model, train/val loss, etc.
But after some time, I noticed major differences. TTS carries a whole class of problems that are:
- easy to miss when you start;
- rarely written about;
- and yet quietly shape your day-to-day work.
This is an article about the biggest of them, which I faced while working on smart assistants and TTS services in different big tech companies. This post is about the parts of that job I didn't see coming.
Metrics you can't trust
This is probably the most crucial part. In short, there is no single metric in TTS that you can rely on and say, "Okay, the model has improved." Well, you can rely on it, but with some caveats.
WER and...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE