Training a multi-speaker VITS speech model
I prepared audio and transcripts, selected speakers using pitch analysis, and trained a multi-speaker English VITS text-to-speech model for a client’s answer robot.
Project at a glance
- Purpose
- Multiple voices for a client’s text-based answer robot.
- My contribution
- Audio segmentation, transcript normalization, pitch-based speaker selection, and ESPnet2 / PyTorch training.
- Status and checks
- Client model training, with data validation and training troubleshooting. Production deployment and measured quality are unconfirmed.
- Evidence
- Processing and training code below; no permitted audio sample available.
The client already had an answer robot. It could reply in text. They wanted it to speak, in more than one voice, from their own recordings. I trained that voice: multi-speaker English, VITS through ESPnet2, which is PyTorch underneath. Inference is the trained model plus a speaker id. Training only worked after their corpus stopped lying to it.
01A podcast is not a training set.
The raw files were long, uneven, and sometimes just an mp3 the decoder did not like. I converted to wav, ran a gentle ffmpeg chain (high-pass, low-pass, loudness, denoise off unless I asked for it), and kept about two of the cleaner minutes per speaker. Denoise stays off by default in this pipeline because it can introduce metallic artifacts into the training audio.
Then I cut on silence. An utterance under 2 seconds was too short for this training setup. I capped segments at 12 seconds to reduce alignment problems. If a quiet stretch is still too long, I split it.
def _split_segment(start, end, max_dur):
duration = end - start
if duration <= max_dur:
return [(start, end)]
n_chunks = math.ceil(duration / max_dur)
segments = []
for idx in range(n_chunks):
seg_start = start + idx * max_dur
seg_end = min(end, seg_start + max_dur)
segments.append((seg_start, seg_end))
return segments
If a file had no transcript, faster-whisper wrote one. Empty text is not allowed through. The train script fails fast on a <NO_TEXT> line. I would rather stop there than train a model to speak silence.
02I did not keep every speaker.
Multi-speaker only helps if the speakers are actually different, and if each one has enough voiced audio to learn. I estimated F0 with pyworld, Dio plus StoneMask, and kept a median, a 10th and a 90th percentile, and a voiced ratio per speaker.
eligible = []
for spk, f0_values in spk_f0_values.items():
voiced_ratio = spk_voiced_frames[spk] / max(spk_total_frames[spk], 1)
row = {
"speaker_id": spk,
"f0_median_hz": round(float(np.median(f0_values)), 2),
"voiced_ratio": round(float(voiced_ratio), 4),
}
if (
spk_sec[spk] >= 60.0
and spk_utts[spk] >= 20
and voiced_ratio >= 0.45
):
eligible.append(row)
deep = sorted(eligible, key=lambda r: r["f0_median_hz"])[:50]
Sixty seconds, twenty utterances, almost half the frames voiced. Then the lowest medians. That list is which of their speakers the robot is allowed to become. A speaker with a pretty recording and no pitch track does not get in.
03The text has to be sayable.
VITS learns a mapping from characters to audio. “2021” is not four digits to a listener. It is “two thousand twenty-one”, or it is a stumble. Dashes and quotes get folded first. Numbers are optional, on purpose, so I can hear the difference.
def normalize_text(text, normalize_numbers=False):
text = _CONTROL_RE.sub(" ", text)
text = _DASH_RE.sub("-", text)
text = _QUOTE_RE.sub("'", text)
if normalize_numbers:
text = re.sub(
r"\b\d{1,4}\b",
lambda m: _number_to_words(int(m.group(0))),
text,
)
return _WS_RE.sub(" ", text).strip()
04Then the GPU, and a speaker id.
Training is ESPnet2’s VITS recipe, on a Vast.ai GPU, with speaker ids turned on so inference can ask for a voice. Sample rate is the constraint I stopped negotiating with. Data prep and the config have to agree, 22050 or 24000, or the run stops. The pipeline validates that the prepared audio and training configuration use the same sample rate.
bash run.sh \
--manifest data/manifest_transcribed.csv \
--out-data data \
--sr 24000 \
--multi-spk
Two failures ate more time than the model. The dump was written as FLAC, and the training loader wanted TorchCodec to read it, so I forced WAV and rebuilt the dump. One utterance dropped out of the stats and the batch came back short: speech had one fewer row than text. Same keys in wav.scp, text, and utt2spk, or it does not train.
Inference is the other side of that same contract. The answer robot sends text and a speaker id, and gets a waveform back. If the id was not in the pitch list, I do not get to be surprised by the voice.