🎙  Interspeech 2026  ·  Anonymous Submission

Lightweight Cross-Lingual Speaker Adaptation
for Indic TTS

Audio samples accompanying the paper

We propose a lightweight speaker adaptation pipeline requiring only a 10-second reference sample. Synthetic speech from IN-F5 is filtered through a four-stage quality-control pipeline and used to fine-tune a FastSpeech2 model with dual-site ECAPA-TDNN speaker conditioning and a cosine consistency loss. CLS-based phoneme unification enables the adapted voice to synthesise speech across Hindi, Marathi, Tamil, and Telugu — running 53× faster than IN-F5 at 4.7× lower parameter count.

Proposed Pipeline

Three stages: pretrain a multilingual base → generate & filter synthetic data → fine-tune to target speaker

1

Multi-Speaker Pretraining

FastSpeech2 + HiFi-GAN trained on ~119 h across Hindi, Marathi, Tamil & Telugu with CLS phoneme unification and dual-site ECAPA-TDNN speaker conditioning.

Base FS2
→
2

Synthetic Data Generation & Filtering

10 s reference → IN-F5 voice cloning → 4-stage pipeline (CER ≤10%, pitch, duration, log-likelihood pruning) → 2.34 h clean corpus.

Quality Control
→
3

Target Speaker Fine-Tuning

Pretrained base model fine-tuned on the filtered corpus. Speaker conditioning layers adapt to the target voice while preserving multilingual phoneme knowledge.

Our Model

Speaker Reference

The only real recording used — a 10-second Hindi clip from PIB India. All systems are conditioned on this identity.

🎤

Hindi Reference (PIB India)

10-second male Hindi speaker · sole adaptation input

Listening Samples

5 utterances per language  ·  Hindi = seen (adaptation language)  ·  Marathi / Tamil / Telugu = cross-lingual (zero adaptation data)

Hindi  हिन्दी seen
# IN-F5 Base FS2 Ours
1
2
3
4
5
Marathi  मराठी cross-lingual
# IN-F5 Base FS2 Ours
1
2
3
4
5
Tamil  தமிழ் cross-lingual
# IN-F5 Base FS2 Ours
1
2
3
4
5
Telugu  తెలుగు cross-lingual
# IN-F5 Base FS2 Ours
1
2
3
4
5