Audio samples accompanying the paper
Three stages: pretrain a multilingual base → generate & filter synthetic data → fine-tune to target speaker
FastSpeech2 + HiFi-GAN trained on ~119 h across Hindi, Marathi, Tamil & Telugu with CLS phoneme unification and dual-site ECAPA-TDNN speaker conditioning.
Base FS210 s reference → IN-F5 voice cloning → 4-stage pipeline (CER ≤10%, pitch, duration, log-likelihood pruning) → 2.34 h clean corpus.
Quality ControlPretrained base model fine-tuned on the filtered corpus. Speaker conditioning layers adapt to the target voice while preserving multilingual phoneme knowledge.
Our ModelThe only real recording used — a 10-second Hindi clip from PIB India. All systems are conditioned on this identity.
10-second male Hindi speaker · sole adaptation input
5 utterances per language · Hindi = seen (adaptation language) · Marathi / Tamil / Telugu = cross-lingual (zero adaptation data)
| # | IN-F5 | Base FS2 | Ours |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 |
| # | IN-F5 | Base FS2 | Ours |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 |
| # | IN-F5 | Base FS2 | Ours |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 |
| # | IN-F5 | Base FS2 | Ours |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 |