Abstract
Flow matching is a powerful generative model and has become a core component of modern text-to-speech (TTS) systems. In order to ensure high-quality speech synthesis, Classifier-Free Guidance (CFG) is widely used during the training and inference stages of flow matching based TTS models. However, in the scenario of zero-shot voice cloning, the models keep human feedback isolated from training and inference, which results in mismatched training objectives and evaluation metrics. In this work, we propose a preference optimization method by separately modeling preferred and dispreferred information with two distinct models, which enables better utilization of preferred and dispreferred data. We further incorporate an improved CFG technique to enhance synthesis adherence to the text and reference audio. The experiments demonstrate that these optimizations significantly enhance intelligibility, speaker similarity and naturalness of synthesized speech, outperforming the baseline TTS model.
Audio Samples
For the pretraining phase, we use WenetSpeech4TTS Premium, a Mandarin dataset contains 945 hours of multi-speaker corpus,
and LibriTTS dataset, a multi-speaker English corpus comprising approximately 585 hours of speech, as the training sets.
For DPO, we construct different preference speech datasets (Data1, Data2 and Data3) for training.
Baseline: Pretrained F5-TTS Small model.
PA-Base-DT1 : Single model DPO training with Data1 (preference pairs constructed by ground truth samples and generated samples of the pretrained model).
PA-Base-DT2 : Single model DPO training with Data2 (preference pairs constructed by ranking regular text-audio pairs with WER and SSIM metrics).
PA-Dual-DT1 : Proposed dual model DPO training with Data1 (preference pairs constructed by ground truth samples and generated samples of the pretrained model).
PA-Dual-DT2 : Proposed dual model DPO training with Data2 (preference pairs constructed by ranking regular text-audio pairs with WER and SSIM metrics).
PA-Dual-DT3 : Proposed dual model DPO training with Data3 (preference pairs constructed by ranking challenging text-audio pairs with WER and SSIM metrics).
Mandarin (Seed-TTS test-zh)
| Text | Prompt | Baseline | PA-Base-DT1 | PA-Base-DT2 | PA-Dual-DT1 | PA-Dual-DT2 | PA-Dual-DT3 |
|---|---|---|---|---|---|---|---|
| 将货物通关时间,从原来的九点二个小时,缩短为九分钟。 | |||||||
| 尽管时代变迁,但这间老店却像激流里,唯一不动的岩石。 | |||||||
| 产品经理其实是随着互联互联网发展应用,而生的。 |
| 将货物通关时间,从原来的九点二个小时,缩短为九分钟。 | |
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
| 尽管时代变迁,但这间老店却像激流里,唯一不动的岩石。 | |
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
| 产品经理其实是随着互联互联网发展应用,而生的。 | |
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
English (Seed-TTS test-en)
| Text | Prompt | Baseline | PA-Base-DT1 | PA-Base-DT2 | PA-Dual-DT1 | PA-Dual-DT2 | PA-Dual-DT3 |
|---|---|---|---|---|---|---|---|
| Outdoor area crowded with jeeps packed with people, seen from behind. |
|||||||
| There is no universal definition of intelligence, but everyone agrees that the ability of learning belongs to it. |
|||||||
This was the strangest of all things that ever came to earth from outer space. |
| Outdoor area crowded with jeeps packed with people, seen from behind. |
|
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
| There is no universal definition of intelligence, but everyone agrees that the ability of learning belongs to it. |
|
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
This was the strangest of all things that ever came to earth from outer space. |
|
|---|---|
| Prompt | |
| Baseline | |
| PA-Base-DT1 | |
| PA-Base-DT2 | |
| PA-Dual-DT1 | |
| PA-Dual-DT2 | |
| PA-Dual-DT3 |
Effect of preference data size
| Text | Prompt | 0 pair | 250 pairs | 500 pairs | 750 pairs | 1000 pairs | 1500 pairs | 2000 pairs |
|---|---|---|---|---|---|---|---|---|
| 厨房师傅使用扫把清扫地面后,将其放入用于做菜的锅内进行漂洗。 |
||||||||
| 而对于话多的大嗓门先生来说,那难受的劲就更别提了。 |
||||||||
| 都需要将车辆的登记身份,从非营非运营换成运营。 |
| 厨房师傅使用扫把清扫地面后,将其放入用于做菜的锅内进行漂洗。 |
|
|---|---|
| Prompt | |
| 0 pair | |
| 250 pairs | |
| 500 pairs | |
| 750 pairs | |
| 1000 pairs | |
| 1500 pairs | |
| 2000 pairs |
| 而对于话多的大嗓门先生来说,那难受的劲就更别提了。 |
|
|---|---|
| 0 pair | |
| 250 pairs | |
| 250 pairs | |
| 500 pairs | |
| 750 pairs | |
| 1000 pairs | |
| 1500 pairs | |
| 2000 pairs |
| 都需要将车辆的登记身份,从非营非运营换成运营。 |
|
|---|---|
| Prompt | |
| 0 pair | |
| 250 pairs | |
| 500 pairs | |
| 750 pairs | |
| 1000 pairs | |
| 1500 pairs | |
| 2000 pairs |
Effect of CFG scale
| Text | Prompt | ω=1.0 | ω=1.5 | ω=2.0 | ω=2.5 | ω=3.0 |
|---|---|---|---|---|---|---|
| 但艾肯尔还是感到害怕,还萌生了打道回府的念头。 |
||||||
| 令人可怕的描述,差不多用了十页稿纸。 |
||||||
| 每天早晨有一个老仆人,来为他打扫房间和跑腿。 |
| 但艾肯尔还是感到害怕,还萌生了打道回府的念头。 |
|
|---|---|
| Prompt | |
| ω=1.0 | |
| ω=1.5 | |
| ω=2.0 | |
| ω=2.5 | |
| ω=3.0 |
令人可怕的描述,差不多用了十页稿纸。 |
|
|---|---|
| Prompt | |
| ω=1.0 | |
| ω=1.5 | |
| ω=2.0 | |
| ω=2.5 | |
| ω=3.0 |
| 每天早晨有一个老仆人,来为他打扫房间和跑腿。 |
|
|---|---|
| Prompt | |
| ω=1.0 | |
| ω=1.5 | |
| ω=2.0 | |
| ω=2.5 | |
| ω=3.0 |
References
[1] Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen, “F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,” in ACL, 2025.
[2] Minghao Fu, Guo-Hua Wang, Liangfu Cao, Qing-Guo Chen,Zhao Xu, Weihua Luo, and Kaifu Zhang, “CHATS: Combining human-aligned optimization and test-time sampling for text-to-image generation,” in ICML, 2025.
[3] Fu-Yun Wang, Yunhao Shui, Piao, et al., “Diffusion-NPO: Negative preference optimization for better preference aligned generation of diffusion models,” in ICLR, 2025.