Abstract

Flow matching is a powerful generative model and has become a core component of modern text-to-speech (TTS) systems. In order to ensure high-quality speech synthesis, Classifier-Free Guidance (CFG) is widely used during the training and inference stages of flow matching based TTS models. However, in the scenario of zero-shot voice cloning, the models keep human feedback isolated from training and inference, which results in mismatched training objectives and evaluation metrics. In this work, we propose a preference optimization method by separately modeling preferred and dispreferred information with two distinct models, which enables better utilization of preferred and dispreferred data. We further incorporate an improved CFG technique to enhance synthesis adherence to the text and reference audio. The experiments demonstrate that these optimizations significantly enhance intelligibility, speaker similarity and naturalness of synthesized speech, outperforming the baseline TTS model.


Audio Samples

For the pretraining phase, we use WenetSpeech4TTS Premium, a Mandarin dataset contains 945 hours of multi-speaker corpus, and LibriTTS dataset, a multi-speaker English corpus comprising approximately 585 hours of speech, as the training sets. For DPO, we construct different preference speech datasets (Data1, Data2 and Data3) for training.

Baseline: Pretrained F5-TTS Small model.
PA-Base-DT1 : Single model DPO training with Data1 (preference pairs constructed by ground truth samples and generated samples of the pretrained model).
PA-Base-DT2 : Single model DPO training with Data2 (preference pairs constructed by ranking regular text-audio pairs with WER and SSIM metrics).
PA-Dual-DT1 : Proposed dual model DPO training with Data1 (preference pairs constructed by ground truth samples and generated samples of the pretrained model).
PA-Dual-DT2 : Proposed dual model DPO training with Data2 (preference pairs constructed by ranking regular text-audio pairs with WER and SSIM metrics).
PA-Dual-DT3 : Proposed dual model DPO training with Data3 (preference pairs constructed by ranking challenging text-audio pairs with WER and SSIM metrics).


Mandarin (Seed-TTS test-zh)

Text Prompt Baseline PA-Base-DT1 PA-Base-DT2 PA-Dual-DT1 PA-Dual-DT2 PA-Dual-DT3
将货物通关时间,从原来的九点二个小时,缩短为九分钟。
尽管时代变迁,但这间老店却像激流里,唯一不动的岩石。
产品经理其实是随着互联互联网发展应用,而生的。
将货物通关时间,从原来的九点二个小时,缩短为九分钟。
Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3
。
尽管时代变迁,但这间老店却像激流里,唯一不动的岩石。
Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3
产品经理其实是随着互联互联网发展应用,而生的。
Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3

English (Seed-TTS test-en)

Text Prompt Baseline PA-Base-DT1 PA-Base-DT2 PA-Dual-DT1 PA-Dual-DT2 PA-Dual-DT3

Outdoor area crowded with jeeps packed with people, seen from behind.

There is no universal definition of intelligence, but everyone agrees that the ability of learning belongs to it.

This was the strangest of all things that ever came to earth from outer space.

Outdoor area crowded with jeeps packed with people, seen from behind.

Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3

There is no universal definition of intelligence, but everyone agrees that the ability of learning belongs to it.

Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3

This was the strangest of all things that ever came to earth from outer space.

Prompt
Baseline
PA-Base-DT1
PA-Base-DT2
PA-Dual-DT1
PA-Dual-DT2
PA-Dual-DT3

Effect of preference data size

Text Prompt 0 pair 250 pairs 500 pairs 750 pairs 1000 pairs 1500 pairs 2000 pairs

厨房师傅使用扫把清扫地面后,将其放入用于做菜的锅内进行漂洗。

而对于话多的大嗓门先生来说,那难受的劲就更别提了。

都需要将车辆的登记身份,从非营非运营换成运营。

厨房师傅使用扫把清扫地面后,将其放入用于做菜的锅内进行漂洗。

Prompt
0 pair
250 pairs
500 pairs
750 pairs
1000 pairs
1500 pairs
2000 pairs

而对于话多的大嗓门先生来说,那难受的劲就更别提了。

0 pair
250 pairs
250 pairs
500 pairs
750 pairs
1000 pairs
1500 pairs
2000 pairs

都需要将车辆的登记身份,从非营非运营换成运营。

Prompt
0 pair
250 pairs
500 pairs
750 pairs
1000 pairs
1500 pairs
2000 pairs

Effect of CFG scale

Text Prompt ω=1.0 ω=1.5 ω=2.0 ω=2.5 ω=3.0

但艾肯尔还是感到害怕,还萌生了打道回府的念头。

令人可怕的描述,差不多用了十页稿纸。

每天早晨有一个老仆人,来为他打扫房间和跑腿。

但艾肯尔还是感到害怕,还萌生了打道回府的念头。

Prompt
ω=1.0
ω=1.5
ω=2.0
ω=2.5
ω=3.0

令人可怕的描述,差不多用了十页稿纸。

Prompt
ω=1.0
ω=1.5
ω=2.0
ω=2.5
ω=3.0

每天早晨有一个老仆人,来为他打扫房间和跑腿。

Prompt
ω=1.0
ω=1.5
ω=2.0
ω=2.5
ω=3.0

References

[1] Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen, “F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,” in ACL, 2025.

[2] Minghao Fu, Guo-Hua Wang, Liangfu Cao, Qing-Guo Chen,Zhao Xu, Weihua Luo, and Kaifu Zhang, “CHATS: Combining human-aligned optimization and test-time sampling for text-to-image generation,” in ICML, 2025.

[3] Fu-Yun Wang, Yunhao Shui, Piao, et al., “Diffusion-NPO: Negative preference optimization for better preference aligned generation of diffusion models,” in ICLR, 2025.