Leverage style diffusion
I really like this model because it's easy to train and insanly fast. However it doesn't have enough styling in the speech.
How about adding something like style diffusion like with StyleTTS2?
https://github.com/yl4579/StyleTTS2
2 条评论