JasonBourne1998
Thanks for your excellent work, Can we use this to achieve singing-to-speech transfer? If so, how can we do that
#7 · closed · 1 comments
Thanks for your excellent work, Can we use this to achieve singing-to-speech transfer? If so, how can we do that
If you utilize the existing model, like the speech-to-singing task, the accompanying auxiliary input information for the speech (such as phonemes and notes) should be provided. You can employ ASR to recognize the words within the speech and then utilize language-specific tools (for example, pypinyin for Chinese) to extract the necessary input phonemes. Additionally, you can use ROSVOT to directly extract MIDI data from the speech - this is merely for the sake of input consistency. If you intend to conduct special processing on speech, such as using paired speech from GTSinger for training, you can omit the note information during training.