Labmem-Zhouyx/CDFSE_FastSpeech2
The Official Implementation of “Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis”
Focus on TTS/Speech/NLP. El Psy Congroo
The Official Implementation of “Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis”
A collection of links and notes on forced alignment tools
A tool for speech dataset to mel-spectrogram.
Claude Code Snapshot for Research. All original source code is the property of Anthropic.
The code of "Dependency Parsing based Semantic Representation Learning with Graph Neural Network for Enhancing Expressiveness of Text-to-Speech"
List of speech synthesis papers.
An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
PyTorch implementation of various methods for continual learning (XdG, EWC, online EWC, SI, LwF, GR, GR+distill, RtF, ER, A-GEM, iCaRL).
A PyTorch inplementation of character-based Tacotron2 for Chinese/Mandarin
A PyTorch inplementation of phoneme-based Tacotron2 for Chinese/Mandarin
Clone a voice in 5 seconds to generate arbitrary speech in real-time
STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech, INTERSPEECH 2021
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
An implementation of Attentron based on FastSpeech2 backbone
label-smooth, amsoftmax, partial-fc, focal-loss, triplet-loss, lovasz-softmax. Maybe useful
Please visit: https://thuhcsi.github.io/interspeech2022-cdfse-tts
Please visit: https://thuhcsi.github.io/interspeech2022-dependency-semantic-tts
Official repository of https://arxiv.org/abs/2111.04040v1
Tacotron 2 - PyTorch implementation with faster-than-realtime inference