** Converted samples coming soon **
A pytorch implementation based on: StarGAN-VC2: https://arxiv.org/pdf/1907.12279.pdf.
- Currently does not implement source-and-target adversarial loss.
- Makes use of gradient penalty.
- Doesnt make use of PS in G.
Tested on Python version 3.6.2 in a linux VM environment
Recommended to use a linux environment - not tested for mac or windows OS
- Create a new environment using Anaconda
conda create -n stargan-vc python=3.6.2- Install conda dependencies
conda install pytorch=1.4.0 torchvision=0.5.0 cudatoolkit=10.1 -c pytorch
conda install pillow=5.4.1
conda install -c conda-forge librosa=0.6.1
conda install -c conda-forge tqdm=4.43.0- Intall dependencies not available through conda using pip
pip install pyworld==0.2.8
pip install mcd==0.4NB: For mac users who cannot install pyworld see: https://github.com/JeremyCCHsu/Python-Wrapper-for-World-Vocoder
- Install binaries
mkdir ../data/VCTK-Data
wget https://datashare.is.ed.ac.uk/bitstream/handle/10283/2651/VCTK-Corpus.zip?sequence=2&isAllowed=y
unzip VCTK-Corpus.zip -d ../data/VCTK-DataIf the downloaded VCTK is in tar.gz, run this:
tar -xzvf VCTK-Corpus.tar.gz -C ../data/VCTK-Data- VCC2016 and 2018 are yet to be included
We will use Mel-Cepstral coefficients(MCEPs) here.
This example script is for the VCTK data which needs resampling to 16kHz, the script allows you to preprocess the data without resampling either. This script assumes the data dir to be ../data/VCTK-Data/
# VCTK-Data
python preprocess.py --perform_data_split y \
--resample_rate 16000 \
--origin_wavpath ../data/VCTK-Data/VCTK-Corpus/wav48 \
--target_wavpath ../data/VCTK-Data/VCTK-Corpus/wav16 \
--mc_dir_train ../data/VCTK-Data/mc/train \
--mc_dir_test ../data/VCTK-Data/mc/test \
--speaker_dirs p262 p272 p229 p232- Currently only tested with conversion between 4 speakers
- Not yet tested with use of tensorboard
Example script:
# example with VCTK
python main.py --train_data_dir ../data/VCTK-Data/mc/train \
--test_data_dir ../data/VCTK-Data/mc/test \
--use_tensorboard False \
--wav_dir ../data/VCTK-Data/VCTK-Corpus/wav16 \
--model_save_dir ../data/VCTK-Data/models \
--sample_dir ../data/VCTK-Data/samples \
--num_iters 200000 \
--batch_size 8 \
--speakers p262 p272 p229 p232 \
--num_speakers 4If you encounter an error such as:
ImportError: /lib64/libstdc++.so.6: version `CXXABI_1.3.9' not foundYou may need to export export LD_LIBRARY_PATH: (See Stack Overflow)
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/<PATH>/<TO>/<YOUR>/.conda/envs/<ENV>/lib/For example: restore model at step 120000 and specify the speakers
# example with VCTK
python convert.py --resume_model 120000 \
--sampling_rate 16000 \
--num_speakers 4 \
--speakers p262 p272 p229 p232 \
--train_data_dir ../data/VCTK-Data/mc/train/ \
--test_data_dir ../data/VCTK-Data/mc/test/ \
--wav_dir ../data/VCTK-Data/VCTK-Corpus/wav16 \
--model_save_dir ../data/VCTK-Data/models \
--convert_dir ../data/VCTK-Data/converted \
--num_converted_wavs 4This saves your converted flies to ../data/VCTK-Data/converted/120000/
Calculate the Mel Cepstral Distortion of the reference speaker vs the synthesized speaker. Use --spk_to_spk tag to define multiple speaker to speaker folders generated with the convert script.
python mel_cep_distance.py --convert_dir ../data/VCTK-Data/converted/120000 \
--spk_to_spk p262_to_p272 \
--output_csv p262_to_p272.csv- Include converted samples
- Include MCD examples
- Include s-t loss like original paper