Train single-speaker RVC v2 models and batch-convert audio to target speaker voices.
Training: Source audio is sliced into ~3.7s segments, then HuBERT extracts content features (768-dim) and RMVPE extracts F0 pitch. The posterior encoder (enc_q) maps mel spectrograms to a VAE latent space z, which is decoded by NSF-HiFiGAN into waveform. A multi-period discriminator (MPD) provides adversarial feedback. A normalizing flow module bridges the prior and posterior distributions.
Inference: Input audio is processed through HuBERT (content) and RMVPE (pitch). FAISS retrieval blends input features with the target speaker's stored features (75/25 ratio). The phone encoder (enc_p) and reverse flow produce the latent z, which NSF-HiFiGAN decodes into the target speaker's voice.
# Clone RVC repo
git clone https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.git /workspace/rvc
# Install dependencies
pip install torch torchaudio fairseq faiss-gpu numpy soundfile librosa \
pyworld wandb pyyaml gcsfs huggingface_hub matplotlib tensorboard
# Pretrained models are auto-downloaded on first run# From HuggingFace dataset
python preprocess.py --config config.yaml --speaker speaker_2 \
--hf-dataset rumik-ai/real-voice-batch-3 --hf-speaker SP001
# From local audio directory
python preprocess.py --config config.yaml --speaker speaker_2 \
--local-audio-dir /path/to/speaker_wavs
# From GCS
python preprocess.py --config config.yaml --speaker speaker_2 \
--gcs-metadata gs://bucket/metadata.jsonl --num-samples 400# Basic training (50 epochs recommended to avoid mode collapse)
python train.py --config config.yaml --speaker speaker_2
# With W&B logging
python train.py --config config.yaml --speaker speaker_2 \
--wandb-project rvc-speaker-training --wandb-run speaker2-v1
# Custom epochs
python train.py --config config.yaml --speaker speaker_2 --epochs 30# Single file conversion
python inference.py --config config.yaml --speaker speaker_2 \
--input source_audio.wav --output converted.wav
# Batch convert a directory
python inference.py --config config.yaml --speaker speaker_2 \
--input-dir /path/to/source_wavs --output-dir /path/to/converted
# Batch from JSONL metadata
python inference.py --config config.yaml --speaker speaker_2 \
--input-jsonl dataset/metadata.jsonl --output-dir /path/to/converted
# With pitch shift (+2 semitones)
python inference.py --config config.yaml --speaker speaker_2 \
--input source.wav --output converted.wav --pitch 2python upload.py --config config.yaml --speaker speaker_2 \
--hf-repo rumik-ai/speaker-2-RVCAll settings are in config.yaml. Key options:
| Setting | Default | Description |
|---|---|---|
training.epochs |
50 | Training epochs (collapse ~60-70, stop early) |
training.fp16 |
false | Disabled for H100 stability |
training.grad_clip_norm_g |
5.0 | Generator gradient norm clip |
training.grad_clip_norm_d |
1.0 | Discriminator gradient norm clip |
preprocessing.slice_duration |
3.7 | Audio segment length in seconds |
inference.index_ratio |
0.75 | FAISS index blend ratio |
- Mode collapse at ~60-70 epochs: RVC's GAN discriminator overpowers the generator. Keep epochs at 50 or lower. Best quality is typically around epoch 25-30.
- fp16 causes NaN on H100: Use
fp16: falsein config. bf16 not supported by RVC's GradScaler. - Synthetic/TTS audio: RVC training collapses quickly on TTS-generated audio. Use real speech recordings only. For TTS audio, use zero-shot VC (Seed-VC) instead.
rvc_pipeline/
├── config.yaml # Configuration
├── preprocess.py # Data preprocessing (download, slice, F0, HuBERT)
├── train.py # Training with W&B logging
├── inference.py # Single file and batch voice conversion
├── upload.py # Push models to HuggingFace
├── utils.py # Shared utilities
└── README.md
/workspace/rvc_models/speaker_2/
├── speaker_2.pth # Extracted model (small, for inference)
├── speaker_2.index # FAISS index for feature retrieval
└── metadata.json # Training metadata
/workspace/rvc/logs/speaker_2/
├── 0_gt_wavs/ # Preprocessed audio segments
├── 2a_f0/ # Coarse F0 (quantized)
├── 2b-f0nsf/ # Continuous F0
├── 3_feature768/ # HuBERT features
├── filelist.txt # Training file list
├── config.json # RVC training config
├── G_*.pth # Generator checkpoints
└── D_*.pth # Discriminator checkpoints