nullHawk/custom-RVC

custom training and inference pipeline for Retrival-Based Voice Conversion models

★ 1Forks 0PythonGitHub ↗Compare
deep-learninghifi-gantts

README

RVC v2 Voice Conversion Pipeline

Train single-speaker RVC v2 models and batch-convert audio to target speaker voices.

Architecture

RVC v2 Architecture

Training: Source audio is sliced into ~3.7s segments, then HuBERT extracts content features (768-dim) and RMVPE extracts F0 pitch. The posterior encoder (enc_q) maps mel spectrograms to a VAE latent space z, which is decoded by NSF-HiFiGAN into waveform. A multi-period discriminator (MPD) provides adversarial feedback. A normalizing flow module bridges the prior and posterior distributions.

Inference: Input audio is processed through HuBERT (content) and RMVPE (pitch). FAISS retrieval blends input features with the target speaker's stored features (75/25 ratio). The phone encoder (enc_p) and reverse flow produce the latent z, which NSF-HiFiGAN decodes into the target speaker's voice.

Setup

# Clone RVC repo
git clone https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.git /workspace/rvc

# Install dependencies
pip install torch torchaudio fairseq faiss-gpu numpy soundfile librosa \
    pyworld wandb pyyaml gcsfs huggingface_hub matplotlib tensorboard

# Pretrained models are auto-downloaded on first run

Quick Start

1. Preprocess

# From HuggingFace dataset
python preprocess.py --config config.yaml --speaker speaker_2 \
    --hf-dataset rumik-ai/real-voice-batch-3 --hf-speaker SP001

# From local audio directory
python preprocess.py --config config.yaml --speaker speaker_2 \
    --local-audio-dir /path/to/speaker_wavs

# From GCS
python preprocess.py --config config.yaml --speaker speaker_2 \
    --gcs-metadata gs://bucket/metadata.jsonl --num-samples 400

2. Train

# Basic training (50 epochs recommended to avoid mode collapse)
python train.py --config config.yaml --speaker speaker_2

# With W&B logging
python train.py --config config.yaml --speaker speaker_2 \
    --wandb-project rvc-speaker-training --wandb-run speaker2-v1

# Custom epochs
python train.py --config config.yaml --speaker speaker_2 --epochs 30

3. Inference

# Single file conversion
python inference.py --config config.yaml --speaker speaker_2 \
    --input source_audio.wav --output converted.wav

# Batch convert a directory
python inference.py --config config.yaml --speaker speaker_2 \
    --input-dir /path/to/source_wavs --output-dir /path/to/converted

# Batch from JSONL metadata
python inference.py --config config.yaml --speaker speaker_2 \
    --input-jsonl dataset/metadata.jsonl --output-dir /path/to/converted

# With pitch shift (+2 semitones)
python inference.py --config config.yaml --speaker speaker_2 \
    --input source.wav --output converted.wav --pitch 2

4. Upload to HuggingFace

python upload.py --config config.yaml --speaker speaker_2 \
    --hf-repo rumik-ai/speaker-2-RVC

Configuration

All settings are in config.yaml. Key options:

Setting Default Description
training.epochs 50 Training epochs (collapse ~60-70, stop early)
training.fp16 false Disabled for H100 stability
training.grad_clip_norm_g 5.0 Generator gradient norm clip
training.grad_clip_norm_d 1.0 Discriminator gradient norm clip
preprocessing.slice_duration 3.7 Audio segment length in seconds
inference.index_ratio 0.75 FAISS index blend ratio

Known Issues

  • Mode collapse at ~60-70 epochs: RVC's GAN discriminator overpowers the generator. Keep epochs at 50 or lower. Best quality is typically around epoch 25-30.
  • fp16 causes NaN on H100: Use fp16: false in config. bf16 not supported by RVC's GradScaler.
  • Synthetic/TTS audio: RVC training collapses quickly on TTS-generated audio. Use real speech recordings only. For TTS audio, use zero-shot VC (Seed-VC) instead.

File Structure

rvc_pipeline/
├── config.yaml      # Configuration
├── preprocess.py    # Data preprocessing (download, slice, F0, HuBERT)
├── train.py         # Training with W&B logging
├── inference.py     # Single file and batch voice conversion
├── upload.py        # Push models to HuggingFace
├── utils.py         # Shared utilities
└── README.md

Output Structure

/workspace/rvc_models/speaker_2/
├── speaker_2.pth      # Extracted model (small, for inference)
├── speaker_2.index    # FAISS index for feature retrieval
└── metadata.json      # Training metadata

/workspace/rvc/logs/speaker_2/
├── 0_gt_wavs/         # Preprocessed audio segments
├── 2a_f0/             # Coarse F0 (quantized)
├── 2b-f0nsf/          # Continuous F0
├── 3_feature768/      # HuBERT features
├── filelist.txt       # Training file list
├── config.json        # RVC training config
├── G_*.pth            # Generator checkpoints
└── D_*.pth            # Discriminator checkpoints

Contributors

nullHawk

Issues