This is ACE-Step paper implementation (Unofficial) from scratch. This paper's goal is to build a model to generate text-to-song & features.
- Linear Attention
- DiT (Diffussion transformer) implemented
- Converted songs into mel-spectrogram
Tags-> mT5 encoderLyrics-> VoiceBPE Tokenizer- Cross Attention
- RoPE implement
- Implement Training pipeline
- Clone the repository:
[email protected]:Imesh7/ace-step.git
- Setup enviornment
conda env create --name ace-step -f environment.yml
- Activate the conda environment:
conda activate ace-step
Install Dependecies
conda install --file requirements.txt
├─── model
│ ├─── autoencoder
| | ├─── autoencoder.py
| | ├─── encoder.py
| | └─── encoder.py
| |
│ ├─── DiT
| | └─── dit.py
| |
│ └─── transformer
| | ├─── attention.py
| | ├─── cross_attention.py
| | └─── mix_feed_forward.py
| |
| ├─── RoPE.py
| └─── m5_encoder.py
|
├─── notebook
├─── tests
├─── train.py
└─── inference.py
- Train the Autoencoder & Upload it to Huggineface
- Implement Encoders for music
- Flash Attention 3