Imesh7/ace-step

ACE-STEP paper implementation (Unoffical)

★ 1Forks 0Jupyter NotebookGitHub ↗Compare
diffusion-modelsditmusic-generation-deep-learning

README

ACE Step

This is ACE-Step paper implementation (Unofficial) from scratch. This paper's goal is to build a model to generate text-to-song & features.

Paper

Key features

  • Linear Attention
  • DiT (Diffussion transformer) implemented
  • Converted songs into mel-spectrogram
  • Tags -> mT5 encoder
  • Lyrics -> VoiceBPE Tokenizer
  • Cross Attention
  • RoPE implement
  • Implement Training pipeline

Environment Setup

  1. Clone the repository:
[email protected]:Imesh7/ace-step.git
  1. Setup enviornment
conda env create --name ace-step -f environment.yml
  1. Activate the conda environment:
conda activate ace-step

Dependency

Install Dependecies

conda install --file requirements.txt

Folder structure

├─── model
│     ├─── autoencoder
|     |        ├─── autoencoder.py
|     |        ├─── encoder.py
|     |        └─── encoder.py
|     |
│     ├─── DiT
|     |      └─── dit.py
|     |
│     └─── transformer
|     |        ├─── attention.py
|     |        ├─── cross_attention.py
|     |        └─── mix_feed_forward.py
|     |
|     ├─── RoPE.py
|     └─── m5_encoder.py
|     
├─── notebook
├─── tests
├─── train.py
└─── inference.py

What's next

  • Train the Autoencoder & Upload it to Huggineface
  • Implement Encoders for music
  • Flash Attention 3

Overall Architecture

image

Contributors

Imesh7imesh-sp

Issues