YuCao16/LiteVSR

★ 6Forks 0PythonGitHub ↗Compare

README

LiteVSR: Unleashing the Potential of Frozen Diffusion Transformers for Video Super-Resolution

Yu Cao1, Ziquan Liu1, Zhensong Zhang2, Jiankang Deng3, Shaogang Gong1, Jifei Song2

1Queen Mary University of London   2Huawei Darwin Research Center   3Imperial College London

Project page  |  Paper


LiteVSR performs Video Super-Resolution using a completely frozen Diffusion Transformer with a lightweight State-Aware Adapter. This repository provides inference code and pretrained weights.

News

  • [2026-07] Inference code and Wan2.2-5B adapter released.
  • [2026-06] Paper accepted to ICML 2026.

TODO

  • Release inference code
  • Release Wan2.2-5B pretrained adapter
  • Release Wan2.1-1B pretrained adapter
  • Release training code
  • Gradio demo

Installation

Requires CUDA 12.x, a GPU with ≥ 24 GB VRAM, and conda to manage the Python interpreter.

1. Create the environment

conda create -n litevsr python=3.11 -y
conda activate litevsr

2. Clone and set up

git clone https://github.com/YuCao16/LiteVSR
cd LiteVSR
bash install.sh

install.sh initializes the Wan2.x submodules, slims their top-level __init__.py files to skip unused upstream re-exports, installs uv into the current conda env, and installs the minimal dependency set with uv pip install ..

Optional FlashAttention (~20–50 % faster, ~30 % less memory; requires nvcc and several minutes to compile):

uv pip install '.[flash]' --extra-index-url https://download.pytorch.org/whl/cu121

Without it LiteVSR falls back to PyTorch's scaled_dot_product_attention automatically.

Optional SciPy backend for the UniPC scheduler:

uv pip install '.[scipy]'

3. Download weights

Base model (Wan2.2-TI2V-5B, ~28 GB):

huggingface-cli download Wan-AI/Wan2.2-TI2V-5B --local-dir checkpoints/Wan2.2-TI2V-5B

LiteVSR adapter (~2.4 GB): download litevsr_wan2_2_5b.ckpt from the weights folder and place it under checkpoints/.

Final layout:

checkpoints/
├── Wan2.2-TI2V-5B/
│   ├── ...
│   └── Wan2.2_VAE.pth
└── litevsr_wan2_2_5b.ckpt

Quick Start

python scripts/infer.py \
  --checkpoint checkpoints/litevsr_wan2_2_5b.ckpt \
  --config     configs/wan2_2_5b.yaml \
  --input_dir  /path/to/lq_videos \
  --output_dir /path/to/sr_outputs \
  --upscale 4 --num_steps 5

Common options: --num_steps (fewer = faster), --tile_size H W (spatial tiling for large inputs), --upscale (SR factor), --crf (output video quality, 0 = lossless, 23 = default). Run python scripts/infer.py --help for the full list.

Reproducing Paper Results

Datasets

We provide the exact LQ / GT splits used in the paper on Google Drive. Download and unpack under datasets/:

datasets/
├── REDS4/    { LQ-Video/, GT-Video/ }
├── UDM10/    { LQ-Video/, GT-Video/ }
├── SPMCS/    { LQ-Video/, GT-Video/ }
├── YouHQ40/  { LQ-Video/, GT-Video/ }
└── VideoLQ/  { LQ-Video/, GT-Video/ }

REDS4 / VideoLQ are subsampled from the original REDS and VideoLQ releases (Nah et al., Chan et al.).

Inference

Run all five datasets:

bash reproduce.sh

Override paths via env vars if needed:

CKPT=path/to/ckpt DATA_ROOT=path/to/datasets OUT_ROOT=path/to/out \
  bash reproduce.sh

Or open reproduce.sh and run a single dataset block by hand.

Evaluation

Evaluation code (PSNR / SSIM / LPIPS / DISTS / CLIPIQA / DOVER / NIQE / MUSIQ) will be released together with the training code.

Citation

@article{cao2026litevsr,
  author  = {Cao, Yu and Liu, Ziquan and Zhang, Zhensong and Deng, Jiankang and Gong, Shaogang and Song, Jifei},
  title   = {LiteVSR: Unleashing the Potential of Frozen Diffusion Transformers for Video Super-Resolution},
  journal = {arXiv preprint arXiv:2606.09250},
  year    = {2026},
}

Acknowledgements

LiteVSR is built on top of Wan2.2. We thank the Wan-AI team for open-sourcing their models.

License

The LiteVSR code and pretrained weights are released for non-commercial research use only under the LiteVSR Non-Commercial Research License. The training method is covered by pending patent applications; no patent license is granted. For commercial licensing, please contact the corresponding author.

Third-party components (Wan2.1, Wan2.2) remain under their original licenses (Apache 2.0).

Contributors

YuCao16

Issues