Yu Cao1, Ziquan Liu1, Zhensong Zhang2, Jiankang Deng3, Shaogang Gong1, Jifei Song2
1Queen Mary University of London 2Huawei Darwin Research Center 3Imperial College London
LiteVSR performs Video Super-Resolution using a completely frozen Diffusion Transformer with a lightweight State-Aware Adapter. This repository provides inference code and pretrained weights.
- [2026-07] Inference code and Wan2.2-5B adapter released.
- [2026-06] Paper accepted to ICML 2026.
- Release inference code
- Release Wan2.2-5B pretrained adapter
- Release Wan2.1-1B pretrained adapter
- Release training code
- Gradio demo
Requires CUDA 12.x, a GPU with ≥ 24 GB VRAM, and conda to manage the Python interpreter.
conda create -n litevsr python=3.11 -y
conda activate litevsrgit clone https://github.com/YuCao16/LiteVSR
cd LiteVSR
bash install.shinstall.sh initializes the Wan2.x submodules, slims their top-level __init__.py files to skip unused upstream re-exports, installs uv into the current conda env, and installs the minimal dependency set with uv pip install ..
Optional FlashAttention (~20–50 % faster, ~30 % less memory; requires nvcc and several minutes to compile):
uv pip install '.[flash]' --extra-index-url https://download.pytorch.org/whl/cu121Without it LiteVSR falls back to PyTorch's scaled_dot_product_attention automatically.
Optional SciPy backend for the UniPC scheduler:
uv pip install '.[scipy]'Base model (Wan2.2-TI2V-5B, ~28 GB):
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B --local-dir checkpoints/Wan2.2-TI2V-5BLiteVSR adapter (~2.4 GB): download litevsr_wan2_2_5b.ckpt from the weights folder and place it under checkpoints/.
Final layout:
checkpoints/
├── Wan2.2-TI2V-5B/
│ ├── ...
│ └── Wan2.2_VAE.pth
└── litevsr_wan2_2_5b.ckpt
python scripts/infer.py \
--checkpoint checkpoints/litevsr_wan2_2_5b.ckpt \
--config configs/wan2_2_5b.yaml \
--input_dir /path/to/lq_videos \
--output_dir /path/to/sr_outputs \
--upscale 4 --num_steps 5Common options: --num_steps (fewer = faster), --tile_size H W (spatial tiling for large inputs), --upscale (SR factor), --crf (output video quality, 0 = lossless, 23 = default). Run python scripts/infer.py --help for the full list.
We provide the exact LQ / GT splits used in the paper on Google Drive. Download and unpack under datasets/:
datasets/
├── REDS4/ { LQ-Video/, GT-Video/ }
├── UDM10/ { LQ-Video/, GT-Video/ }
├── SPMCS/ { LQ-Video/, GT-Video/ }
├── YouHQ40/ { LQ-Video/, GT-Video/ }
└── VideoLQ/ { LQ-Video/, GT-Video/ }
REDS4 / VideoLQ are subsampled from the original REDS and VideoLQ releases (Nah et al., Chan et al.).
Run all five datasets:
bash reproduce.shOverride paths via env vars if needed:
CKPT=path/to/ckpt DATA_ROOT=path/to/datasets OUT_ROOT=path/to/out \
bash reproduce.shOr open reproduce.sh and run a single dataset block by hand.
Evaluation code (PSNR / SSIM / LPIPS / DISTS / CLIPIQA / DOVER / NIQE / MUSIQ) will be released together with the training code.
@article{cao2026litevsr,
author = {Cao, Yu and Liu, Ziquan and Zhang, Zhensong and Deng, Jiankang and Gong, Shaogang and Song, Jifei},
title = {LiteVSR: Unleashing the Potential of Frozen Diffusion Transformers for Video Super-Resolution},
journal = {arXiv preprint arXiv:2606.09250},
year = {2026},
}LiteVSR is built on top of Wan2.2. We thank the Wan-AI team for open-sourcing their models.
The LiteVSR code and pretrained weights are released for non-commercial research use only under the LiteVSR Non-Commercial Research License. The training method is covered by pending patent applications; no patent license is granted. For commercial licensing, please contact the corresponding author.
Third-party components (Wan2.1, Wan2.2) remain under their original licenses (Apache 2.0).