[๐ Paper] [๐ Project Page] [๐ค Model Weights]
worlddirector_demo.mp4
Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen
WorldDirector is a controllable video world model framework for persistent dynamic object memory and unrestricted viewpoint exploration. It decouples semantic motion orchestration from visual generation: an LLM plans 3D object and camera trajectories, these plans are projected into 2D location controls, and appearance binding preserves object identity when dynamic entities leave the view and later re-enter. This enables long-horizon, causally generated world simulation with controllable object actions, camera motion, object permanence, and stable visual appearance.
Strongly recommend seeing our demo page.
If you enjoyed the videos we created, please consider giving us a star ๐.
- Full inference code
WorldDirector-14B
This codebase follows an environment setup similar to LingBot-World.
Clone the repo:
git clone https://github.com/pPetrichor/WorldDirector
cd WorldDirectorInstall dependencies:
# Ensure torch >= 2.4.0
pip install -r generate/requirements.txtInstall FlashAttention:
pip install flash-attn --no-build-isolation| Model | Resolution | Download Links |
|---|---|---|
| WorldDirector-14B | 480P | ๐ค HuggingFace |
Download models using huggingface-cli:
pip install "huggingface_hub[cli]"
huggingface-cli download hlwang06/WorldDirector --local-dir ./WorldDirector-14BWorldDirector first uses an LLM to plan the motion of dynamic objects and the camera trajectory. This planning stage can be customized by users, as long as the final output follows a JSON format similar to LLM_plan.json, which describes all dynamic object motions and camera movements throughout the video.
Below is our inference pipeline.
First, choose an input image to operate on. If the image already contains the dynamic objects you want to control, estimate their initial 3D bounding boxes before LLM planning. We provide a simple SAM- and DepthAnything-based tool for this step; please see 3D_bbox_infer for details. If the input image does not contain the dynamic objects to be designed, this step can be skipped.
Next, use an LLM to generate the JSON file that describes the full-video motion of all dynamic objects and the camera. We provide our complete prompt template in prompt.txt. After running the code generated by Gemini 3.1 Pro, you should obtain an output similar to LLM_plan.json, together with a 3D box visualization video.
At this stage, users can freely adjust the prompt, including the video duration, camera motion, and object behaviors. The LLM can also design 3D bounding-box trajectories for new dynamic objects, even if those objects do not appear in the input image. The planned video duration should be a multiple of 5 seconds, and the total frame count must satisfy 80*n+1 (for example, 81, 161, or 241 frames at 16 FPS).
After LLM planning, run the following command to extract video latents and compute the box condition and appearance condition used for generation. Please refer to the paper for more details.
cd worldlatent_prepare
python extract_games_hacking_from_json.py \
--input_json <PATH_TO_LLM_GENERATED_JSON> \
--output_dir <OUTPUT_DIR> \
--vae_path <PATH_TO_Wan2.1_VAE.pth>The output directory should contain several segment_{n} folders, where each folder represents a 5-second video segment. WorldDirector then causally generates the full video chunk by chunk.
Before generation, configure eval_dataset.json by setting the text prompt for each segment. The text field should be a list: the first item is the global prompt, and the following items are local prompt descriptions for each dynamic object. Use the generated text_subject_map.json to determine the order of the local object prompts. You also need to fill in the initial image_path in the first entry of eval_dataset.json with the path to the input image. We provide three example prompt settings on Hugging Face for reference. If you want to directly test these examples, please update image_path and all file paths in the corresponding eval_dataset.json files to your local paths.
Finally, run causal video generation:
cd ../generate
torchrun --nproc_per_node=8 generate_channel.py \
--task i2v-A14B \
--size 480*832 \
--ckpt_dir <PATH_TO_WEIGHTS> \
--eval_dataset_path <PATH_TO_EVAL_DATASET_JSON> \
--save_dir <OUTPUT_DIR> \
--frame_num <GENERATED_VIDEO_FRAME_NUM> \
--sample_steps 50 \
--sample_shift 10.0 \
--sample_guide_scale 1.0 \
--fps 16 \
--height 480 \
--width 832 \
--box_attn_bias_value 3.5 \
--prefix_num_total 10 \
--ulysses_size 8 \
--dit_fsdp \
--t5_fsdpHere, <PATH_TO_WEIGHTS> is the local directory containing the downloaded WorldDirector model weights, and <PATH_TO_EVAL_DATASET_JSON> is the path to the configured eval_dataset.json generated in the previous step.
We provide example generation results below. Each video shows the generated result on the left and the corresponding 3D plan on the right.
demo_woman_with_plan.mp4
demo_walk_with_plan.mp4
demo_fire_with_plan.mp4
The generated videos preserve object permanence and autonomous motion for dynamic entities, while maintaining consistent dynamic object memory.
Users are not limited to objects that appear in the input image. The LLM can also design new objects and use text prompts to control their generation, as shown below:
text_control_with_plan.mp4
We would like to express our gratitude to the Wan Team for open-sourcing their code and models. Their contributions have been instrumental to the development of this project.
We also thank LingBot-World for its open-source world-model codebase and environment setup, which helped support this project.
If you find this work useful, please consider citing our paper:
@article{wang2026worlddirector,
title={WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory},
author={Hanlin Wang and Hao Ouyang and Qiuyu Wang and Wen Wang and Qingyan Bai and Ka Leong Cheng and Yue Yu and Yixuan Li and Yihao Meng and Zichen Liu and Yanhong Zeng and Yujun Shen and Qifeng Chen},
journal={arXiv preprint arXiv:2607.02517},
year={2026}
}This project is licensed under the CC BY-NC-SA 4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License).
The code is provided for academic research purposes only.
For any questions, please contact [email protected].