MTDickens/WorldDirector

โ˜… 0Forks 0GitHub โ†—Compare

README

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

[๐Ÿ“„ Paper] [๐ŸŒ Project Page] [๐Ÿค— Model Weights]

worlddirector_demo.mp4

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen

TLDR

WorldDirector is a controllable video world model framework for persistent dynamic object memory and unrestricted viewpoint exploration. It decouples semantic motion orchestration from visual generation: an LLM plans 3D object and camera trajectories, these plans are projected into 2D location controls, and appearance binding preserves object identity when dynamic entities leave the view and later re-enter. This enables long-horizon, causally generated world simulation with controllable object actions, camera motion, object permanence, and stable visual appearance.

Strongly recommend seeing our demo page.

If you enjoyed the videos we created, please consider giving us a star ๐ŸŒŸ.

๐Ÿš€ Open-Source Plan

โœ… Released

  • Full inference code
  • WorldDirector-14B

Setup

This codebase follows an environment setup similar to LingBot-World.

Installation

Clone the repo:

git clone https://github.com/pPetrichor/WorldDirector
cd WorldDirector

Install dependencies:

# Ensure torch >= 2.4.0
pip install -r generate/requirements.txt

Install FlashAttention:

pip install flash-attn --no-build-isolation

Model Download

Model Resolution Download Links
WorldDirector-14B 480P ๐Ÿค— HuggingFace

Download models using huggingface-cli:

pip install "huggingface_hub[cli]"
huggingface-cli download hlwang06/WorldDirector  --local-dir ./WorldDirector-14B

Inference

WorldDirector first uses an LLM to plan the motion of dynamic objects and the camera trajectory. This planning stage can be customized by users, as long as the final output follows a JSON format similar to LLM_plan.json, which describes all dynamic object motions and camera movements throughout the video.

Below is our inference pipeline.

1. Plan Dynamic Objects and Camera Motion

First, choose an input image to operate on. If the image already contains the dynamic objects you want to control, estimate their initial 3D bounding boxes before LLM planning. We provide a simple SAM- and DepthAnything-based tool for this step; please see 3D_bbox_infer for details. If the input image does not contain the dynamic objects to be designed, this step can be skipped.

Next, use an LLM to generate the JSON file that describes the full-video motion of all dynamic objects and the camera. We provide our complete prompt template in prompt.txt. After running the code generated by Gemini 3.1 Pro, you should obtain an output similar to LLM_plan.json, together with a 3D box visualization video.

At this stage, users can freely adjust the prompt, including the video duration, camera motion, and object behaviors. The LLM can also design 3D bounding-box trajectories for new dynamic objects, even if those objects do not appear in the input image. The planned video duration should be a multiple of 5 seconds, and the total frame count must satisfy 80*n+1 (for example, 81, 161, or 241 frames at 16 FPS).

2. Prepare Latents and Conditions

After LLM planning, run the following command to extract video latents and compute the box condition and appearance condition used for generation. Please refer to the paper for more details.

cd worldlatent_prepare
python extract_games_hacking_from_json.py \
  --input_json <PATH_TO_LLM_GENERATED_JSON> \
  --output_dir <OUTPUT_DIR> \
  --vae_path <PATH_TO_Wan2.1_VAE.pth>

The output directory should contain several segment_{n} folders, where each folder represents a 5-second video segment. WorldDirector then causally generates the full video chunk by chunk.

Before generation, configure eval_dataset.json by setting the text prompt for each segment. The text field should be a list: the first item is the global prompt, and the following items are local prompt descriptions for each dynamic object. Use the generated text_subject_map.json to determine the order of the local object prompts. You also need to fill in the initial image_path in the first entry of eval_dataset.json with the path to the input image. We provide three example prompt settings on Hugging Face for reference. If you want to directly test these examples, please update image_path and all file paths in the corresponding eval_dataset.json files to your local paths.

3. Causal Video Generation

Finally, run causal video generation:

cd ../generate
torchrun --nproc_per_node=8 generate_channel.py \
  --task i2v-A14B \
  --size 480*832 \
  --ckpt_dir <PATH_TO_WEIGHTS> \
  --eval_dataset_path <PATH_TO_EVAL_DATASET_JSON> \
  --save_dir <OUTPUT_DIR> \
  --frame_num <GENERATED_VIDEO_FRAME_NUM> \
  --sample_steps 50 \
  --sample_shift 10.0 \
  --sample_guide_scale 1.0 \
  --fps 16 \
  --height 480 \
  --width 832 \
  --box_attn_bias_value 3.5 \
  --prefix_num_total 10 \
  --ulysses_size 8 \
  --dit_fsdp \
  --t5_fsdp

Here, <PATH_TO_WEIGHTS> is the local directory containing the downloaded WorldDirector model weights, and <PATH_TO_EVAL_DATASET_JSON> is the path to the configured eval_dataset.json generated in the previous step.

Demo Results

We provide example generation results below. Each video shows the generated result on the left and the corresponding 3D plan on the right.

demo_woman_with_plan.mp4

demo_walk_with_plan.mp4

demo_fire_with_plan.mp4

The generated videos preserve object permanence and autonomous motion for dynamic entities, while maintaining consistent dynamic object memory.

Users are not limited to objects that appear in the input image. The LLM can also design new objects and use text prompts to control their generation, as shown below:

text_control_with_plan.mp4

๐Ÿ“š Related Projects

โœจ Acknowledgement

We would like to express our gratitude to the Wan Team for open-sourcing their code and models. Their contributions have been instrumental to the development of this project.

We also thank LingBot-World for its open-source world-model codebase and environment setup, which helped support this project.

Citation

If you find this work useful, please consider citing our paper:

@article{wang2026worlddirector,
  title={WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory},
  author={Hanlin Wang and Hao Ouyang and Qiuyu Wang and Wen Wang and Qingyan Bai and Ka Leong Cheng and Yue Yu and Yixuan Li and Yihao Meng and Zichen Liu and Yanhong Zeng and Yujun Shen and Qifeng Chen},
  journal={arXiv preprint arXiv:2607.02517},
  year={2026}
}

License

This project is licensed under the CC BY-NC-SA 4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License).

The code is provided for academic research purposes only.

For any questions, please contact [email protected].

Contributors

pPetrichor

Issues