Wild4D builds language and question-answering models over dynamic 4D point clouds reconstructed from monocular video. It combines VGGT-Omega geometry, DynamicVerse RGB/depth/masks, Utonia point features, an LLM/VLM alignment stack, and versioned grounded-QA generation pipelines.
| Area | Purpose |
|---|---|
vggt-omega/ |
VGGT-Omega camera/depth prediction and static/dynamic 4D reconstruction |
third_party/Utonia/ |
vendored Utonia encoder plus Wild4D feature/PCA integrations |
wild4d_common/ |
portable filesystem configuration shared by Python entry points |
wild4d_llm/ |
Utonia-token → MLP/Q-Former → decoder-only LLM training and evaluation |
wild4d_vlm/ |
Qwen2.5-VL RGB-frame baseline |
wild4d_qa/ |
text QA v3/v4/v5 and binary geometry QA construction |
scripts/reconstruction/ |
DynamicVerse batch reconstruction launchers |
scripts/cache/ |
six-way token-sampling cache extraction |
scripts/training/ |
caption/QA LLM and VLM train/eval launchers |
scripts/qa/ |
versioned production and pilot QA drivers |
scripts/serving/ |
point-cloud viewers and Qwen3.6 vLLM workers |
scripts/setup/ |
reproducible component-specific environments |
scripts/hf/, assets/ |
gated Hugging Face manifest and upload workflow |
Legacy SegAnyMo/GeoMotion prototype scripts with unavailable external dependencies are isolated
under scripts/legacy/ and are not part of the supported DynamicVerse pipeline.
DynamicVerse RGB + masks
│
├── VGGT-Omega camera/depth ──▶ per-frame 4D point clouds
│ │
│ └── Utonia encoder ──▶ [frames, tokens, features]
│ │
│ ┌─────────────────────────────────────┴─────────────┐
│ │ │
│ MLP / Q-Former QA construction
│ │ text v3/v4/v5 + geometry
│ ▼ │
└─────────────────▶ decoder LLM ◀─────────────────────────────────────────┘
RGB frames ──▶ Qwen2.5-VL LoRA baseline (same splits and metrics)
The sampling ablation keeps 256 visual tokens per scene and compares:
- 16 frames × 16 tokens vs. 8 frames × 32 tokens;
- random vs. FPS vs. FPS + curvature saliency;
- MLP vs. 32-query Q-Former alignment.
Large files are intentionally outside Git. Portable defaults resolve through ignored local links:
Wild4D/
├── data/dynamicVerse -> source DynamicVerse dataset
├── outputs/dynamicVerse -> generated datasets, experiments, checkpoints
├── checkpoints/ -> optional local links to gated third-party weights
└── .venvs/ -> component environments
outputs/dynamicVerse/
├── datasets/ # reconstructed point-cloud datasets
├── experiments/ # captioning, ablation, QA, profiling, tracking
└── checkpoints/ # Wild4D checkpoints + manifest; third_party is excluded from upload
Copy .env.example or export the path variables before running on another server:
export WILD4D_DATA_ROOT=/data/dynamicVerse
export WILD4D_OUTPUT_ROOT=/data/4d_outputs/dynamicVerse
export WILD4D_ENV_ROOT=/data/wild4d_envs
export HF_HOME=/data/huggingface_cache
source scripts/lib/wild4d_env.shThe model families use incompatible Torch/Transformers stacks, so environments are separated.
See environment/README.md for details.
git clone https://github.com/Ever2after/Wild4D.git
cd Wild4D
bash scripts/setup/bootstrap.sh core # LLM, QA, analysis, HF tools
bash scripts/setup/bootstrap.sh vggt # VGGT-Omega reconstruction
bash scripts/setup/bootstrap.sh utonia # point feature extraction
bash scripts/setup/bootstrap.sh vlm # Qwen2.5-VL baseline
bash scripts/setup/bootstrap.sh qwen # Qwen3.6 vLLM servingVGGT-Omega weights must be requested from the official gated Hugging Face repository and placed
at checkpoints/vggt_omega_1b_512.pt (or provided through VGGT_CHECKPOINT). They are never
included in Git or Wild4D uploads.
Reconstruct one scene or run a batch:
bash scripts/reconstruction/run_4d_dynamicverse.sh DAVIS/soccerball
GPUS=0,1,2,3,4,5,6,7 bash scripts/reconstruction/run_4d_dynamicverse_batch.shExtract the six sampling caches and run the 12 alignment ablations:
GPUS=0,1,2,3,4,5,6,7 bash scripts/cache/run_wild4d_llm_ablation_cache_10k.sh
GPUS=0,1,2,3 bash scripts/training/run_wild4d_llm_ablation_12.shTrain/evaluate caption and QA models:
GPUS=0,1,2,3 bash scripts/training/run_wild4d_llm_train.sh
GPU=0 bash scripts/training/run_wild4d_llm_eval.sh
SPLIT_DIR=/path/to/splits GPUS=0,1,2,3 bash scripts/training/run_wild4d_llm_qa_train.sh
SPLIT_DIR=/path/to/splits GPU=0 BLIND=1 CKPT=/path/to/model.pt \\
bash scripts/training/run_wild4d_llm_qa_eval.shGenerate versioned QA:
bash scripts/qa/v3/run_generation.sh
bash scripts/qa/v4/run_10k.sh
bash scripts/qa/v5/run_pilot.sh
bash scripts/qa/v5/run_v51_pilot.shSee wild4d_llm/README.md, wild4d_vlm/README.md, and wild4d_qa/README.md for model and QA
implementation details.
Asset collections are declared in assets/collections.json. Always build and inspect a manifest
before upload. Upload is a dry-run unless --execute is explicitly supplied.
python scripts/hf/build_asset_manifest.py checkpoints --hash sha256
python scripts/hf/upload_assets.py checkpoints --repo-id ORG/Wild4D-Checkpoints
python scripts/hf/upload_assets.py checkpoints --repo-id ORG/Wild4D-Checkpoints --executeRaw DynamicVerse data and gated VGGT-Omega checkpoints are blocked from re-upload. See
assets/README.md and THIRD_PARTY_NOTICES.md for the collection-level policy.
python -m compileall -q wild4d_common wild4d_llm wild4d_vlm wild4d_qa scripts
bash -n scripts/{reconstruction,training,cache,serving,setup}/*.sh
python scripts/cache/extract_wild4d_llm_ablation_caches.py --help
python scripts/hf/upload_assets.py third_party_weights --repo-id example/blockedA project-level Wild4D license has not yet been selected. Third-party code, source data, model
weights, and derived artifacts retain their own terms; consult THIRD_PARTY_NOTICES.md before
publishing code or assets. Add the Wild4D paper citation here before the public release.