A comprehensive Real2Sim2Real system featuring two independent pipelines for converting videos into 3D assets:
This project provides two separate, independent pipelines:
Generates 3D models of tracked objects from video:
- Video โ SAM2 Object Tracking: Extract and track objects using Grounded-SAM-2
- TRELLIS 3D Reconstruction: Convert tracked object images into 3D meshes (PLY, GLB, MP4)
- Output: Individual 3D object models ready for simulation
Generates environment mesh from video (no object tracking):
- Video โ OpenPano: Generate initial panorama from video frames
- Diffusion360 Enhancement: Enhance panorama quality using diffusion models
- Pano2Room Mesh Generation: Convert panorama into 3D room mesh
- Output: Complete environment mesh with textures ready for simulation
- Object Pipeline: Focus on specific objects in the scene (people, cars, furniture, etc.)
- Environment Pipeline: Focus on the overall scene/room structure (walls, floors, ceilings)
- Combined: Use both pipelines together to create complete simulated environments with both scene structure and individual objects
real2sim2real_sjtu/
โโโ src/ # Core source files
โ โโโ object_focused_tracking.py # Grounded-SAM-2 object tracking
โ โโโ example_multi_image.py # TRELLIS 3D generation
โโโ scripts/ # Automation scripts
โโโ configs/ # Configuration files
โโโ data/ # Data directory
โ โโโ input/ # Input videos
โ โโโ output/ # Generated results
โโโ docs/ # Documentation
โโโ examples/ # Example usage
โโโ environments/ # Conda environment exports
โ โโโ diffusion360.yml
โ โโโ Pano2room.yml
โ โโโ discoverse.yml
โโโ README.md
You need the following conda environments installed:
diffusion360Pano2roomdiscoverse
- Clone required repositories:
# Clone Grounded-SAM-2
git clone https://github.com/IDEA-Research/Grounded-SAM-2.git
cd Grounded-SAM-2
# Follow their installation instructions
# Clone TRELLIS
git clone https://github.com/microsoft/TRELLIS.git
cd TRELLIS
# Follow their installation instructions- Setup environments:
# Use the provided environment files to recreate conda environments
conda env create -f environments/diffusion360.yml
conda env create -f environments/Pano2room.yml
conda env create -f environments/discoverse.ymlYou can run either pipeline independently or both together depending on your needs.
# Run the complete object pipeline: SAM2 tracking + TRELLIS 3D reconstruction
bash scripts/run_object_pipeline.sh \
"data/input/your_video.mp4" \
"data/output/object_results" \
"Objects."What it does:
- Stage 1: Tracks objects in video using SAM2 (discoverse env)
- Stage 2: Generates 3D models using TRELLIS (diffusion360 env)
- Output: 3D object models (PLY, GLB, MP4 videos)
When to use: You want to extract and reconstruct specific objects from the video
# Run the complete environment pipeline: OpenPano + Diffusion360 + Pano2Room
bash scripts/run_environment_pipeline.sh \
"data/input/your_video.mp4" \
"data/output/environment_results"What it does:
- Stage 1: Generates panorama using OpenPano (diffusion360 env)
- Stage 2: Enhances panorama using Diffusion360 (diffusion360 env)
- Stage 3: Creates 3D room mesh using Pano2Room (Pano2room env)
- Stage 4: Converts and processes mesh formats (Pano2room env)
- Output: Environment mesh with textures (OBJ, PLY, PNG)
When to use: You want to reconstruct the overall scene/room structure
# Launch the web interface
bash start_gradio.sh
# Or directly:
python gradio_app_dual.pyFeatures:
- ๐ฏ Object Pipeline Tab: Upload video, set detection prompt, track objects, generate 3D models
- ๐ Environment Pipeline Tab: Upload video, generate panorama, create environment mesh
- Real-time logs and progress tracking
- Visual results preview (tracking images, 3D videos, panoramas, meshes)
- Download generated assets
Access at: http://127.0.0.1:7860
# Run the complete pipeline: BOTH object + environment pipelines
bash scripts/run_example.sh \
"data/input/your_video.mp4" \
"data/output/complete_results" \
"Objects." \
"false"What it does:
- Runs BOTH Object Pipeline (Stages 1-2) AND Environment Pipeline (Stages 3-6)
- Optional: Simulation environment import (Stage 7, set last param to "true")
- Output: Complete scene with both 3D objects and environment mesh
When to use: You want everything - both individual objects and the environment structure
| Script | Purpose | Pipelines |
|---|---|---|
run_object_pipeline.sh |
Object tracking + 3D reconstruction | Object only (SAM2 + TRELLIS) |
run_environment_pipeline.sh |
Scene/room reconstruction | Environment only (OpenPano + Diffusion360 + Pano2Room) |
run_example.sh |
Complete pipeline | Both pipelines + optional simulation |
gradio_app_dual.py |
Web UI | Both pipelines with visual interface |
If you prefer fine-grained control, you can run individual stages manually. See the detailed pipeline documentation:
- Object Pipeline: See
docs/COMPLETE_PIPELINE_USAGE.md(Stages 1-2) - Environment Pipeline: See
docs/COMPLETE_PIPELINE_USAGE.md(Stages 3-6)
This script handles:
- Video frame extraction
- Object detection using Grounding DINO
- Object tracking using SAM2
- Isolation of tracked objects with transparent backgrounds
- Generation of object-specific videos
Key Features:
- Supports multiple prompt types (point, box, mask)
- Automatic object isolation with transparency
- Batch processing of video frames
- Organized output structure per tracked object
This script handles:
- Loading multiple images of the same object
- 3D asset generation (Gaussian, Radiance Field, Mesh)
- Rendering and visualization
- Export to various 3D formats (.glb, .ply)
Key Features:
- Multi-image input for better 3D reconstruction
- Multiple output formats (Gaussian Splatting, Mesh, Radiance Field)
- Configurable generation parameters
- Video output for visualization
Key parameters to adjust in object_focused_tracking.py:
# Model paths
sam2_checkpoint = "./checkpoints/sam2.1_hiera_large.pt"
model_cfg = "configs/sam2.1/sam2.1_hiera_l.yaml"
grounding_dino_model = "./grounding-dino-base"
# Detection thresholds
box_threshold = 0.3
text_threshold = 0.3
# Tracking parameters
prompt_type = "box" # or "point", "mask"Key parameters in example_multi_image.py:
# Generation parameters
sparse_structure_sampler_params = {
"steps": 12,
"cfg_strength": 7.5,
}
slat_sampler_params = {
"steps": 12,
"cfg_strength": 3,
}- Input: Video file (MP4, AVI, etc.)
- Stage 1: Object detection and tracking
- Extract video frames
- Detect objects in first frame
- Track objects throughout video
- Generate isolated object image sequences
- Stage 2: 3D reconstruction
- Select representative frames from each tracked object
- Generate 3D assets using TRELLIS
- Export meshes, Gaussian splats, and radiance fields
- Output: 3D models ready for simulation environments
- Robotics Simulation: Convert real-world objects to simulation assets
- Digital Twins: Create 3D models from video observations
- AR/VR Content: Generate 3D assets from video footage
- Game Development: Convert real objects to game assets
- Research: Study object geometry and motion
This repository includes conda environment exports for:
- diffusion360.yml: Environment for TRELLIS 3D generation
- Pano2room.yml: Environment for panoramic image processing
- discoverse.yml: Environment for object discovery and segmentation
These can be used to recreate the exact conda environments used in development.
- Grounded-SAM-2: Object tracking and segmentation
- TRELLIS: 3D generation from images
- PyTorch
- OpenCV
- PIL/Pillow
- NumPy
- Transformers
- Supervision
- CUDA Memory Issues: Reduce batch size or use smaller models
- Model Download Failures: Ensure internet connection and sufficient disk space
- Environment Conflicts: Use separate conda environments for each component
- Path Issues: Use absolute paths for model checkpoints and data
- Use GPU acceleration for both SAM2 and TRELLIS
- Preprocess videos to reduce resolution if needed
- Filter objects by confidence scores to reduce noise
- Use appropriate model sizes based on available VRAM
This project combines multiple components with different licenses:
- Grounded-SAM-2: Apache 2.0
- TRELLIS: MIT
- Project-specific code: MIT