MINT-SJTU/SIM2REAL2SIM

โ˜… 2Forks 0PythonGitHub โ†—Compare

README

Real2Sim2Real SJTU - Dual Pipeline System

A comprehensive Real2Sim2Real system featuring two independent pipelines for converting videos into 3D assets:

๐ŸŽฏ Project Overview

Dual Pipeline Architecture

This project provides two separate, independent pipelines:

๐ŸŽฏ Object Pipeline (SAM2 + TRELLIS)

Generates 3D models of tracked objects from video:

  1. Video โ†’ SAM2 Object Tracking: Extract and track objects using Grounded-SAM-2
  2. TRELLIS 3D Reconstruction: Convert tracked object images into 3D meshes (PLY, GLB, MP4)
  3. Output: Individual 3D object models ready for simulation

๐ŸŒ Environment Pipeline (OpenPano + Diffusion360 + Pano2Room)

Generates environment mesh from video (no object tracking):

  1. Video โ†’ OpenPano: Generate initial panorama from video frames
  2. Diffusion360 Enhancement: Enhance panorama quality using diffusion models
  3. Pano2Room Mesh Generation: Convert panorama into 3D room mesh
  4. Output: Complete environment mesh with textures ready for simulation

Why Two Pipelines?

  • Object Pipeline: Focus on specific objects in the scene (people, cars, furniture, etc.)
  • Environment Pipeline: Focus on the overall scene/room structure (walls, floors, ceilings)
  • Combined: Use both pipelines together to create complete simulated environments with both scene structure and individual objects

๐Ÿ“ Project Structure

real2sim2real_sjtu/
โ”œโ”€โ”€ src/                          # Core source files
โ”‚   โ”œโ”€โ”€ object_focused_tracking.py # Grounded-SAM-2 object tracking
โ”‚   โ””โ”€โ”€ example_multi_image.py     # TRELLIS 3D generation
โ”œโ”€โ”€ scripts/                      # Automation scripts
โ”œโ”€โ”€ configs/                      # Configuration files
โ”œโ”€โ”€ data/                        # Data directory
โ”‚   โ”œโ”€โ”€ input/                   # Input videos
โ”‚   โ””โ”€โ”€ output/                  # Generated results
โ”œโ”€โ”€ docs/                        # Documentation
โ”œโ”€โ”€ examples/                    # Example usage
โ”œโ”€โ”€ environments/                # Conda environment exports
โ”‚   โ”œโ”€โ”€ diffusion360.yml
โ”‚   โ”œโ”€โ”€ Pano2room.yml
โ”‚   โ””โ”€โ”€ discoverse.yml
โ””โ”€โ”€ README.md

๐Ÿš€ Quick Start

Prerequisites

You need the following conda environments installed:

  • diffusion360
  • Pano2room
  • discoverse

Installation

  1. Clone required repositories:
# Clone Grounded-SAM-2
git clone https://github.com/IDEA-Research/Grounded-SAM-2.git
cd Grounded-SAM-2
# Follow their installation instructions

# Clone TRELLIS
git clone https://github.com/microsoft/TRELLIS.git
cd TRELLIS
# Follow their installation instructions
  1. Setup environments:
# Use the provided environment files to recreate conda environments
conda env create -f environments/diffusion360.yml
conda env create -f environments/Pano2room.yml
conda env create -f environments/discoverse.yml

Usage

You can run either pipeline independently or both together depending on your needs.

Option 1: Run Object Pipeline (Recommended for Object Tracking)

# Run the complete object pipeline: SAM2 tracking + TRELLIS 3D reconstruction
bash scripts/run_object_pipeline.sh \
    "data/input/your_video.mp4" \
    "data/output/object_results" \
    "Objects."

What it does:

  • Stage 1: Tracks objects in video using SAM2 (discoverse env)
  • Stage 2: Generates 3D models using TRELLIS (diffusion360 env)
  • Output: 3D object models (PLY, GLB, MP4 videos)

When to use: You want to extract and reconstruct specific objects from the video

Option 2: Run Environment Pipeline (Recommended for Scene Reconstruction)

# Run the complete environment pipeline: OpenPano + Diffusion360 + Pano2Room
bash scripts/run_environment_pipeline.sh \
    "data/input/your_video.mp4" \
    "data/output/environment_results"

What it does:

  • Stage 1: Generates panorama using OpenPano (diffusion360 env)
  • Stage 2: Enhances panorama using Diffusion360 (diffusion360 env)
  • Stage 3: Creates 3D room mesh using Pano2Room (Pano2room env)
  • Stage 4: Converts and processes mesh formats (Pano2room env)
  • Output: Environment mesh with textures (OBJ, PLY, PNG)

When to use: You want to reconstruct the overall scene/room structure

Option 3: Use Gradio Web Interface (Recommended for Beginners)

# Launch the web interface
bash start_gradio.sh

# Or directly:
python gradio_app_dual.py

Features:

  • ๐ŸŽฏ Object Pipeline Tab: Upload video, set detection prompt, track objects, generate 3D models
  • ๐ŸŒ Environment Pipeline Tab: Upload video, generate panorama, create environment mesh
  • Real-time logs and progress tracking
  • Visual results preview (tracking images, 3D videos, panoramas, meshes)
  • Download generated assets

Access at: http://127.0.0.1:7860

Option 4: Run Both Pipelines Together (Complete Pipeline)

# Run the complete pipeline: BOTH object + environment pipelines
bash scripts/run_example.sh \
    "data/input/your_video.mp4" \
    "data/output/complete_results" \
    "Objects." \
    "false"

What it does:

  • Runs BOTH Object Pipeline (Stages 1-2) AND Environment Pipeline (Stages 3-6)
  • Optional: Simulation environment import (Stage 7, set last param to "true")
  • Output: Complete scene with both 3D objects and environment mesh

When to use: You want everything - both individual objects and the environment structure

Pipeline Scripts Summary

Script Purpose Pipelines
run_object_pipeline.sh Object tracking + 3D reconstruction Object only (SAM2 + TRELLIS)
run_environment_pipeline.sh Scene/room reconstruction Environment only (OpenPano + Diffusion360 + Pano2Room)
run_example.sh Complete pipeline Both pipelines + optional simulation
gradio_app_dual.py Web UI Both pipelines with visual interface

Option 5: Manual Step-by-Step Execution

If you prefer fine-grained control, you can run individual stages manually. See the detailed pipeline documentation:

  • Object Pipeline: See docs/COMPLETE_PIPELINE_USAGE.md (Stages 1-2)
  • Environment Pipeline: See docs/COMPLETE_PIPELINE_USAGE.md (Stages 3-6)

๐Ÿ› ๏ธ Core Components

object_focused_tracking.py

This script handles:

  • Video frame extraction
  • Object detection using Grounding DINO
  • Object tracking using SAM2
  • Isolation of tracked objects with transparent backgrounds
  • Generation of object-specific videos

Key Features:

  • Supports multiple prompt types (point, box, mask)
  • Automatic object isolation with transparency
  • Batch processing of video frames
  • Organized output structure per tracked object

example_multi_image.py

This script handles:

  • Loading multiple images of the same object
  • 3D asset generation (Gaussian, Radiance Field, Mesh)
  • Rendering and visualization
  • Export to various 3D formats (.glb, .ply)

Key Features:

  • Multi-image input for better 3D reconstruction
  • Multiple output formats (Gaussian Splatting, Mesh, Radiance Field)
  • Configurable generation parameters
  • Video output for visualization

๐Ÿ”ง Configuration

Grounded-SAM-2 Configuration

Key parameters to adjust in object_focused_tracking.py:

# Model paths
sam2_checkpoint = "./checkpoints/sam2.1_hiera_large.pt"
model_cfg = "configs/sam2.1/sam2.1_hiera_l.yaml"
grounding_dino_model = "./grounding-dino-base"

# Detection thresholds
box_threshold = 0.3
text_threshold = 0.3

# Tracking parameters
prompt_type = "box"  # or "point", "mask"

TRELLIS Configuration

Key parameters in example_multi_image.py:

# Generation parameters
sparse_structure_sampler_params = {
    "steps": 12,
    "cfg_strength": 7.5,
}
slat_sampler_params = {
    "steps": 12,
    "cfg_strength": 3,
}

๐Ÿ“Š Expected Workflow

  1. Input: Video file (MP4, AVI, etc.)
  2. Stage 1: Object detection and tracking
    • Extract video frames
    • Detect objects in first frame
    • Track objects throughout video
    • Generate isolated object image sequences
  3. Stage 2: 3D reconstruction
    • Select representative frames from each tracked object
    • Generate 3D assets using TRELLIS
    • Export meshes, Gaussian splats, and radiance fields
  4. Output: 3D models ready for simulation environments

๐ŸŽฎ Applications

  • Robotics Simulation: Convert real-world objects to simulation assets
  • Digital Twins: Create 3D models from video observations
  • AR/VR Content: Generate 3D assets from video footage
  • Game Development: Convert real objects to game assets
  • Research: Study object geometry and motion

๐Ÿ“ Environment Files

This repository includes conda environment exports for:

  • diffusion360.yml: Environment for TRELLIS 3D generation
  • Pano2room.yml: Environment for panoramic image processing
  • discoverse.yml: Environment for object discovery and segmentation

These can be used to recreate the exact conda environments used in development.

๐Ÿค Dependencies

External Repositories Required:

Key Python Libraries:

  • PyTorch
  • OpenCV
  • PIL/Pillow
  • NumPy
  • Transformers
  • Supervision

๐Ÿ› Troubleshooting

Common Issues:

  1. CUDA Memory Issues: Reduce batch size or use smaller models
  2. Model Download Failures: Ensure internet connection and sufficient disk space
  3. Environment Conflicts: Use separate conda environments for each component
  4. Path Issues: Use absolute paths for model checkpoints and data

Performance Tips:

  1. Use GPU acceleration for both SAM2 and TRELLIS
  2. Preprocess videos to reduce resolution if needed
  3. Filter objects by confidence scores to reduce noise
  4. Use appropriate model sizes based on available VRAM

๐Ÿ“„ License

This project combines multiple components with different licenses:

  • Grounded-SAM-2: Apache 2.0
  • TRELLIS: MIT
  • Project-specific code: MIT

Contributors

XcloudFance

Issues