idejie/ChinaTextbook
所有小初高、大学PDF教材。
The world is so noisy that let's listen to ourselves more.
所有小初高、大学PDF教材。
Hierarchical Sub-action Tree for Continuous Sign Language Recognition,ICME2025
The third assignment for Multi-Lodal Learning@PKU2024-Autumn
A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
Vue+FastAPI application for my personal pages and posts
verl: Volcano Engine Reinforcement Learning for LLMs
Collections for Violence Detection
Layout Supervised Image Generation with HunyuanDiT
[WIP] PyTorch & Hydra Template
Unified Vision-Language-Action Model
Open-source unified multimodal model
An official implementation for " UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation"
pretrained_model from Shan et. al . “ Understanding Human Hands in Contact at Internet Scale (CVPR 2020, Oral).”
Online RL with Simple Reward Enables Training VLA Models with Only One Trajectory
Code for ICML 2025 Paper "Highly Compressed Tokenizer Can Generate Without Training"
A collection of tabletop tasks in Mujoco
DisTime: Distribution-based Time Representation for Video Large Language Models.
Code for the paper "Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos" [NeurIPS (spotlight), 2024]
Long Context Transfer from Language to Vision
[CVPR2025] Number it: Temporal Grounding Videos like Flipping Manga