hhaAndroid/xtuner
XTuner is a toolkit for efficiently fine-tuning LLM
LLM&MLLM Infra, RL
XTuner is a toolkit for efficiently fine-tuning LLM
Unified Efficient Fine-Tuning of 100+ LLMs (ACL 2024)
Efficient Triton Kernels for LLM Training
Janus-Series: Unified Multimodal Understanding and Generation Models
OpenMMLab YOLO series toolbox and benchmark
Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks
Train transformer language models with reinforcement learning.
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference
Ring attention implementation with flash attention
A GPipe implementation in PyTorch
Python资源大全中文版,内容包括:Web框架、网络爬虫、网络内容提取、模板引擎、数据库、数据可视化、图片处理、文本处理、自然语言处理、机器学习、日志、代码分析等
The official Meta Llama 3 GitHub site
This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.
u2net
Reformatted Alignment
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Accelerating the development of large multimodal models (LMMs) with lmms-eval
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
Implementing a ChatGPT-like LLM from scratch, step by step
Visual Instruction Tuning: Large Language-and-Vision Assistant built towards multimodal GPT-4 level capabilities.
LAVIS - A One-stop Library for Language-Vision Intelligence
GLEE: General Object Foundation Model for Images and Videos at Scale
Recent LLM-based CV and related works. Welcome to comment/contribute!
Grounded Language-Image Pre-training
A paper list of object detection using deep learning.
Flickr30K Entities Dataset
A technical report on convolution arithmetic in the context of deep learning