ZSL98/openhands-aci
Agent computer interface for AI software engineer.
www.shulai.org
Agent computer interface for AI software engineer.
Shulai Zhang's Homepage
The open source coding agent.
A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresses it with AI (using Claude's agent-sdk), and injects relevant context back into future sessions.
veRL: Volcano Engine Reinforcement Learning for LLM
Checkpoint-engine is a simple middleware to update model weights in LLM inference engines
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
Ongoing research training transformer models at scale
A collection of hands-on demos and code snippets for learning and experimenting with various programming concepts, frameworks, and tools. Ideal for self-study and quick references.
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
A high-throughput and memory-efficient inference and serving engine for LLMs
Fast Embodied AI
[RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
Hackable and optimized Transformers building blocks, supporting a composable construction.
A simple but complete full-attention transformer with a set of promising experimental features from various papers
A fast MoE impl for PyTorch
Tutel MoE: An Optimized Mixture-of-Experts Implementation
An unofficial cuda assembler, for all generations of SASS, hopefully :)
Artifacts for our NSDI'23 paper TGS
An interference-aware scheduler for fine-grained GPU sharing
Early Exits of DNN Networks with TensorRT
⏰ Collaboratively track deadlines of conferences recommended by CCF (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~
Linux kernel modules for secure sharing of memory buffers
Hooked CUDA-related dynamic libraries by using automated code generation tools.
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
First-Class GPU Resource Management: Device Drivers, Runtimes, and CUDA Compilers for Nouveau.
First-Class GPU Resource Management: Device Drivers, Runtimes, and CUDA Compilers for Nouveau.