Shulai Zhang

@ZSL98 · User

GitHub profile ↗ · Compare

www.shulai.org

58 followers60 repositories

Repositories

ZSL98/claude-mem

A Claude Code plugin that automatically captures everything Claude does during your coding sessions, compresses it with AI (using Claude's agent-sdk), and injects relevant context back into future sessions.

★ 0Forks 0

ZSL98/verl

veRL: Volcano Engine Reinforcement Learning for LLM

★ 4PythonForks 0

ZSL98/mini-swe-agent

The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!

★ 0Forks 0

ZSL98/Megatron-LM

Ongoing research training transformer models at scale

★ 2PythonForks 0

ZSL98/study-demo

A collection of hands-on demos and code snippets for learning and experimenting with various programming concepts, frameworks, and tools. Ideal for self-study and quick references.

★ 0Forks 0

ZSL98/flux

A fast communication-overlapping library for tensor/expert parallelism on GPUs.

★ 0C++Forks 0

ZSL98/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

ZSL98/diffusion_policy

[RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion

★ 0PythonForks 0

ZSL98/xformers

Hackable and optimized Transformers building blocks, supporting a composable construction.

★ 1PythonForks 1

ZSL98/x-transformers

A simple but complete full-attention transformer with a set of promising experimental features from various papers

★ 1PythonForks 1

ZSL98/tutel

Tutel MoE: An Optimized Mixture-of-Experts Implementation

★ 0PythonForks 0

ZSL98/CuAssembler

An unofficial cuda assembler, for all generations of SASS, hopefully :)

★ 0Forks 0

ZSL98/TGS

Artifacts for our NSDI'23 paper TGS

★ 0Forks 0

ZSL98/orion

An interference-aware scheduler for fine-grained GPU sharing

★ 0PythonForks 0

ZSL98/PAME

Early Exits of DNN Networks with TensorRT

★ 2PythonForks 1

ZSL98/ccf-deadlines

⏰ Collaboratively track deadlines of conferences recommended by CCF (Website, Python Cli, Wechat Applet) / If you find it useful, please star this project, thanks~

★ 0Forks 0

ZSL98/nvsci

Linux kernel modules for secure sharing of memory buffers

★ 0Forks 0

ZSL98/cuda_hook

Hooked CUDA-related dynamic libraries by using automated code generation tools.

★ 0Forks 0

ZSL98/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.

★ 0C++Forks 0

ZSL98/gdev_new

First-Class GPU Resource Management: Device Drivers, Runtimes, and CUDA Compilers for Nouveau.

★ 0Forks 0

ZSL98/gdev

First-Class GPU Resource Management: Device Drivers, Runtimes, and CUDA Compilers for Nouveau.

★ 0CForks 0