Chauncey

@chaunceyjiang · User

GitHub profile ↗ · Compare

Karmada/OpenELB/HAMi/vLLM approver cilium/istio member

@DaoCloud 201 followers87 repositories

Repositories

chaunceyjiang/humming

Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

★ 0Forks 0

chaunceyjiang/fake-gpu

This project is designed to simulate GPU information, making it easier to test scenarios where a GPU is not available.

★ 65C++Forks 4

chaunceyjiang/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

chaunceyjiang/scuda

SCUDA is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

★ 0Forks 0

chaunceyjiang/velero

Backup and migrate Kubernetes applications and their persistent volumes

★ 0Forks 0

chaunceyjiang/MinivLLM

Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation

★ 1PythonForks 0

chaunceyjiang/mpu

A shim driver allows in-docker nvidia-smi showing correct process list without modify anything

★ 0Forks 0

chaunceyjiang/cuda_hook

Hooked CUDA-related dynamic libraries by using automated code generation tools.

★ 0CForks 0