Zhihui Du

@zhihuidu-amd · User

GitHub profile ↗ · Compare

AMDCA4 followers19 repositories

Repositories

zhihuidu-amd/llvm-project

The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

★ 0Forks 0

zhihuidu-amd/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 0Forks 0

zhihuidu-amd/hipgraph-ms

HIP Multi-Stream Graph Capture — upstream-compatible fix for AMD ROCm hipStreamCaptureModeGlobal

★ 0C++Forks 0

zhihuidu-amd/warp

A Python framework for accelerated simulation, data generation and spatial computing.

★ 0Forks 0

zhihuidu-amd/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 0Forks 0

zhihuidu-amd/Speech

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

★ 0Forks 0

zhihuidu-amd/llm-awq

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

★ 0Forks 0

zhihuidu-amd/fp6_llm

An efficient GPU support for LLM inference with x-bit quantization (e.g. FP6,FP5).

★ 0Forks 0

zhihuidu-amd/ROCm-unsloth

ROCm port of unsloth (unslothai/unsloth) for AMD MI300X (gfx942). LoRA/QLoRA fine-tuning with ROCm 7.0+ support.

★ 0PythonForks 0

zhihuidu-amd/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0Forks 0

zhihuidu-amd/onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

★ 0Forks 0