zhihuidu-amd/llvm-project
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
AMD's graph optimization engine.
GTP engine and self-play learning in Go
GPU-optimized version of the MuJoCo physics simulator.
HIP Multi-Stream Graph Capture — upstream-compatible fix for AMD ROCm hipStreamCaptureModeGlobal
The Triton backend for the ONNX Runtime.
A Python framework for accelerated simulation, data generation and spatial computing.
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
An efficient GPU support for LLM inference with x-bit quantization (e.g. FP6,FP5).
Utils for Unsloth https://github.com/unslothai/unsloth
ROCm port of unsloth (unslothai/unsloth) for AMD MI300X (gfx942). LoRA/QLoRA fine-tuning with ROCm 7.0+ support.
super repo for rocm libraries
Flash-Attention 2 SDPA engine for hipDNN (ROCm/rocm-libraries PR)
A high-throughput and memory-efficient inference and serving engine for LLMs
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator