saienduri/sparc
test sparc
test sparc
Temporary UTTR driver stack packaging images
Flexible GPU fractionalization — run more workloads per GPU.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Becnhmark kernels performance through IREE
vgpu.rs is the fractional GPU & vgpu-hypervisor implementation written in Rust
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
empirical study
Device Metrics Exporter exports metrics from AMD devices (GPUs) to collectors like Prometheus.
Caching scripts
A collection of postmortem templates
SGLang is a fast serving framework for large language models and vision language models.
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Documentation for bringing up an Azure Kubernetes cluster integrated with GitHub Actions Runner Controller for IREE Project
The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm
testing mi300 cluster
testing k8s mounting functionality
Development repository for the Triton language and compiler
The "leetcode"-like platform for GPU hackers
NanoGPT (124M) in 3.4 minutes
The goal of the OSSCI Fleet is to provide a central mechanism to enable test automation, batch job scheduling, and developer access to a federated set of limited GPU resources
IREE's PyTorch Frontend, based on Torch Dynamo.
Ongoing research training transformer models at scale
Efficient Triton Kernels for LLM Training
A high-throughput and memory-efficient inference and serving engine for LLMs
Reference implementations of MLPerf™ inference benchmarks
Tracks SHARK SDXL Performance