Xinyu302/rocm-systems
super repo for rocm systems projects
Student of Beihang University
super repo for rocm systems projects
[DEPRECATED] Moved to ROCm/rocm-systems repo
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
A task benchmark
The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.
A machine learning compiler for GPUs, CPUs, and ML accelerators
BladeDISC is an end-to-end DynamIc Shape Compiler project for machine learning workloads.
CUDA Templates for Linear Algebra Subroutines
飞桨护航计划集训营
ByteIR
Backward compatible ML compute opset inspired by HLO/MHLO
how to optimize some algorithm in cuda.
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
A code translator, translate openearth mlir code into artemis dsl.
ncnn is a high-performance neural network inference framework optimized for the mobile platform
PaddlePaddle Developer Community
development repository for the open earth compiler
OpenMMLab Foundational Library for Training Deep Learning Models
To test vectorized functions in NCNN gemm kernel
A simple high performance CUDA GEMM implementation.
MegCC是一个运行时超轻量,高效,移植简单的深度学习模型编译器
Optimize gemm on riscv