airMeng/sycl-tla
SYCL* Templates for Linear Algebra (SYCL*TLA) - SYCL based CUTLASS implementation for Intel GPUs
SYCL* Templates for Linear Algebra (SYCL*TLA) - SYCL based CUTLASS implementation for Intel GPUs
SGLang is a fast serving framework for large language models and vision language models.
OpenAI Triton backend for Intel® GPUs
oneAPI Deep Neural Network Library (oneDNN)
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of Facebook's LLaMA model in C/C++
Stable Diffusion and Flux in pure C/C++
PyTorch native quantization and sparsity for training and inference
GEMM performance kernels for Intel GPUs, Nvidia GPUs, and Intel CPUs, written using SYCL joint matrix extension
Tensors and Dynamic neural networks in Python with strong GPU acceleration
A Python package for extending the official PyTorch that can easily obtain performance on Intel platform
Representation and Reference Lowering of ONNX Models in MLIR Compiler Infrastructure
Fast sparse deep learning on CPUs
MLIR Sample dialect
a JIT assembler for x86(IA-32)/x64(AMD64, x86-64) MMX/SSE/SSE2/SSE3/SSSE3/SSE4/FPU/AVX/AVX2/AVX-512 by C++ header
Library for specialized dense and sparse matrix operations, and deep learning primitives.
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Intel staging area for llvm.org contribution. Home for Intel LLVM-based projects.
model mid state comparison tools for pytorch
Assembler for NVIDIA Maxwell architecture
This fork of BVLC/Caffe is dedicated to improving performance of this deep learning framework when running on CPU, in particular Intel® Xeon processors.
Models and examples built with TensorFlow
Example code for the guide
2019年4-22日-bilibili-干杯站后端源码(原包删除前最后一版170M)
profiling tools for pytorch