Meng, Hengyu

@airMeng · User

GitHub profile ↗ · Compare

Intel40 followers42 repositories

Repositories

airMeng/sycl-tla

SYCL* Templates for Linear Algebra (SYCL*TLA) - SYCL based CUTLASS implementation for Intel GPUs

★ 0C++Forks 0

airMeng/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0PythonForks 0

airMeng/oneDNN

oneAPI Deep Neural Network Library (oneDNN)

★ 0C++Forks 0

airMeng/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

airMeng/ao

PyTorch native quantization and sparsity for training and inference

★ 0PythonForks 0

airMeng/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 0PythonForks 0

airMeng/onnx-mlir

Representation and Reference Lowering of ONNX Models in MLIR Compiler Infrastructure

★ 0C++Forks 0

airMeng/xbyak

a JIT assembler for x86(IA-32)/x64(AMD64, x86-64) MMX/SSE/SSE2/SSE3/SSSE3/SSE4/FPU/AVX/AVX2/AVX-512 by C++ header

★ 0C++Forks 0

airMeng/libxsmm

Library for specialized dense and sparse matrix operations, and deep learning primitives.

★ 0Forks 0

airMeng/onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

★ 0Forks 0

airMeng/llvm

Intel staging area for llvm.org contribution. Home for Intel LLVM-based projects.

★ 0Forks 0

airMeng/maxas

Assembler for NVIDIA Maxwell architecture

★ 0Forks 0

airMeng/caffe

This fork of BVLC/Caffe is dedicated to improving performance of this deep learning framework when running on CPU, in particular Intel® Xeon processors.

★ 0Forks 0