PerkzZheng/prims-ts-examples
Standalone examples for FlashInfer PrimTS attention APIs
Currently work as an AI Technology Developer Engineer @ Nvidia
Standalone examples for FlashInfer PrimTS attention APIs
Reproducible accuracy and performance regression suite for FlashInfer PrimTS attention kernels
Private FlashInfer K3 MLA development repository
FlashMLA: Efficient Multi-head Latent Attention Kernels
A high-throughput and memory-efficient inference and serving engine for LLMs
FlashInfer: Kernel Library for LLM Serving
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
CUDA Templates for Linear Algebra Subroutines
A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
Transformer related optimization, including BERT, GPT
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Run MNIST inference in Apache Flink
This repository contains compilation of some chosen topics.