Tom-Zheng/flashinfer
FlashInfer: Kernel Library for LLM Serving
I am currently a software engineer at NVIDIA.
FlashInfer: Kernel Library for LLM Serving
The source of LMSYS website and blogs
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
This is a basic project that takes video stream from OV7670 camera and displays via VGA on Artix7 FPGA.
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in both training and inference.
👑 Easy-to-use and powerful NLP and LLM library with 🤗 Awesome model zoo, supporting wide-range of NLP tasks from research to industrial applications, including 🗂Text Classification, 🔍 Neural Search, ❓ Question Answering, ℹ️ Information Extraction, 📄 Document Intelligence, 💌 Sentiment Analysis etc.
A JSBox app that prints notes from Bear.app using Memobird printer.
Realtek RTL8811CU/RTL8821CU USB Wi-Fi adapter driver for Linux
a plugin for showdown to be able to use Bear.app's syntax to generate html like the html it generates when you "export as html"
Re-implementation of SparseConvNet & 3D UNet using pytorch C++ API.
RV-Debugger-BL702 Project, an opensource debugger implement
Xv6 for RISC-V
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
Debian disk image builder for the Sipeed Lichee RV Risc-v Single Board Computer
Ongoing research training transformer language models at scale, including: BERT & GPT-2
LPIPS metric. pip install lpips
Autoencoder for Point Clouds
Tangent Convolutions for Dense Prediction in 3D
An simple hash table that works inside cuda kernel. Based on Robin Hood Hashing.