zyongye/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
mts @Inferact
A high-throughput and memory-efficient inference and serving engine for LLMs
The official repository for the gem5 computer-system architecture simulator.
Reference implementation and examples of the CuTe Layout representation and algebra.
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
FlashInfer: Kernel Library for LLM Serving
Development repository for the Triton language and compiler
Fast and memory-efficient exact attention
Student version of Assignment 2 for Stanford CS336 - Language Modeling From Scratch
Student version of Assignment 1 for Stanford CS336 - Language Modeling From Scratch
Common recipes to run vLLM
Verilog Realization of Little Computer 3, a processor used for book "Intro to Computing Systems"
Repo to create custom RISCV test to run on boom
WasmEdge is a lightweight, high-performance, and extensible WebAssembly runtime for cloud native, edge, and decentralized applications. It powers serverless apps, embedded functions, microservices, smart contracts, and IoT devices.
Guide to securing and improving privacy on macOS
Online Graduate Virtualization Class. UT Austin CS Dept. Instructor: Vijay Chidambaram. Copyright held by Vijay Chidambaram and UT Austin.
SwiftUI Cheat Sheet