Sunt-ing/SUSTech-CS-Notes
一名南科大CS本科生的课程笔记、考点预测和论文,包含CS和非CS的内容 :rocket::rocket:
一名南科大CS本科生的课程笔记、考点预测和论文,包含CS和非CS的内容 :rocket::rocket:
Parse LaTeX math expressions
DeepGEMM: clean and efficient BLAS kernel library on GPU
Ting Sun — personal homepage
A computer algebra system written in pure Python
Keep working while jobs run. Wake your Codex session with results, without model-side polling.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
:yum: A curated reading list about database systems
:innocent: A PyTorch-like deep learning framework. Just for fun.
Ongoing research training transformer models at scale
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
🚀 Efficient implementations for emerging model architectures
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
TokenSpeed is a speed-of-light LLM inference engine.
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM Quantization plugin for GGUF
[NSDI26] Pilot Execution: Simulating Failure Recovery In Situ for Production Distributed Systems
A scalable, end-to-end training pipeline for general-purpose agents
LLM inference in C/C++
LLM-Merging: Building LLMs Efficiently through Merging
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
UP-TO-DATE LLM Watermark paper. 🔥🔥🔥
ai4db and db4ai work