RomaA2000/dynamo
A Datacenter Scale Distributed Inference Serving Framework
A Datacenter Scale Distributed Inference Serving Framework
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
TokenSpeed is a speed-of-light LLM inference engine.
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
sources for daal4py - a convenient Python API to DAAL
solutions for the MountainCar-v0 and MountainCarContinuous-v0
💬 Chatbot web app + HTTP and Websocket endpoints for LLM inference with the Petals client
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
Easily build, customize and control your own LLMs
homework of software development course
haskell course hw
🏅State-of-the-art learned data structure that enables fast lookup, predecessor, range searches and updates in arrays of billions of items using orders of magnitude less space than traditional indexes
cpp programs
Benchmark for optimizations to scikit-learn in the Intel Distribution for Python*
ios game
naural_face
exam list
vector with move
vector with copy on write and small object optimization
simple tuple implementation
huffman encoder and decoder
really big numbers with small object and copy on write optimizations
really big numbers
my java homeworks