Roman Ageev

@RomaA2000 · User

GitHub profile ↗ · Compare

NvidiaAmsterdam23 followers31 repositories

Repositories

RomaA2000/dynamo

A Datacenter Scale Distributed Inference Serving Framework

★ 0RustForks 0

RomaA2000/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

★ 0Forks 0

RomaA2000/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.

★ 0Forks 0

RomaA2000/chat.petals.dev

💬 Chatbot web app + HTTP and Websocket endpoints for LLM inference with the Petals client

★ 0Forks 0

RomaA2000/petals

🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading

★ 0Forks 0

RomaA2000/PGM-index

🏅State-of-the-art learned data structure that enables fast lookup, predecessor, range searches and updates in arrays of billions of items using orders of magnitude less space than traditional indexes

★ 0C++Forks 0

RomaA2000/vector

vector with copy on write and small object optimization

★ 0C++Forks 0