Julien Debache

@jdebache · User

GitHub profile ↗ · Compare

Zurich13 followers42 repositories

Repositories

jdebache/TensorRT-Model-Optimizer

nvidia-modelopt is a unified library of state-of-the-art model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM or TensorRT to optimize inference speed.

★ 0PythonForks 0

jdebache/DeepEP

DeepEP: an efficient expert-parallel communication library

★ 0Forks 0

jdebache/dynamo

A Datacenter Scale Distributed Inference Serving Framework

★ 0RustForks 0

jdebache/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.

★ 0PythonForks 0

jdebache/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

jdebache/tokenizers

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

★ 0Forks 0

jdebache/transformers

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

★ 0Forks 0

jdebache/numba

NumPy aware dynamic Python compiler using LLVM

★ 0Forks 0

jdebache/llvmlite

A lightweight LLVM python binding for writing JIT compilers

★ 0Forks 0

jdebache/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 0PythonForks 0

jdebache/mscclpp

MSCCL++: A GPU-driven communication stack for scalable AI applications

★ 0Forks 0

jdebache/TensorRT

NVIDIA® TensorRT™, an SDK for high-performance deep learning inference, includes a deep learning inference optimizer and runtime that delivers low latency and high throughput for inference applications.

★ 0C++Forks 0