jthomson04

@jthomson04 · User

GitHub profile ↗ · Compare

NVIDIA Dynamo

9 followers16 repositories

Repositories

jthomson04/aiperf

AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.

★ 0Forks 0

jthomson04/axum

HTTP routing and request-handling library for Rust that focuses on ergonomics and modularity

★ 0Forks 0

jthomson04/srt-slurm

NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.

★ 0Forks 0

jthomson04/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0Forks 0

jthomson04/RL

Scalable toolkit for efficient model reinforcement

★ 0Forks 0

jthomson04/srt-slurm-trtllm

Collection of SLURM deployment scripts for various hardware + trtllm disaggregation variants

★ 2PythonForks 0

jthomson04/InferenceMAX

Open Source Continuous Inference Benchmarking - GB200 NVL72 vs MI355X vs B200 vs H200 vs MI325X & soon™ TPUv6e/v7/Trainium2/3/GB300 NVL72 - DeepSeek 670B MoE, GPTOSS

★ 0Forks 0

jthomson04/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0Forks 0

jthomson04/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.

★ 0C++Forks 0