CatherineSue/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Shepherd Model Gateway
A high-throughput and memory-efficient inference and serving engine for LLMs
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
SGLang is a fast serving framework for large language models and vision language models.
Code for RL experiments in "Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks"
A framework for reproducible reinforcement learning research
rllab is a framework for developing and evaluating reinforcement learning algorithms, fully compatible with OpenAI Gym.
This is a project to make a commend to people seeking for jobs, which contributed by my National University Student Innovation Program team