Greg Pereira

@Gregory-Pereira · User

GitHub profile ↗ · Compare

Sr. Machine Learning Engineer @ Red Hat | Inference Engineering | Building llm-d: distributed inference for LLMs on Kubernetes

@RedHatOfficial @llm-dSan Francisco84 followers198 repositories

Repositories

Gregory-Pereira/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

Gregory-Pereira/llm-d

llm-d is a Kubernetes-native high-performance distributed LLM inference framework

★ 0PythonForks 0

Gregory-Pereira/speculators

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

★ 0Forks 0

Gregory-Pereira/MemOS

Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings

★ 0Forks 0

Gregory-Pereira/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

★ 0Forks 0

Gregory-Pereira/konflux-central

Central repository for managing Konflux resource files. This streamlines maintenance by consolidating configurations and leveraging GitHub Actions for automated syncing.

★ 0PythonForks 0

Gregory-Pereira/nvshmem

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.

★ 0Forks 0