Roberto L. Castro

@LopezCastroRoberto · User

GitHub profile ↗ · Compare

Senior ML Engineer at RedHat AI

@RedHatOfficialSpain37 followers5 repositories

Repositories

LopezCastroRoberto/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

LopezCastroRoberto/humming

Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

★ 0Forks 0