Keyang Ru

@key4ng · User

GitHub profile ↗ · Compare

Nebulizing..

15 followers25 repositories

Repositories

key4ng/lite-cc

A tiny Claude-Code-style CLI that chats with multiple AI models, runs bash with guardrails like a brave little engineer, and occasionally learns new tricks through plugins.

★ 3PythonForks 1

key4ng/lite-bb

Because sometimes you want gh, but your company uses Bitbucket. lite-bb is a gh-style minimal CLI for pull request creation, diffs, reviews, and comments — designed for humans, scripts, and LLM agents.

★ 6RustForks 0

key4ng/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

key4ng/claude-skill-bb

Claude Code skill for Bitbucket CLI (lite-bb) — the gh equivalent for Bitbucket

★ 1Forks 0

key4ng/datadog-auto-ack

For when your wife/girlfriend's company uses Datadog, write access to API keys is disabled, and someone configured 100+ stupid/noisy SMS alerts that mostly just need an ack.

★ 0PythonForks 0

key4ng/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

★ 0Forks 0

key4ng/minimind

Train a 64M-parameter LLM from scratch in just 2h!

★ 0Forks 0

key4ng/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

★ 0Forks 0

key4ng/ome-howto

A developer journey through using and building with OME

★ 0Forks 0

key4ng/ome

OME is a Kubernetes operator for enterprise-grade management and serving of Large Language Models (LLMs)

★ 0Forks 0

key4ng/genai-bench-rs

A Rust reimplementation of genai-bench for benchmarking LLM serving systems at high concurrency with accurate timing and industry-standard metrics.

★ 0RustForks 0

key4ng/lite-grpc-router

A minimal learning-focused implementation of SMG's gRPC router — proxies OpenAI-compatible requests to SGLang inference backends in ~1000 lines of Rust.

★ 0RustForks 0

key4ng/KAI-Scheduler

KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale

★ 0Forks 0

key4ng/cli

GitHub’s official command line tool

★ 0Forks 0

key4ng/mini-sglang

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

★ 0Forks 0

key4ng/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0PythonForks 0

key4ng/dynamo

A Datacenter Scale Distributed Inference Serving Framework

★ 0Forks 0

key4ng/tensorzero

TensorZero is an open-source stack for industrial-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluation, and experimentation.

★ 0Forks 0

key4ng/genai-bench

Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.

★ 0PythonForks 0

key4ng/production-stack

vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

★ 0Forks 0