Phil-OSophy-42/llm-d-inference-sim
A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual heavy models.
A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual heavy models.
Inference scheduler for llm-d
NVSentinel detects and remediates GPU faults on Kubernetes nodes
helm repo add daocloud https://daocloud.github.io/dce-charts-repackage/
A benchmark to evaluate LLMs on Kubernetes tasks
Self-hosted, open-source agent skill registry for enterprises. Publish & version skill packages, govern with RBAC and audit logs, deploy on-premise with Docker or Kubernetes.
Achieve state of the art inference performance with modern accelerators on Kubernetes
HAMi-core compiles libvgpu.so, which ensures hard limit on GPU in container
很多镜像都在国外。比如 gcr 。国内下载很慢,需要加速。致力于提供连接全世界的稳定可靠安全的容器镜像服务。
d.run website
Variant optimization autoscaler for distributed inference workloads
DaoCloud Enterprise 5.0 Documentation
Heterogeneous AI Computing Virtualization Middleware(Project under CNCF)