wbigat/xllm
A high-performance inference engine for LLMs, optimized for diverse AI accelerators.
A high-performance inference engine for LLMs, optimized for diverse AI accelerators.
A community-driven pypto implementation
An independent Python feature port of Claude Code, entirely rewritting from scratch using oh-my-codex. Educational Purpose only.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
SGLang is a high-performance serving framework for large language models and multimodal models.
A PyTorch native platform for training generative AI models
On-device AI across mobile, embedded and edge for PyTorch
Community maintained hardware plugin for vLLM on Ascend
A torch compile backend for multi-targets
Ongoing research training transformer models at scale