ved1beta/ved1beta
Config files for my GitHub profile.
Hi, I'm Ved! ๐ฃ , AI Infra | LLM Inference | Systems Performance
Config files for my GitHub profile.
ML papers Implementation
MLX: An array framework for Apple silicon
inference engine with speculative decoding.
Go ahead and axolotl questions
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
SGLang is a high-performance serving framework for large language models and multimodal models.
kimi4all
A high-throughput and memory-efficient inference and serving engine for LLMs
JAX backend for SGL
writingjaxtobekool
wrote v0.dev backend and inspired by bolt.new in TS
Train transformer language models with reinforcement learning.
Open ABI and FFI for Machine Learning Systems
Tensors and Dynamic neural networks in Python with strong GPU acceleration
common in-memory tensor structure
Distributed training
FlashInfer: Kernel Library for LLM Serving
Efficient Triton Kernels for LLM Training
Build compute kernels and load them from the Hub.
Small scale distributed training of sequential deep learning models, built on Numpy and MPI.
"Efficient and scalable solutions for PyTorch, enabling large language model quantization with k-bit precision for enhanced accessibility.
๐ค PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
Composable building blocks to build Llama Apps
ARC Relay โ WebSocket relay server for agent remote control by Axolotl AI
E2E pipeline 4 inference
Tenstorrent MLIR compiler