zaristei/FlashMLA
FlashMLA: Efficient Multi-head Latent Attention Kernels
Zachary Aristei
FlashMLA: Efficient Multi-head Latent Attention Kernels
A high-throughput and memory-efficient inference and serving engine for LLMs
OpenShell is the safe, private runtime for autonomous AI agents.
Run OpenClaw more securely inside NVIDIA OpenShell with managed inference
Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents
Bonsplit is a custom tab bar and layout split library for macOS apps. Out of the box 120fps animations, drag-and-drop reordering, SwiftUI support & keyboard navigation.
🙌 OpenHands: Code Less, Make More
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A Datacenter Scale Distributed Inference Serving Framework
Low-Level Graph Neural Network Operators for PyG
Graph Neural Network Library for PyTorch
Me experimenting with making a sudoku solver to see if it makes me a better player
A collection of tools for manipulating and analyzing SVG Path objects and Bezier curves.
Project leaderboard and per-team scores.