WindChimeRan/speculators
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Interested in Desktop LLM inference (vllm-metal & DGX-Spark) and speculative decoding.
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
This repository includes code and materials for the paper "SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops"
Play whack-a-mole using local llm (vllm-metal or vllm dgx-spark)
A high-throughput and memory-efficient inference and serving engine for LLMs
Community maintained hardware plugin for vLLM on Apple Silicon
BERT + reproduce "Joint entity recognition and relation extraction as a multi-head selection problem" for Chinese and English IE
A TurboQuant inference server
deep research agent for codebase
Apple Silicon GPU/CPU/Memory monitoring CLI — like gpustat, but for Metal
This repo covers almost all the papers (35) related to Neural Relation Extraction in ACL, EMNLP, COLING, NAACL, AAAI, IJCAI in 2018.
EMNLP2020 findings paper: Minimize Exposure Bias of Seq2Seq Models in Joint Entity and Relation Extraction
This repo will cover almost all the papers related to Neural Relation Extraction in ACL, EMNLP, COLING, NAACL, AAAI, IJCAI in 2019.
AAAI20 "CopyMTL: Copy Mechanism for Joint Extraction of Entities and Relations with Multi-Task Learning"
Agent-friendly GPU profile-query CLI
One-handed vibe coding using Apple TV Remote, customizable buttons and gestures.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
MLX: An array framework for Apple silicon
Online problem-driven learning system
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Run frontier AI locally.