Ranran

@WindChimeRan · User

GitHub profile ↗ · Compare

Interested in Desktop LLM inference (vllm-metal & DGX-Spark) and speculative decoding.

@psunlpgroupUS115 followers82 repositories

Repositories

WindChimeRan/speculators

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

★ 0PythonForks 0

WindChimeRan/SiliconBench

This repository includes code and materials for the paper "SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops"

★ 8HTMLForks 1

WindChimeRan/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

WindChimeRan/metalstat

Apple Silicon GPU/CPU/Memory monitoring CLI — like gpustat, but for Metal

★ 3PythonForks 0

WindChimeRan/NREPapers2018

This repo covers almost all the papers (35) related to Neural Relation Extraction in ACL, EMNLP, COLING, NAACL, AAAI, IJCAI in 2018.

★ 24Forks 5

WindChimeRan/OpenJERE

EMNLP2020 findings paper: Minimize Exposure Bias of Seq2Seq Models in Joint Entity and Relation Extraction

★ 51PythonForks 7

WindChimeRan/NREPapers2019

This repo will cover almost all the papers related to Neural Relation Extraction in ACL, EMNLP, COLING, NAACL, AAAI, IJCAI in 2019.

★ 103Forks 10

WindChimeRan/CopyMTL

AAAI20 "CopyMTL: Copy Mechanism for Joint Extraction of Entities and Relations with Multi-Task Learning"

★ 128PythonForks 21

WindChimeRan/hypervibe

One-handed vibe coding using Apple TV Remote, customizable buttons and gestures.

★ 0SwiftForks 0

WindChimeRan/lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

★ 0Forks 0

WindChimeRan/transformers

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

★ 0Forks 0