anish

@anishesg · User

GitHub profile ↗ · Compare

cs @ princeton | formerly microsoft ai, aws, nasa

princetonsf7 followers108 repositories

Repositories

anishesg/revive

Phone swarm collective intelligence — distributed LLM inference via Mixture of Agents on iOS

★ 1PythonForks 0

anishesg/flash-evict-prefill

Flash-Evict: fused prefill KV cache eviction kernel — importance scoring inside FlashAttention blockwise computation, single O(n) pass

★ 0PythonForks 0

anishesg/kv-distill-training

KV Cache Distillation: training models to produce compression-friendly KV representations via nuclear norm, entropy, and gumbel losses

★ 0PythonForks 0

anishesg/rlvr-fdc

RLVR-FDC: Forward Density Credit for RL post-training — dense per-step rewards from frozen reference model + GFlowNet proportional sampling

★ 0Forks 0

anishesg/tele-attn

Fused telescoping attention: progressive multi-resolution KV-cache with importance-weighted representative selection and multiplicity-correct single-pass decode

★ 0CudaForks 0

anishesg/linear-decode

Fused recurrent linear attention decode: shared-memory-resident matrix state with gated rank-1 updates and warp-cooperative query contraction

★ 0CudaForks 0

anishesg/slab-attn

GPU-resident KV-cache slab allocator: lock-free warp-cooperative page management fused with paged attention eliminating CPU round-trip allocation

★ 0CudaForks 0

anishesg/spec-decode

Fused GPU-resident speculative decoding verification: single-kernel acceptance chain with warp-cooperative correction sampling and tree-structured multi-draft path selection

★ 0CudaForks 0

anishesg/cache-evict

Fused attention-eviction kernel: online importance scoring during tiled attention with warp-cooperative stream compaction for bounded-memory long-context inference

★ 0CudaForks 0

anishesg/dfa-sample

Fused grammar-constrained decode: GPU-resident DFA token masking with register-heap top-k sampling in a single kernel

★ 0CudaForks 0

anishesg/dyn-quant

Fused dynamic-precision GEMV: warp-cooperative outlier detection with runtime INT4/FP16 channel splitting in a single kernel

★ 0CudaForks 0

anishesg/chunk-scan

Fused chunk-wise state space dual kernel: tensor-core intra-chunk matmul with warp-shuffle inter-chunk parallel scan

★ 0CudaForks 0

anishesg/latent-attn

Fused multi-head latent attention: absorbed Q/K projections with online softmax in compressed latent space

★ 0CudaForks 0

anishesg/lora-fuse

Fused LoRA decode inference: base weight matvec with register-resident low-rank accumulation and multi-adapter batched dispatch

★ 0CudaForks 0

anishesg/topk-head

Fused LM-head top-k decode: tiled vocabulary projection with register-heap online selection and Cauchy-Schwarz tile pruning

★ 0CudaForks 0

anishesg/persist-decode

Persistent-kernel transformer decode: full-layer fusion with shared-memory-resident hidden state and double-buffered weight streaming

★ 0CudaForks 0

anishesg/sparse-ffn

Decode-time sparse SwiGLU: skip SiLU-dead neurons in up/down projections with fused dynamic-sparse matvec kernels

★ 0CudaForks 0

anishesg/warp-route

Fused decode-time MoE dispatch: warp-cooperative gating, expert selection, and FFN in a single kernel launch

★ 0CudaForks 0

anishesg/bitcache

Binary KV-cache attention via XNOR-popcount with precision-preserving residual correction

★ 0CudaForks 0

anishesg/fork-attn

Fused tree-structured decode attention: shared-prefix KV computation with single-pass multi-branch output

★ 0CudaForks 0

anishesg/hydra

Hydra is a framework for elegantly configuring complex applications

★ 0Forks 0

anishesg/kv-compress

Online KV-cache compression via per-head learned codebooks with age-based mixed-precision attention

★ 0CudaForks 0

anishesg/scout-attention

Predictive block-skipping attention kernel: fused importance scoring within FlashAttention tiled computation

★ 0CudaForks 0

anishesg/neurovoice

EEG-driven communication system for non-verbal users — NeuralRL + OpenAI + ElevenLabs

★ 0PythonForks 0

anishesg/axiom

Brain-computer interface for macOS — control your computer with EEG (Muse S headband)

★ 0PythonForks 0