cs @ princeton | formerly microsoft ai, aws, nasa
Repositories
anishesg/qcb-research-dashboard
QCB research dashboard
anishesg/revive
Phone swarm collective intelligence — distributed LLM inference via Mixture of Agents on iOS
anishesg/flash-evict-prefill
Flash-Evict: fused prefill KV cache eviction kernel — importance scoring inside FlashAttention blockwise computation, single O(n) pass
anishesg/kv-distill-training
KV Cache Distillation: training models to produce compression-friendly KV representations via nuclear norm, entropy, and gumbel losses
anishesg/rlvr-fdc
RLVR-FDC: Forward Density Credit for RL post-training — dense per-step rewards from frozen reference model + GFlowNet proportional sampling
anishesg/tele-attn
Fused telescoping attention: progressive multi-resolution KV-cache with importance-weighted representative selection and multiplicity-correct single-pass decode
anishesg/linear-decode
Fused recurrent linear attention decode: shared-memory-resident matrix state with gated rank-1 updates and warp-cooperative query contraction
anishesg/slab-attn
GPU-resident KV-cache slab allocator: lock-free warp-cooperative page management fused with paged attention eliminating CPU round-trip allocation
anishesg/spec-decode
Fused GPU-resident speculative decoding verification: single-kernel acceptance chain with warp-cooperative correction sampling and tree-structured multi-draft path selection
anishesg/cache-evict
Fused attention-eviction kernel: online importance scoring during tiled attention with warp-cooperative stream compaction for bounded-memory long-context inference
anishesg/dfa-sample
Fused grammar-constrained decode: GPU-resident DFA token masking with register-heap top-k sampling in a single kernel
anishesg/dyn-quant
Fused dynamic-precision GEMV: warp-cooperative outlier detection with runtime INT4/FP16 channel splitting in a single kernel
anishesg/chunk-scan
Fused chunk-wise state space dual kernel: tensor-core intra-chunk matmul with warp-shuffle inter-chunk parallel scan
anishesg/latent-attn
Fused multi-head latent attention: absorbed Q/K projections with online softmax in compressed latent space
anishesg/lora-fuse
Fused LoRA decode inference: base weight matvec with register-resident low-rank accumulation and multi-adapter batched dispatch
anishesg/topk-head
Fused LM-head top-k decode: tiled vocabulary projection with register-heap online selection and Cauchy-Schwarz tile pruning
anishesg/persist-decode
Persistent-kernel transformer decode: full-layer fusion with shared-memory-resident hidden state and double-buffered weight streaming
anishesg/sparse-ffn
Decode-time sparse SwiGLU: skip SiLU-dead neurons in up/down projections with fused dynamic-sparse matvec kernels
anishesg/warp-route
Fused decode-time MoE dispatch: warp-cooperative gating, expert selection, and FFN in a single kernel launch
anishesg/bitcache
Binary KV-cache attention via XNOR-popcount with precision-preserving residual correction
anishesg/fork-attn
Fused tree-structured decode attention: shared-prefix KV computation with single-pass multi-branch output
anishesg/hydra
Hydra is a framework for elegantly configuring complex applications
anishesg/kv-compress
Online KV-cache compression via per-head learned codebooks with age-based mixed-precision attention
anishesg/scout-attention
Predictive block-skipping attention kernel: fused importance scoring within FlashAttention tiled computation
anishesg/david
openllm research
anishesg/axiom-weavehacks
AXIOM: Neural BCI + Multi-Agent Rover Control | WeaveHacks 4
anishesg/axiom-weave
AXIOM: Neural BCI + Multi-Agent Rover Control | WeaveHacks 4
anishesg/neurovoice
EEG-driven communication system for non-verbal users — NeuralRL + OpenAI + ElevenLabs
anishesg/axiom
Brain-computer interface for macOS — control your computer with EEG (Muse S headband)