Syed Azeez

@syedazeez337 · User

GitHub profile ↗ · Compare

Working on Open Source projects #Cilium #CoreDNS

63 followers240 repositories

Repositories

syedazeez337/tiny-dino

A 21,699-parameter MLP that plays a deterministic Chromium Dino runner with no planner, solver, or safety shield.

★ 0PythonForks 0

syedazeez337/emoji-pile

Semantic emoji search visualized as physics. Type a phrase, matching emoji rise out of the pile. Qwen3 Embedding at Cloudflare's edge, cosine in the browser.

★ 0TypeScriptForks 0

syedazeez337/from-bert-to-deepseek-v41

Interactive Slidev explainer: DeepSeek-V4.1-Flash Causal Encoder-Decoder vs BERT and classic encoder-decoder. Live: https://syedazeez337.github.io/from-bert-to-deepseek-v41/

★ 0VueForks 0

syedazeez337/agentfoundry

An experiment platform that makes agent-architecture claims falsifiable. Runs coding agents as experimental subjects in isolated environments, records immutable evidence, and grades both outcomes and their trustworthiness with cost matching, FDR control, and equivalence testing.

★ 0PythonForks 0

syedazeez337/Macaw

Macaw — on-device AI agent for macOS. 2.7B LFM2.5 fine-tune, MLX 4-bit, 97 macOS tools, no cloud.

★ 0Forks 0

syedazeez337/skill-recorder

Desktop app that records your on-screen work session and uses the GitHub Copilot CLI to reconstruct it as an intent + ordered steps, then builds a reusable Skill or Automation for Microsoft Scout, Microsoft Copilot Cowork, or Copilot Studio.

★ 0Forks 0

syedazeez337/Under-the-hood

Under the Hood - Build Every Layer of a Large Language Model from Scratch. A 36-project, build-it/break-it/measure-it manual by Ramchand Kumaresan.

★ 0PythonForks 0

syedazeez337/core-dumped

Learn Modern C++ by Debugging It. A workbook of broken C++26 programs, for Linux and native Windows.

★ 0C++Forks 0

syedazeez337/mesh-lab

KV-cache-aware LLM inference mesh on a single 6 GB GPU. 2× vLLM + LMCache shared KV + prefix-aware router on k3s, observed with Cilium/Hubble eBPF. Every number measured, every breakage documented.

★ 1PythonForks 0

syedazeez337/gpu-passport

A Go binary that checks whether a GPU is the card it claims to be. Measures bandwidth, FP32 throughput, VRAM integrity, SM count and bus width, compares them against an embedded spec table, and writes an ed25519-signed scorecard. Driver-only: no CUDA toolkit needed.

★ 0GoForks 0

syedazeez337/speculative-decoding-explained

A beginner-friendly, from-scratch course on speculative decoding: plain-English lessons, an interactive explainer, designed PDF notes, and runnable cross-platform code (CUDA / Apple Metal / CPU).

★ 0TypstForks 0

syedazeez337/raglearn

Local hybrid RAG over the vLLM codebase — BGE-M3 + Qdrant (dense+sparse, RRF) + cross-encoder rerank + Qwen3.5-4B, with a source-grounded evaluation harness.

★ 0PythonForks 0

syedazeez337/data-movement-roofline

Verifying Reiner Pope's 'data-movement tax' chip-design claim on an RTX 3050 (6GB, sm_86) with measured roofline + Triton kernels

★ 0PythonForks 0

syedazeez337/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0Forks 0

syedazeez337/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0Forks 0

syedazeez337/attention-drift-repro

Small-scale single-GPU (6GB) reproduction of Attention Drift / EAGLE-3.1 (arXiv:2605.09992) + a probe of its untested §4.4 verifier-quantization hypothesis

★ 0PythonForks 0

syedazeez337/slm-toolkit

Complete SLM Toolkit — build, train, align, serve and operate Small Language Models from first principles. Includes internal-docs RAG assistant.

★ 0Forks 0

syedazeez337/opam

opam is a source-based package manager. It supports multiple simultaneous compiler installations, flexible package constraints, and a Git-friendly development workflow.

★ 1Forks 0

syedazeez337/llm-from-scratch

A GPT-2 style LLM built from scratch in 8 stages — tokenizer → embeddings → attention → transformer → pretraining → generation — with pytest coverage and guided CodeTours.

★ 0PythonForks 0

syedazeez337/vllm-lab

Measuring vLLM's engine internals on a 6GB GPU — 8 self-contained demos with live dashboards

★ 0PythonForks 0