syedazeez337/tiny-dino
A 21,699-parameter MLP that plays a deterministic Chromium Dino runner with no planner, solver, or safety shield.
Working on Open Source projects #Cilium #CoreDNS
A 21,699-parameter MLP that plays a deterministic Chromium Dino runner with no planner, solver, or safety shield.
Semantic emoji search visualized as physics. Type a phrase, matching emoji rise out of the pile. Qwen3 Embedding at Cloudflare's edge, cosine in the browser.
JOB
Interactive Slidev explainer: DeepSeek-V4.1-Flash Causal Encoder-Decoder vs BERT and classic encoder-decoder. Live: https://syedazeez337.github.io/from-bert-to-deepseek-v41/
An experiment platform that makes agent-architecture claims falsifiable. Runs coding agents as experimental subjects in isolated environments, records immutable evidence, and grades both outcomes and their trustworthiness with cost matching, FDR control, and equivalence testing.
Macaw — on-device AI agent for macOS. 2.7B LFM2.5 fine-tune, MLX 4-bit, 97 macOS tools, no cloud.
Desktop app that records your on-screen work session and uses the GitHub Copilot CLI to reconstruct it as an intent + ordered steps, then builds a reusable Skill or Automation for Microsoft Scout, Microsoft Copilot Cowork, or Copilot Studio.
Under the Hood - Build Every Layer of a Large Language Model from Scratch. A 36-project, build-it/break-it/measure-it manual by Ramchand Kumaresan.
Learn Modern C++ by Debugging It. A workbook of broken C++26 programs, for Linux and native Windows.
KV-cache-aware LLM inference mesh on a single 6 GB GPU. 2× vLLM + LMCache shared KV + prefix-aware router on k3s, observed with Cilium/Hubble eBPF. Every number measured, every breakage documented.
A Go binary that checks whether a GPU is the card it claims to be. Measures bandwidth, FP32 throughput, VRAM integrity, SM count and bus width, compares them against an embedded spec table, and writes an ed25519-signed scorecard. Driver-only: no CUDA toolkit needed.
A beginner-friendly, from-scratch course on speculative decoding: plain-English lessons, an interactive explainer, designed PDF notes, and runnable cross-platform code (CUDA / Apple Metal / CPU).
Local hybrid RAG over the vLLM codebase — BGE-M3 + Qdrant (dense+sparse, RRF) + cross-encoder rerank + Qwen3.5-4B, with a source-grounded evaluation harness.
Verifying Reiner Pope's 'data-movement tax' chip-design claim on an RTX 3050 (6GB, sm_86) with measured roofline + Triton kernels
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
Small-scale single-GPU (6GB) reproduction of Attention Drift / EAGLE-3.1 (arXiv:2605.09992) + a probe of its untested §4.4 verifier-quantization hypothesis
Complete SLM Toolkit — build, train, align, serve and operate Small Language Models from first principles. Includes internal-docs RAG assistant.
opam is a source-based package manager. It supports multiple simultaneous compiler installations, flexible package constraints, and a Git-friendly development workflow.
A GPT-2 style LLM built from scratch in 8 stages — tokenizer → embeddings → attention → transformer → pretraining → generation — with pytest coverage and guided CodeTours.
Measuring vLLM's engine internals on a 6GB GPU — 8 self-contained demos with live dashboards
Learn GPU Programming in Mojo🔥 by Solving Puzzles
tiny Go library to normalize URLs
Config files for my GitHub profile.