Chebaleomkar/ZeroHost-vLLM
Production-style LLM inference on free compute: vLLM + Qwen2.5-7B on a Kaggle T4, OpenAI-compatible API, tool calling, full tracing and benchmarks.
Inference | AI/ML Engineer
Production-style LLM inference on free compute: vLLM + Qwen2.5-7B on a Kaggle T4, OpenAI-compatible API, tool calling, full tracing and benchmarks.
Single GPU vs data parallel vs tensor parallel on Kaggle's free 2x T4: vLLM + Qwen2.5-7B, load-balancing proxy, in-session benchmarks, OpenAI-compatible API.
Open-source AI coworkers that each get a computer of their own: a browser, files and tools, with every action decided before it happens and recorded after. Bring any AG-UI agent.
Supermemory plugin for OpenCode
Open-sourcing our company brain - A teammate in your Slack that remembers everything your team says, and can go do the work.
A mini Subconscious Cache: pruning an agent's context with prefix + suffix KV reuse, measured with a real coding agent on a free Kaggle T4
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Native Windows desktop notifications for Claude Code — works with any terminal
Fix desktop notifications for Claude Code's Warp plugin on Windows
Official Warp terminal integration for Claude Code - native notifications and more
My solutions to TensorTonic problems
Built a parameter-efficient domain adaptation pipeline for legal clause simplification using QLoRA on Gemma 2B.
Outcome driven agent development framework that evolves
Full stack developer assesment of cooltrack organization
An agent-based AI application that generates, reviews, and refines educational content through a transparent multi-step workflow. It uses specialized generator and reviewer agents, structured JSON outputs, and a simple UI to demonstrate practical AI agent pipelines for real-world use cases.