Kolade Fajimi

@koladefaj · User

GitHub profile ↗ · Compare

Backend & AI Engineer | FastAPI • Celery • Redis • gRPC | Building Production RAG, MLOps & Distributed Systems at Scale

Nigeria.12 followers17 repositories

Repositories

koladefaj/JamAIBase

The collaborative spreadsheet for AI. Chain cells into powerful pipelines, experiment with prompts and models, and evaluate LLM responses in real-time. Work together seamlessly to build and iterate on AI applications.

★ 0PythonForks 0

koladefaj/litellm

Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]

★ 0PythonForks 0

koladefaj/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

koladefaj/AReaL

Lightning-Fast RL for LLM Reasoning and Agents. Made Simple & Flexible.

★ 0PythonForks 0

koladefaj/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0PythonForks 0

koladefaj/StoryCanon

A local, privacy-first writing tool that catches continuity errors in fiction and lets the author decide what's canon.

★ 1TypeScriptForks 0

koladefaj/autograder

Flexible, Consistent and Powerful Autograding Tool that grades and generates reports on your students submissions.

★ 0PythonForks 0

koladefaj/embeddedllm

EmbeddedLLM: API server for Embedded Device Deployment. Currently support CUDA/OpenVINO/IpexLLM/DirectML/CPU

★ 0Forks 0

koladefaj/gia

Conversational voice AI with streaming STT/TTS, speculative execution, and a distilled routing cascade. Features background memory consolidation using Weaviate hybrid retrieval and mood inference pipelines via Celery.

★ 0PythonForks 0

koladefaj/Phalanx

Microsecond-latency ML fraud orchestration engine featuring gRPC microservices, sub-millisecond ONNX inference, and self-healing MLOps workflows.

★ 1PythonForks 0

koladefaj/Engram

Hardware-accelerated RAG pipeline and document intelligence engine. Features pgvector-native storage, asynchronous OCR workers, and local LLM orchestration via Ollama. Built for high-concurrency knowledge retrieval.

★ 0PythonForks 0

koladefaj/Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

★ 0PythonForks 0

koladefaj/e-pharmacy-backend

A full-featured e-pharmacy backend API built with FastAPI, featuring role-based access control (User/Pharmacist/Admin), prescription verification workflow, Stripe payment integration and asynchronous task processing with FastAPI background task.

★ 0PythonForks 0

koladefaj/Telco-churn-prediction

Machine learning project predicting telecom customer churn using a tuned XGBoost model with custom threshold optimization, EDA, and deployment-ready pipeline.

★ 0Jupyter NotebookForks 0

koladefaj/data-wrangling

Beginner-friendly data wrangling and visualization projects using Python (Pandas, NumPy, Matplotlib)

★ 0Jupyter NotebookForks 0