yossiovadia/hindsight
Hindsight: Agent Memory That Learns
Hindsight: Agent Memory That Learns
Development metering service for the AI Inference Gateway — simulates OpenMeter, Monetize360, and other metering backends for testing per-user token usage tracking and cost attribution
Side-Eye — on-demand session review: draft cheap, review expensive
Collection of Architectural Decision Records
AI features and capabilities for Praxis
Live demo/education dashboard for MaaS (Models as a Service) platform
One tiny model, every LLM API. Drop-in test server for OpenAI, Anthropic, Bedrock, and Vertex. Real inference, no API keys, no GPU needed.
Inference payload processor for llm-d
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Enterprise context compression gateway — deploy headroom as a centralized proxy service on OpenShift/Kubernetes
AI Gateway documentation, gap analysis, and design docs
A framework for efficient model inference with omni-modality models
That's just, like, your opinion, man. — AI-powered infrastructure-level bug reproduction and validation.
That's just like, literally what it is, man. One word at a time. — RSVP speed reader
model-as-a-service
Gateway API Inference Extension
Vision/Mission https://ambient-code.ai : Virtual team management and collaboration platform. User guides: https://ambient-code.github.io/platform/
The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools, not merely storing the archive of leaked Claude Code but make real things done. Now rewriting in Rust using oh-my-codex.
ReproBot — AI-powered infrastructure-level bug reproduction demo
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
AI-powered SRE monitoring using local LLMs (Ollama + Qwen/Llama/etc)
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
LLM inference in C/C++
Configure OpenClaw to use Ollama + Claude Code CLI
Intelligent Mixture-of-Models Router for Efficient LLM Inference
Custom scripts and configs for Claude Code CLI
Demo project 2 - testing PM/Dev agents