Context compressor - tool for improving vibe-coding experience, preventing LLMs from forgetting useful things from your chat session.
Context Compressor is a FastAPI-based web application that compresses chat conversations to help maintain context in LLM interactions. It uses OpenAI's API to summarize and condense chat histories by 80% or more, preserving only the most important information about user preferences and assistant decisions.
- Conversation Compression: Reduce chat history by 80%+ while preserving key information
- FastAPI Backend: RESTful API with health check endpoint
- Jinja2 Templating: Flexible template system for conversation formatting
- OpenAI Integration: Uses configurable LLM models for compression
- Easy Setup: Simple configuration with environment variables
- Docker Support: Containerized deployment with Dockerfile
- Python 3.12 or higher
- OpenAI API key
- Access to an OpenAI-compatible API endpoint
- Clone the repository:
git clone https://github.com/finettt/context-compressor
cd context-compressor- Install dependencies:
pip install -e .- Configure environment variables:
touch .envEdit .env file with your configuration:
BASE_URL=your-openai-compatible-api-url
API_KEY=your-api-key
MODEL=your-model-name # e.g., gpt-3.5-turbo, gpt-4uv run python main.pydocker build -t context-compressor .
docker run -p 8000:8000 --env-file .env context-compressorThe application will start on http://localhost:8000
curl http://localhost:8000/healthResponse:
{"status": "healthy"}curl -X POST "http://localhost:8000/chat/completion" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello"}, {"role": "assistant", "content": "Hi there!"}]}'Response:
{
"content": "Compressed conversation summary...",
"message": "Now, you can use this message instead of previous history"
}You can integrate the context compressor into your chat application by making HTTP requests to the /chat/completion endpoint. Send your conversation messages and receive a compressed summary that can be used to maintain context in LLM interactions.
context-compressor/
โโโ main.py # Application entry point
โโโ src/
โ โโโ app.py # FastAPI application setup
โ โโโ routes/
โ โ โโโ api.py # API endpoints
โ โโโ client/
โ โโโ main.py # Core compression logic
โโโ assets/
โ โโโ task_template.jinja2 # Compression task template
โ โโโ chat_template.jinja2 # Conversation formatting template
โโโ pyproject.toml # Project configuration
โโโ app.dockerfile # Docker configuration
โโโ .dockerignore # Docker ignore file
โโโ README.md # This file
| Variable | Description | Required |
|---|---|---|
BASE_URL |
OpenAI-compatible API endpoint | Yes |
API_KEY |
API key for authentication | Yes |
MODEL |
Model name to use for compression | Yes |
You can customize the compression behavior by modifying the templates in the assets/ directory:
task_template.jinja2: Controls how the compression task is presented to the LLMchat_template.jinja2: Controls how conversation messages are formatted
This project uses several tools to maintain code quality:
- Ruff: Fast Python linter and code formatter
- Ty: Modern task runner
- Pytest: Testing framework
- Bandit: Security linter
- Safety: Security vulnerability checker
# Install development dependencies
pip install -e ".[dev]"
# Run tests
pytest
# Run linting
ruff check .
# Format code
ruff format .- Fork the repository
- Create a feature branch
- Make your changes
- Add tests if applicable
- Submit a pull request
This project is licensed under the APGL-3.0 License - see the LICENSE file for details.
- Built with FastAPI
- Uses OpenAI API
- Template engine powered by Jinja2
- Package management with uv