v587d/multimodal-skill
Give text-only LLMs eyes. A Pi Agent skill + zero-dependency Python CLI that adds image understanding and document parsing (OCR, tables, formulas, PDF → Markdown) to any text-only model such as DeepSeek, using free-tier third-party multimodal APIs.