The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 4,100TypeScriptForks 126
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 1,216PythonForks 48
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
★ 887TypeScriptForks 47
DSH 插件:图片与文件直达纯文本模型——图片保留原生附件体验,PDF/Office/压缩包/视频/音频显示为附件栏方块,点击发送时自动转为工作区路径,配合 dsh-vision-toolkit 粘贴即看图。A DSH plugin that delivers images AND files to text-only models as workspace paths: images keep the native attachment UI, other files show as square chips in the rail, paths append on send — pairs with dsh-vision-toolkit.
★ 11JavaScriptForks 2
Community MCP vision bridge for Xiaomi MiMo Vision, enabling image understanding for text-only LLM agents.
★ 8PythonForks 1
Give text-only LLMs eyes. A Pi Agent skill + zero-dependency Python CLI that adds image understanding and document parsing (OCR, tables, formulas, PDF → Markdown) to any text-only model such as DeepSeek, using free-tier third-party multimodal APIs.
★ 6PythonForks 0
DSH Relay 让你可以通过 Telegram 远程与 DeepSeek Harness 对话,并接收通知。DSH Relay turns Telegram into a remote conversation and notification channel for DeepSeek Harness.
★ 4TypeScriptForks 0
带上她的眼睛 · Give a text-only LLM eyes — a single-binary MCP tool that lets agents like Claude Code / Codex call a vision model to extract structured key information from images.
★ 2GoForks 0
CLI-first vision sidecar for text-only coding agents. Analyze screenshots, diagrams, charts, UI diffs, and videos with OpenAI-compatible multimodal models.
★ 2TypeScriptForks 0
Local OCR for DeepSeek Harness — read text from screenshots on your Mac with Apple's Vision framework. No API key, no upload.
★ 1JavaScriptForks 0
DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH 插件集:qwen-image 让纯文本模型借千问 VL 读图,返回文本;纯 ESM,安装无需构建授权。
★ 1JavaScriptForks 0
Most LLMs see images. With squint-mcp, the rest imagine seeing them.
★ 1TypeScriptForks 0
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, and GUI automation, through layered evidence memory and cache.
★ 1TypeScriptForks 0
DeepSeek Harness Vision Helper/DeepSeek Harness 视觉辅助方案
★ 1JavaScriptForks 0
零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。
★ 0JavaScriptForks 0