Topic: text-only-llm

15 repositories

liustack/modlens

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

★ 4,100TypeScriptForks 126

Anionex/agent-vision-toolkit

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

★ 1,216PythonForks 48

Anionex/dsh-vision-toolkit

[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

★ 887TypeScriptForks 47

loudMore/dsh-drop-to-path

DSH 插件:图片与文件直达纯文本模型——图片保留原生附件体验,PDF/Office/压缩包/视频/音频显示为附件栏方块,点击发送时自动转为工作区路径,配合 dsh-vision-toolkit 粘贴即看图。A DSH plugin that delivers images AND files to text-only models as workspace paths: images keep the native attachment UI, other files show as square chips in the rail, paths append on send — pairs with dsh-vision-toolkit.

★ 11JavaScriptForks 2

SeverinQuan/mimo-vision-mcp

Community MCP vision bridge for Xiaomi MiMo Vision, enabling image understanding for text-only LLM agents.

★ 8PythonForks 1

v587d/multimodal-skill

Give text-only LLMs eyes. A Pi Agent skill + zero-dependency Python CLI that adds image understanding and document parsing (OCR, tables, formulas, PDF → Markdown) to any text-only model such as DeepSeek, using free-tier third-party multimodal APIs.

★ 6PythonForks 0

congchuanling-dot/DSH-Telegram-Relay

DSH Relay 让你可以通过 Telegram 远程与 DeepSeek Harness 对话,并接收通知。DSH Relay turns Telegram into a remote conversation and notification channel for DeepSeek Harness.

★ 4TypeScriptForks 0

Penty-d/her-eyes

带上她的眼睛 · Give a text-only LLM eyes — a single-binary MCP tool that lets agents like Claude Code / Codex call a vision model to extract structured key information from images.

★ 2GoForks 0

ZhuXinAI/sidesight

CLI-first vision sidecar for text-only coding agents. Analyze screenshots, diagrams, charts, UI diffs, and videos with OpenAI-compatible multimodal models.

★ 2TypeScriptForks 0

uknowmyface/locallens

Local OCR for DeepSeek Harness — read text from screenshots on your Mac with Apple's Vision framework. No API key, no upload.

★ 1JavaScriptForks 0

zjcdkj/dsh-plugins

DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH 插件集:qwen-image 让纯文本模型借千问 VL 读图,返回文本;纯 ESM,安装无需构建授权。

★ 1JavaScriptForks 0

theJian/squint-mcp

Most LLMs see images. With squint-mcp, the rest imagine seeing them.

★ 1TypeScriptForks 0

kanchengw/dsh-mindseye

Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, and GUI automation, through layered evidence memory and cache.

★ 1TypeScriptForks 0

junhongchashui/dsh-vision-relay

零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。

★ 0JavaScriptForks 0