NVIDIA-NeMo/Automodel
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
40 repositories
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
macOS menubar app for fast local DeepSeek V4.1, with 1M context.
⚡️ A community driven PHP client for DeepSeek AI, designed to bring clean API access, fluent developer experience, and framework-friendly integration to PHP applications.
Fixes missing reasoning_content for DeepSeek V4
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support
基于 FastAPI 的 DeepSeek Chat 反向代理,将 DeepSeek 网页版的 API 转换为 OpenAI 兼容格式。 支持流式/非流式对话、专家模式、深度思考(reasoning_content)、工具调用(DSML prompt injection)。 自动处理 PoW 鉴权挑战,无需官方 API Key。
DeepSeek-V4-Flash-0731 284B inference in ~25 GB of RAM / Qwen3.8-Next-Flash-FP8 inference in ~18 GB of RAM on any M-series MacBook
DeepSeek V4 Flash CPU/NVMe research fork: 78.62 GiB GGUF validated on 7.7 GiB RAM, CPU-only, using demand paging.
My public AI learning journey — sharing what I learn, build, test, and discover across AI. Traffic: 7K+ monthly clicks · 150K+ AI citations/month · 5.5 average search position across search engines. local-ai-zone.github.io
A tool to have multiple claude-code instance with deepseek, minimax, and z.ai glm models
Production-ready, reproducible Ansible for DeepSeek V4 Flash on 128 GiB AMD Strix Halo, with two qualified Vulkan/ROCmFPX stacks, matched quality and throughput benchmarks, and 512K context validation.
Guide for development with DeepSeek Harness. Building plugin for DeepSeek Harness Project.
Reproducible kit to deploy DeepSeek-V4-Flash-DSpark on a 2× NVIDIA DGX Spark (GB10) cluster: vLLM TP=2 over QSFP 200GbE, NVFP4 KV, DSpark speculative decoding, 1M context, systemd self-heal. Apache-2.0.
A collection of recipes/notebooks showcasing use-cases of open-source models with Qubrid AI.
Codex 桌面版为 DeepSeek-V4-Flash 开启 Max 推理档位的完整排障与配置指南 / Complete guide to enable Max reasoning effort for DeepSeek-V4-Flash in Codex desktop
Codex vision bridge for DeepSeek V4 Flash: give text-only DeepSeek image capability in Codex. Local proxy turns pasted images and view_image into text via free GLM-4V-Flash or any OpenAI-compatible vision API. No GPU, no Ollama.
Dynamic Agent-to-Agent (A2A) task graph generation, subtask independence verification, and parallel multi-agent orchestration
Supercharge your daily workflows! Autonomously browse pages, scrape content, switch tabs, fill out inputs, and stream Chain-of-Thought (CoT) reasoning under granular human-in-the-loop safety switches and glassmorphic UI controls.
Unlock Claude Code with DeepSeek V4. Get Anthropic's agent tools with 95% lower costs and local vision.
Golang wrapper for DwarfStar4 (ds4)
A Rust proxy that exposes an **OpenAI Responses API** interface, translating to various upstream backends (DeepSeek, OpenAI, Anthropic) with full streaming SSE support. Built for Codex CLI and similar tools hardcoded to OpenAI.
A browser-side analytics dashboard for DeepSeek API usage. Drag your monthly CSV exports onto the page and get instant cost charts, per-key breakdowns, cache analysis, and usage trends — all processed locally in your browser. no upload, no signup. 一款纯浏览器端的 DeepSeek API 用量分析仪表盘。将月度 CSV 导出文件拖拽到页面,即刻获取费用图表、各 Key 用量明细、缓存分析和用量趋势 — 所有数据均在浏览器本地处理。
AI编程助手 | AI coding agent
LLM-as-a-Verifier (arXiv:2607.05391) as a dsh plugin — Best-of-N conversation mode: give DeepSeek V4 Flash test-time scaling. Bo5 self-verification hits 88% on Terminal-Bench 2.1, beating some frontier models at a fraction of the cost. Fine-grained logprob-expectation scoring, PPT tournament, zero-config.
DeepSeek Harness 超高自由度的个性化插件,支持液态玻璃效果、自定义图片、视频背景、字体颜色,同时支持批准提示音和账户余额查询。
Most LLMs see images. With squint-mcp, the rest imagine seeing them.
DSH Web composer stats strip: official CNY pricing with auto-synced peak/off-peak tiers, per-model accounting, live LLM/tool timers, agent-team tree merge, DeepSeek balance, budget alerts, streaming cost estimates
Run llama.cpp across two GPUs of different vendors at once (AMD ROCm/HIP or Vulkan + NVIDIA CUDA) in one llama-server process, no RPC server. Windows (PowerShell) and Linux (bash) implementations, shared docs on ROCm memory bugs and benchmarks.
DeepSeek-V4-Flash 推理引擎 — 纯 Rust / 纯 CPU / 零依赖 / MoE 284B 满血模型本地推理
A high-performance API Gateway that standardizes DeepSeek endpoints for AI agents and developer tools (fully compatible with OpenAI, Claude, and Gemini protocols).