kaixinbaba / dsh-vision-recognizer

Listed

DeepSeek Harness 识图插件:保持 DeepSeek 对话,15+ 供应商视觉模型把图片转译为文字,可在 设置→插件 配置

mainTool View source

Installation

npx -y @deepseek-ai/dsh plugin --profile web add github:kaixinbaba/dsh-vision-recognizer

This installation command is an unverified starting point generated from the GitHub repository address.

README

Maintainer-authored documentation snapshot.

View on GitHub ↗
Commit e25e5b7Synced Aug 18, 2026

dsh-vision-recognizer

English | 简体中文

Keep DeepSeek as the conversation brain, attach images anyway, and switch the image-recognition provider any time from Settings → Plugins. A vision plugin for DeepSeek Harness.

It registers a new provider route (default vision-recognizer, shown as DeepSeek + 识图 in the model picker) that wraps the real DeepSeek adapter: it declares image input (so the attachment preflight and the read_image gate admit images) and, in the request stream, transcribes every attached image to text through the vision model you select, then delegates the text-only conversation to DeepSeek. DeepSeek still answers; recognition is an add-on.

attached image ──▶ vision-recognizer route ──▶ vision-model transcription (OCR + layout + detail)
                     │                          │
                     ▼                          ▼
              DeepSeek answers ◀── text-only conversation (image replaced by [图片转译] text)

Features

  • One-click install: dsh plugin --profile web add dsh-vision-recognizer — no build scripts, no sharp approval (no native dependencies at all).
  • Configure from Settings → Plugins → Vision: pick a provider, enter an API key, override model / endpoint / token cap / timeout / marker. Saved changes take effect immediately, no restart.
  • 15+ providers, domestic and international: OpenAI, Anthropic Claude, Google Gemini, OpenRouter, Azure OpenAI, Ollama (local), plus Alibaba DashScope, QwenCloud (Intl), Zhipu GLM, Baidu Qianfan, iFlytek Spark, Moonshot Kimi, Tencent Hunyuan, Volcengine Doubao, SiliconFlow. Any OpenAI-compatible endpoint works via the custom provider.
  • Two wire protocols: OpenAI-compatible (/chat/completions) and native Anthropic Messages — Claude works out of the box.
  • No hangs: local/anonymous endpoints get a hard 20s timeout cap, HTTP 429 fails fast, failed endpoints cool down for 60s; without a key and without local Ollama it fails fast with actionable guidance.
  • Fallback chain: after the primary model fails, each fallbackModels entry is tried in order (each may target a different vendor); only after all fail does the request fail, listing every attempt.
  • Content-hash cache: the same image is transcribed at most once per process (in-process, capped at 200).
  • Zero-config local path: autoLocalOllama (default on) probes http://localhost:11434 and prepends a running Ollama to the chain — images never leave your machine.

Quick start

dsh plugin --profile web add dsh-vision-recognizer

Slow npm registry? dsh plugin --profile web add dsh-vision-recognizer --registry=https://registry.npmmirror.com

Install from a local checkout (development):

dsh plugin --profile web add file:/path/to/dsh-vision-recognizer

Use the file: prefix (copies the package into node_modules). A bare add . or add link:… makes pnpm symlink the package, in which case the plugin's schemastery dependency resolves from the source checkout and is not found — a general pnpm symlink-install gotcha, not a bug in the plugin.

Restart dsh web, then:

  1. Pick DeepSeek + 识图 in the model selector;
  2. Open Settings → Plugins → Vision, choose a provider, enter an API key, save;
  3. Paste an image into any conversation → you should see the [图片转译] marker followed by a DeepSeek answer.

With no key and no local Ollama, a turn fails fast in a few seconds with guidance — that is the intended anti-hang behavior.

Supported providers

ProviderbaseURLDefault modelKey env varProtocol
OpenAIhttps://api.openai.com/v1gpt-4o-miniOPENAI_API_KEYOpenAI
Anthropic Claudehttps://api.anthropic.com/v1claude-3-5-sonnet-latestANTHROPIC_API_KEYAnthropic
Google Geminihttps://generativelanguage.googleapis.com/v1beta/openaigemini-2.0-flashGEMINI_API_KEYOpenAI
OpenRouterhttps://openrouter.ai/api/v1qwen/qwen-2.5-vl-72b-instructOPENROUTER_API_KEYOpenAI
Azure OpenAIuser-supplied (…/openai/deployments/<deployment>)gpt-4o-miniAZURE_OPENAI_API_KEYOpenAI
Ollama (local)http://localhost:11434/v1auto-detectednoneOpenAI
Alibaba DashScopehttps://dashscope.aliyuncs.com/compatible-mode/v1qwen-vl-maxDASHSCOPE_API_KEYOpenAI
QwenCloud (Intl)https://dashscope-intl.aliyuncs.com/compatible-mode/v1qwen-vl-plusDASHSCOPE_API_KEYOpenAI
Zhipu GLMhttps://open.bigmodel.cn/api/paas/v4glm-4v-flashZHIPU_API_KEYOpenAI
Baidu Qianfanhttps://qianfan.baidubce.com/v2ernie-4.5-vl-8kQIANFAN_API_KEYOpenAI
iFlytek Sparkhttps://spark-api-open.xf-yun.com/v1generalv3.5SPARK_API_KEYOpenAI
Moonshot Kimihttps://api.moonshot.cn/v1moonshot-v1-8k-vision-previewMOONSHOT_API_KEYOpenAI
Tencent Hunyuanhttps://api.hunyuan.cloud.tencent.com/v1hunyuan-visionHUNYUAN_API_KEYOpenAI
Volcengine Doubaohttps://ark.cn-beijing.volces.com/api/v3doubao-1.5-vision-pro-32k-250115ARK_API_KEYOpenAI
SiliconFlowhttps://api.siliconflow.cn/v1Qwen/Qwen2.5-VL-72B-InstructSILICONFLOW_API_KEYOpenAI

Model ids drift over time; the defaults are starting points — override Model in the settings UI. Key resolution order: key entered in the UI → the provider env var → $VISION_API_KEY / $DASHSCOPE_API_KEY.

Configuration storage

Config saved from the UI is written to $DSH_HOME/vision-recognizer.json and merged over the bundle defaults at startup. cordis.patch.yml only carries factory defaults; a user cordis.patch.yml override still works as the composition-time fallback.

⚠️ patch semantics: the bundle's - insert: appends this row to the entry list. Writing a second - insert: with the same id in your own cordis.patch.yml would register the adapter twice (undefined behavior). To override individual keys, write a single top-level - id: dsh-vision-recognizer entry; better yet, use the Settings UI.

Implementation notes (for plugin authors)

Only stable rc.6 public interfaces are used:

  • ctx.llm.registration(innerProvider).adapter — fetch the wrapped adapter;
  • ctx.llm.registerAdapter([providerId], proxyAdapter) — register the new route;
  • proxying resolveModel overrides inputModalities to ['text', 'image'];
  • proxying stream transcribes image blocks ({ type: 'image', attachment }, bytes read via ctx.get('attachments').readImage(ref)), then yield* forwards the inner adapter's stream;
  • the settings UI rides the settings.plugins.tab slot plus custom webServer routes, persisting config to its own JSON file (independent of the api-proxy settings allowlist).

Privacy

Transcription sends image bytes (base64, HTTPS) to the vision endpoint you configure — image data leaves your machine unless the endpoint is local (e.g. Ollama). Nothing beyond the harness's own attachment storage persists any image. For sensitive images, use your own endpoint or a local model, or don't install this plugin.

License

MIT

Project files and signals

Shown items are public repository signals detected in the directory snapshot.

TestsDetected

Repository information

Language
JavaScript
License
MIT
Latest release
v0.1.2
Last updated
Aug 16, 2026, 7:20 AM

Install deliberately

Review source code, permissions, lifecycle hooks, dependencies and network access. Test untrusted plugins in an isolated environment.