ferstar / dsh-tool-ocr

목록에 있음

本地 OCR 插件:让纯文本生成 LLM 也能读懂图片 | Local OCR plugin: give text-only generative LLMs the ability to read images

master모델도구 소스 보기

설치

npx -y @deepseek-ai/dsh plugin --profile web add github:ferstar/dsh-tool-ocr

이 설치 명령은 GitHub 저장소 주소에서 생성된 확인되지 않은 시작점입니다.

README

유지 관리자가 작성한 문서 스냅샷입니다.

GitHub에서 보기 ↗
커밋 cc5f3ac동기화 2026. 8. 18.

dsh-tool-ocr

English | 中文

Local image text recognition for DeepSeek Harness models without vision input, backed by the standalone newbee-ocr (nbocr) engine over PP-OCRv6 models. Fully out-of-tree: depends only on published dsh base packages.

Features

ItemDescription
ocr toolrecognize / status / install / check actions
Image inputsLocal path, or attachment_id (an image already in the conversation)
OutputReading-ordered text, per-block bounding boxes, review flags (low confidence / amounts / numbers / dates / quantities), heuristic Markdown tables
RobustnessSurvives MNN diagnostics in engine stdout; honors cancellation and timeout
Settings cardOn dsh ≥ 0.1.0-rc.7 the plugin registers the tool-ocr namespace: the web Settings → Plugins page renders a card that edits the engine command, model tier, and run bounds live — no composition edit needed

Quick start

1. Install the engine — grab nbocr from the newbee-ocr-cli releases (one-line installer or prebuilt archive):

curl -LsSf https://github.com/zibo-chen/newbee-ocr-cli/releases/latest/download/newbee_ocr_cli-installer.sh | sh
# Windows PowerShell: irm https://github.com/zibo-chen/newbee-ocr-cli/releases/latest/download/newbee_ocr_cli-installer.ps1 | iex

2. Install the plugin

dsh plugin --profile web add dsh-tool-ocr
# or: cd ~/.dsh/profiles/web && pnpm add dsh-tool-ocr

3. Mount it in your profile's cordis.patch.yml:

- insert:
    - id: ocr
      name: 'dsh-tool-ocr'
      inject: [tools, subprocess, systemPrompt]
      config:
        command: 'C:/path/to/nbocr.exe'   # required — the nbocr executable

That's it — restart dsh, then point the model at an image: ocr { path: "C:/screenshot.png" }. To tune recognition (language, model tier, limits), open Settings → Plugins → Plugin config and edit the OCR card; changes apply on save, no restart needed. command is the only field that must come from the composition.

Image inputs

  • path works on every dsh build: the model reads the file from disk.
  • attachment_id needs a build that admits images for text-only models (e.g. the dsh fork with api-gateway.allowImagePlaceholder: true): images enter the session, the llm layer substitutes [image attachment <id>] text, and the model hands the id to ocr.

Configuration

FieldDefaultMeaning
command— (required)The nbocr executable: absolute path or PATH-resolved name.
args[]Extra arguments before the nbocr subcommand (no shell).
env{}Extra environment entries.
languagechineseRecognition model/language alias.
detModelv6-tinyDetection model tier.
modelsDir''Model directory for non-embedded models; empty uses embedded models.
maxImageBytes26214400Largest accepted image in bytes.
maxOutputBytes2000000Largest collected engine stdout in bytes.
maxTextChars12000Largest recognized text returned in text.
timeoutMs600000Tool-call timeout budget in ms.

Tool usage

ocr { path | attachment_id, action?, include_boxes?, table?, max_text_chars? }

status probes readiness without engine work; check runs a real end-to-end 1x1 probe.

Development

pnpm typecheck   # tsc --noEmit
pnpm test        # vitest
pnpm lint        # oxlint

Real-engine smoke (skipped unless both env vars are set):

NBOCR_BIN=/path/to/nbocr OCR_E2E_IMAGE=/path/to/image.png pnpm test

License

MIT

프로젝트 파일 및 신호

표시된 항목은 디렉터리 스냅샷에서 감지된 공개 저장소 신호입니다.

테스트감지됨

저장소 정보

언어
TypeScript
라이선스
MIT
최신 릴리스
v0.1.0
마지막 업데이트
2026. 8. 18. 오후 1:40

신중하게 설치하기

소스 코드, 권한, 수명 주기 스크립트, 의존성 및 네트워크 접근을 검토하고 신뢰하지 않는 플러그인은 격리 환경에서 테스트하세요.