renjianguojinqianfan / dsh-skill-eval

목록에 있음

DSH 插件:用 LLM judge 评测技能 description 的触发准确率(欠触发/过触发)

main모델스킬 소스 보기

설치

npx -y @deepseek-ai/dsh plugin --profile web add github:renjianguojinqianfan/dsh-skill-eval

이 설치 명령은 GitHub 저장소 주소에서 생성된 확인되지 않은 시작점입니다.

README

유지 관리자가 작성한 문서 스냅샷입니다.

GitHub에서 보기 ↗
커밋 8288582동기화 2026. 8. 18.

dsh-skill-eval

CI

Skill-trigger evaluation plugin for DeepSeek Harness (DSH).

An LLM judge recreates the exact DSH skill catalog prompt and decides, for each test query, whether the target skill should be triggered. The plugin reports accuracy, precision, recall, false-positive/negative rates, and a confusion matrix — a reproducible measure of how reliably a skill description routes matching queries (and how often it over- or under-triggers).

Install

From the repo root:

dsh plugin --profile web add ./dsh-skill-eval

Then configure the judge model route in your profile/overlay cordis.patch.yml:

- id: skill-eval
  config:
    provider: <provider-id>
    model: <model-name>

The provider must be registered in the DSH LLM runtime (the same one your profile uses for chat). The plugin validates the route at startup and warns if the provider is not yet registered.

Usage

Slash command:

/skill-eval <skill-name> [test-file]

Model-callable tool:

run_skill_eval(skill_name="<skill-name>", test_file="examples/dsh-plugin-eval.json")

test-file is optional; it defaults to examples/dsh-plugin-eval.json inside the plugin package. Relative paths resolve against the plugin package directory.

Test-case format

A JSON array of { query, should_trigger } objects:

[
  { "query": "add a tool to the harness that persists across restarts", "should_trigger": true },
  { "query": "help me write a Python script for this CSV", "should_trigger": false }
]

category is optional and reserved for future use.

How it works

  1. Enumerate the session's model-invocable skills (ctx.skills.snapshot).
  2. Recreate the official catalog message verbatim (<system-reminder> + <available_skills> + normalized/truncated/escaped descriptions).
  3. For each query, ask the judge model whether the target skill should be triggered, forcing a one-line YES/NO answer.
  4. Compare against the expected label and aggregate metrics.

The evaluation measures the judge model's routing accuracy for the given skill description. Swap provider/model in the config to test other judges.

Development and tests

npm run check     # syntax check for every JS file
npm test          # node:test, including official catalog fidelity and mock ctx tests
npm run smoke     # 51 pure-function smoke assertions
npm pack --dry-run  # published file list check
bash scripts/mount-smoke.sh  # real DSH mount smoke in a scratch home

The catalog fidelity fixture pins the official dsh-tool-skill@0.1.0-rc.6 template. After a DSH upgrade, refresh the fixture from a local official install and review the diff:

node scripts/refresh-catalog-fixture.mjs <path-to-dsh-tool-skill/lib/index.js>

Files

  • index.js — plugin entry: registers the run_skill_eval tool and the /skill-eval command.
  • runner.js — catalog reproduction, judge LLM call, verdict parsing.
  • catalog.js — pure functions: catalog message rendering and verdict parsing.
  • llm-helpers.js — dependency-free BlockAssembler, createUserMessage, deepFreeze (mirrors the official dsh-llm pattern).
  • parser.js — test-case JSON loading and validation.
  • metrics.js — confusion matrix, metrics, and markdown report formatting.
  • examples/ — default test cases.

프로젝트 파일 및 신호

표시된 항목은 디렉터리 스냅샷에서 감지된 공개 저장소 신호입니다.

테스트감지됨
기여 가이드감지됨
예제감지됨

저장소 정보

언어
JavaScript
라이선스
MIT
마지막 업데이트
2026. 8. 17. PM 6:35

신중하게 설치하기

소스 코드, 권한, 수명 주기 스크립트, 의존성 및 네트워크 접근을 검토하고 신뢰하지 않는 플러그인은 격리 환경에서 테스트하세요.