dsh-plugin-evaluation / dsh-plugin-evaluation-standards

목록에 있음

Open evaluation datasets, test cases, and metrics for DSH plugins.

main모델 소스 보기

설치

npx -y @deepseek-ai/dsh plugin --profile web add github:dsh-plugin-evaluation/dsh-plugin-evaluation-standards

이 설치 명령은 GitHub 저장소 주소에서 생성된 확인되지 않은 시작점입니다.

README

유지 관리자가 작성한 문서 스냅샷입니다.

GitHub에서 보기 ↗
커밋 e1eb1db동기화 2026. 8. 18.

DSH Plugin Evaluation Datasets

English | 中文 | 日本語

A growing collection of evaluation datasets for DSH plugins.

Each dataset is a profile (which metrics to use) and a cases file (test prompts and expected answers). Pick one that fits your plugin, run its cases, and use the results to understand how your plugin behaves.

Start here

  1. Browse the datasets.
  2. Choose one that matches your plugin and the scenarios you want to cover.
  3. Open its profile and cases files.
  4. Run the cases against your plugin and review the results.

Need a dataset that is not here yet? Use the AI-assisted authoring guide to draft one, then contribute it.

Build this collection with us

Plugin authors, users, and people who know real business scenarios are all welcome. You do not need a finished JSON dataset to participate:

  • Have a real scenario? Open an issue with how a user would ask, what the plugin should do, and the supporting facts or setup conditions.
  • Have a small set of cases? Submit a profile and cases following the contribution guide.
  • Maintain a dataset long term? Keep it in your own repository and add it to this catalog using the external dataset listing guide.

Common tasks, tricky conditions, and cases where a plugin should avoid making things up are all valuable. Do not submit private business material, personal data, or secrets.

Datasets

DatasetPlugin typeCoversCasesMetrics
Knowledge Query Basicsknowledge-queryRefunds, shipping, invoices3answer-matches-expected, duration

Knowledge Query Basics

A small starting set for plugins that look up clear facts from an installed knowledge source.

Included cases
CaseWhat it checksExpected answer
Refund request windowThe default deadline for a refund request30 days
Standard shipping SLAThe promised time for standard shipping3 business days
Electronic invoice channelWhere an electronic invoice is sentEmail address bound to the order
Metrics
  • answer-matches-expected checks whether the final answer matches the expected answer. It decides whether a case passes.
  • duration records how long the case takes. It does not change the pass/fail result.

Dataset files

Each dataset has two files:

profiles/<id>.json  Which metrics to use and where to find the cases
cases/<id>.json     Plugin types and test cases

A test case looks like this:

{
  "id": "case-id",
  "title": "A short name for the case",
  "prompt": "The input sent to the plugin",
  "expected": "The answer you expect"
}

Supported metrics

Metric typeAvailable nowChanges pass/fail
llm_judgeYesYes
observationYesNo
tool_traceNot yetNo
thresholdNot yetNo

Add a dataset

You can contribute a small dataset directly to this repository, or keep a larger dataset in its own repository and add it to the catalog.

npm run validate
npm test

프로젝트 파일 및 신호

표시된 항목은 디렉터리 스냅샷에서 감지된 공개 저장소 신호입니다.

테스트감지됨
보안 정책감지됨
기여 가이드감지됨

저장소 정보

언어
JavaScript
라이선스
CC0-1.0
최신 릴리스
v1.0.0
마지막 업데이트
2026. 8. 18. AM 5:17

신중하게 설치하기

소스 코드, 권한, 수명 주기 스크립트, 의존성 및 네트워크 접근을 검토하고 신뢰하지 않는 플러그인은 격리 환경에서 테스트하세요.