dshbase

插件目录 / Developer / dsh-skill-eval

dsh-skill-eval

未验证 renjianguojinqianfan

✓ 持续维护

查看 GitHub ↗ ← 返回插件目录

1Stars
0Forks
0未关闭 issue
语言
2026-08-17最近推送
跨平台平台

功能简介

DSH 插件:用 LLM judge 评测技能 description 的触发准确率(欠触发/过触发)

我们的评价
未验证 — 尚未实测

DSH 插件:用 LLM judge 评测技能 description 的触发准确率(欠触发/过触发) 尚未验证——请自行安装测试。

「未验证」表示我们的自动化 CI 尚未安装过该插件。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

你是插件作者? 想拿到「已验证」标签——提交你自己的验证证据(截图、日志或短视频),我们审核通过后即改为「已验证」。

提交验证证据 ↗

README

dsh-skill-eval

CI

Skill-trigger evaluation plugin for DeepSeek Harness (DSH).

An LLM judge recreates the exact DSH skill catalog prompt and decides, for each
test query, whether the target skill should be triggered. The plugin reports
accuracy, precision, recall, false-positive/negative rates, and a confusion
matrix — a reproducible measure of how reliably a skill description routes
matching queries (and how often it over- or under-triggers).

Install

From the repo root:

dsh plugin --profile web add ./dsh-skill-eval

Then configure the judge model route in your profile/overlay cordis.patch.yml:

- id: skill-eval
  config:
    provider: <provider-id>
    model: <model-name>

The provider must be registered in the DSH LLM runtime (the same one your
profile uses for chat). The plugin validates the route at startup and warns if
the provider is not yet registered.

Usage

Slash command:

/skill-eval <skill-name> [test-file]

Model-callable tool:

run_skill_eval(skill_name="<skill-name>", test_file="examples/dsh-plugin-eval.json")

test-file is optional; it defaults to examples/dsh-plugin-eval.json inside
the plugin package. Relative paths resolve against the plugin package directory.

Test-case format

A JSON array of { query, should_trigger } objects:

[
  { "query": "add a tool to the harness that persists across restarts", "should_trigger": true },
  { "query": "help me write a Python script for this CSV", "should_trigger": false }
]

category is optional and reserved for future use.

How it works

  1. Enumerate the session's model-invocable skills (ctx.skills.snapshot).
  2. Recreate the official catalog message verbatim (<system-reminder> +
    <available_skills> + normalized/truncated/escaped descriptions).
  3. For each query, ask the judge model whether the target skill should be
    triggered, forcing a one-line YES/NO answer.
  4. Compare against the expected label and aggregate metrics.

The evaluation measures the judge model's routing accuracy for the given
skill description. Swap provider/model in the config to test other judges.

Development and tests

npm run check     # syntax check for every JS file
npm test          # node:test, including official catalog fidelity and mock ctx tests
npm run smoke     # 51 pure-function smoke assertions
npm pack --dry-run  # published file list check
bash scripts/mount-smoke.sh  # real DSH mount smoke in a scratch home

The catalog fidelity fixture pins the official [email protected]
template. After a DSH upgrade, refresh the fixture from a local official
install and review the diff:

node scripts/refresh-catalog-fixture.mjs <path-to-dsh-tool-skill/lib/index.js>

Files

  • index.js — plugin entry: registers the run_skill_eval tool and the
    /skill-eval command.
  • runner.js — catalog reproduction, judge LLM call, verdict parsing.
  • catalog.js — pure functions: catalog message rendering and verdict parsing.
  • llm-helpers.js — dependency-free BlockAssembler, createUserMessage,
    deepFreeze (mirrors the official dsh-llm pattern).
  • parser.js — test-case JSON loading and validation.
  • metrics.js — confusion matrix, metrics, and markdown report formatting.
  • examples/ — default test cases.

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-skill-eval」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:renjianguojinqianfan/dsh-skill-eval

Headless(CLI)profile:

dsh plugin --profile headless add github:renjianguojinqianfan/dsh-skill-eval

实测报告

尚未 L3 验证——若已跑过,见下方失败备注。

状态:pending · 最近测试 2026-08-26
备注:验证: runtime-fail 浏览全部待验证失败 →
安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7795 个插件 →