Plugin directory / Developer / dsh-skill-eval
dsh-skill-eval
Unverified renjianguojinqianfan
What it does
DSH 插件:用 LLM judge 评测技能 description 的触发准确率(欠触发/过触发)
Unverified — not yet verified
DSH 插件:用 LLM judge 评测技能 description 的触发准确率(欠触发/过触发) Not yet verified — install and test it yourself.
“Unverified” means our automated CI has not yet installed this plugin. Feature descriptions and version compatibility are the author’s claims. This is not a security audit and not an endorsement of third-party code.
README
dsh-skill-eval
Skill-trigger evaluation plugin for DeepSeek Harness (DSH).
An LLM judge recreates the exact DSH skill catalog prompt and decides, for each
test query, whether the target skill should be triggered. The plugin reports
accuracy, precision, recall, false-positive/negative rates, and a confusion
matrix — a reproducible measure of how reliably a skill description routes
matching queries (and how often it over- or under-triggers).
Install
From the repo root:
dsh plugin --profile web add ./dsh-skill-eval
Then configure the judge model route in your profile/overlay cordis.patch.yml:
- id: skill-eval
config:
provider: <provider-id>
model: <model-name>
The provider must be registered in the DSH LLM runtime (the same one your
profile uses for chat). The plugin validates the route at startup and warns if
the provider is not yet registered.
Usage
Slash command:
/skill-eval <skill-name> [test-file]
Model-callable tool:
run_skill_eval(skill_name="<skill-name>", test_file="examples/dsh-plugin-eval.json")
test-file is optional; it defaults to examples/dsh-plugin-eval.json inside
the plugin package. Relative paths resolve against the plugin package directory.
Test-case format
A JSON array of { query, should_trigger } objects:
[
{ "query": "add a tool to the harness that persists across restarts", "should_trigger": true },
{ "query": "help me write a Python script for this CSV", "should_trigger": false }
]
category is optional and reserved for future use.
How it works
- Enumerate the session's model-invocable skills (
ctx.skills.snapshot). - Recreate the official catalog message verbatim (
<system-reminder>+<available_skills>+ normalized/truncated/escaped descriptions). - For each query, ask the judge model whether the target skill should be
triggered, forcing a one-lineYES/NOanswer. - Compare against the expected label and aggregate metrics.
The evaluation measures the judge model's routing accuracy for the given
skill description. Swap provider/model in the config to test other judges.
Development and tests
npm run check # syntax check for every JS file
npm test # node:test, including official catalog fidelity and mock ctx tests
npm run smoke # 51 pure-function smoke assertions
npm pack --dry-run # published file list check
bash scripts/mount-smoke.sh # real DSH mount smoke in a scratch home
The catalog fidelity fixture pins the official [email protected]
template. After a DSH upgrade, refresh the fixture from a local official
install and review the diff:
node scripts/refresh-catalog-fixture.mjs <path-to-dsh-tool-skill/lib/index.js>
Files
index.js— plugin entry: registers therun_skill_evaltool and the/skill-evalcommand.runner.js— catalog reproduction, judge LLM call, verdict parsing.catalog.js— pure functions: catalog message rendering and verdict parsing.llm-helpers.js— dependency-freeBlockAssembler,createUserMessage,deepFreeze(mirrors the official dsh-llm pattern).parser.js— test-case JSON loading and validation.metrics.js— confusion matrix, metrics, and markdown report formatting.examples/— default test cases.
Install
Install the catalog once, then DeepSeek Harness can find and install any plugin from this site automatically:
dsh plugin add dshbase-catalog Then say "install dsh-skill-eval for me" — your agent finds it in the directory and installs it. Docs: dshbase-catalog · verified packs.
This plugin is GitHub source (not published to npm) — install it straight from the repo:
Web profile:
dsh plugin --profile web add github:renjianguojinqianfan/dsh-skill-eval Headless (CLI) profile:
dsh plugin --profile headless add github:renjianguojinqianfan/dsh-skill-eval Test report
Not yet L3-verified — see failure note below if we already ran it.