dshbase

插件目录 / Developer / dsh-tool-accurate-vision

dsh-tool-accurate-vision

未验证 imkingjh999

✓ 持续维护 基于 4 个官方 DSH 包

查看 GitHub ↗ ← 返回插件目录

0Stars
0Forks
0未关闭 issue
语言
2026-08-23最近推送
跨平台平台

功能简介

Model-facing accurate_vision tool for DeepSeek Harness: precise spatial reasoning via any OpenAI-compatible vision model (0-1000 bbox primitives + annotated SVG)

我们的评价
未验证 — 尚未实测

Model-facing accurate_vision tool for DeepSeek Harness: precise spatial reasoning via any OpenAI-compatible vision model (0-1000 bbox primitives + annotated SVG) 尚未验证——请自行安装测试。

「未验证」表示我们的自动化 CI 尚未安装过该插件。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

你是插件作者? 想拿到「已验证」标签——提交你自己的验证证据(截图、日志或短视频),我们审核通过后即改为「已验证」。

提交验证证据 ↗

README

dsh-tool-accurate-vision

Awesome DSH Plugin

Model-facing accurate_vision tool for
DeepSeek Harness:
precise spatial reasoning over an image file via an OpenAI-compatible vision model.
Ported from pi-accurate-vision.

A vision model reads the image and returns a structured note plus
bounding-box primitives normalised to 0–1000; this tool formats them as a
<vision-context> block the next model turn reads — giving a text-only agent
exact object positions, layout, and OCR without losing spatial fidelity.

English | 中文

Install

dsh plugin --profile web add dsh-tool-accurate-vision

Or from source:

dsh plugin --profile web add github:your-username/dsh-tool-accurate-vision

Set the vision API key (separate from DEEPSEEK_API_KEY):

export VISION_API_KEY=sk-...

How it works

image file ──► base64 data URL ──► vision chat/completions ──► JSON note + primitives
                                                                      │
                                                          <vision-context> XML ──► next model turn

The pure vision core (src/bridge.ts) is provider-agnostic:
any OpenAI-compatible multimodal chat/completions endpoint works.
The Cordis host (src/index.ts) owns config, credential
resolution, and the registered tool.

Every call also writes a self-contained SVG — the original image with every
bounding box and label drawn on it — returned as the annotatedImage path,
so the boxes can be eyeballed instead of trusted blind
(set annotate: false to skip it).

Case study: rigorous distance computation

Ask an image question with a checkable answer — in this hand-drawn physicists
network, which node sits physically closest to 居里夫人 (Marie Curie), ignoring
the connecting lines?
— and the gap between plain vision and this tool becomes
measurable. The test image is the aged network diagram below:

The test image: a hand-drawn physicists network

  1. Asking a multimodal model directly yields a visual impression, not a
    measurement: "郎之万, at the lower left, looks closest" — nothing to verify,
    and as it turns out, wrong.

    A plain VLM answers by intuition

  2. Vision text without structured primitives can be worse than no numbers
    at all: the model invents plausible-looking coordinates in prose, then
    contradicts itself — a claimed ~15-unit gap while its own two boxes imply
    59 — and returns the same wrong answer.

    Unstructured output hallucinates coordinates

  3. With this tool's normalised primitives, every node carries a checkable
    0–1000 bounding box, so the agent computes real edge-to-edge distances in
    code: 皮卡尔德 25.96 vs 郎之万 58.00. The correct answer — 皮卡尔德
    (Piccard) — arrives with the numbers that prove it.

    Structured primitives enable exact distances

That is the core advantage: bounding-box primitives turn visual impressions
into geometry. Positions, distances, and layout become facts a text-only agent
can compute and verify, not guesses it has to trust. For distance questions the
canonical edge-to-edge computation pairs the facing edges per axis
(dx = max(a.x1 - b.x2, b.x1 - a.x2, 0), same for y, then hypot); the
tested helper bboxEdgeDistance(a, b) ships with this package so downstream
agents never pair the wrong edges.

Configuration

Override in your profile's cordis.patch.yml:

- id: tool-accurate-vision
  config:
    model: gpt-4o              # any OpenAI-compatible multimodal model
    baseURL: https://api.openai.com/v1
    apiKeyEnv: VISION_API_KEY  # credential reference
    primitives: true           # request bounding-box primitives
    annotate: true             # also write an SVG with boxes drawn on the image
    maxTokens: 8192
    timeoutSecs: 120
    temperature: 0
    disableThinking: true     # skip the reasoning phase (MiniMax): faster & steadier

Origin

Faithful port of pi-accurate-vision (which itself extracted DeepSeek-TUI's
crates/tui/src/vision/bridge.rs). The parsing, prompt, and formatting logic
is preserved verbatim; only the host integration targets the Cordis ctx.tools
registry with schemastery config and the credentials seam.

License

MIT

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-tool-accurate-vision」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:imkingjh999/dsh-tool-accurate-vision

Headless(CLI)profile:

dsh plugin --profile headless add github:imkingjh999/dsh-tool-accurate-vision

实测报告

尚未 L3 验证——若已跑过,见下方失败备注。

状态:pending · 最近测试 2026-08-27
备注:验证: install-fail 浏览全部待验证失败 →
安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7795 个插件 →