dshbase

插件目录 / Developer / kubemd

kubemd

已验证 · 实测可装 guiyi-labs

✓ 持续维护 2 位贡献者

查看 GitHub ↗ ← 返回插件目录

0Stars
0Forks
0未关闭 issue
Shell语言
2026-08-16最近推送
跨平台平台

功能简介

证据优先的 Kubernetes 运行时故障诊断 skill:7 阶段流程(上下文→反馈环→信号→可证伪假设→dry-run 修复→案例记录→输出契约),已针对真实故障注入 kind 集群验证(5 场景:CrashLoopBackOff / OOMKilled / ImagePullBackOff / Pending / NetworkPolicy deny)。自带 cases.yaml 案例记忆库 + scripts 采集脚本 + 可 go install 的 CLI 孪生(aiops diagnose / aiops cases)。

我们的评价
可用 — 实测通过,早期项目

证据优先的 Kubernetes 运行时故障诊断 skill:7 阶段流程(上下文→反馈环→信号→可证伪假设→dry-run 修复→案例记录→输出契约),已针对真实故障注入 kind 集群验证(5 场景:CrashLoopBackOff / OOMKilled / ImagePullBackOff / Pending / NetworkPolicy deny)。自带 cases.yaml 案例记忆库 + scripts 采集脚本 + 可 go install 的 CLI 孪生(aiops diagnose / aiops cases)。 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

KubeMD — The Kubernetes Surge Doctor

**Evidence-first runtime diagnosis for Kubernetes failures — with case memory.** A DSH (DeepSeek Harness) skill that diagnoses **live, broken clusters**, not manifests. When a pod is CrashLooping, a node goes NotReady, or a Service stops answering, KubeMD runs a disciplined loop: capture context → build a red-capable feedback loop → collect signals → rank falsifiable hypotheses → fix with dry-run semantics → **record the case for instant recall next time**. > Different from KubeShark-style skills: they prevent hallucinations *while writing YAML*. KubeMD finds out *why your running workload is broken* — and never forgets a fix. https://img.shields.io/badge/powered_by-dsh-4D6BFE?style=flat-square&logo=deepseek&logoColor=white](https://github.com/deepseek-ai/deepseek-harness) https://img.shields.io/badge/License-Apache%202.0-yellow.svg](LICENSE) https://img.shields.io/badge/listed_in-awesome--deepseek--harness-4D6BFE](https://github.com/Dominic789654/awesome-deepseek-harness) https://img.shields.io/badge/listed_in-Awesome--DeepSeek--Harness-Plugins-4D6BFE](https://github.com/Zhiyuan-Fan/Awesome-DeepSeek-Harness-Plugins) https://img.shields.io/badge/listed_in-0xsline__awesome--dsh-4D6BFE](https://github.com/0xsline/awesome-deepseek-harness) https://img.shields.io/badge/listed_in-dshbase.com-4D6BFE](https://dshbase.com/plugins/kubemd/)

Install (30 seconds)

``bash git clone https://github.com/guiyi-labs/kubemd ~/.dsh/skills/dsh-k8s-diagnosis ` That's it. DSH auto-discovers skills in ~/.dsh/skills/. No restart needed. > DSH (DeepSeek Harness) — everything is a plugin. Skills are instruction bundles + scripts that agents load on demand. > 💡 **Install as a skill directory:** the repo layout we ship is exactly a DSH skill bundle. Or copy the folder and rename to dsh-k8s-diagnosis under ~/.dsh/skills/.

Demo

CLI running against a real fault-injected kind cluster (diagnose → 4 findings → case recall): ![KubeMD demo — aiops CLI on a real fault cluster](assets/demo-cli.gif) Reproduce it yourself in ~60s (needs Docker, kind, and the aiops CLI — or just the skill):
`bash

1) a real broken cluster

kind create cluster --name kubemd-demo kubectl run crash-app --image=nginx:1.25 --command -- sleep 10 # crashes on purpose kubectl rollout status deployment/crash-app 2>/dev/null || true

2) diagnose it (CLI twin of the skill, same deterministic engine)

go install github.com/guiyi-labs/aiops-platform/cmd/aiops@latest aiops diagnose --namespace default --pod crash-app --period 5 # signals → root cause

3) recall the case next time

aiops cases --query crash-loop
` Same loop the skill runs: signals first, hypotheses ranked, fix suggested dry-run.

What it does

` Symptom: "pod CrashLoopBackOff after image update to :latest" │ ├─ Phase 1 capture context (cluster, scope, recent changes) ├─ Phase 2 build feedback loop (kubectl events/logs → 10s red-capable signal) ├─ Phase 3 collect signals (events → status → --previous logs → node) ├─ Phase 4 rank 3-5 falsifiable hypotheses (predictions, not vibes) ├─ Phase 5 verify, dry-run (kubectl diff / rollout undo) ├─ Phase 6 record the case (cases.yaml → recall next time) └─ Phase 7 output contract (ROOT_CAUSE / EVIDENCE / FIX / CASE_RECORDED) `

Included

| Path | Purpose | |---|---| |
SKILL.md | The 7-phase procedure (short, token-efficient) | | references/signal-map.md | Symptom → signal → command cheatsheet | | references/playbooks/ | Deep playbooks: crashloop, oom, network, pending, node-not-ready | | scripts/collect-signals.sh | One-shot signal collection for Phase 3 | | scripts/record-case.sh | Append a resolved diagnosis to cases.yaml | | cases.yaml | Your growing case library (starts with examples; grows with your fleet) |

Case memory (the differentiator)

Every resolved diagnosis becomes a record. Next time the same symptom appears, search first:
`bash grep -i "crashloop" ~/.dsh/skills/dsh-k8s-diagnosis/cases.yaml ` A recalled past case is the fastest diagnosis: reproduction loop + remembered fix + re-verify. This is a local MVP of a broader AIOps knowledge loop — the same "distill resolved diagnoses into a searchable library" idea that powers LLM-assisted root-cause analysis at platform scale.

Also: the aiops CLI

Prefer a terminal? The same deterministic diagnosis rules ship as a
go install-able CLI: `bash go install github.com/guiyi-labs/aiops-platform/cmd/aiops@latest aiops diagnose --namespace demo --pod web-0 # rule-based root cause aiops cases --query "crashloop" # historical case recall ` No server. No database. One binary. Same engine, two doors: KubeMD (agent guidance) ↔ aiops CLI (terminal automation).

Design principles (borrowed from the best)

- **Feedback loop first** (mattpocock/diagnosing-bugs): no hypothesis before a red-capable loop exists - **Token-efficient progressive disclosure** (KubeShark): SKILL.md stays short; playbooks load on demand - **Truthfulness**: every step marks verified vs unverified; never claim what you didn't run - **Dry-run semantics**:
kubectl diff before apply, rollout undo` over live edits

Roadmap

- [x] SKILL.md + signal-map + 5 playbooks + scripts - [x] cases.yaml examples + LICENSE + branding - [ ] Verified against kind cluster (real fault injection: crashloop / oom / netpol deny) - [ ] MCP tooling for DSH diagnosis hints - [ ] Sync cases.yaml ↔ aiops-platform knowledge base (RAG)

License

Apache-2.0

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 kubemd」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:guiyi-labs/kubemd

Headless(CLI)profile:

dsh plugin --profile headless add github:guiyi-labs/kubemd

实测报告

验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7795 个插件 →