DSH Marketplace

The catalogue

Vision & Multimodal

Synced from the community registry and the dsh-plugin GitHub topic. Star counts and last-push dates come straight from GitHub.

306 plugins

99% install-verified

7.6k

A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

Installed cleanly when we ran it

npm packageTypeScript15d ago

AI review

dsh-tool-describe-image

What it is — dsh-tool-describe-image is an image description tool for text-only models. It sends local paths, URLs or attachments to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

Who it is for — When using text-only models in DeepSeek Harness on tasks involving image inputs, this plugin is suitable for installation. Users who rely on built-in vision capabilities in their models can choose not to install it, as they have native image support.

Watch out — Sandbox testing in a new profile showed successful installation and registration with no obvious pitfalls found. No issues were detected in static checks.

The verdict — I would install it because it allows text-only models to handle image inputs in the session.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source

modlens

liustack

4.1k

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

Installed cleanly when we ran it

npm packageTypeScript3d ago

AI review

modlens

What it is — ModLens bridges text-only models with vision, allowing pasted images to obtain structured JSON evidence including OCR, layout, and semantics.

Who it is for — When using DeepSeek Harness with text-only models like DeepSeek-V4-Flash to analyze pasted images, it is suitable. For users of native multimodal models, this plugin is not necessary because native models retain their built-in image pasting capability.

Watch out — Sandbox testing passed, with successful installation in a fresh profile and registration with Harness. No obvious issues found.

The verdict — I would install it because it outputs structured visual evidence without saving the image to a file path first.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Details

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

Installed cleanly when we ran it

GitHub sourcePython1mo ago

AI review

agent-vision-toolkit

What it is — agent-vision-toolkit provides a vision toolkit and skills for text-only models.

Who it is for — When text models in DeepSeek Harness have image tool attempts blocked by the system, this plugin is suitable to install. Users not needing it are those whose system already allows native image tools.

Watch out — From source installation, manually allow the build script first. The installation will execute shell commands. Sandbox test passed: installed in a new profile and registered into the profile. No obvious issues found.

The verdict — I would install it because it has been verified in real sessions with DeepSeek and Claude Code.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

Installed cleanly when we ran it

npm packageJavaScriptyesterday

AI review

dsh-vision-router

What it is — dsh-vision-router provides free vision capabilities for text-only agents in DeepSeek Harness with pixel tools.

Who it is for — When your agent needs vision grounding or vision crop operations, this plugin is suitable. If your agent supports multimodal input natively instead of being text-only, this plugin is not needed.

Watch out — No obvious issues found. Sandbox test passed in new profile: the plugin was installed and registered by harness. The plugin executes shell commands but no issues occurred.

The verdict — I would install it because it supports continuous multi-step image processing until the task is complete.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Details

Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.

Installed cleanly when we ran it

npm packageTypeScript4d ago

AI review

dsh-vision-toolkit

What it is — DSH Vision Toolkit enables text-only models to process images for Q&A, OCR, UI reproduction and related visual tasks.

Who it is for — When using text-only models in DeepSeek Harness web for screenshot analysis, element grounding or multi-image comparison, this plugin is suitable. Users with native vision models do not need this to enhance text models.

Watch out — No obvious issues found in sandbox testing. Images are processed by the author-hosted free service with 100 per machine per day by default, configurable to your own provider.

The verdict — I would install it because it integrates vision tools natively into DeepSeek Harness profiles and sessions.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Details

dsh-browser

omdsh-dev

705

Chrome sidebar extension that lets DeepSeek Harness operate your browser directly, no vision capabilities required. 一款 Chrome 侧边栏扩展程序,可让 DeepSeek Harness 直接操控您的浏览器,无需视觉能力。

Installed cleanly when we ran it

GitHub sourceTypeScript10d ago

AI review

dsh-browser

What it is — This is a Chrome sidebar extension that lets DeepSeek Harness directly control your browser without vision capabilities.

Who it is for — Users who need to navigate browser tabs, fill forms and manage tabs with deepseek-v4-flash are suitable for this. Users who do not require browser interaction do not need to install it.

Watch out — It requires executing shell commands during install and providing an API key. Sandbox testing passed in a fresh profile. It showed lower latency than Playwright in benchmarks.

The verdict — I would install it because it operates on real browser sessions instead of headless simulations, preserving logins and cookies.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
2Source
396

DeepWatch, powered by Watch Skill: a local-first perception and verification stack for AI agents. Understand video, audio and screen activity, keep timestamped evidence, and independently verify what a tool or workflow actually did. MCP, CLI, REST, Web and DeepSeek Harness.

Installed cleanly when we ran it

GitHub sourcePython16d ago

AI review

watch-skill

What it is — watch-skill is a local-first perception and verification plugin for AI agents that enables understanding of video, audio and screen activity with timestamped evidence and independent verification of agent actions.

Who it is for — When handling browser flows or generated videos in DeepSeek Harness that require verifying agent effects, watch-skill is suitable. For users not involving dynamic screen or video input, watch-skill is not necessary.

Watch out — Sandbox testing in a fresh profile succeeded with registration by DeepSeek Harness. The plugin will execute shell commands. No obvious issues were found.

The verdict — I would install it because it can drive browser and prove the effect of each action, which is necessary in agent workflows.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

Photo-to-editorial Skill with Original (Codex) and V3 Adaptive editions. Scene-aware layouts, creative controls, Strict Fidelity composition, structured QA, and a validated DeepSeek Harness capability path.

Installed cleanly when we ran it

GitHub sourcePython24d ago

AI review

photo-abstract-editorial

What it is — Converts a photograph into an editorial composition with original photo, abstract memory panel, and poetic English title.

Who it is for — When using DeepSeek Harness for photo editing tasks requiring scene-aware layouts and creative controls, this plugin is suitable. If you want the fixed Codex-only workflow without adaptive features, you do not need this plugin.

Watch out — Sandbox test passed: installed in a new profile and registered by Harness. No obvious issues found.

The verdict — I would install it because it includes a validated DeepSeek Harness capability path for photo-to-editorial tasks.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source
90

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

Installed cleanly when we ran it

npm packageTypeScript5d ago

AI review

dsh-imagegen

What it is — Integrated AI image generation plugin for DeepSeek Harness Web GUI, supporting text-to-image and image-to-image.

Who it is for — Users requiring image generation in DeepSeek Harness Web GUI are suitable to install this plugin. Users not using the DSH host process for API proxying need not install this plugin.

Watch out — Sandbox test passed in a fresh profile after installation and registration. API URL and key must be configured in the settings card. No obvious issues found.

The verdict — I would install it because it supports cross-device sharing of image history records.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

dsh-comfyui

fandc520

80

Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.

Installed cleanly when we ran it

npm packageTypeScript17d ago

AI review

dsh-comfyui

What it is — The dsh-comfyui plugin enables DeepSeek Harness agents to drive local or remote ComfyUI servers for generating images and videos, with workflow library and skill package management.

Who it is for — If you are using DeepSeek Harness for image or video generation tasks that require an agent to directly execute ComfyUI workflows, this plugin fits your scenario. If your tasks do not involve a ComfyUI backend, you can skip installing it.

Watch out — No additional setup is required. It was successfully tested in a fresh profile and registered by Harness. No obvious issues were found.

The verdict — I would install it because the skill package feature lets agents record and reuse experiences across sessions, reducing repeated configuration for complex workflows.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source

Community Docker and Kubernetes packaging for DeepSeek Harness (@deepseek-ai/dsh), with a hardened image, Compose stack, Helm chart, Web UI, and headless CLI.

Installed cleanly when we ran it

GitHub sourceShell22d ago

AI review

deepseek-harness-docker

What it is — DeepSeek Harness Docker packaging provides hardened image, Compose stack, Helm chart, Web UI and headless CLI for containerized deployment using official npm release in Node.js 24.

Who it is for — It is suitable when running DeepSeek Harness Web UI in Docker needs persistent profile, sessions and workspace. Users requiring only headless mode or without container boundary needs are not suitable.

Watch out — Manual build script approval is required from source installation. Shell commands are executed. Sandbox testing passed in a new profile and registered by harness.

The verdict — I would install it because it handles official release container boundaries like PTY and host binding without forking upstream.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

dsh-design-qa

sunxin-ai

44

Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.

Installed cleanly when we ran it

npm packageJavaScript23d ago

AI review

dsh-design-qa

What it is — The dsh-design-qa plugin enables text-only models in DeepSeek Harness to read images using the deepseek_vision tool.

Who it is for — If you need to have the model judge whether an implementation matches its mock in DeepSeek Harness, this plugin is suitable. If you do not have an OpenAI-compatible vision route, you cannot use this plugin.

Watch out — Sandbox testing installed and registered in a new profile. The plugin will execute shell commands.

The verdict — Provided that the model supports tool calling, I would install it because it includes the benchmark and questioning discipline needed for the judgment.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source
36

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

Installed cleanly when we ran it

npm packageJavaScript12d ago

AI review

picturereader

What it is — picturereader gives text-only models the ability to read images by converting them into text grids via downscaling and color reduction.

Who it is for — When using text-only models such as DeepSeek to analyze images, picturereader is suitable. Users performing purely text tasks do not need it because they require no visual information input.

Watch out — Sandbox testing passed, with successful installation and registration into the profile by DeepSeek Harness. ZCode version has additional overhead from MCP process communication. No other obvious issues found.

The verdict — I would install it because it provides an image-reading methodology skill verified through extensive real image iterations, enabling text models to autonomously verify image content.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

DSH-vison

hisence999

35

DSH-vison is a DeepSeek Harness plugin that adds image understanding for text-only models by converting images into descriptions from a configured vision-capable model.

Installed cleanly when we ran it

GitHub sourceJavaScript1mo ago

AI review

DSH-vison

What it is — DSH-vison provides image understanding for text-only models by converting images into descriptions from a configured vision model.

Who it is for — When you use a text-only model to process user messages containing images, you need image understanding capability. If you use a multimodal model, you can skip installing this plugin.

Watch out — Sandbox test passed: it was installed in a fresh profile and registered by harness. No obvious pitfalls found.

The verdict — I would install it because the sandbox test passed and it enables global image handling for text models.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
LINUX DO1Details

dsh-vision

william-jin-cmu

30

dsh 插件:给纯文本 DeepSeek 加视觉——view_image 工具桥接任意 OpenAI 兼容 VLM(默认智谱免费档,实测 4 厂商 10 模型)

Installed cleanly when we ran it

GitHub sourceTypeScript1mo ago

AI review

dsh-vision

What it is — dsh-vision plugin registers view_image tool to bridge any OpenAI compatible VLM for text-only DeepSeek.

Who it is for — When using DeepSeek-V4-Flash to describe desktop images in DSH web, install dsh-vision. Users without need for visual analysis can skip this plugin.

Watch out — Sandbox test passed: installed in new profile and registered by harness. Source install requires manual linking of host dependencies. No obvious pitfalls found.

The verdict — I would install it because it enables DeepSeek-V4-Flash image description, but only after configuring the backend apiKey.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

dsh-directorx

LaplaceYoung

28

DirectorX as a DeepSeek Harness plugin: AI video/image/audio skills, knowledge corpus, and configurable vision/image/video/audio model tools.

Installed cleanly when we ran it

GitHub sourceJavaScript27d ago

AI review

dsh-directorx

What it is — DirectorX is a DeepSeek Harness plugin integrating AI video/image/audio skills, knowledge corpus, and configurable vision/image/video/audio model tools.

Who it is for — When making multi-shot video projects in DeepSeek Harness that require storyboard planning and editing, this plugin provides the necessary tools. For users who only need simple one-time generation without canvas or editing workflows, this plugin is not required.

Watch out — Sandbox test passed: Installed in a new profile and registered into the profile. No obvious pitfalls found.

The verdict — I would install it because it enables local ffmpeg-based editing that can be iterated without re-generating, but only if confirmation is done before placing on canvas.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

dsh-clawrouter

BlockRunAI

21

A safety gate for DeepSeek Harness: a stronger model reviews dangerous tool calls before they run. Plus vision and 67 models from one wallet, paid per request over x402.

Installed cleanly when we ran it

npm packageTypeScript25d ago

AI review

dsh-clawrouter

What it is — DeepSeek Harness plugin reviews dangerous tool calls by stronger model before execution.

Who it is for — When DeepSeek Harness agent proposes rm -rf ~ for high risk commands, use anthropic/claude-opus-5 to review the decision. Users who do not want to pay model requests via x402 are not suitable for this plugin.

Watch out — Tested successfully in new profile, registered into DeepSeek Harness, passed sandbox test. No obvious issues found.

The verdict — I would not install it, because it only reviews dangerous commands and does not review ordinary work.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

Installed cleanly when we ran it

GitHub sourcePython29d ago

AI review

dsh-media-skills

What it is — Provides free image reading and generation skills for DeepSeek Harness, supporting image pasting into text-only sessions.

Who it is for — When you lack image capabilities in text-only sessions, this plugin is suitable for installation. If you only want to use paid models, you can choose not to install this plugin.

Watch out — Sandbox test passed. Installed successfully in a new profile and registered by harness. No obvious issues found.

The verdict — I will install it because it provides free image processing and structured evidence output.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source

dsh-media-skills

akqwpeter-prog

18

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

Installed cleanly when we ran it

GitHub sourcePython10d ago

AI review

dsh-media-skills

What it is — Provides free image reading and generation for DeepSeek Harness.

Who it is for — When working with text-only models on image analysis or content creation tasks, this plugin adds paste-image functionality and auto-routes vision models. If your tasks are purely text-based with no image input needs, no installation is required.

Watch out — Installation requires manual approval of the build script from source. Sandbox test passed in a new profile with successful registration. No obvious issues found.

The verdict — I would install it because it bridges free vision and generation for text models, but only if you can obtain the required API keys from Zhipu and SiliconFlow.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source

Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.

No one-line install — this plugin lives inside a larger repository and publishes no npm package.

GitHub source1mo ago

AI review

dsh-tool-describe-image

What it is — Provides image understanding for text-only models through a vision-language model, exposed as the describe_image tool.

Who it is for — If you are using text-only models in DeepSeek Harness conversations and need to analyze images. If you prefer to use vision-language models directly for image processing, this plugin serves as a supplement.

Watch out — No obvious issues found. Sandbox testing has not been performed.

The verdict — I would install it if image understanding is required, as it extends text models with visual capabilities.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source
16

Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.

Installed cleanly when we ran it

npm packageTypeScript28d ago

AI review

dsh-visual-plugin

What it is — This plugin enables text-only models to process images by forwarding uploaded images to an OpenAI-compatible vision model and displaying descriptions in the right panel of the Web UI.

Who it is for — It suits users who employ text models for image analysis or visual description tasks. Users who stick to pure text interactions without image uploads do not need it.

Watch out — Sandbox testing passed with successful installation and registration in a fresh profile. Configuration of the vision endpoint, model, and key completed without observed issues.

The verdict — I would install it because it provides a bridge for text models to handle image inputs, but only if a compatible vision API backend is available.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source
14

DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.

Installed cleanly when we ran it

npm packageJavaScript1mo ago

AI review

dsh-vision-proxy

What it is — This plugin registers the deepseek-vision route to transcribe GUI-attached images to text before sending them to text-only DeepSeek models.

Who it is for — Users who attach images in the DeepSeek Harness GUI can install this plugin to enable image recognition. Users of official vision models do not need to install it because the official already supports multimodal.

Watch out — It passes when installed in a fresh profile and registered by harness. No obvious issues found.

The verdict — I would install it for text-only DeepSeek models because it bridges image recognition, but not for official vision model users who have native support.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

dsh-docs

Sqhao-O

14

Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR. | DeepSeek Harness 全本地文档智能插件,支持 PDF、Office、图片与离线 OCR

Installed cleanly when we ran it

npm packageTypeScript1mo ago

AI review

dsh-docs

What it is — DeepSeek Harness plugin providing local document intelligence with offline OCR for PDF, Office files, and images.

Who it is for — If you need to process local PDF or Office files in a DeepSeek Harness project, install this plugin. If you require remote HTTP service support, skip this plugin.

Watch out — Installation executes shell commands. Sandbox test passed in a new profile with successful registration. No obvious pitfalls found.

The verdict — I would install it if offline OCR is needed on Windows x64, as it passed sandbox verification.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

Installed cleanly when we ran it

GitHub sourceJavaScript21d ago

AI review

dsh-vision-opencode

What it is — This DSH plugin automatically converts images to text via vision models for pure-text LLMs.

Who it is for — Users of pure-text models like DeepSeek who need to handle images in chats are suitable to install this plugin. Users of native multi-modal models can use their original image capabilities without this plugin.

Watch out — Sandbox test passed: installed in a new profile and registered into the profile by harness. The installation script will execute shell commands and uninstall requires backing up image sessions. No obvious issues found.

The verdict — I would install it because it allows pure-text models to handle images without switching the main model and the sandbox test passed.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source