DSH Marketplace

dsh-vision-router

ysr666/dsh-vision-router

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

1,12951JavaScriptMITSource

Install

Add dsh-vision-router to DeepSeek Harness

via npm

Resolves a published tarball rather than cloning the repository, and installs without any extra setup. Swap `web` for your profile name if you run another one.

via GitHub · npm package

Installing from GitHub runs the project's build script, which pnpm blocks until you allowlist it — run the command once and pnpm prints the exact key to add under `allowBuilds` in ~/.dsh/profiles/web/pnpm-workspace.yaml.

What happened when we ran it

Installed cleanly when we ran it

Every command here is run in a throwaway container against a clean profile, and the result is whatever the harness recorded — not a guess from the source. Last run 11d ago.

Show it in your README

install verified — dshmarketplace

For maintainers: the badge serves this listing's latest sandbox verdict, so a re-run updates it on its own — and it links readers to the full result here.

Due diligence

Before you install dsh-vision-router

  • Source of record: ysr666/dsh-vision-router — present in the community registry that DSH's own plugin market installs from.
  • Licensed under MIT.
  • Detected: terminal surface. Read the source before granting these.
  • A listing here is not a security review. Plugins run with your agent's permissions.

The AI take

What it is — dsh-vision-router provides free vision capabilities for text-only agents in DeepSeek Harness with pixel tools.

Who it is for — When your agent needs vision grounding or vision crop operations, this plugin is suitable. If your agent supports multimodal input natively instead of being text-only, this plugin is not needed.

Watch out — No obvious issues found. Sandbox test passed in new profile: the plugin was installed and registered by harness. The plugin executes shell commands but no issues occurred.

The verdict — I would install it because it supports continuous multi-step image processing until the task is complete.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.

What dsh-vision-router does

Image tiles pass through model gates and a lens, producing crops, pixel comparisons, traced contours and screenshots.

dsh-vision-router is a DeepSeek Harness plugin that routes image turns to vision models while leaving DeepSeek responsible for reasoning. It keeps the original image available to the vision side instead of first converting it into a text description. The plugin uses agent/request waterfalls and the bundled dsh.bundle.patch to add a composition row, an admission wrapper and attachment limits. dsh plugin add applies that wiring automatically; taking over the official DeepSeek route is an optional stealth-mode setting, disabled by default.

Vision requests use user-provided vision models first, followed by a built-in keyless OVHcloud fallback chain. The README describes pixel tools for Q&A, grounding, cropping, pixel comparison, colour extraction, OCR, SVG tracing, cutout and screenshots. Named tools include vision_ground, vision_crop, vision_describe and vision_pixel_diff. Image answers are cached by image content, while uploaded images remain visible in the conversation UI and the routing rewrite stays inside the model call.

This is intended for text-only agents that need iterative image inspection without switching their main model. It is a poor fit where the existing harness already provides the required vision route, or where external anonymous endpoints are unsuitable: the fallback is rate-limited to 2 requests per minute per IP per model. The plugin also reaches the terminal surface because its local pipeline uses sharp, potrace, tesseract and system Chrome for image processing and HTML screenshots.

dsh-vision-router documentation

How it behaves

Image handling is performed as an ordinary tool-calling turn. A multi-step sequence can call vision_ground, then vision_crop, vision_describe or vision_pixel_diff, apply a change, and inspect a new screenshot. The uploaded image remains an image in the session UI; the internal rewrite that directs the model to vision tools is not added to the session log. Text turns continue using the normal model, cost and context path.

Tools

The README describes a pixel-tool set covering visual Q&A, grounding, cropping, pixel diffs, colour sampling, OCR, SVG tracing, cutout and screenshots. It explicitly shows these tool names:

Tool Documented use
vision_ground Locate an element in an image
vision_crop Crop an image region
vision_describe Produce an image description
vision_pixel_diff Compare image pixels

Routing and fallback

User-provided vision models are tried first. If they are unavailable, the built-in OVHcloud anonymous fallback chain provides five models without an account or key. The README states a limit of 2 requests/minute per IP per model, with roughly 10 RPM in theory across independent buckets. Answers are cached by image content.

Requirements and local access

The package requires Node.js >=22 and does not require Python. Its processing pipeline uses sharp, potrace, tesseract and system Chrome; system Chrome is used for HTML screenshots. The README identifies dsh.bundle.patch as the composition patch applied by the plugin. A bundled DSH Web profile is also referenced by cordis.patch.yml.

Routing option

Taking over the official DeepSeek route is available through an optional stealth-mode setting and is off by default. The README does not specify a configuration key or command-line flag for that setting.

Written from the project's own documentation and kept in sync with it. Where the two disagree, the source is authoritative — read the README on GitHub

Same category

Alternatives to dsh-vision-router

7.6k

A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

Installed cleanly when we ran it

npm packageTypeScript15d ago

AI review

dsh-tool-describe-image

What it is — dsh-tool-describe-image is an image description tool for text-only models. It sends local paths, URLs or attachments to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

Who it is for — When using text-only models in DeepSeek Harness on tasks involving image inputs, this plugin is suitable for installation. Users who rely on built-in vision capabilities in their models can choose not to install it, as they have native image support.

Watch out — Sandbox testing in a new profile showed successful installation and registration with no obvious pitfalls found. No issues were detected in static checks.

The verdict — I would install it because it allows text-only models to handle image inputs in the session.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Source

modlens

liustack

4.1k

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

Installed cleanly when we ran it

npm packageTypeScript3d ago

AI review

modlens

What it is — ModLens bridges text-only models with vision, allowing pasted images to obtain structured JSON evidence including OCR, layout, and semantics.

Who it is for — When using DeepSeek Harness with text-only models like DeepSeek-V4-Flash to analyze pasted images, it is suitable. For users of native multimodal models, this plugin is not necessary because native models retain their built-in image pasting capability.

Watch out — Sandbox testing passed, with successful installation in a fresh profile and registration with Harness. No obvious issues found.

The verdict — I would install it because it outputs structured visual evidence without saving the image to a file path first.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Details

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

Installed cleanly when we ran it

GitHub sourcePython1mo ago

AI review

agent-vision-toolkit

What it is — agent-vision-toolkit provides a vision toolkit and skills for text-only models.

Who it is for — When text models in DeepSeek Harness have image tool attempts blocked by the system, this plugin is suitable to install. Users not needing it are those whose system already allows native image tools.

Watch out — From source installation, manually allow the build script first. The installation will execute shell commands. Sandbox test passed: installed in a new profile and registered into the profile. No obvious issues found.

The verdict — I would install it because it has been verified in real sessions with DeepSeek and Claude Code.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source

Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.

Installed cleanly when we ran it

npm packageTypeScript4d ago

AI review

dsh-vision-toolkit

What it is — DSH Vision Toolkit enables text-only models to process images for Q&A, OCR, UI reproduction and related visual tasks.

Who it is for — When using text-only models in DeepSeek Harness web for screenshot analysis, element grounding or multi-image comparison, this plugin is suitable. Users with native vision models do not need this to enhance text models.

Watch out — No obvious issues found in sandbox testing. Images are processed by the author-hosted free service with 100 per machine per day by default, configurable to your own provider.

The verdict — I would install it because it integrates vision tools natively into DeepSeek Harness profiles and sessions.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
Details

dsh-browser

omdsh-dev

705

Chrome sidebar extension that lets DeepSeek Harness operate your browser directly, no vision capabilities required. 一款 Chrome 侧边栏扩展程序,可让 DeepSeek Harness 直接操控您的浏览器,无需视觉能力。

Installed cleanly when we ran it

GitHub sourceTypeScript11d ago

AI review

dsh-browser

What it is — This is a Chrome sidebar extension that lets DeepSeek Harness directly control your browser without vision capabilities.

Who it is for — Users who need to navigate browser tabs, fill forms and manage tabs with deepseek-v4-flash are suitable for this. Users who do not require browser interaction do not need to install it.

Watch out — It requires executing shell commands during install and providing an API key. Sandbox testing passed in a fresh profile. It showed lower latency than Playwright in benchmarks.

The verdict — I would install it because it operates on real browser sessions instead of headless simulations, preserving logins and cookies.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
2Source
396

DeepWatch, powered by Watch Skill: a local-first perception and verification stack for AI agents. Understand video, audio and screen activity, keep timestamped evidence, and independently verify what a tool or workflow actually did. MCP, CLI, REST, Web and DeepSeek Harness.

Installed cleanly when we ran it

GitHub sourcePython16d ago

AI review

watch-skill

What it is — watch-skill is a local-first perception and verification plugin for AI agents that enables understanding of video, audio and screen activity with timestamped evidence and independent verification of agent actions.

Who it is for — When handling browser flows or generated videos in DeepSeek Harness that require verifying agent effects, watch-skill is suitable. For users not involving dynamic screen or video input, watch-skill is not necessary.

Watch out — Sandbox testing in a fresh profile succeeded with registration by DeepSeek Harness. The plugin will execute shell commands. No obvious issues were found.

The verdict — I would install it because it can drive browser and prove the effect of each action, which is necessary in agent workflows.

Generated by grok-4.6, and a starting point rather than a verdict. Where it says a plugin installs or does not, that is from a real run in a clean profile — everything else is read off the repository. Trust the source over this.Read the source
1Source