How it behaves
The plugin integrates agent-vision-toolkit into DSH. In DSH Web, pasting an image causes text-only routes to use a Vision Toolkit variant and preserve a workspace path so the image can be reused. The request includes the inspection reason, allowing answers to focus on questions such as an error location or a button colour instead of returning only a general caption.
Capabilities and outputs
The documented workflows cover intent-aware image Q&A, element grounding with original-image pixel coordinates, long-screenshot OCR, UI reconstruction, asset extraction and screenshot comparison. The toolkit can produce:
- OCR Markdown, chunks, manifests and resumable run state;
- crops, transparent PNGs, colour palettes and editable SVG traces;
- HTML screenshots, difference percentages, ranked regions, heatmaps and JSON.
Service and local processing
New installations use the built-in Gemma 4 vision service and do not require an API key. Cropping, pixel diffing, colour analysis, foreground extraction, SVG tracing and HTML screenshots run locally rather than consuming vision API requests. A real image-request test is available in Settings; checking /models alone is not sufficient to verify image support.
Limits and platform notes
The README documents free-service limits of 100/day per client, 400/day globally and 20/minute for bursts. Windows first-time isolated-runtime setup supports Microsoft Store Python. The provided excerpt does not specify additional configuration keys or standalone tool command names.
Written from the project's own documentation and kept in sync with it. Where the two disagree, the source is authoritative — read the README on GitHub