How it behaves
Image handling is performed as an ordinary tool-calling turn. A multi-step sequence can call vision_ground, then vision_crop, vision_describe or vision_pixel_diff, apply a change, and inspect a new screenshot. The uploaded image remains an image in the session UI; the internal rewrite that directs the model to vision tools is not added to the session log. Text turns continue using the normal model, cost and context path.
Tools
The README describes a pixel-tool set covering visual Q&A, grounding, cropping, pixel diffs, colour sampling, OCR, SVG tracing, cutout and screenshots. It explicitly shows these tool names:
| Tool |
Documented use |
vision_ground |
Locate an element in an image |
vision_crop |
Crop an image region |
vision_describe |
Produce an image description |
vision_pixel_diff |
Compare image pixels |
Routing and fallback
User-provided vision models are tried first. If they are unavailable, the built-in OVHcloud anonymous fallback chain provides five models without an account or key. The README states a limit of 2 requests/minute per IP per model, with roughly 10 RPM in theory across independent buckets. Answers are cached by image content.
Requirements and local access
The package requires Node.js >=22 and does not require Python. Its processing pipeline uses sharp, potrace, tesseract and system Chrome; system Chrome is used for HTML screenshots. The README identifies dsh.bundle.patch as the composition patch applied by the plugin. A bundled DSH Web profile is also referenced by cordis.patch.yml.
Routing option
Taking over the official DeepSeek route is available through an optional stealth-mode setting and is off by default. The README does not specify a configuration key or command-line flag for that setting.
Written from the project's own documentation and kept in sync with it. Where the two disagree, the source is authoritative — read the README on GitHub