agent-vision-toolkit — tool profile and history

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

Active

Canonical source ↗

Language
Python
License
MIT
Created
Last activity
Topics
  • agent
  • agent-skills
  • claude-code
  • codex
  • computer-use
  • deepseek
  • dsh-plugin
  • glm
  • harness-engineering
  • multimodal
  • opencode
  • text-only-llm
  • vision
  • vision-language-model

Install git clone https://github.com/Anionex/agent-vision-toolkit ~/.claude/skills/agent-vision-toolkit

Current measurement

1 146GitHub stars+28 (+2.5 %)
measured growth

Momentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.

Reading history

90 days · 90 maximum readings
DateValueMetric
1 146GitHub stars
1 133GitHub stars
1 125GitHub stars
1 119GitHub stars
1 118GitHub stars
1 114GitHub stars
1 110GitHub stars
1 107GitHub stars
1 102GitHub stars
1 098GitHub stars
1 095GitHub stars
1 095GitHub stars
1 084GitHub stars
1 072GitHub stars
1 049GitHub stars
1 005GitHub stars
989GitHub stars
911GitHub stars
845GitHub stars
830GitHub stars
409GitHub stars
397GitHub stars
392GitHub stars
384GitHub stars

Classifications

Domains
Type
Agent skills
Use cases