ocr
OCR skill for extracting text from images and PDFs. Use when you need to read text from screenshots, photos, scanned documents, or any image file. Supports Chinese, English, and 100+ languages.
How do I install this agent skill?
npx skills add https://github.com/mr-shaper/opencode-skill-hybrid-ocr --skill ocrIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The ocr skill is a privacy-focused local tool for extracting text from images and PDFs using DeepSeek-OCR or PaddleOCR. It operates entirely on the local machine by communicating with a local Ollama instance. The code follows secure development practices, with no evidence of data exfiltration, malicious command execution, or obfuscation.
- Socketwarn
1 alert: gptSecurity
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
OCR Skill
Usage
To extract text from an image or PDF, run:
python3 "/Users/mrshaper/Library/Application Support/com.differentai.openwork/workspaces/starter/.opencode/skills/paddle-ocr/scripts/ocr.py" "/path/to/image.png"
Options
| Option | Description |
|---|---|
--prompt "text" | Custom prompt (e.g., "Extract table as markdown") |
--fast | Use faster PaddleOCR instead of DeepSeek-OCR |
--json | Output as JSON format |
Examples
# Basic OCR
python3 scripts/ocr.py image.png
# Extract table as markdown
python3 scripts/ocr.py table.png --prompt "Extract this table as markdown"
# Fast mode
python3 scripts/ocr.py image.png --fast
# PDF OCR
python3 scripts/ocr.py document.pdf
Supported Formats
Images: PNG, JPG, JPEG, BMP, GIF, WEBP, TIFF Documents: PDF
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/mr-shaper/opencode-skill-hybrid-ocr/ocr">View ocr on skillZs</a>