web-scraper
Search the web and read page contents without API keys. Use when you need to search via DuckDuckGo/Brave/Google (multi-page), extract readable text from URLs, browse interactively with a persistent visible browser (with tabs, click, screenshot, text search), download files/PDFs, or dismiss cookie banners. Supports JSON/markdown/text output. Powered by Playwright + Chromium.
How do I install this agent skill?
npx skills add https://github.com/liranudi/openclaw-web-scraper --skill web-scraperIs this agent skill safe to install?
- Gen Agent Trust Hubwarn
This skill provides powerful web scraping and browser automation capabilities. While the functionality is consistent with its stated purpose, it includes features such as arbitrary JavaScript execution within the browser and the ability to bypass SSL certificate verification, which could be misused if the agent is directed to malicious websites.
- Socketpass
No alerts
- Snykwarn
Risk: MEDIUM · No issues
- Runlayerpass
3/7 files flagged
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
Web Scraper
Four scripts, zero API keys. All output is JSON by default.
Dependencies: requests, beautifulsoup4, playwright (with Chromium).
Optional: pdfplumber or PyPDF2 for PDF text extraction.
Install: pip install requests beautifulsoup4 playwright && playwright install chromium
1. Search the Web
python3 scripts/google_search.py "query" --pages N --engine ENGINE
--engine—duckduckgo(default),brave, orgoogle- Returns
[{title, url, snippet}, ...]
2. Read a Page (one-shot)
python3 scripts/read_page.py "https://url" [--max-chars N] [--visible] [--format json|markdown|text] [--no-dismiss]
--format—json(default),markdown, ortext- Auto-dismisses cookie consent banners (skip with
--no-dismiss)
3. Persistent Browser Session
python3 scripts/browser_session.py open "https://url" # Open + extract
python3 scripts/browser_session.py navigate "https://other" # Go to new URL
python3 scripts/browser_session.py extract [--format FMT] # Re-read page
python3 scripts/browser_session.py screenshot [path] [--full] # Save screenshot
python3 scripts/browser_session.py click "Submit" # Click by text/selector
python3 scripts/browser_session.py search "keyword" # Search text in page
python3 scripts/browser_session.py tab new "https://url" # Open new tab
python3 scripts/browser_session.py tab list # List all tabs
python3 scripts/browser_session.py tab switch 1 # Switch to tab index
python3 scripts/browser_session.py tab close [index] # Close tab
python3 scripts/browser_session.py dismiss-cookies # Manually dismiss cookies
python3 scripts/browser_session.py close # Close browser
- Cookie consent auto-dismissed on open/navigate
- Multiple tabs supported — open, switch, close independently
- Search returns matching lines with line numbers
- Extract supports json/markdown/text output
4. Download Files
python3 scripts/download_file.py "https://example.com/doc.pdf" [--output DIR] [--filename NAME]
- Auto-detects filename from URL/headers
- PDFs: extracts text if pdfplumber/PyPDF2 installed
- Returns
{status, path, filename, size_bytes, content_type, extracted_text}
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/liranudi/openclaw-web-scraper/web-scraper">View web-scraper on skillZs</a>