pdf-text-extractor
Download PDFs (when available) and extract plain text to support full-text evidence, writing `papers/fulltext_index.jsonl` and `papers/fulltext/*.txt`.
How do I install this agent skill?
npx skills add https://github.com/willoscar/research-units-pipeline-skills --skill pdf-text-extractorIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is generally safe for its intended purpose of downloading research papers and extracting text. It contains standard risks associated with downloading content from external URLs and parsing untrusted PDF files, which could potentially serve as a vector for indirect prompt injection or parser exploitation.
- Socketwarn
1 alert: gptAnomaly
- Snykwarn
Risk: MEDIUM · 1 issue
- Runlayerwarn
2/2 files flagged
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
PDF Text Extractor
Triggers & routing
- Trigger: PDF download, fulltext, extract text, papers/pdfs, 全文抽取, 下载PDF.
- Use when:
queries.md设置evidence_mode: fulltext(或你明确需要全文证据)并希望为 paper notes/claims 提供更强 evidence。
Optionally collect full-text snippets to deepen evidence beyond abstracts.
This skill is intentionally conservative: in many survey runs, abstract/snippet mode is enough and avoids heavy downloads.
Inputs
papers/core_set.csv(expectspaper_id,title, and ideallypdf_url/arxiv_id/url)- Optional:
outline/mapping.tsv(to prioritize mapped papers)
Outputs
papers/fulltext_index.jsonl(one record per attempted paper)- Side artifacts:
papers/pdfs/<paper_id>.pdf(cached downloads)papers/fulltext/<paper_id>.txt(extracted text)
Decision: evidence mode
queries.mdcan setevidence_mode: "abstract" | "fulltext".abstract(default template): do not download; write an index that clearly records skipping.fulltext: download PDFs (when possible) and extract text topapers/fulltext/.
Local PDFs Mode
When you cannot/should not download PDFs (restricted network, rate limits, no permission), provide PDFs manually and run in “local PDFs only” mode.
- PDF naming convention:
papers/pdfs/<paper_id>.pdfwhere<paper_id>matchespapers/core_set.csv. - Set
- evidence_mode: "fulltext"inqueries.md. - Run:
uv run python .codex/skills/pdf-text-extractor/scripts/run.py --workspace <workspace> --local-pdfs-only
If PDFs are missing, the script writes a to-do list:
output/MISSING_PDFS.md(human-readable summary)papers/missing_pdfs.csv(machine-readable list)
Workflow (heuristic)
- Read
papers/core_set.csv. - If
outline/mapping.tsvexists, prioritize mapped papers first. - For each selected paper (fulltext mode):
- resolve
pdf_url(usepdf_url, else derive fromarxiv_id/urlwhen possible) - download to
papers/pdfs/<paper_id>.pdfif missing - extract a reasonable prefix of text to
papers/fulltext/<paper_id>.txt - append/update a JSONL record in
papers/fulltext_index.jsonlwith status + stats
- resolve
- Never overwrite existing extracted text unless explicitly requested (delete the
.txtto re-extract).
Quality checklist
-
papers/fulltext_index.jsonlexists and is non-empty. - If
evidence_mode: "fulltext": at least a small but non-trivial subset has extracted text (strict mode blocks if extraction coverage is near-zero). - If
evidence_mode: "abstract": the index covers everypapers/core_set.csvpaper and every record clearly reflectsskip_mode_abstract(no downloads attempted).fulltext_max_papersdoes not truncate this zero-download index.
Script
Quick Start
uv run python .codex/skills/pdf-text-extractor/scripts/run.py --helpuv run python .codex/skills/pdf-text-extractor/scripts/run.py --workspace <workspace>
All Options
--max-papers <n>: cap number of papers processed (can be overridden byqueries.md)--max-pages <n>: extract at most N pages per PDF--min-chars <n>: minimum extracted chars to count as OK--sleep <sec>: delay between downloads--local-pdfs-only: do not download; only usepapers/pdfs/<paper_id>.pdfif presentqueries.mdsupports:evidence_mode,fulltext_max_papers,fulltext_max_pages,fulltext_min_chars
Examples
- Abstract mode (no downloads):
- Set
- evidence_mode: "abstract"inqueries.md, then run the script (it will emitpapers/fulltext_index.jsonlwith skip statuses)
- Set
- Fulltext mode with local PDFs only:
- Set
- evidence_mode: "fulltext"inqueries.md, put PDFs underpapers/pdfs/, then run:uv run python .codex/skills/pdf-text-extractor/scripts/run.py --workspace <workspace> --local-pdfs-only
- Set
- Fulltext mode with smaller budget:
uv run python .codex/skills/pdf-text-extractor/scripts/run.py --workspace <workspace> --max-papers 20 --max-pages 4 --min-chars 1200
Notes
- Downloads are cached under
papers/pdfs/; extracted text is cached underpapers/fulltext/. - The script does not overwrite existing extracted text unless you delete the
.txtfile.
Troubleshooting
Issue: no PDFs are available to download
Fix:
- Use
evidence_mode: abstract(default) or provide local PDFs underpapers/pdfs/and rerun with--local-pdfs-only.
Issue: extracted text is empty/garbled
Fix:
- Try a different extraction backend if supported; otherwise mark the paper as
abstractevidence level and avoid strong fulltext claims.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/willoscar/research-units-pipeline-skills/pdf-text-extractor">View pdf-text-extractor on skillZs</a>