document-granular-decompose
Upload local documents to TianGong AI Unstructure `/mineru_with_images` API for fine-grained parsing and return only plain fulltext content. Use when a task needs document fulltext extraction with `return_txt=true`, strict file-type allowlist validation, API base URL/auth token from environment variables, and optional provider/model overrides.
How do I install this agent skill?
npx skills add https://github.com/tiangong-ai/agent-skills --skill document-granular-decomposeIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill extracts text from local documents by uploading them to a configurable remote API. It includes a Python script to handle document parsing and returns the resulting fulltext to the agent.
- Socketwarn
1 alert: gptAnomaly
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Document Granular Decompose
Core Goal
- Parse a local document through
POST /mineru_with_images. - Always force
return_txt=true. - Read environment variables for endpoint, request identity, and model routing:
UNSTRUCTURED_API_BASE_URL(example:https://your-unstructured-host:7770)UNSTRUCTURED_AUTH_TOKENUNSTRUCTURED_PROVIDER(optional)UNSTRUCTURED_MODEL(optional)
- Return only plain fulltext (prefer API
txt; fallback to joinedresult[].text).
Triggering Conditions
- Need robust document fulltext extraction for PDF/Office/image files.
- Need image-aware MinerU parsing but only textual output for downstream chunking/search/summarization.
- Need to standardize provider/model/token input via environment variables instead of ad-hoc command parameters.
Workflow
- Prepare environment variables.
export UNSTRUCTURED_AUTH_TOKEN="your-fastapi-bearer-token"
export UNSTRUCTURED_API_BASE_URL="https://your-unstructured-host:7770"
# Optional routing overrides. Omit them to let the server choose its defaults.
export UNSTRUCTURED_PROVIDER="vllm"
export UNSTRUCTURED_MODEL="Qwen/Qwen3.5-122B-A10B-FP8"
- Run extraction and print fulltext to stdout.
python3 scripts/mineru_fulltext_extract.py \
--file "/absolute/path/to/document.pdf"
- Save fulltext to a local file when needed.
python3 scripts/mineru_fulltext_extract.py \
--file "/absolute/path/to/document.pdf" \
--output "/absolute/path/to/fulltext.txt"
Request Contract
- Endpoint resolution:
--api-urlif provided- else
UNSTRUCTURED_API_BASE_URL + /mineru_with_images - else fail fast with missing environment variable error
- Method:
POSTmultipart form. - Query params:
- Force
return_txt=true(always set by script).
- Force
- Form fields sent:
file(required)provider(optional, fromUNSTRUCTURED_PROVIDERwhen set)model(optional, fromUNSTRUCTURED_MODELwhen set)
- Header sent:
Authorization: Bearer $UNSTRUCTURED_AUTH_TOKEN
Supported File Types (Strict)
- Supported file types:
.bmp, .doc, .docm, .docx, .dot, .dotx, .gif, .jp2, .jpeg, .jpg, .odp, .odt, .pdf, .png, .pot, .potx, .pps, .ppsx, .ppt, .pptm, .pptx, .tiff, .webp, .xls, .xlsm, .xlsx, .xlt, .xltx
- Office formats:
.doc, .docm, .docx, .dot, .dotx, .odp, .odt, .pot, .potx, .pps, .ppsx, .ppt, .pptm, .pptx, .xls, .xlsm, .xlsx, .xlt, .xltx
- Any other extension is rejected before sending API requests.
Output Rules
- Success output must be plain text fulltext only.
- Normalize the confirmed upstream Markdown underscore escape (
\_to_) in both supported response paths; preserve other backslashes and escapes. - Fulltext source priority:
response.txt- join non-empty
response.result[].textby blank lines
- Do not output chunk metadata/json unless the user explicitly requests debugging.
Error Handling
- Missing required env vars (
UNSTRUCTURED_API_BASE_URL,UNSTRUCTURED_AUTH_TOKEN): fail fast with actionable message. - Missing
UNSTRUCTURED_PROVIDERorUNSTRUCTURED_MODEL: omit the form field and let the service choose its default. - HTTP 401/403: report token/auth issue.
- HTTP 4xx/5xx: print status and API error body if available.
- Missing text in response: fail with explicit schema mismatch error.
References
references/env.mdreferences/request-response.md
Assets
assets/config.example.env
Scripts
scripts/mineru_fulltext_extract.py
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/tiangong-ai/agent-skills/document-granular-decompose">View document-granular-decompose on skillZs</a>