doc-to-markdown
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing. Fixes pandoc grid tables, simple tables, image paths, CJK bold spacing, attribute noise, and code blocks; for PDFs also strips OCR garbage blocks, repeated headers/footers/watermarks, and absolute image paths from pymupdf4llm output. Benchmarked best-in-class (7.6/10) against Docling, MarkItDown, Pandoc raw, and Mammoth. Trigger on "convert document", "docx to markdown", "parse word", "doc to markdown", "解析word", "转换文档", "HTML to Markdown".
How do I install this agent skill?
npx skills add https://github.com/daymade/claude-code-skills --skill doc-to-markdownIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill is a document conversion utility that transforms DOCX, PDF, and PPTX files into Markdown. It orchestrates well-known tools like Pandoc and MarkItDown using Python scripts. No malicious behavior or security risks were identified beyond the inherent risk of processing untrusted document content.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
Doc to Markdown
Convert documents to high-quality markdown with intelligent multi-tool orchestration and automatic DOCX post-processing.
Architecture: Pandoc (best-in-class extraction) + 8 post-processing fixes (our value-add).
Quick Start
# DOCX → Markdown (one command, zero manual fixes)
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.docx -o output.md --assets-dir ./media
# PDF → Markdown
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.pdf -o output.md
# Saved HTML → Markdown (Pandoc required; no remote fetching)
uv run scripts/convert.py page.html -o page.md
# Run tests
uv run --with pytest pytest scripts/test_convert.py -v
Dual Mode
| Mode | Speed | Quality | Use Case |
|---|---|---|---|
| Quick (default) | Fast | Good | Drafts, simple documents |
| Heavy | Slower | Best | Final documents, complex layouts |
Tool Selection
| Format | Quick Mode | Heavy Mode |
|---|---|---|
| pymupdf4llm | pymupdf4llm + markitdown | |
| DOCX | pandoc + post-processing | pandoc + markitdown |
| PPTX | markitdown | markitdown + pandoc |
| XLSX | markitdown | markitdown |
| HTML/HTM | pandoc + source-href retention check | unsupported; use the HTML quick path |
Saved HTML And Website Manuals
For HTML/HTM, read references/html-conversion.md
before converting. Use scripts/convert.py; the default converts the whole body
and does not trim navigation. An explicit --html-selector selects exactly one
tag, #id or .class. --html-heading-offset shifts parsed headings for assembly
and rejects overflow beyond H6. Relative assets are not downloaded or copied.
Verify source href occurrences against the emitted Markdown AST, then rerun
scripts/html_to_markdown.py after cleanup or merging. For manuals, map the book,
chapters, lessons and internal headings before assembly; preserve code fences and
reconcile rewritten anchors. Conversion success does not certify that figures
are readable inside the recipient's actual Markdown reader.
DOCX Post-Processing (automatic)
When converting DOCX via pandoc, 8 cleanups are applied automatically:
| Problem | Fix | Test coverage |
|---|---|---|
Grid tables (+:---+) | Single-column → blockquote, multi-column → pipe table | TestPostprocessPipeline |
Simple tables ( ---- ----) | Multi-column images → pipe table with captions | TestSimpleTable |
Image path nesting (media/media/) | Flatten to media/, absolute → relative | test_stats_tracking |
Pandoc attributes ({width="..."}) | Removed | test_pandoc_attributes_removed |
CJK bold spacing (**粗体**中文) | Add space around ** for CJK bold spans | TestCjkBoldSpacing (15 cases) |
| Indented dashed code blocks | → fenced ``` with language detection | test_code_block_with_language |
Escaped brackets (\[...\]) | → [...] | test_escaped_brackets_fixed |
Double-bracket links ([[text]](url)) | → [text](url) | test_double_bracket_links_fixed |
PDF Post-Processing (automatic, 2026-08-30 起)
When converting PDF via pymupdf4llm, 3 cleanups are applied automatically (skip with --no-postprocess):
| Problem | Fix | Test coverage |
|---|---|---|
Tesseract OCR garbage on image regions (<!-- Start of picture text -->...) | Block removed; images themselves kept | TestStripOcrPictureText |
| Repeated header/footer/watermark lines (same normalized line on ≥60% of pages, incl. diagonal watermarks) | Detected via pymupdf cross-page scan, removed from markdown; bold-wrapped and merged-with-page-number variants also caught | TestRepeatingLines |
Absolute image paths () | Rewritten relative to the output markdown file (portable output) | TestImagePathsRelative |
Heavy mode additionally prints a loud ⚠️ HEAVY MODE DEGRADED warning on stderr when one engine fails and the merge would otherwise silently degrade to single-engine output.
Known limits (learned from a 62-page Chinese research-report conversion, 2026-08-30):
- pymupdf4llm may emit duplicated paragraphs (source text layer has only one copy) — not auto-fixed; spot-check.
- Dotted TOC pages get detected as tables — rewrite the TOC manually if it matters.
- Cross-page tables are NOT merged (each page's fragment keeps its own header row) — merge manually.
- Table cells overlapped by diagonal watermarks can contain watermark character shards (
dn,uFE, ...); the repeating-line stripper removes full lines only, not intra-cell shards. Watermark-heavy PDFs need cell-level rebuild (collect non-watermark spans per cell bbox). - Complex infographics (dense in-image text) come out as images only; transcribing in-image text needs a VLM pass, not this tool.
CJK Bold Spacing — why and how
DOCX uses run-level styling (no spaces between bold/normal runs in CJK text). Markdown renderers need whitespace around ** to recognize bold boundaries.
Rule: if a **content** span contains any CJK character, ensure both sides have a space — unless already spaced or at line boundary. This handles CJK punctuation, emoji adjacency, and mixed content.
Before: 打开**飞书**,就可以 → some renderers fail to bold
After: 打开 **飞书** ,就可以 → universally renders correctly
Heavy Mode Workflow
Heavy Mode runs multiple tools in parallel and selects the best segments:
- Parallel Execution: Run all applicable tools simultaneously
- Segment Analysis: Parse each output into segments (tables, headings, images, paragraphs)
- Quality Scoring: Score each segment based on completeness and structure
- Intelligent Merge: Select best version of each segment across tools
Merge Criteria
| Segment Type | Selection Criteria |
|---|---|
| Tables | More rows/columns, proper header separator |
| Images | Alt text present, local paths preferred |
| Headings | Proper hierarchy, appropriate length |
| Lists | More items, nested structure preserved |
| Paragraphs | Content completeness |
Image Extraction
# Extract images with metadata
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf -o ./extracted-images
# Generate markdown references file
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf --markdown refs.md
Output:
- Images:
extracted-images/img_page1_1.png,extracted-images/img_page2_1.jpg - Metadata:
extracted-images/images_metadata.json(page, position, dimensions)
Quality Validation
# Validate conversion quality
uv run --with pymupdf scripts/validate_output.py document.pdf output.md
# Generate HTML report
uv run --with pymupdf scripts/validate_output.py document.pdf output.md --report report.html
Quality Metrics
| Metric | Pass | Warn | Fail |
|---|---|---|---|
| Text Retention | >95% | 85-95% | <85% |
| Table Retention | 100% | 90-99% | <90% |
| Image Retention | 100% | 80-99% | <80% |
Merge Outputs Manually
# Merge multiple markdown files
python scripts/merge_outputs.py output1.md output2.md -o merged.md
# Show segment attribution
python scripts/merge_outputs.py output1.md output2.md -o merged.md --verbose
Path Conversion (Windows/WSL)
# Windows to WSL conversion
python scripts/convert_path.py "C:\Users\<windows-user>\Documents\file.pdf"
# Output: /mnt/c/Users/<windows-user>/Documents/file.pdf
Common Issues
"No conversion tools available"
# Install all tools
pip install pymupdf4llm
uv tool install "markitdown[pdf]"
brew install pandoc
FontBBox warnings during PDF conversion
- Harmless font parsing warnings, output is still correct
Images missing from output
- Use Heavy Mode for better image preservation
- Or extract separately with
scripts/extract_pdf_images.py
Tables broken in output
- Use Heavy Mode - it selects the most complete table version
- Or validate with
scripts/validate_output.py
Bundled Scripts
| Script | Purpose |
|---|---|
convert.py | Main orchestrator with Quick/Heavy mode + DOCX post-processing |
html_to_markdown.py | Pandoc HTML adapter and saved-output source-href retention verifier |
test_convert.py | 31 tests covering all post-processing functions |
merge_outputs.py | Merge multiple markdown outputs |
validate_output.py | Quality validation with HTML report |
extract_pdf_images.py | PDF image extraction with metadata |
convert_path.py | Windows to WSL path converter |
References
references/benchmark-2026-03-22.md- 5-tool benchmark (Docling/MarkItDown/Pandoc/Mammoth/ours)references/heavy-mode-guide.md- Detailed Heavy Mode documentationreferences/tool-comparison.md- Tool capabilities comparisonreferences/conversion-examples.md- Batch operation examplesreferences/html-conversion.md- Saved HTML scope, link retention, assets and manual heading assembly
Next Step: Clean Up Converted Content
After converting documents to markdown, suggest cleanup:
Conversion complete: [N] files converted to markdown.
Options:
A) Clean up docs — run /daymade-docs:docs-cleaner to consolidate redundant content (Recommended if multiple files)
B) Check facts — run /fact-checker to verify claims in the converted content
C) No thanks — the markdown conversion is sufficient
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/daymade/claude-code-skills/doc-to-markdown">View doc-to-markdown on skillZs</a>