Detalle del Skill
pdf-to-html
Converts PDFs to readable HTML, but is a specialized conversion workflow.
Revisar antes de usar
La revisión automática comprueba relevancia, no seguridad ni respaldo. Lee las instrucciones de la fuente antes de usar este Skill.
SKILL.md
Este extracto es una copia guardada durante la revisión. La fuente externa contiene la versión completa y actual.
--- name: pdf-to-html description: Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure. Uses structured extraction (PyMuPDF), font-size-driven layout, compressed base64-inlined images (a single portable file), and mandatory headless-Chrome visual verification. Use whenever someone wants to READ a PDF as a web page or clean document, turn a PDF into HTML, or translate a PDF into another language while keeping its images/tables/charts intact — e.g. "PDF 转 HTML", "把这个 PDF 转成中文网页版", "make this report readable", "translate this PDF but don't lose the charts", "I just want to read this PDF on my phone". Distinct from doc-to-markdown (plain Markdown text) and pdf-creator (Markdown→PDF) — this one produces a styled, image-faithful HTML reading experience. --- # PDF to HTML Turn a PDF into a single, self-contained, readable HTML file — images, tables, charts and reading order preserved — and optionally translate it, keeping every figure in place. The pipeline is **extract → look → (translate) → build → verify**. The middle "look" and final "verify" steps are where faithfulness actually comes from: a PDF is a layout, not just a text stream, so you read the rendered pages before building and the rendered HTML before delivering. This skill runs **inline** (no `context: fork`): translation orchestrates a Dynamic Workflow, and a subagent cannot spawn one. ## When to use / not use - **Use** when the goal is to *read* a PDF as HTML/web page, to convert a PDF to a styled HTML document, or to translate a PDF into another language while keeping its figures and tables. - **doc-to-markdown** instead if they want plain Markdown text (no styling, figures optional). - **pdf-creator** instead for the reverse direction (Markdown → PDF). ## What it does NOT do - **Scanned/image-only PDFs** (no text layer): OCR first (e.g. `ocrmypdf`), then use this. - **Complex multi-column tables**: cell *text* is preserved and readable, but column alignment can flatten into a text flow — PyMuPDF reads a table as text blocks, not a grid, so the grid lines are gone. Tables that are *images* in the PDF survive as images. If the table's grid structure is essential, use **doc-to-markdown** (pandoc rebuilds real tables) or convert that page separately. - **Pixel-perfect facsimile**: output is a clean *re-flow* that keeps images and reading order, not a 1:1 copy of the original page layout. - **Rewriting**: it translates and re-lays-out; it does not summarize, add a TL;DR, or editorialize. Faithfulness is the point (see Fidelity below). ## Dependencies `uv` (runs Python with inline deps), Google Chrome or Chromium (visual verification). Python packages come via `uv run --with`: PyMuPDF, Pillow, numpy. Nothing to pre-install beyond Chrome and uv. ## Workflow Copy this checklist and tick as you go: ``` - [ ] 1. Extract structure + render pages (extract_pdf.py) - [ ] 2. Read pages/*.png — SEE the layout, find content vs decorative images - [ ] 3. (only if translating) run the translation workflow - [ ] 4. Build the single-file HTML (build_html.py) - [ ] 5. Verify visually (verify_render.py → Read every segment) - [ ] 6. Deliver the .html ``` ### 1. Extract ```bash uv run --with pymupdf python scripts/extract_pdf.py input.pdf ``` Writes `input-build/` with `structure.json` (text blocks with font sizes + image blocks flagged `decorative`), `images/`, and `pages/` (one PNG per page). ### 2. Look before you build Read `input-build/pages/*.png`. This is not optional: you need to see the real layout, confirm which images are content vs decoration, and spot tables/charts. For a long PDF, read every page; for a short one it's quick. This is also where you understand the document well enough to translate it well. ### 3. Translate (optional) Only if the user asked for another languagLeer la fuente completa en GitHub (abre una página externa)