Editing a document you only have as a PDF
Turn a PDF notice someone sent you into a .docx, edit the wording and send it back.
Guide
PDF to Word turns a PDF that carries a text layer into a .docx you can keep editing in Word or WPS Office: it reads the text coordinates on each page, joins character blocks into lines, then merges lines into paragraphs using three signals — a sudden jump in line spacing, a font-size change and a first-line indent — keeping an approximate size and bold state per paragraph. The goal is to get the text out together with its paragraph structure, not to reproduce the layout.
Updated 2026-09-093 min read
PDF to Word turns a PDF that carries a text layer into a .docx you can keep editing in Word or WPS Office: it reads the text coordinates on each page, joins character blocks into lines, then merges lines into paragraphs using three signals — a sudden jump in line spacing, a font-size change and a first-line indent — keeping an approximate size and bold state per paragraph. The goal is to get the text out together with its paragraph structure, not to reproduce the layout.
It is currently in Beta, with the paragraph rules still being tuned. Unlike Extract text, it outputs a .docx and tries to segment it properly; unlike PDF OCR it does no OCR at all — a scanned file yields almost no text here, and the page says so and offers a “Go to PDF OCR” button.
| Input | Output | Notes |
|---|---|---|
20-page single-column paper PDF (text-based) |
paper.docx: bold headings, split paragraphs, page breaks between pages |
single-column body text works best |
a two-column journal page |
lines from the left and right columns interleave by height |
columns are a known weak spot |
a scanned PDF |
a “may be a scan” warning plus Go to PDF OCR |
triggered below 30 characters |
Turn a PDF notice someone sent you into a .docx, edit the wording and send it back.
Drop long passages from papers or reports into your own Word file — copying from the converted text brings fewer broken lines than copying from a PDF reader.
See how many characters come out first; almost none means the file is a scan, and the page points you to OCR.
This version only handles the text layer, so images, table rules and layout elements are not written into the .docx. For table data use PDF to Excel; for images use Extract images.
Paragraphs are inferred from line spacing, font size and indents, so a PDF with very tight spacing or an unusual layout can be misjudged. Turn off “Page breaks as in the original PDF” and merge paragraphs in Word with find-and-replace.
This tool reads the text layer directly, which is fast and character-accurate and suits text-based PDFs; PDF OCR runs OCR over page images, which suits scans but is slower and can misread characters.
The PDF is parsed, its text extracted and the .docx built in your browser; the file contents are never uploaded to any server, and the data is released from memory when you close the page.
Updated 2026-09-09
Extract PDF text and basic paragraphs into a .docx; complex layouts are not preserved
Extracts the text and basic paragraph structure into a .docx; complex layout (images, columns, headers and footers) is not kept