PDF to Word: what survives the conversion and what does not
Every converter promises a perfect copy. None can give you one, and it helps to know why before you judge the result: the two formats are built on opposite ideas.
Two opposite ideas
A Word document is a flow: paragraphs one after another, and Word decides where the page ends when you open it. A PDF is the opposite — it stores where every fragment of text was printed on the page, in coordinates, precisely so it looks the same everywhere.
Converting from PDF to Word means reading those coordinates and guessing the flow back: which fragments form one line, which lines form one paragraph, which is a heading, where a table starts. A good converter guesses well; none of them knows for certain, because the information was deliberately thrown away when the PDF was made.
What comes back well
The text, with its size, its bold and italic, its underlines and its colours. Headings come back as headings, so Word's navigation pane works and you can regenerate a table of contents.
Tables drawn with lines come back as real Word tables, cell by cell, which is what lets you sort them or add a column. Tables without borders are deduced from how the columns line up, and honest tools mark that as a guess.
Images, the page size and the margins. Lists keep their bullets or numbers. Repeated headers and footers are detected and written as headers and footers, not as text stuck at the top of every page.
What does not, and why
The exact line breaks, when the original used a font your computer does not have. Word substitutes the font, the substitute is a hair wider or narrower, and the last word of a line moves. Nothing can prevent that except installing the original font.
Text boxes, WordArt, footnotes, charts and drawings made with shape tools: a PDF does not store them as objects, only as drawing instructions. What you get is their appearance, not an editable chart.
A scanned PDF has no text at all. Its pages are photographs, so there is nothing to move into Word until somebody reads the words out of the picture.
Reading a picture needs OCR, which CLEKARA does not do yet: it detects that case and stops with an explanation instead of handing you an empty document.
Get a better result
Start from the original PDF, not from a photocopy of a printout. Every extra generation makes the letters fuzzier and the guessing worse.
If the document has two columns, check the reading order first. Columns are where converters go wrong most often, and it is much faster to notice it in the first page than after editing twenty.
Convert, then fix in Word — not the other way round. Regenerate the table of contents, check the tables, and save as .docx. Ten minutes of tidying on a good conversion beats an hour of retyping.
You can try it with your own file in PDF to Word; the file is read in your browser and never uploaded.