Build1 publisher2 min readPublished
MarkItDown reports success on scanned PDFs that yield only a newline
Microsoft's MarkItDown returns a single newline and exit code 0 for a scanned PDF, according to a dev.to test of version 0.1.8 on 14 files. Any pipeline that feeds a retrieval index from it has to check the returned text itself before writing a file.
The Engineer · Build desk
What happened
- Without a guard, three of the author's fixtures would have been saved as files holding one newline, and nothing in the pipeline would have complained.
- The project README says OCR is not part of the default pipeline, and a scanned PDF has no text layer for the converter to read.
- Text PDFs with tables kept correct numbers but gained empty columns between value pairs, a layout artifact that downstream parsers see as phantom cells.
- Installing markitdown[all] at 0.1.8 pulled a pre-release Azure dependency, and the resolver fell back to 0.1.5 with little warning.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams building ingestion on MarkItDown have to decide for themselves what counts as a failed conversion, because the library's success signal also covers empty output.
- exposure An index built without a guard counts one-newline documents as ingested, and nothing in the run flags them for anyone to review.
- cost Table-heavy PDFs need cleanup that the author puts at about two minutes per table, or a column-normalizing step before indexing.
- constraint Projects that install the all extra can end up on 0.1.5 without noticing, so builds that rely on 0.1.8 behavior need to pin specific extras or assert the version.
The failure arrives as a normal return value. Run from the command line on a scanned PDF, `markitdown` writes one newline and exits 0 [2]. Called from Python, `convert()` returns without raising an error, and `result.markdown` holds next to nothing [2][11]. A try/except around the call has nothing to catch. The author's batch loop wraps every conversion in exactly that block, and without a further check nothing in the run would have complained [8][4].
The loop handles it with a minimum output length. It strips the returned text, skips anything under 20 characters, and prints a note to check for a text layer or add OCR [8]. The author wrote that the `len(text.strip()) < 20` guard is not paranoia [9]. The check measures length and nothing else. It does not open the PDF to ask whether a text layer exists. Any output of 20 characters or more gets written, whatever produced it [8].
Three one-newline files out of 14 fixtures is about 21 percent [14]. The 14 were public samples chosen to find the tool's edges, with scanned PDFs and two PNGs in the mix [3]. The failure follows from the missing text layer, so the rate depends on how many image-only PDFs a given archive holds [5]. A corpus with no scans would sit near zero. The 21 percent figure applies only to an archive whose share of image-only files is close to that test set's.
For its stated job, the tool is well built. The Markdown keeps headings, lists, tables and links instead of flattening them to plain text [10]. A chunker splitting documents for a retrieval index needs that structure [10]. The README says plainly that the output is meant for machines and does not aim for faithful layout for human readers [12]. The basic API is three lines, and the author's examples switch plugins off with `enable_plugins=False` [11].
I would rather the library warned when a non-empty file converts to empty text. For now the check belongs in the calling code, next to a version check. The author's one-liner, `python -c "import importlib.metadata as m; print(m.version('markitdown'))"`, shows which release the resolver actually installed [13].
What to watch
- Whether a later MarkItDown release raises an error or a warning when a non-empty input converts to empty text.
- Whether the Azure extra's pre-release dependency is fixed so markitdown[all] installs 0.1.8 or later without --pre.
- Whether OCR moves into the default pipeline for PDFs that lack a text layer.