The Page Was Never Unreadable: Thaana PDF Text Recall From 23% to 99.7%
The Claim, Up Front
Dhivehi documents exported to PDF from Microsoft Word extract, using the ordinary route through PyMuPDF, at 23% to 27% word agreement with the Word file they were exported from. Repairing two specific defects takes the same pages to between 90.8% and 99.9% on eleven real documents, and to 99.7% or better on the minutes. No OCR, no vision model, no machine learning of any kind. The information was in the file the whole time.
This came out of two projects. rihaPDF is a browser-based PDF editor for Dhivehi documents, and it found the first defect because it had to: an editor that loses a vowel when you click a word is not an editor. Ilma is a document management system I am building for a large Dhivehi archive, and it needed the second one, because an archive that indexes text nobody can read has not archived anything. The refinement in this post is mostly Ilma’s. The discovery is rihaPDF’s.
What a Dhivehi PDF Actually Is
Thaana is written right to left. Its consonants carry their vowels as combining marks called fili, and every consonant must carry either a fili or a sukun, which marks the absence of a vowel. In Unicode the marks live at U+07A6 through U+07B0, they have no advance width, and they are painted over the letter in front of them. A Thaana word without its fili is not a word with a typo. It is a consonant skeleton that a reader has to guess at.
Two things about how these documents are produced matter more than anything else.
The first is that a large share of them are set in legacy Thaana fonts. Faruma, A_Waheed, MV Boli, and about two hundred relatives predate any usable Thaana Unicode support and map Thaana glyphs onto Latin code points. A PDF set in one of those extracts as fluent-looking ASCII garbage. That failure is at least honest: it is visibly wrong, and any quality check catches it.
The second is that the documents in this archive are authored in Word. Word’s PDF exporter handles Unicode Thaana, embeds the font, resolves the bidi layout itself, and produces a page that is typographically correct. Those are the documents this post is about, and they are the dangerous ones, because nothing about them looks broken until you read the text back.
The Baseline Was 23%, and Getting a Baseline Was the Hard Part
Every measurement before this one had been a proxy. Does the page look right. Does the extracted string contain a phrase I know is on it. Does a spot check of a heading come back sane. Those checks all share a defect: a phrase you check is a phrase you already knew to look for, so they report on the parts of the page you were already thinking about and say nothing about the rest.
The ground truth was sitting in a folder the whole time. These PDFs were exported from Word
documents, and the Word documents still exist. The .docx holds the text as clean Unicode in
word/document.xml. So the measurement is not a judgement at all: extract the PDF, extract the
source, and count what share of the source’s Thaana words appear in the extraction.
Measured that way, across eleven real documents:
before after
minutes, 5 documents 23%-27% 99.7%-99.9% of their Thaana words
agendas, 5 documents 23%-27% 90.8%-98.2%
malformed fili 2.9%-24.7% 0.00%
The 99.7% in the title is the floor of the minutes bucket, not a single uniform figure. Agendas
land lower and I will come back to why. The before column is the same for both, which is the
first useful signal: whatever is wrong is systematic, not per-document.
Two defects account for the gap. They are independent, they compound, and they need completely different fixes.
Fault One: Word Writes a Space Where a Vowel Belongs
A PDF does not store text. It stores instructions to paint glyphs, addressed by an index into a
font. To recover text from that, an extractor consults the font’s /ToUnicode CMap, which maps
each glyph index back to the character it stands for.
Word’s exporter writes one entry of a fili bfrange as a literal <0020>. A space.
<006D> <006E> [<07A6> <0020>] CID 0x6E should be U+07A7, not a space
<006F> <0077> <07A8> CID 0x6F..0x77 carry on correctly at U+07A8
In this archive it is always the same character: U+07A7, aabaafili, the long aa vowel and one of
the commonest characters in the language.
The reason this survived for years is the part worth sitting with. Nothing consults /ToUnicode
while painting a page. The renderer takes the glyph index and draws the glyph, and the glyph is
correct, so the document is perfect on screen, perfect in print, and perfect to a human reader.
The map is consulted only when someone tries to read the text back out. Every extractor
downstream, faithfully doing exactly what the PDF told it to do, silently drops the vowel and
gains a space in its place.
Measured over the archive, 247 of the 248 pages that carry a text layer come back with between 2.9% and 24.7% of their fili unattached to any letter.
The fix is inference. A fili bfrange maps consecutive glyph indices to consecutive code points,
so a gap in the middle of a block is recoverable arithmetically: look at what the neighbours map
to, check that they advance in step, and interpolate. rihaPDF found the defect and does this by
looking at the immediate neighbours, cid - 1 and cid + 1.
That repair alone lifts agreement from 23%-27% to 33%-64%.
Which is a large improvement and still, obviously, nowhere near a usable page. The characters were now right. The order was not.
The Refinement: Arithmetic Over a Window
Before the order, the smaller of the two improvements Ilma made.
rihaPDF’s rule reads the CIDs on either side of the gap. That is enough for a single broken entry, and a single broken entry is what these documents contain, so it was the right rule for the file in front of it. It cannot repair two broken CIDs in a row, because then both neighbours are themselves spaces and there is nothing to interpolate from.
Ilma’s version is arithmetic over a window instead of a neighbour lookup: find the nearest mapped
fili on each side within four positions, confirm that the code points advance in step with the
CIDs, and interpolate across the whole run. One-sided inference at the end of a block falls out of
the same rule for free, which matters because the last entry of a bfrange has no right-hand
neighbour by construction.
Both versions confine every inferred value to U+07A6 through U+07B0. That guard is not decoration. Plenty of fonts map plenty of CIDs to a genuine space, and a rule that rewrites a space because a fili happens to sit near it in CID space would be a much worse bug than the one being fixed. The repair is allowed to produce a Thaana combining mark and nothing else.
Fault Two: The Order Is Wrong in Three Ways at Once
MuPDF linearises a page in the order the content stream paints glyphs, and does not run the Unicode bidi algorithm over the result. A span carries a bidi level, but that level is metadata, not an instruction to reorder. For a pure-Thaana span MuPDF reports level 0 and leaves the paint order alone.
Word resolves bidi itself before emitting the page, which is entirely correct of it. The consequence is that a Word-exported Thaana page extracts in visual order, and it is wrong in three distinct ways at the same time:
- A run’s trailing mark is emitted at the head of the run, so the sukun that closes a word arrives attached to the word before it.
- The space glyphs land inside words rather than between them, so word boundaries in the extracted string are simply not where the words are.
- Digits inside a Thaana line read backwards.
2026extracts as6202.
That third one is the tell that made the rest tractable. A date reading backwards is not corruption. It is a correct right-to-left traversal of characters that happen to want to be read left to right. Nothing has been lost. The page was laid out by a typesetter that knew exactly where every glyph belonged, and every glyph in the trace still carries its own position.
The geometry is the account of the page that is not damaged.
Reading the Page Back From Where Its Glyphs Were Painted
So the second repair discards the extracted string entirely and rebuilds the page from glyph positions. Five rules, all of them small:
Lines come from the baseline. Glyphs sharing a text-origin y are one line, within a tolerance of 1.5 points, because marks and the odd superscript sit a hair off the baseline of the text they belong to.
A line is read in its own direction. Right to left when its strong characters are Thaana or
Arabic, left to right when they are Latin. Digits and punctuation get no vote, because they have
no direction of their own. This matters more than it sounds: a heading like 2026 vana aharu is
half digits and is read right to left all the same, so counting the weak characters would flip it.
And for a pure RTL line, right to left is logical order, which is why the geometry alone
recovers the text.
A mark belongs to the base nearest it. A fili has no advance width, so it is painted over its consonant. Word positions fili with their own offsets, though, and a mark can land slightly to the right of its base and sort ahead of it. So the rule is attach, not sort: find the nearest base and make the mark follow it. The base is never in doubt even when the ordering is.
Word boundaries come from the gaps, not the space glyphs. A horizontal gap wider than 0.14 of the font size is a word boundary. That threshold was chosen against the archive rather than reasoned about, and the useful part of the measurement is that it was flat between 0.14 and 0.18: Thaana letter spacing is nowhere near that wide and a real space is far beyond it, so the threshold sits in an empty gap between two populations rather than on a slope. The space glyphs in these documents are in the wrong place. The gaps are what a reader actually sees.
Embedded left-to-right runs are turned back. Reading an RTL line right to left reverses any
digits or Latin inside it, which is the one thing the geometry gets wrong on its own. A run of
digits or Latin letters, including internal punctuation when it sits between two such characters,
is reversed back. The punctuation condition is what keeps 2/2/2026 together without swallowing
the brackets around a Thaana phrase.
That is the whole of it. No dictionary, no language model, no Dhivehi-specific knowledge beyond which Unicode ranges are which. On the eleven documents it returns 90.8% to 99.9% of the source’s Thaana words, and the pages stop needing OCR at all.
The agendas land at the bottom of that range because they are short, heavily tabulated documents where a table cell’s contents and its neighbour’s share a baseline; the minutes are long runs of body prose, which is the case the baseline rule handles cleanly. That is a line-finding limit, not a character-recovery one.
Two Repairs That Must Never Meet
Ilma manufactures some of its own PDFs. Office files are converted to a PDF master through
Gotenberg, which means LibreOffice for documents and Chromium for HTML, and both of those
producers emit visually-ordered Thaana too. So there was already a repair pass in the pipeline
before the Word work started, and it works on a different principle: it reads shaped glyph
clusters and normalises them, because LibreOffice wraps every glyph of a cluster in its own span
carrying the whole cluster as ActualText, so a consonant plus a fili extracts as four
characters instead of two.
Forced onto a Word page, that cluster repair scores 0.1% to 5.4%. It tears marks off their consonants and produces a different kind of wrong.
So the two paths have to be chosen apart, and the choice cannot be a guess:
if the page has no RTL text leave it alone
if the producer is LibreOffice or Chromium,
or the text is duplicated outright cluster repair
else if fili are attached to nothing positional recovery
else leave it alone
Producer metadata decides the first branch. The second is decided by measurement on the page itself, and the positional pass is additionally only kept if it leaves fewer fili stranded than the flat text did. If the recovery is not demonstrably better than what it replaces, it is discarded and the page is failed for OCR instead. Churning text without fixing it is worse than leaving it alone, because the next person to look at the page cannot tell which happened.
The Third Case: A Converted Master Is Not the Document
The nastiest thing I found in this whole exercise was not either defect above. It was what happens when a Word document is uploaded as a Word document.
Gotenberg converts it to a PDF master for display, pagination and citation, and until recently its text came from that master. LibreOffice, converting Thaana, substitutes the vowels. Not displaces them, not drops them. Substitutes them, so one fili becomes a different fili, and the result is well-formed Thaana.
That is the whole problem. Well-formed text passes every check I had. Every fili sits properly
behind a letter, the character distribution is sane, the density is plausible, so the page scores
good and is indexed as though it were sound. Measured against the documents’ own sources, a
converted master recovers 0.9% to 3.9% of their Thaana words. One document was sitting in the
archive at 0.4% agreement, marked good, with nothing anywhere objecting.
Reordering cannot help, and this is the one case where the geometry is no use: the code points themselves are wrong. The positional reader scores 0.5% to 3.9% here, against 99.7% on a Word PDF.
So the text comes from the document instead. Read word/document.xml, xl/sharedStrings.xml and
the slide parts, and join the runs with nothing at all, because a run boundary is not a word
boundary. Word splits a word across runs whenever formatting or spell-check state changes, so a
fili can end up in a run of its own, and joining with a space orphans it. That detail cost me an
afternoon and looks completely correct in a diff.
Then the pages, which are the real work. Citations are page-and-line, and page boundaries only exist once something is paginated, so the source’s text has to be cut the way the master cuts it. What survives an office conversion is the alignment key: digits, Latin and Arabic all come through intact, and so does the order. Reduce both texts to that skeleton, align them, and each of the master’s page boundaries lands at an offset in the source.
Thaana is deliberately excluded from the alignment key, for the reason this section is about: matching on it would align the document onto nonsense.
Verifying that alignment needed a check that does not require per-page ground truth, since the source has no pages to compare against. The property used is that every number printed on master page k must appear in the slice assigned to page k:
minutes 49 pages 100.0% (248/248)
minutes 30 pages 100.0% (218/218)
minutes 31 pages 99.5% (199/200)
agenda 4 pages 80.0% (16/20)
The 49-page document that had been sitting in the archive at 0.4% now stores its own text exactly: 100.0% recall, every page marked as sourced from the document rather than the master.
Why the Quality Score Is Orthography and Not a Threshold
All of this routing depends on being able to ask one question about a page of Dhivehi: is this text something a person could read?
The answer that worked is not a heuristic. It is orthography. A fili is a combining mark and belongs to the consonant in front of it, so in readable Dhivehi every fili has a Thaana letter immediately before it. Any other arrangement is text no reader can read, whatever put it there: a run painted right to left, a cluster whose mark was emitted before its base, or a legacy font whose glyphs came back mapped to the wrong code points. One measurement, three unrelated defects.
What makes it usable as a gate is the distribution. Across the archive, all 68 pages that had been transcribed by OCR score exactly zero. Of the 248 pages carrying an embedded text layer, 247 score between 2.9% and 24.7%. Nothing falls in between. The threshold sits at 2%, inside an empty gap, which means it is choosing between two populations rather than tuning a dial. I have written a lot of thresholds that were really just a number I liked, and the difference in how they behave later is not subtle.
The one guard it needs is a minimum: fewer than 60 fili on a page and the score is forced to zero, because a single stray mark where a word is split across a page break proves nothing and the verdict costs an OCR pass.
What This Replaced
The verdict before all of this was: the text layer is damaged, fail the page, read it again with OCR. That verdict was a correct response to a real problem and it was still the wrong answer, because these pages were never unreadable. They were misread.
It also would not have worked, which I only know because the OCR side of the same project was
measured just as carefully. Transcribing these pages is worse than extracting them in a specific
and unrecoverable way: OCR reads them broadly correctly and corrupts a number in about a third of
passes, 2020 for 2026 and 1441 for 1448 among them. Those are the years, document numbers
and reference numbers that the documents are searched by. A commercial vision model, meanwhile,
invents Dhivehi proper nouns outright, fluently, differently on each run, and every substituted
word is ordinary Dhivehi, so no dictionary check can catch one.
The comparison is the part I would want someone else to take from this. A page with a damaged text layer and a scan look like the same problem, and they are not. The scan has genuinely lost information and something has to guess. The Word page has lost nothing at all, and reaching for a model there does not fill a gap, it replaces recoverable data with a guess, and the guess is plausible enough that nothing downstream can tell.
The cheapest 77% I have ever recovered came from reading the file more carefully.
What Was Not Documented
I looked for prior art on the /ToUnicode defect before writing any of this, and again before
writing this post. There is a great deal about ToUnicode CMaps in general, and about PDFs that
omit them entirely, and there is real work on Dhivehi OCR. I could find nothing on this specific
failure: a producer writing one entry of an otherwise correct combining-mark bfrange as U+0020,
in a document that renders perfectly, so that the defect is invisible to everyone except whoever
tries to extract the text. Nothing in a bug tracker, nothing in a paper, nothing in a mailing
list thread.
That is unsurprising, and it is the reason for this post. The defect needs three conditions to be visible at all: a script whose vowels are combining marks, a producer that writes the broken entry, and somebody who cares what the extracted text says rather than what the page looks like. For most of the world’s documents the second and third conditions never meet.
If you work with a script whose vowels or tone marks are combining marks, and text extraction from
a PDF that looks perfect is coming back subtly wrong, check the /ToUnicode CMap for a mark
mapped to a space, and check whether your extractor is handing you paint order rather than logical
order. Both are cheap to test and neither is visible from the rendered page.
Technical Notes
Ilma’s extraction path is Python: PyMuPDF for the document, processing/tounicode.py for the CMap
repair, processing/positional.py for the geometric recovery, processing/thaana.py for the
cluster repair and the orthographic quality measure, and processing/office.py for reading an
office document’s own text and cutting it into the master’s pages. The pure logic is separated from
PyMuPDF deliberately, so every rule above is unit-testable against synthetic glyphs, and the
fixture word-dhivehi-3pages.pdf holds the whole path as a regression test. One of its tests
asserts that the flat extraction is still damaged, so that if a future PyMuPDF starts resolving
this itself, the test fails loudly and the repair can be deleted.
The CMap defect and the shape of its fix were found by rihaPDF, in TypeScript, over pdf.js and pdf-lib: github.com/yashau/rihaPDF (Apache-2.0). The build log for that project is here. Ilma’s repository is not public, so the file paths above are references rather than links.