Docs / Guides

OCR, auto-filing and the Identity Card

What happens when you add a document

Keepsake runs OCR on your device — PurpleOCR on Windows, ML Kit on Android, WebAssembly in the browser — and then a shared rules engine reads the text:

  • Category — passport, insurance, contract, certificate, bill… scored from keywords, including regional terms (CNIC, NADRA, iqama).
  • Expiry date — found from labels like "Date of expiry" in 40+ formats and normalised properly.
  • Document number, issuing authority, tags — extracted with per-type patterns.
  • Passport MRZ — the machine-readable zone is parsed with every ICAO check digit verified. If the arithmetic doesn't prove the read exact, nothing is auto-filled from it; if it does, the verified passport number, expiry, date of birth and nationality beat any ordinary text match.

You always see the suggestions before saving, and your corrections improve future suggestions on that device only.

Tips for best results

  • Photograph documents flat, filling the frame, in decent light.
  • 300-dpi scans read essentially perfectly; phone photos are what the engine is tuned for.
  • Text PDFs (e.g. e-tickets, statements) are read directly — no OCR pass needed.

Wondering how accurate it actually is? We publish the numbers — including a reproducible benchmark against raw Tesseract — on the OCR feature page.

High-accuracy mode (Premium, Windows): Settings → OCR & Document Intelligence → Enable high-accuracy OCR downloads a second, deep-learning engine (PaddleOCR PP-OCRv4, ~12 MB, checksum-verified, runs fully on your device) that works beside Tesseract. Keepsake keeps whichever engine reads a page better — lifting overall accuracy on the published benchmark to 0.1% character error and 99.3% field accuracy, with the biggest gains on shadows, low light and phone photos.

Table extraction (Premium, Windows)

Bills, statements and invoices are really tables — but OCR normally flattens them into a jumble of text. On Windows, open a document and press 📋 Table: Keepsake walks the word geometry instead, rebuilding the rows and columns, and gives you Copy for Excel (paste straight into a spreadsheet) or CSV. If no aligned columns exist, it says so rather than guessing.

The Identity Card (Premium)

Settings → Identity Card aggregates every identity field found across all your documents — passport number, national ID, licence, IBAN — deduplicated into one list with tap-to-copy. Filling government forms stops being a scavenger hunt.

Copied values clear from the clipboard after 45 seconds on every platform, and Android marks them sensitive so previews stay hidden.

Languages

Latin-script documents work everywhere out of the box. Urdu (اردو), Arabic (العربية) and Hindi (हिन्दी) language packs shipped in every current build:

  • Windows — Settings → OCR & Document Intelligence → Language packs: download the pack once (checksum-verified against the official Tesseract models — nothing unverified is ever installed) and tick it on. New scans then read English plus your selected languages in one pass, fully offline.
  • Web app — Settings → OCR languages: tick a language; the recognition model downloads on first use and is cached for offline scanning after that.
  • Android — Settings → OCR Script: switch between Latin and Devanagari (Hindi). Urdu and Arabic aren't available on Android yet — Google's on-device recogniser has no Arabic-script model — so use the Windows or web app for those documents.

How well Urdu actually works

Honestly, and with numbers: it depends entirely on how the Urdu is set.

We measured 120 rendered lines of realistic Urdu document text — ID card fields, dates, addresses, amounts — where the correct answer is known exactly because we rendered it, and ran them through the same pipeline the app uses.

| Urdu style | Where you see it | Median character error | Lines read perfectly | |---|---|---|---| | Naskh | printed forms, ID card fields, bills | 4.2% | 38 of 80 | | Nastaliq | letters, notices, most Urdu prose | 20.9% | 1 of 40 |

So: printed Urdu works reasonably well — on clean scans, most lines come back exactly right. Nastaliq is harder, and the paragraph below revises this table's Nastaliq figure in both directions.

Two corrections, and a correction to one of the corrections. The 20.9% above is measured on ten short sentences in one typeface, which is too small a set to describe Urdu. In July 2026 we built a larger held-out set and published 43.2% character error, not one line of 540 read exactly. That figure was wrong — not dishonest, but drawn badly. Our sample was taken from the front of a corpus that emits one document template at a time, so "thirty held-out lines" turned out to be thirty variations of the same ID-card sentence. Sampled properly, across all thirty templates, and with a Nastaliq-specific model we have since trained:

| Nastaliq, measured properly | Median character error | Lines read perfectly | |---|---|---| | Stock Urdu model | 33.3% | 0 of 540 | | Our first fine-tune (July 2026) | 12.8% | 77 of 540 | | Our current model | 4.2% | 156 of 540 | | Raw Tesseract, same model | 10.0% | 0 of 540 |

Which of these you get, and what it costs. The stock Urdu model is free, on every plan, and always will be — being able to read a language is not something we will sell you. Our Nastaliq model is the top row of that table, it took three training runs to make, and it is Premium (£19/yr). That is the line we draw everywhere: capability is free, accuracy is paid. If you are on the free plan, Urdu still works — it reads at the 33.3% row rather than the 4.2% one.

That table does not carry over to ordinary Urdu writing, and we would rather say so than let you find out. Measured on Urdu Wikipedia prose — text sharing nothing at all with what the model was trained on — every version of our model reads at about 21%, against the stock model's 27%. Our first fine-tune, our second and our current one are statistically indistinguishable on it. All the improvement from 12.8% to 4.2% above is improvement at reading documents of the kind we trained on: ID cards, bills, certificates, policies. That is what the app is for, so it is the number that matters here — but if you point it at a page of Urdu literature, expect roughly one character in five to be wrong.

We were wrong about ID numbers, and it is fixed. This page previously said that identity and account numbers with hyphens — CNIC, mobile, bank account — did not work at all, scoring zero, and that "the digits themselves are not recovered, so this is a real gap rather than a formatting quirk". That was our bug, not a limit of the recogniser. Urdu lays a hyphenated number out with its groups right-to-left; our training data had them labelled the other way round, and the recognition engine did not put them back afterwards. Both halves are corrected:

| Held over 2,268 field checks | Before | Now | |---|---|---| | CNIC | 0% | 96% | | Mobile / bank account | 0% | 92% | | Dates | 85% | 90% | | Every value on the line | 36% | 60% |

What this means for you. Nastaliq now works for the things a vault is actually for: expiry dates, ID numbers, account numbers. What still does not read reliably is names (49%) and running prose (34%) — Urdu names in particular are short, varied and unforgiving of a single wrong letter.

So: trust Urdu OCR to find a document, to suggest a renewal date, and to pick up an ID number — but check any name it fills in. Photograph in good light; low light roughly doubles the error.

English and other Latin-script documents are unaffected and score far better; see the published benchmark.

English always stays active alongside your packs, so dates, document numbers and other Latin fields keep auto-filling.


Was this page helpful?

If something here is missing, wrong, or just unclear, say so — corrections to these pages usually start as a comment.

Leave a comment Ask the community →

Comments

No comments yet — be the first.

Sign in to comment — website account only; your vault never touches it.