On-device OCR that files for you
Drop in a photo of a passport, policy or bill. Keepsake reads it locally, understands what it is, and files it — category, expiry date, document number, tags — without one byte leaving your device.
The problem
Filing is why document apps fail. If every upload means typing a title, picking a category and copying an expiry date from tiny print, people stop after a week — and an unfiled vault protects nothing.
Competitors solve this with cloud AI: your passport is uploaded to their servers to be read. That is precisely the document you least want on someone else's machine.
How Keepsake solves it
📄Text recognition on the device
PurpleOCR on Windows, ML Kit on Android, WebAssembly OCR in the browser. The image is processed where it already is — never uploaded for reading.
🧠One rules engine, everywhere
A shared extraction engine turns raw text into structure: document type, expiry date (40+ date formats, ISO-normalised), document numbers, issuing authority, tags.
🌍Documents from everywhere
Recognises regional formats competitors ignore — CNIC, IBAN, iqama — and reads the passport MRZ with full ICAO check-digit validation, plus checksum-verified Urdu, Arabic and Hindi language packs on Windows and the web app — with the Urdu accuracy measured and published rather than assumed.
📷Scan paper with your phone camera
On Android, tap Scan with camera: automatic edge detection and deskew, multi-page capture straight to PDF, then the same on-device OCR files it. The scanner runs on your phone — the pages never leave it.
Under the hood — shared extraction rules
- All three platforms interpret the same
extraction-rules.json: keyword-scored categorisation, labelled-date patterns with declared component order, per-field regex with explicit flags. - Dates are normalised to ISO with a two-digit-year pivot and impossible-day/month swap detection — "03/04/26" resolves correctly per document context.
- Passport and ID-card machine-readable zones are parsed per ICAO 9303 with every check digit verified — a field is auto-filled from the MRZ only when the arithmetic proves the read exact, and a verified MRZ overrides a mere label match.
- Identical behaviour is enforced by a shared fixture suite: 8 realistic documents (passport, insurance, utility bill, CNIC…) asserted byte-for-byte on C#, Kotlin and TypeScript.
- Corrections you make re-rank suggestions locally. Nothing is ever contributed back to us — there is no telemetry channel to contribute through.
Trust through specificity: the full crypto design is documented on the security page.
How accurate is it? Measured, published, reproducible
"It has OCR" tells you nothing. So we grade ourselves on a synthetic test set where the ground truth is exact by construction — four realistic document layouts (passport, insurance policy, utility bill, official letter) rendered across four fonts, then degraded the way phones degrade paper: sensor noise, skew, a shadow band, low light, phone-photo compression. 96 cases per engine.
Two numbers matter. CER (character error rate) is raw engine quality — lower is better. Field accuracy is the number you actually feel: did the expiry date, document number and amount come out readable? Below is Keepsake's on-device pipeline against raw Tesseract 5 with no preprocessing.
| Degradation | Keepsake CER | Raw Tesseract CER | Keepsake field | Raw field |
|---|---|---|---|---|
| Clean scan | 0.1% | 0.1% | 100% | 100% |
| Sensor noise | 0.6% | 0.6% | 94.8% | 94.8% |
| Skewed / rotated | 0.2% | 12.3% | 97.9% | 85.4% |
| Shadow band | 1.7% | 100% | 97.9% | 0% |
| Low light | 0.6% | 2.4% | 97.9% | 88.5% |
| Phone photo | 0.3% | 0.4% | 97.9% | 95.8% |
| Overall | 0.6% | 19.3% | 97.7% | 77.4% |
Two rows repay a close look. On the shadow band — a bright streak across the page, the classic phone-photo failure — raw Tesseract reads essentially nothing (100% error, zero fields); Keepsake detects the failed read and re-runs with heavier binarisation, recovering it to 1.7% and 97.9% of fields. On skew, the pipeline now reads a rotated page at 0.2% against raw Tesseract's 12.3% — but that row used to say the two were tied, and the reason is worth knowing: our deskewer had never actually run. Its angle detector sheared each row sideways and then counted the dark pixels in that row, which cannot change no matter how far you shift it, so every candidate angle scored the same and the answer was always "no skew". The tie we published so proudly was measuring raw Tesseract twice.
High-accuracy mode (Premium) adds a second, deep-learning engine — PaddleOCR PP-OCRv4, running fully on-device via ONNX Runtime — beside Tesseract, and keeps whichever read is stronger per page. On the same 96-case set it lifts the overall numbers to 0.1% character error and 99.3% field accuracy, winning most on exactly the hard cases — shadows, low light and sensor noise — while a skew guard keeps the deskewing pipeline in charge of rotated pages. That guard is doing real work: the deep engine alone scores 89.1% error on skewed pages and 50.4% on phone photos, because its text boxes are axis-aligned and a rotated line does not fit one. Keeping the better of two engines is only an improvement if you can tell which is better, so we publish what each engine scores on its own. (These read 2.1% / 96.9% until July 2026 — measured, like everything else, against a deskewer that turned out never to run.)
Reproducing these requires an idle machine. We first published 0.5% from a run that had been sharing the processor with two other benchmark processes; a repeat gave 3.6%, and the raw-Tesseract control column moved too — which it cannot legitimately do, since that arm has no adaptive logic to destabilise. The figures above come from an exclusive run of the shipping build; two further exclusive runs of that same build put overall character error at 0.6–0.7% and field accuracy at 96.7–97.7%. The clean, skew, shadow and phone rows come out identical every time. The degraded rows — sensor noise and low light — move by a few points in either direction from run to run, and so does raw Tesseract's control column beside them, because the recognition model is threaded and its output depends on how the work happened to be scheduled. Treat a two-point difference on those rows as weather, not signal. If you run PurpleOCR.Benchmark accuracy on a loaded machine, expect considerably worse than that.
These numbers moved twice in July 2026. First, work on Urdu exposed a way the engine could be fooled: a preprocessing pass that turns speckle into glyphs scores well on our internal quality measure precisely because it produces more text. A plausibility check — a pass may not multiply the text it claims to recover, and a pass replacing a blank read must itself be confident — fixed that. Then chasing the skew row uncovered the dead deskewer described above. Fixing it took overall character error from 2.6% to 0.6%, word error from 4.1% to 2.0%, and field accuracy to 97.7%. Both of those were bugs we found by insisting on publishing a row we were losing.
Reproduce it yourself from source:
dotnet run -c Release --project PurpleOCR/PurpleOCR.Benchmark -- accuracy
Questions
Does OCR work offline?
Yes — it is the same engine whether you are online or not. On the web app the OCR module is cached by the service worker after first use.
What file types can Keepsake read?
Photos and scans (JPG/PNG), and PDFs — both text PDFs (read directly) and scanned PDFs (OCR per page).
Can I scan paper documents directly?
Yes — the Android app has a built-in camera scan flow with automatic edge detection and deskew; multiple pages become a single PDF. On Windows and the web, import a photo or scan and the same pipeline takes over. Capture and recognition both happen on your device.
How do I know these accuracy numbers are real?
The whole benchmark is in the source tree and runs with one command — the test set, the ground truth and the scoring are all reproducible. The table above is the exact output; we also publish the methodology in detail on our blog.
Last verified — by dotnet run --project PurpleOCR/PurpleOCR.Benchmark. Every claim on this site is listed, with its evidence, in the claim ledger.