How do I search inside scanned documents without uploading them anywhere?
How do I search inside scanned documents without uploading them anywhere?
Run OCR on your own machine to add an invisible text layer to each scanned PDF, then use the search built into your operating system. Free tools do this well: OCRmyPDF wraps Tesseract on any platform, and both Windows and macOS index the text inside PDFs automatically once it is there. Nothing needs to leave the computer.
A scan is a photograph of a page. Your computer can see the picture and cannot read a word of it, which is why a folder of scans is unsearchable however carefully you named the files. The fix is a step called OCR, and the useful thing to know is that it runs perfectly well on an ordinary laptop — the cloud is a convenience here, not a requirement.
What the OCR step actually does
OCR reads the picture and produces text. For documents the important part is where it puts that text, and there are two very different answers.
- A separate text file — useful for processing, useless for filing. You now have two files and one of them looks like nonsense.
- An invisible layer inside the PDF, positioned exactly over the words in the image. The page looks identical, you can select and copy text off it, and every search tool on your machine can read it. This is what you want, and it is what "searchable PDF" means.
The second kind costs nothing extra to produce and is what the tools below do by default. It also means the searchability travels with the document: copy the file to a new machine in ten years and it is still searchable, with no index to rebuild and no application required.
Doing it locally, free, on any platform
OCRmyPDF is the tool most people end up on. It wraps Tesseract, runs on Windows, macOS and Linux, adds the invisible text layer, and leaves the original image untouched.
ocrmypdf --skip-text in.pdf out.pdf is the whole thing for one file. --skip-text matters more than it looks: it leaves alone any page that already has text, so you can run it over a folder repeatedly without doing damage or wasting an afternoon.
Other routes to the same place, depending on what you already have:
- Your scanner's own software. Most bundled drivers have a "searchable PDF" output option that is off by default. If yours does, this is the least work available — the OCR happens as the page comes through.
- Adobe Acrobat, and most paid PDF editors. "Recognise text" does the same job locally.
- NAPS2 on Windows, which is free, scans and OCRs in one step, and is considerably friendlier than a command line.
- Preview and Quick Actions on macOS, and the Files app on iPhone, both of which now recognise text in images without asking anyone.
Accuracy on a flat, well-lit scan of printed text is high enough that you will rarely think about it. It falls off on handwriting, on faint thermal receipts, and on anything photographed at an angle — so if a document matters, scan it properly once rather than fixing it later.
Searching the folder afterwards
Once the text is inside the files, the search you already have works.
- Windows. File Explorer searches inside PDFs, but only in indexed locations and only with a PDF iFilter installed — Adobe Reader installs one, and so do several free PDF readers. Add your documents folder under Indexing Options and the search box starts returning hits from inside the pages.
- macOS. Spotlight reads PDF text with no setup at all. It is the one place this simply works.
- Linux.
pdfgrep -r "term" ~/Documentsneeds no index and no daemon, and is the fastest way to answer "which file was that in". - Anywhere. Everything on Windows for instant filename search, or DocFetcher, which builds a portable full-text index that can live on the same drive as the documents.
Worth knowing before you install anything: good file names solve most of this without OCR at all. If a file is called 2026-04-12 Passport Ali Muhammad HMPO.pdf, you will find it by name in a second. Full-text search earns its place for the question you cannot phrase as a file name — the policy number you half remember, the letter that mentioned a name.
Why people reach for a server, and when they should
The standard advice for this problem is paperless-ngx or Docspell: excellent software, genuinely better than a folder for high volumes, and both of them expect you to run a server. That means Docker, a machine that stays on, a database, backups of that database, and a security update cadence — a hobby, in other words, and a good one, but a hobby.
The honest test is volume and shape. A household filing perhaps two hundred documents a year does not need a document management system; it needs consistent names and OCR. A small practice filing several thousand, with correspondents, tags and retention rules, is exactly what those systems are for and will outgrow a folder quickly.
The middle case — more than a folder, less than a server — is where people get stuck, and it is mostly filled by desktop applications rather than by self-hosted ones. That is a real gap in the free tooling and it is worth saying so plainly rather than pretending a folder scales further than it does.
Keeping the text private once it exists
One consequence people miss: OCR makes a document readable by everything, including things you did not intend. Before OCR, a passport scan in a synced folder was an image. After it, the passport number is plain text inside the file — indexed by your desktop search, visible to any process that can read the folder, and searchable by whoever holds the account if the folder syncs somewhere.
That is not a reason to skip the step. It is a reason to do it before deciding where the file lives:
- Keep identity and financial documents in an encrypted container rather than a plain synced folder — the free options are in the guide on organising without a subscription.
- Check what your desktop search indexes. Indexes are ordinary files, they are rarely encrypted, and they are frequently backed up.
- If a tool offers to do the OCR "in the cloud for better accuracy", that is a full copy of the document leaving your machine. Sometimes the right trade; never an invisible one.
Step by step
- Scan to PDF, not to photo. Use a scanner or a phone scanning app so a multi-page document stays one file. Flat, even light, no shadow.
- Name the file before anything else. Date first: 2026-04-12 Passport Ali Muhammad HMPO.pdf. This alone answers most searches.
- Add the text layer locally. ocrmypdf --skip-text in.pdf out.pdf, or your scanner software's "searchable PDF" option, or NAPS2 on Windows.
- Let your operating system index it. On Windows, add the folder under Indexing Options and install a PDF iFilter. On macOS, Spotlight already has it. On Linux, pdfgrep needs no index at all.
- Verify before you trust it. Search for a word you know is in the middle of a scanned page. If it does not come back, the text layer or the index is missing — find out which now rather than in two years.
- Decide where the searchable copy lives. The text is now readable by anything with access to the file. Sensitive documents belong in an encrypted container before they reach a synced folder.
Questions
Is local OCR as accurate as a cloud service?
On clean, printed, well-lit pages the difference is small enough not to matter — both will read a typed letter or a bill essentially perfectly. Cloud services pull ahead on hard inputs: handwriting, unusual layouts, photographs taken at an angle, and low-resource languages. For a household archive of printed documents, local OCR is not a compromise. For a box of handwritten letters, it is.
Does OCR change the scan itself?
No, and this is the reassuring part. The text layer is added on top of the image, which is left exactly as it was. The page still looks like the scan, prints like the scan, and can be re-OCRed later with better software without any loss. If a tool offers to "clean up" or re-compress the image, that is a separate option and one to be careful with — aggressive compression on a scanned document can turn digits into other digits.
How long does it take?
Roughly a second or two per page on an ordinary laptop, so a hundred-page backlog is a coffee break rather than a project. It parallelises across cores, and it is the kind of job to point at a folder and leave running.
Can I do this on a phone?
Increasingly yes. Both major mobile platforms now recognise text in images on the device, and several free scanning apps produce a searchable PDF directly. The weak point on mobile is not the OCR, it is that phone file management makes a consistent archive harder to maintain — most people scan on the phone and file on a computer.
What about searching inside Word documents and emails?
Those already contain text, so they need no OCR and every desktop search tool reads them. The gap is almost always scans and photographs. Worth checking that your index actually covers the folder they live in, because the default indexed locations on Windows do not include every drive.
Where Keepsake fits
Keepsake is our product, so read this part with that in mind. Everything above is true whether or not you use it, and most of it you can do with a folder and an afternoon.
Keepsake does the OCR step on the device as documents come in, so scans are searchable without a separate tool or a separate pass — and because it runs locally, the same is true with no network at all. It also reads expiry dates and document numbers off the page and fills the tracker in from them, which is the part a text layer alone does not give you. The search index stays inside the encrypted vault rather than in an operating-system index, which is the privacy consequence described above, handled.
If what you want is searchable scans in a folder, OCRmyPDF plus the search you already have is a complete answer and costs nothing. This page would recommend it. Keepsake is aimed at the case where the documents also need to expire, be shared with a family, or survive you.