How do I get my documents out of paperless-ngx?
How do I get my documents out of paperless-ngx?
paperless-ngx includes a document exporter that writes your original files into a folder along with a manifest.json describing every document, tag, correspondent and document type. That folder is a complete, self-contained copy of your archive: it can be restored into another paperless-ngx instance, read by another program, or simply kept as the backup that proves you are not locked in.
paperless-ngx is self-hosted, open source, and the strongest answer available to the question "what happens to my documents if the software goes away". This guide is about the part that makes that true: the exporter, what it writes, and what another program can do with it.
If you are running paperless-ngx happily, you do not need to move anywhere. Understanding the export is still worth twenty minutes, because an export you have never produced is a plan rather than a backup.
The exporter
paperless-ngx ships a management command, document_exporter, which writes a target directory containing your document files and a manifest.json beside them. It is run through the same mechanism as the rest of the project's management commands — inside the container for a Docker installation, or through the project's script for a bare-metal one. The project's own documentation is the place to confirm the exact invocation for your setup, because it depends on how you installed it and that changes over time.
Two options are worth knowing about whatever your setup: one that splits the manifest into per-document files, and one that uses the documents' original file names rather than internal identifiers. Both make the export easier for a human and for another program to read.
What the manifest actually contains
The manifest is a list of Django fixtures — the same format the project uses to load a database. Each entry is an object with a model naming what kind of record it is and a fields object holding the record itself:
[
{ "model": "documents.tag", "pk": 3,
"fields": { "name": "insurance" } },
{ "model": "documents.correspondent", "pk": 7,
"fields": { "name": "Aviva" } },
{ "model": "documents.document", "pk": 42,
"fields": { "title": "Car insurance 2026",
"filename": "2026/car-insurance.pdf",
"created": "2026-03-31",
"correspondent": 7, "tags": [3],
"content": "...the full OCR text..." } }
]The important structural point is the last one: documents refer to their tags, correspondents and document types by number, not by name. A program reading the manifest has to collect the tag and correspondent records first and then resolve the references, or it ends up importing documents tagged "3" and "7". That is the single thing most hand-written scripts against this format get wrong.
The content field, and why we leave it alone
Every document record carries a content field holding the full OCR text of the document — often several pages of it. It is genuinely useful data and it is the wrong thing to copy into a notes field.
Notes are a place a person writes a sentence they will want later. Four pages of machine-read text pasted into one turns every document's notes into something nobody will ever open, and pushes the one sentence that mattered off the bottom. So Keepsake's importer deliberately does not copy it: the document itself is read on import, so the text is searchable anyway, and the notes stay a place for a sentence.
Worth knowing regardless of what you import into: that field means your manifest contains the readable text of every document you own. It is not a file to leave in a shared folder or attach to a support ticket.
How the fields land
| paperless-ngx | Becomes |
|---|---|
| title | The document title, unchanged. |
| created | The issue date. |
| correspondent | The issuing authority, resolved from its record. |
| document type | The category, where the name is one we hold; otherwise the document is filed under Other and the type is kept as a tag. |
| tags | Tags, resolved from their records. |
| archive serial number | The document number. |
| content | Nothing. See above. |
paperless-ngx has no expiry-date concept, so nothing fills that in and no import invents one. Expiry dates are the thing you will be adding by hand afterwards, and they are the reason to bother: they are what a reminder is built from.
Step by step
- Run the exporter. Follow the paperless-ngx documentation for your installation type, exporting into an empty directory. Use the option that writes the original file names if your version offers it.
- Check the manifest is there and the files are beside it. A folder with a manifest.json and your documents in it is a complete export. Open the manifest in a text editor and confirm it starts with a list of records.
- Keep a copy of the export as a backup. Whatever you do next, you now have a portable copy of the archive that does not need paperless-ngx running to be readable. Treat it as sensitive: the manifest holds the OCR text of every document.
- Import the folder. Point the importer at the folder. The manifest is recognised by its contents rather than its name, because manifest.json is one of the commonest file names in software.
- Add expiry dates to the handful that need them. Nothing in the export carries one. Passports, licences, insurance and visas are the four groups worth doing, and it is a ten-minute job for most archives.
Questions
Is the export a real backup?
It is a portable copy of your documents and their metadata, which is most of what people mean. It is not a byte-for-byte backup of the paperless-ngx installation — the project documents a separate procedure for that, including the database and the search index. For the purpose of "could I read these files if the software stopped existing", the export is the thing that answers yes.
Do I have to leave paperless-ngx to read the export?
No, and if paperless-ngx suits you, staying is a perfectly good decision — it is one of the few products in this space that genuinely cannot lock you in. Producing an export occasionally is worth doing anyway, and being able to read it somewhere else is the point of it.
What happens to documents whose type Keepsake does not have?
They are filed under Other with the paperless-ngx document type kept as a tag. No guessing: a document filed under an invented category is one you cannot find, because you will not think to look for it there.
Does anything get uploaded during the import?
No. The folder is read on the device you are using, against a vault already unlocked there. Keepsake never connects to a paperless-ngx server and asks for no credential of yours.
What if the manifest names files that are not in the folder?
Each missing file is listed at the end of the import with that reason. A partial export is a thing that happens — a disk filling up mid-run, a copy interrupted — and the useful response is a named list, not silence.
Where Keepsake fits
Keepsake is our product, so read this part with that in mind. Everything above is true whether or not you use it, and most of it you can do with a folder and an afternoon.
Keepsake reads a paperless-ngx export folder directly: the manifest is parsed, the tag and correspondent references are resolved, and the documents are filed with their titles, dates and tags intact. On Windows, Android and the web app, from one shared definition of the mapping.
We are a different shape of product and it is worth saying plainly: paperless-ngx is self-hosted, open source and free, and if you want a server you control that OCRs everything you feed it, it is better at that than we are. Keepsake is for the household case — nothing to host, works on a phone, tracks what expires, and has a plan for somebody else opening the vault when you cannot. Some people run both, with paperless-ngx as the archive and Keepsake holding the documents that get carried and renewed.