OCR in n8n: Extracting Receipts Without a Cloud Service

n8n OCR without the cloud: integrating Tesseract or Ollama Vision for receipts, including common issues from real-world practice.

Hand-drawn sketch: a magnifying glass hovers over a receipt with scan lines, the magnifying glass is colored teal.

Automatically extract receipts, invoices, and delivery notes without sending them to a cloud API: this is possible with OCR in n8n, either using the classic Tesseract engine or a local AI vision model via Ollama. For structured documents with a clean print layout, Tesseract delivers usable results; for poor scans or complex layouts, a vision model usually comes out ahead, but requires significantly more computing power. As of: August 2026.

What ways are there to do OCR in n8n?

n8n does not come with a built-in OCR node, so the path leads either through community nodes or through the Execute Command node combined with a command-line tool. The second option is the most common for self-hosted instances, because it doesn't require additional npm packages and can be built into your own Docker image.

How do you set up the Tesseract variant?

The core is the Execute Command Node, which runs shell commands on the machine where n8n is running. According to the documentation, this node is disabled by default for security reasons as of version 2.0 and is not available at all on n8n Cloud, because it poses a risk in environments with untrusted users. For a self-hosted instance with a fixed set of users, however, it can be enabled deliberately.

  • Write binary data to disk: The Read/Write Files from Disk Node saves the incoming PDF or image as a file before Tesseract can access it. On n8n Cloud, access is restricted to `/home/node/`; self-hosted, this area can be restricted or extended via an environment variable.
  • Convert PDF to image: For multi-page PDFs, Tesseract first needs an image format; `pdftoppm` from the poppler-utils package is suitable for this, installed into the n8n image via a custom Dockerfile.
  • Recognize text: The actual `tesseract` command reads the image and returns the recognized text. Tesseract is an open-source OCR engine that, according to the project description on GitHub recognizes text in more than 100 languages.
  • Read the result back in: A second Read/Write Files node reads the text file back in and passes it on to the next workflow steps.

What are the practical limits of Tesseract?

Tesseract reaches its limits with poorly scanned or complex-layout receipts. This is openly discussed in the n8n community: one user reports in the thread Local OCR in n8n with Ollama that Tesseract quality was simply poor for their documents, while vision models like minicpm-v or llava-llama3 perform better on dense, structured text. This observation comes from individual users' practical experience and is not a universally valid benchmark, but it matches the known weakness of classic OCR engines with unclean templates.

What is the Ollama vision alternative, and where does it struggle?

Instead of Tesseract, a local vision model can be used via Ollama, which can describe images directly and extract text from them. According to community reports, this approach also has practical pitfalls that are rarely documented: large multi-page PDFs exceed the model's context window and cause the Ollama container to stall when all pages are sent at once. The common solution is to process each page individually and insert a Wait node between requests so the local container isn't overloaded. For both variants: anyone who frequently uses `/tmp` or similar paths for temporary files should check whether the environment variable restricting file access allows this.

Is the effort worth it compared to a cloud OCR API?

The effort is worthwhile above all if receipts should not be sent to an external provider for GDPR or confidentiality reasons. If you don't have this requirement, a cloud OCR API often gets you to stable results faster, since model maintenance and scaling are no longer your concern. For companies that already self-host their automation and want to keep control over their data, the local solution is the logical next step. If you want an n8n instance set up on a clean foundation for workflows like this, you can find background information at /leistungen/n8n.

Frequently asked questions about OCR in n8n

Do I necessarily need the Execute Command node for OCR in n8n?

No, there are also community nodes like n8n-nodes-tesseractjs that wrap Tesseract directly as a node, without requiring shell commands. However, the Execute Command approach is more flexible, since it can be combined with any command-line tool, not just Tesseract.

Does OCR with Tesseract work on n8n Cloud?

No, the Execute Command node is fundamentally unavailable on n8n Cloud, and file access is limited to a restricted path. For this use case, you need a self-hosted n8n instance with your own Docker image.

Is a vision model via Ollama always the better choice over Tesseract?

Not always, because vision models require significantly more computing power and are also not error-free with very poor scans. For clearly printed, clean documents, Tesseract is often sufficient; for dense or unclean templates, practitioners report better results with vision models.

How do I handle multi-page PDFs without overloading the local server?

The proven approach in practice is to process each page individually instead of sending the entire document at once, and to insert short wait times between processing steps. This prevents the Ollama container or the Tesseract process from being blocked by too many simultaneous requests.

Simon Glowik, founder of NordFlux
About the author

Founder of NordFlux. Spent four years automating processes at enterprise scale at Dräger, and now brings that depth to the mid-market — pragmatic and with full data sovereignty.

Certifications

  • Microsoft certified — PL-900 and AZ-900
  • UiPath certified — Automation Developer Associate
  • UiPath zertifiziert — Automation Developer Associate
All articles
Free initial call

Concrete questions about automation or AI?

In a free 30-minute initial call we discuss your case directly. No strings attached.