ENSLČlanek je na voljo samo v angleščini.Slovenske novice →Book a free audit

AI & automation

Run Local LLMs in n8n with Ollama (Llama 3 and Newer): 2026 Guide

Updated

Diagram titled Local LLM n8n Workflow: four connected workflow boxes, the third holding an n8n node icon.

To run a local LLM with n8n, install Ollama on a machine with enough memory, pull an open model such as Llama 3 (ollama pull llama3), and point n8n’s Ollama Chat Model node, or a plain HTTP Request node, at Ollama’s API on port 11434. Prompts and documents then never leave your own infrastructure. In exchange you take on hardware, maintenance and, for complex reasoning, a quality gap compared with the best cloud models.

This guide covers the setup, the Docker networking issue almost everyone hits, how to size the model, and how we decide between local and cloud models in client workflows.

Why run a local LLM with n8n?

Three reasons come up in practice:

  • The data can’t leave. Contracts, health or financial records, internal research: if a document shouldn’t go to a third-party model provider, a local model keeps it on your server.
  • Predictable cost. You pay for hardware and electricity instead of per-token fees. For steady, high-volume jobs such as classifying or summarising thousands of documents, that can be cheaper.
  • No dependency on an outside API. The workflow keeps running when a provider has an outage or changes its pricing.

The trade-off is quality and effort. Small open models are good at classification, extraction and summarising; they are weaker than frontier cloud models at long, multi-step reasoning. You also maintain the server yourself.

What do you need before you start?

  • Self-hosted n8n. n8n Cloud can only reach Ollama if you expose Ollama to the internet, which we don’t recommend. Run n8n on the same server or private network instead. Our n8n hosting guide compares the options.
  • Ollama, installed natively or in Docker.
  • Enough memory for the model. Ollama’s own guidance is roughly 8 GB of RAM for 7B models, 16 GB for 13B models and 32 GB for 33B models. A GPU with enough memory makes responses much faster; CPU-only works for small models and light loads, and Apple Silicon Macs run small models well.
  • Docker and Docker Compose, if you run n8n in a container (most self-hosted setups do).

n8n also publishes a self-hosted AI starter kit, a Docker Compose setup that bundles n8n with Ollama and a vector store, if you’d rather start from a template.

How do you connect n8n to Ollama, step by step?

  1. Install Ollama and pull a model. Run ollama pull llama3 (or a newer model; see below), then check it answers with ollama run llama3 "Say hello". The API listens on http://localhost:11434.
  2. Make Ollama reachable from n8n. This is where most setups fail:
    • n8n installed natively on the same machine: use http://localhost:11434.
    • n8n in Docker, Ollama on the host: use http://host.docker.internal:11434. On Linux, add extra_hosts: ["host.docker.internal:host-gateway"] to the n8n service, and start Ollama with OLLAMA_HOST=0.0.0.0 so it accepts connections from the container, not only from localhost.
    • Both in the same Docker Compose file: use the service name, for example http://ollama:11434.
  3. Add the credential in n8n. Create an Ollama credential with that base URL.
  4. Build the workflow. Add an AI Agent or a Basic LLM Chain node and attach the Ollama Chat Model sub-node, choosing your model from the list. For full control, use an HTTP Request node instead: POST /api/generate with a JSON body such as {"model": "llama3", "prompt": "Summarise this document: {{ $json.text }}", "stream": false}, or POST /api/chat with a messages array.
  5. Test on real input. Run a handful of real documents, check the answers and the response time, and only then connect the workflow to anything downstream.

For retrieval over your own documents, Ollama also serves embedding models (for example nomic-embed-text), so the whole retrieval pipeline can stay local. We explain how we keep retrieval answers honest in how we use RAG to stop AI hallucinations.

Which local model should you choose?

Start with the smallest model that does the job well on your own test set, because smaller models are faster and cheaper to run. Newer open models appear every few months, so treat any list as a starting point and test on your data.

Model family Typical sizes Good for Licence note
Llama 3 (Meta) 8B, 70B and newer variants General text, summarising, chat Meta’s own licence; read the terms for commercial use
Mistral 7B / Mixtral 8x7B 7B; 8x7B needs much more memory Fast general tasks; stronger reasoning with Mixtral Apache 2.0
Code models (e.g. DeepSeek Coder) Small to mid sizes Code generation and review Check the model’s licence
Embedding models (e.g. nomic-embed-text) Small Search and retrieval (RAG) Check the model’s licence

Quantised versions (the default for most models in Ollama) use far less memory with a modest quality loss, which is usually the right trade for automation jobs.

How do you make a local LLM fast enough for production?

  • Keep the model loaded. Ollama unloads idle models after a few minutes; set OLLAMA_KEEP_ALIVE longer for workflows that run all day, so the first request isn’t slow.
  • Control concurrency. Process items in batches with n8n’s Loop Over Items node, and tune OLLAMA_NUM_PARALLEL to what your GPU can handle, rather than firing hundreds of requests at once.
  • Keep prompts short. Long context windows cost memory and time. Send the part of the document the task needs, not the whole archive.
  • Put n8n and Ollama close together. Same machine or same private network.
  • Add retries and timeouts. Local models can stall under load; a retry with a timeout and an alert keeps a stuck item from blocking the queue.

When should you use a cloud model instead?

Most of our client workflows end up hybrid: a local model where the data is sensitive, a cloud model (OpenAI, Claude or Gemini, chosen per task) where quality matters more and the data is allowed to leave.

Situation Our default
Personal, health, financial or confidential data that shouldn’t go to a third party Local model
Fine-tuned or domain-specific model you trained yourself Local model
Complex multi-step reasoning, long documents, high-quality writing Cloud model
Prototyping before you know the volume Cloud model, then move steady jobs local
Customer-facing output Either, but always through a human review step

In n8n this is simple to build: an IF or Switch node checks a sensitivity flag and routes the item to the local or the cloud model. Whatever the model, the output goes to a review queue that a person empties, the rule behind all our AI automation work.

How do you keep a local LLM setup secure?

Ollama’s API has no authentication by default, so never expose port 11434 to the internet. Keep it on localhost or a private network, put n8n behind HTTPS with proper user management, and keep both updated. Our security guide for self-hosted n8n covers the rest.

FAQ

Can n8n Cloud use Ollama?

Only if Ollama is reachable from the internet, which means putting authentication and HTTPS in front of it yourself. For sensitive data, self-hosted n8n on the same server or network as Ollama is the cleaner setup.

Do I need a GPU to run Ollama with n8n?

No, but it helps a lot. Small models run on CPU or Apple Silicon for light workloads. For steady production volume or larger models, a GPU with enough memory to hold the model is the difference between seconds and minutes per item.

Does running a local LLM make us GDPR compliant?

It removes one data transfer, to the model provider, which helps. It doesn’t cover everything else: legal basis, retention, access control and logging still apply to the workflow as a whole.

Is Llama 3 still the right model in 2026?

It is a solid, well-supported default, but newer open models are released regularly. Keep a small test set of real inputs and expected outputs, and re-test when a new model appears; switching in Ollama is one pull and a dropdown change in n8n.

If you want help deciding which workflows belong on a local model, talk to us.

Related reading