IPv4 & IPv6 Leasing - Any RIR, Any LocationOrder Now
Hostperl

Debian Private AI Inference with Ollama and Nginx

By Raman Kumar

Share:

Updated on Oct 1, 2026

Debian Private AI Inference with Ollama and Nginx

Why this setup makes sense on Debian

Private AI inference on Debian works well when you want tight access control, readable logs, and a stable base for support. This tutorial builds a small but real deployment: Ollama listens only on localhost, Nginx exposes a protected HTTPS endpoint, and UFW limits inbound traffic to SSH, HTTP, and HTTPS.

The setup suits teams that need a private model endpoint for internal tools, support workflows, or app backends without sending prompts to a third-party API. If you are comparing infrastructure, a Hostperl VPS is a sensible place to start for light inference and staging, while heavier model work may need a dedicated server or GPU-capable platform.

This guide uses Debian 12 or Debian 13 because Debian keeps package behavior predictable, systemd service management straightforward, and networking tools familiar. It also matches the kind of setup many Hostperl customers run after a migration or a fresh VPS turnup.

What you will build

  • Ollama running as a systemd service on Debian
  • Nginx as a reverse proxy with request size and timeout controls
  • UFW rules that keep the model endpoint private
  • A non-root administrator account for day-to-day access
  • Verification commands, smoke tests, and rollback steps

This is not a broad AI app overview. The goal is a working endpoint you can hand to a team, monitor, and support without exposing the model port to the internet.

Prerequisites and server assumptions

Start with a fresh Debian 12 or Debian 13 VPS. The examples use 203.0.113.10 as the server IP, server.example.com as the hostname, and deploy as the non-root administrator. Replace 203.0.113.10 with your real public IP.

Model size matters. A small quantized model can run acceptably on modest VPS RAM, but larger models quickly need more memory and CPU headroom. If your workload includes concurrent users, long prompts, or larger context windows, plan for a dedicated server rather than a small VPS.

Connect to the server and identify the OS

On your local computer, connect with the documented example below.

ssh root@203.0.113.10

On the VPS as root, confirm the operating system before you touch packages or firewall rules.

cat /etc/os-release

You should see Debian 12 or Debian 13 in the output. If you use a non-standard SSH port, connect with the correct port first and keep the same server IP convention throughout the guide.

Create a non-root administrator and keep root open

Do not close your root session yet. Create the daily-use account, add sudo access, and verify that SSH keys work before you remove any root convenience.

On the VPS as root, create the user and grant sudo membership.

adduser deploy
usermod -aG sudo deploy

Now create the SSH directory and key file for the new account. If you already have a public key on your local computer, paste it into the file below.

mkdir -p /home/deploy/.ssh
chmod 700 /home/deploy/.ssh
nano /home/deploy/.ssh/authorized_keys

Paste one public key line, save the file, and exit nano. Then secure the ownership and permissions.

chown -R deploy:deploy /home/deploy/.ssh
chmod 600 /home/deploy/.ssh/authorized_keys

Open a second terminal from your local computer and test the new login.

ssh deploy@203.0.113.10

After login, verify sudo.

sudo -v

If this works, keep both sessions open. That gives you a safe fallback while you change network exposure later.

Update packages and set the hostname

On the VPS as root, refresh the package index and install the tools used in this tutorial.

apt update
apt -y upgrade
apt -y install curl ca-certificates gnupg ufw nginx

Set the hostname to something recognizable in logs and prompts.

hostnamectl set-hostname server.example.com

Check the result.

hostnamectl

You should see server.example.com in the static hostname field.

Install Ollama on Debian

Ollama does not ship from the standard Debian archive, so install it with the upstream script and then keep it behind localhost. That keeps the model port off the public interface, which matters on any public VPS.

On the VPS as root, install Ollama.

curl -fsSL https://ollama.com/install.sh | sh

Confirm the service exists and is running.

systemctl status ollama --no-pager

Check the listening address. On a correct setup, Ollama should bind locally, not to the public IP.

ss -lntp | grep 11434

If you see 127.0.0.1:11434 or localhost:11434, that is the safe state you want.

Pull a small model for testing. Pick one that suits your CPU and RAM budget.

ollama pull llama3.2:3b

Verify the model list.

ollama list

If the download stalls or fails, check network reachability and disk space before retrying.

Run Ollama as a controlled local service

In production, you want the service to start automatically and stay private. On Debian, Ollama usually installs its own service file, but you should still confirm the environment and binding behavior.

On the VPS as root, inspect the service details.

systemctl cat ollama

If the service is not already bound to localhost, create an override.

systemctl edit ollama

In the editor, add this content:

[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"

Save and exit, then reload systemd and restart Ollama.

systemctl daemon-reload
systemctl restart ollama
systemctl enable ollama

Confirm the service state and local binding again.

systemctl status ollama --no-pager
ss -lntp | grep 11434

Configure Nginx as the public entry point

Nginx will handle external traffic and proxy it to Ollama on localhost. That gives you TLS, request limits, and logs without exposing the model service directly.

On the VPS as root, create a dedicated site file.

nano /etc/nginx/sites-available/ollama.conf

Paste this configuration. Replace only the server name if your real hostname differs.

server {
    listen 80;
    listen [::]:80;
    server_name server.example.com;

    access_log /var/log/nginx/ollama.access.log;
    error_log /var/log/nginx/ollama.error.log;

    client_max_body_size 20m;
    proxy_read_timeout 300s;
    proxy_connect_timeout 60s;
    proxy_send_timeout 300s;

    location / {
        proxy_pass http://127.0.0.1:11434;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

Enable the site and disable the default if it is not needed.

ln -s /etc/nginx/sites-available/ollama.conf /etc/nginx/sites-enabled/ollama.conf
rm -f /etc/nginx/sites-enabled/default

Test the syntax before reload. This is where many support tickets begin, and it is cheap to catch here.

nginx -t

If the syntax is correct, reload Nginx.

systemctl reload nginx

Open the firewall safely

Open SSH first, then HTTP. If you later add TLS, HTTPS can be opened as well. Do not close your working SSH session until you have tested the new rules.

On the VPS as root, allow the ports.

ufw allow OpenSSH
ufw allow 80/tcp
ufw enable

Check the rules.

ufw status verbose

You should see SSH and HTTP allowed. If you are adding TLS after certificate issuance, add 443/tcp in the same controlled way.

Add HTTPS with Let’s Encrypt

For a private AI endpoint, TLS is still worth doing. It protects credentials, session cookies, and any proxy-auth headers you may add later.

On the VPS as root, install Certbot for Nginx.

apt -y install certbot python3-certbot-nginx

Issue the certificate.

certbot --nginx -d server.example.com

During the prompts, choose the redirect option so port 80 sends traffic to HTTPS. After the certificate is issued, test renewal logic.

certbot renew --dry-run

Check that Nginx now listens on 443.

ss -lntp | grep nginx

Test the API from the server and your workstation

First confirm local access on the VPS. This isolates reverse-proxy issues from network or DNS problems.

On the VPS as root, run a basic prompt against the local Ollama port.

curl http://127.0.0.1:11434/api/generate -d '{"model":"llama3.2:3b","prompt":"Say hello in one short sentence.","stream":false}'

Then test the public endpoint through Nginx. Use your hostname, not the raw model port.

curl -k https://server.example.com/api/generate -d '{"model":"llama3.2:3b","prompt":"Return the word ready.","stream":false}'

The -k flag is only for quick testing if your certificate chain is not yet trusted on the client. A normal browser or production client should verify TLS properly.

Check logs and confirm reboot persistence

Support teams need logs that separate proxy failures from model failures. Nginx access and error logs will tell you whether requests reached the proxy layer.

On the VPS as root, watch the logs.

tail -f /var/log/nginx/ollama.access.log /var/log/nginx/ollama.error.log

In another terminal, make a test request again and confirm you see a 200 response in the access log.

Next, confirm persistence across reboot. A clean production service should survive a restart without manual intervention.

systemctl is-enabled ollama
systemctl is-enabled nginx
reboot

After the server comes back, reconnect and verify the services again.

systemctl status ollama --no-pager
systemctl status nginx --no-pager
ufw status verbose

Rollback and recovery

If something breaks, revert in the same order you deployed it. First stop exposing the endpoint, then return to the local service state.

On the VPS as root, disable the Nginx site and reload safely.

rm -f /etc/nginx/sites-enabled/ollama.conf
nginx -t
systemctl reload nginx

If the proxy is not the problem, stop Ollama, fix the service override, and start again.

systemctl stop ollama
systemctl edit ollama

Remove the override content if you need to return to the package default, save, and then run:

systemctl daemon-reload
systemctl start ollama

For a full retreat, you can remove the firewall changes you added.

ufw delete allow 80/tcp
ufw delete allow 443/tcp
ufw status verbose

Troubleshooting the most likely failures

1. Ollama will not start
Diagnostic command:

journalctl -u ollama -n 50 --no-pager

Expected clue: a bad model path, missing permissions, or a port binding conflict. Fix the configuration, then restart with systemctl restart ollama.

2. Nginx returns 502 Bad Gateway
Diagnostic command:

ss -lntp | grep 11434
journalctl -u nginx -n 50 --no-pager

Expected clue: Ollama is down or not listening on localhost. Correct the Ollama service and re-test the local curl command before reloading Nginx.

3. HTTPS works locally but not from outside
Diagnostic command:

ufw status verbose
ss -lntp | grep ':443\|:80'

Expected clue: 80 or 443 is not allowed, or Nginx is only listening on localhost. Add the missing firewall rule and confirm the listen sockets.

4. The model request is too slow or times out
Diagnostic command:

tail -f /var/log/nginx/ollama.error.log

Expected clue: the response exceeds the proxy timeout or the model is too large for the available CPU and RAM. Reduce the model size, increase timeout values, or move the workload to stronger hardware.

When to move from VPS to dedicated hardware

Light internal tooling can work on a VPS, but repeated model calls, larger context windows, or several users at once can saturate memory and CPU quickly. That is the point where a dedicated server makes support easier and reduces noisy-neighbor risk. Hostperl’s dedicated server hosting is the better fit once your private AI endpoint becomes a business system instead of a test.

For teams building customer-facing AI features, keep the model port private, terminate TLS at the proxy, and monitor latency at the application edge. If your workload grows into GPU inference, the same access pattern still applies, but the underlying machine needs more capacity than a small VPS can provide.

If you want a private AI endpoint with predictable support and a clean handoff, Hostperl can help you size the platform before launch. Start with a Hostperl VPS for a small internal service, or move to dedicated server hosting when model load or concurrency increases.

That keeps the deployment simple to run, easier to secure, and easier for your team to support over time.

FAQ

Can I expose Ollama directly on the internet?
Technically yes, but you should not. Put Nginx in front of it, keep Ollama on localhost, and use TLS.

Does this work on Debian 13?
Yes, the package and systemd workflow remains the same for this setup, though you should always confirm repository behavior after a major release upgrade.

What model should I start with?
Start with a smaller quantized model such as llama3.2:3b so you can validate the stack before using heavier workloads.

How do I know if the VPS is too small?
If requests are slow, swap pressure rises, or the model fails under moderate concurrency, move to stronger CPU, more RAM, or dedicated hardware.

What should I monitor first?
Watch Ollama service status, Nginx error logs, request latency, and memory use. Those four signals usually explain early failures.