Host an LLM API on Ubuntu Server with Nginx and UFW

What you are building
This tutorial shows you how to place a private LLM API on Ubuntu Server behind Nginx, with UFW, systemd, and a small model-serving process that listens only on localhost. The result is a setup you can use for internal tools, client portals, or a controlled AI endpoint on a Hostperl VPS.
The backend stays off the public internet. Nginx handles TLS termination and request limits, UFW exposes only the web ports, and systemd keeps the service running after reboots. If the workload grows beyond a small test model, you can move the same layout to a larger Hostperl VPS or a dedicated server for more RAM, faster NVMe, and steadier latency.
For teams comparing hosting options, this pattern fits well on Hostperl VPS hosting for a pilot deployment, or on dedicated server hosting when the model, cache, and logs need more memory headroom.
Before you start
This guide assumes Ubuntu Server 24.04 LTS or a current Ubuntu Server release that still uses apt, systemd, UFW, AppArmor, and netplan. It is written for a fresh server. You will start as root, create a non-root administrator, then continue from that account. Keep the original root session open until the new login is verified.
You will also need a domain name pointed at the server. In the examples below, the hostname is server.example.com and the domain is example.com. Replace them with your own values.
Connect to the server and confirm the OS
On your local computer:
ssh root@203.0.113.10203.0.113.10 is a reserved documentation address. Replace it with the real public IP assigned to your server. If your provider gives you a different default SSH user, connect with that account instead and keep the same IP.
After you log in, confirm the operating system.
On the VPS as root:
cat /etc/os-releaseYou should see Ubuntu in the output. If the server is not Ubuntu, stop here and follow an Ubuntu-compatible path only.
Create a non-root administrator
Do not manage the server as root after the first login. Create a sudo user named deploy, give it SSH key access, and test it in a second terminal before you change any login rules.
On the VPS as root:
adduser deploySet a strong password when prompted. Then grant sudo access.
usermod -aG sudo deployCreate the SSH directory and copy your public key. Replace the example key path with the key on your local computer.
mkdir -p /home/deploy/.ssh
chmod 700 /home/deploy/.ssh
cat /root/.ssh/authorized_keys > /home/deploy/.ssh/authorized_keys
chmod 600 /home/deploy/.ssh/authorized_keys
chown -R deploy:deploy /home/deploy/.sshIf your root account does not already have the key, add it first from your local machine with ssh-copy-id deploy@203.0.113.10 after the user exists. Then verify the new account in a second terminal.
On your local computer:
ssh deploy@203.0.113.10Then confirm sudo works.
On the VPS as the non-root sudo user:
sudo -v
whoami
pwdYou should see deploy from whoami and no errors from sudo -v. Keep the root session open until this step succeeds.
Update the server and install the required packages
Use apt to refresh package lists and install the web server, firewall tools, and Python runtime needed for the example model API wrapper.
On the VPS as the non-root sudo user:
sudo apt update
sudo apt -y upgrade
sudo apt -y install nginx python3 python3-venv python3-pip ufw curlCheck that the main packages are present.
nginx -v
python3 --version
ufw statusAt this point, Nginx should be installed but not yet serving your domain, and UFW will likely still be inactive.
Set hostname, time sync, and a working directory
Give the server a clear name so logs and prompts are easier to read. Then create the application directory.
On the VPS as the non-root sudo user:
sudo hostnamectl set-hostname server.example.com
sudo mkdir -p /opt/myapp
sudo chown -R deploy:deploy /opt/myapp
cd /opt/myapp
pwdReplace server.example.com with your real hostname. The final pwd should print /opt/myapp.
Ubuntu usually keeps time with systemd-timesyncd or a similar service. Check it now.
timedatectl statusYou want NTP to be active. Accurate time matters for TLS certificates, logs, and token expiry.
Build a private LLM API wrapper
For this tutorial, the model server is a local Python service that exposes a simple chat endpoint. In production, you would connect this same reverse-proxy pattern to a real inference engine such as Ollama, vLLM, llama.cpp server mode, or a vendor runtime that binds to localhost.
Here, you will create a minimal service that behaves like a private LLM API. It is enough to prove the network, firewall, Nginx, and systemd flow without exposing a real model port to the internet.
On the VPS as the non-root sudo user:
cd /opt/myapp
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pipCreate the application file.
On the VPS as the non-root sudo user:
cat > /opt/myapp/app.py <<'EOF'
from http.server import BaseHTTPRequestHandler, HTTPServer
import json
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path == "/health":
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(b'{"status":"ok"}')
return
self.send_response(404)
self.end_headers()
def do_POST(self):
if self.path != "/v1/chat":
self.send_response(404)
self.end_headers()
return
length = int(self.headers.get("Content-Length", 0))
body = self.rfile.read(length)
try:
payload = json.loads(body.decode("utf-8"))
except Exception:
self.send_response(400)
self.end_headers()
return
prompt = payload.get("prompt", "")
reply = {
"model": "local-private-demo",
"reply": f"You said: {prompt}",
}
data = json.dumps(reply).encode("utf-8")
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(data)))
self.end_headers()
self.wfile.write(data)
httpd = HTTPServer(("127.0.0.1", 8000), Handler)
httpd.serve_forever()
EOFThis example binds to 127.0.0.1:8000 only. That is the right behavior for a private API behind Nginx.
Create a systemd service for the API
Run the service automatically at boot and keep its logs in journald. Create the unit file carefully, then validate syntax before starting it.
On the VPS as the non-root sudo user:
sudo tee /etc/systemd/system/private-llm-api.service > /dev/null <<'EOF'
[Unit]
Description=Private LLM API Demo
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=deploy
Group=deploy
WorkingDirectory=/opt/myapp
Environment="PATH=/opt/myapp/.venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
ExecStart=/opt/myapp/.venv/bin/python /opt/myapp/app.py
Restart=on-failure
RestartSec=3
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=full
ProtectHome=true
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl start private-llm-api
sudo systemctl enable private-llm-apiCheck the unit status and listening socket.
systemctl status private-llm-api --no-pager
ss -lntp | grep 8000You should see the service active and bound only to 127.0.0.1:8000.
Test the local API before adding Nginx
Always test the backend directly first. That keeps troubleshooting simple if the reverse proxy misbehaves later.
On the VPS as the non-root sudo user:
curl -s http://127.0.0.1:8000/health
curl -s -X POST http://127.0.0.1:8000/v1/chat -H 'Content-Type: application/json' -d '{"prompt":"hello"}'The first command should return JSON with status equal to ok. The second should return a reply that echoes your prompt.
Configure Nginx as the public front end
Next, place Nginx in front of the service. It will answer on ports 80 and 443, forward API traffic to the local process, and keep the backend private.
Create a server block for your domain.
On the VPS as the non-root sudo user:
sudo tee /etc/nginx/sites-available/private-llm-api > /dev/null <<'EOF'
server {
listen 80;
server_name example.com server.example.com;
access_log /var/log/nginx/private-llm-api.access.log;
error_log /var/log/nginx/private-llm-api.error.log;
location /health {
proxy_pass http://127.0.0.1:8000/health;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
location /v1/ {
client_max_body_size 2m;
limit_req zone=api burst=10 nodelay;
proxy_pass http://127.0.0.1:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
location / {
return 404;
}
}
EOF
sudo ln -sfn /etc/nginx/sites-available/private-llm-api /etc/nginx/sites-enabled/private-llm-apiNginx needs a small rate-limit zone in the main configuration. Open the file and add it inside the http { } block if it is not already present.
On the VPS as the non-root sudo user:
sudo nano /etc/nginx/nginx.confAdd this line inside the http block, then save and exit with Ctrl+O, Enter, and Ctrl+X.
limit_req_zone $binary_remote_addr zone=api:10m rate=5r/s;Test the configuration before reloading Nginx.
sudo nginx -t
sudo systemctl reload nginxIf nginx -t reports success, Nginx will reload cleanly. If it fails, fix the error before continuing.
Open only the required firewall ports
UFW should allow SSH, HTTP, and HTTPS. Add the rules before enabling the firewall so you do not lock yourself out.
On the VPS as the non-root sudo user:
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw --force enable
sudo ufw status verboseExpect to see rules for OpenSSH, 80/tcp, and 443/tcp. Do not expose port 8000. That port stays local only.
Add TLS with Let's Encrypt
Once DNS points example.com and server.example.com at the server, issue a certificate with Certbot. If DNS is not correct yet, stop and fix the records first.
On the VPS as the non-root sudo user:
sudo apt -y install certbot python3-certbot-nginx
sudo certbot --nginx -d example.com -d server.example.comChoose the redirect option when prompted so HTTP goes to HTTPS automatically. After issuance, verify the certificate and the new Nginx config.
sudo nginx -t
sudo systemctl reload nginx
sudo certbot renew --dry-runA successful dry run confirms renewal will work later without manual intervention.
Verify the full path from client to backend
Now test the public endpoint from your local computer. Use HTTPS, not the local loopback address.
On your local computer:
curl -s https://example.com/health
curl -s -X POST https://example.com/v1/chat -H 'Content-Type: application/json' -d '{"prompt":"production check"}'You should get the health JSON and a reply that includes production check. If that works, the proxy chain is healthy.
Check service and port status on the server too.
On the VPS as the non-root sudo user:
systemctl status private-llm-api --no-pager
systemctl status nginx --no-pager
ss -lntp | grep -E ':(80|443|8000)'
journalctl -u private-llm-api -n 20 --no-pagerThis should show the backend on 8000, Nginx on 80/443, and no unexpected errors in the journal.
Make the deployment restart safely
Reboot persistence is part of production readiness. Reboot the server and confirm the API returns after boot.
On the VPS as the non-root sudo user:
sudo rebootAfter the server comes back, reconnect and check both services.
On your local computer:
ssh deploy@203.0.113.10Then run:
On the VPS as the non-root sudo user:
systemctl is-active private-llm-api
systemctl is-active nginx
curl -s https://example.com/healthAll three commands should succeed. If the service is inactive, review the journal first.
Rollback and recovery
If you need to back out the deployment, stop the service and remove the Nginx site before changing anything else.
On the VPS as the non-root sudo user:
sudo systemctl stop private-llm-api
sudo systemctl disable private-llm-api
sudo rm -f /etc/nginx/sites-enabled/private-llm-api
sudo nginx -t
sudo systemctl reload nginxThis returns the server to a state where Nginx is still running, but the private API path is gone. Your data in /opt/myapp remains intact unless you remove it manually.
If you need to restore quickly, reverse the commands: recreate the symlink, start the systemd unit, and retest the backend locally before exposing it again.
Common problems and fixes
Nginx returns 502 Bad Gateway
Check whether the backend is listening on 127.0.0.1:8000.
systemctl status private-llm-api --no-pager
ss -lntp | grep 8000
journalctl -u private-llm-api -n 50 --no-pagerIf the service crashed, fix the Python file or the systemd unit, then restart it with sudo systemctl restart private-llm-api.
Certbot fails to issue a certificate
Check DNS and port 80 reachability.
dig +short example.com
curl -I http://example.comIf DNS is wrong, update the A record. If port 80 is blocked, confirm UFW and any upstream firewall allow it.
UFW is active but SSH stopped working
This usually means the SSH rule was missing or removed too early.
sudo ufw status numbered
sudo ufw allow OpenSSHKeep the original root session open until the new user and firewall rules are confirmed.
Why this pattern works for private AI hosting
A private inference endpoint does not need to expose its model port. Nginx gives you TLS, logging, request limits, and a stable public URL. systemd gives you restart control and boot persistence. UFW keeps the server attack surface small.
For teams that later move from a small VPS to larger hardware, the same layout applies. You can keep Nginx, UFW, and systemd unchanged while swapping the backend to a higher-throughput engine on a Hostperl VPS or a self-managed dedicated server when memory, CPU, or GPU demand increases.
For related operational reading, see RAG application hosting on Ubuntu Server with PostgreSQL, private AI API security on VPS, and Nginx logging for faster app troubleshooting in 2026 for more operational context.
If you want to run a private LLM API on infrastructure you can support, Hostperl VPS hosting is a practical starting point for small teams, pilots, and internal tools. When the workload grows, Hostperl VPS hosting and dedicated server hosting give you room to scale without changing the deployment pattern.
For launch planning, monitoring, and migration help, Hostperl’s support-led hosting model is a better fit than a DIY stack that leaves you on your own after the first deploy.
FAQ
Can I use a real model server instead of the demo API?
Yes. Replace app.py with your actual inference process, as long as it binds to 127.0.0.1 or a private socket and Nginx remains the public entry point.
Why keep port 8000 closed to the internet?
Because the model server should not be directly exposed. Nginx adds TLS, logging, and traffic control, and UFW keeps the backend private.
Can this run on a small VPS?
It can run for testing or light internal use. Real model serving usually needs more memory, faster storage, and sometimes GPU capacity, which is where a larger VPS or dedicated server becomes a better fit.
What should I check first if the API stops working after a reboot?
Run systemctl status private-llm-api, then journalctl -u private-llm-api -n 50. If Nginx loads but the backend does not, the journal usually shows the cause.
