IPv4 & IPv6 Leasing - Any RIR, Any LocationOrder Now
Hostperl

Windows Server RAG Hosting with pgvector and IIS

By Raman Kumar

Share:

Updated on Oct 6, 2026

Windows Server RAG Hosting with pgvector and IIS

What this tutorial builds

Windows Server RAG hosting works well when you need a private chatbot or internal knowledge assistant on a Windows stack, but want the retrieval layer in PostgreSQL with pgvector. In this guide, you will set up Windows Server 2022 or Windows Server 2025, install PostgreSQL with vector search, publish a FastAPI app on IIS as a reverse proxy, and verify the full request path from browser to database.

This tutorial assumes a fresh server. You will start at first login, create a non-administrator account, tighten the firewall, deploy the app, and test recovery before production. If you run Hostperl VPS or Hostperl dedicated server infrastructure, this is the kind of controlled rollout our support teams expect before a cutover.

Hostperl VPS is a practical choice for smaller private RAG deployments, while a dedicated server makes more sense once your document corpus, worker queue, and database need more memory or steadier IOPS.

Architecture and prerequisites

The stack in this tutorial stays deliberately simple:

  • Windows Server 2022 or Windows Server 2025
  • PostgreSQL 17 on Windows
  • pgvector extension for similarity search
  • Python 3.12 for the FastAPI service
  • IIS with URL Rewrite and Application Request Routing acting as a reverse proxy

This split keeps the LLM or embedding model out of the web tier. It also gives you a cleaner rollback path if the API needs a restart or the database needs a restore. For a broader deployment pattern, compare this with our RAG application hosting on Ubuntu Server with PostgreSQL and our RAG vector database hosting guide.

1) Connect to the server

On your local computer, open PowerShell or Terminal and connect with the default documentation IP shown below:

ssh root@203.0.113.10

Replace 203.0.113.10 with the real public IP assigned to your Hostperl server. If your access model uses a non-admin account first, connect with that account instead, but keep the same server IP.

After you log in, open a second terminal on your computer and keep the original root or administrator session open until the new account is verified.

2) Detect the operating system and update Windows

On the VPS as root, confirm the platform and build level before you change anything else. On Windows Server, use PowerShell rather than Linux commands.

Get-ComputerInfo | Select-Object WindowsProductName, WindowsVersion, OsHardwareAbstractionLayer, OsBuildNumber

This confirms whether you are on Windows Server 2022 or 2025. You should see the product name and build number in the output.

Install-WindowsUpdate -MicrosoftUpdate -AcceptAll -AutoReboot

This command works if the PSWindowsUpdate module is already present. If it is not, install it first with:

Set-ExecutionPolicy RemoteSigned -Scope Process -Force; Install-Module PSWindowsUpdate -Force; Import-Module PSWindowsUpdate

Then rerun the update command. Reboot if Windows asks for it, and reconnect before continuing. A current patch level matters here because IIS, TLS, and PowerShell modules are easier to support when the OS is fully updated.

3) Create a non-administrator deployment account

Do not run the service day to day from the built-in administrator. Use a dedicated local account for the app and a separate admin account for maintenance.

On the VPS as root, create a local user called deploy and add it to the Administrators group:

New-LocalUser -Name deploy -Password (Read-Host -AsSecureString "Enter a strong password for deploy")

That command prompts you for a password without echoing it. Then add the account to the admin group:

Add-LocalGroupMember -Group Administrators -Member deploy

Open a second remote session and test the account before you close the original one.

On your local computer, open a new PowerShell session and connect as deploy using your normal remote access method, such as RDP or your provider console. Then verify group membership on the server:

whoami /groups

You should see the Administrators group listed. Keep the original root or admin session open until this succeeds.

4) Prepare directories and a working layout

On the VPS as the non-root sudo user, create a predictable layout for the RAG service, database assets, and logs. In Windows Server, that means normal NTFS paths rather than Linux-style directories.

New-Item -ItemType Directory -Force -Path C:\opt\rag-app, C:\opt\rag-app\logs, C:\opt\rag-app\data, C:\opt\rag-app\models

This creates a clean application tree under C:\opt\rag-app. You will place the FastAPI app, a local configuration file, and any cached document artifacts there.

Check the paths exist:

Get-ChildItem C:\opt\rag-app

You should see the four directories you just created.

5) Install PostgreSQL and pgvector

For Windows Server RAG hosting, PostgreSQL stores your document chunks and embeddings. pgvector adds similarity search without forcing you into a separate vector service.

On the VPS as the non-root sudo user, install PostgreSQL 17 from the official Windows installer or your standard software deployment method. After installation, verify the service and version:

Get-Service postgresql* | Format-Table -AutoSize

You should see the PostgreSQL service in the Running state once setup completes.

"C:\Program Files\PostgreSQL\17\bin\psql.exe" --version

This confirms the client tools are installed. Adjust the path if your PostgreSQL version differs.

Next, open postgresql.conf. The file path is usually under the PostgreSQL data directory, for example C:\Program Files\PostgreSQL\17\data\postgresql.conf. Add or confirm these settings:

listen_addresses = 'localhost'
shared_buffers = 1GB
max_connections = 100

These values keep the database local to the machine, which is the safer default for a private RAG app. Save the file, then open pg_hba.conf and make sure local trust is not used for production. A safer minimum is:

# TYPE  DATABASE  USER  ADDRESS  METHOD
local   all       all            scram-sha-256
host    all       all   127.0.0.1/32  scram-sha-256

Restart PostgreSQL after editing the files:

Restart-Service postgresql-x64-17

Now create the database and extension.

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d postgres -c "CREATE DATABASE ragdb;"

Then enable pgvector in that database:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE EXTENSION IF NOT EXISTS vector;"

Confirm the extension is present:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "\dx"

You should see vector in the extension list. For a deeper database performance path, our PostgreSQL tuning article covers maintenance tradeoffs that matter once your chunk table grows.

6) Create the schema for documents and embeddings

Now create tables for document chunks and metadata. Keep them small and explicit.

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE TABLE documents (id bigserial PRIMARY KEY, source text NOT NULL, chunk text NOT NULL, embedding vector(1536) NOT NULL, created_at timestamptz NOT NULL DEFAULT now());"

That example uses a 1536-dimensional embedding, which matches many common embedding models. If your model uses a different size, change the column definition before you load data.

Add an index for similarity search:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE INDEX documents_embedding_idx ON documents USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);"

Check the table and index exist:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "\d documents"

If you see the column list and the ivfflat index, the retrieval layer is ready.

7) Install Python and build the RAG API

On the VPS as the non-root sudo user, install Python 3.12 and the runtime packages you need. Use the Windows installer if Python is not already present, then confirm it works:

python --version

You should see Python 3.12.x or a compatible version. Next, create a virtual environment and install the app dependencies.

cd C:\opt\rag-app
python -m venv .venv
C:\opt\rag-app\.venv\Scripts\pip.exe install fastapi uvicorn psycopg[binary] pgvector python-dotenv

That command installs the web framework, ASGI server, PostgreSQL driver, pgvector integration, and environment file support.

Create the application file:

notepad C:\opt\rag-app\app.py

Paste the following content, then save and close Notepad:

from fastapi import FastAPI
from pydantic import BaseModel
import os
import psycopg
from pgvector.psycopg import register_vector

app = FastAPI()
DB_DSN = os.getenv("DB_DSN", "dbname=ragdb user=postgres password=CHANGE_ME host=127.0.0.1 port=5432")

class Query(BaseModel):
    question: str

@app.get("/health")
def health():
    return {"status": "ok"}

@app.post("/ask")
def ask(payload: Query):
    with psycopg.connect(DB_DSN) as conn:
        register_vector(conn)
        with conn.cursor() as cur:
            cur.execute(
                "SELECT source, chunk FROM documents ORDER BY embedding <=> %s::vector LIMIT 3",
                ([0.0] * 1536,)
            )
            rows = cur.fetchall()
    return {"question": payload.question, "matches": [{"source": r[0], "chunk": r[1]} for r in rows]}

This is a minimal retrieval service. In production, you would replace the zero vector placeholder with an actual embedding generated by your model pipeline.

Test the syntax before you go further:

C:\opt\rag-app\.venv\Scripts\python.exe -m py_compile C:\opt\rag-app\app.py

If the command returns no output, the file compiles cleanly.

8) Add environment settings and secure the file

Create a local environment file that stores the database connection string outside the source code.

notepad C:\opt\rag-app\.env

Use this content, then save and close the editor:

DB_DSN=dbname=ragdb user=postgres password=CHANGE_ME host=127.0.0.1 port=5432

Replace CHANGE_ME with the real database password you set for PostgreSQL. Then restrict access to the app directory so only the service account and administrators can read it.

icacls C:\opt\rag-app /inheritance:r
icacls C:\opt\rag-app /grant deploy:(OI)(CI)F Administrators:(OI)(CI)F

Those ACLs remove inherited permissions and grant full control only to deploy and Administrators. That matters because the database password should not be readable by casual users or shared service accounts.

9) Create the Windows service for the API

You need the RAG API to start on boot and restart on failure. One straightforward production approach on Windows Server is to run the FastAPI app under a service wrapper.

On the VPS as the non-root sudo user, create a startup script:

notepad C:\opt\rag-app\start-api.ps1

Paste this content and save it:

Set-Location C:\opt\rag-app
$env:PYTHONUTF8 = "1"
.\.venv\Scripts\uvicorn.exe app:app --host 127.0.0.1 --port 8000

Check the script file exists:

Get-Content C:\opt\rag-app\start-api.ps1

Now create a scheduled task or service wrapper. If you use NSSM, the service can be installed like this:

nssm install rag-api C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe

Set the arguments to:

-ExecutionPolicy Bypass -File C:\opt\rag-app\start-api.ps1

Then start the service:

Start-Service rag-api

Verify it is running:

Get-Service rag-api

If you prefer Task Scheduler instead of NSSM, keep the same script and schedule it at startup. The important part is unchanged: the API must start automatically after a reboot.

10) Configure IIS as a reverse proxy

IIS will accept external HTTPS traffic and forward it to the local FastAPI service on port 8000. That keeps the Python app off the public interface.

On the VPS as the non-root sudo user, install the IIS role and management components if they are not already present:

Install-WindowsFeature Web-Server, Web-Http-Redirect, Web-Filtering, Web-Request-Monitor, Web-Mgmt-Tools

Then add URL Rewrite and Application Request Routing. After installation, confirm IIS responds locally:

Invoke-WebRequest http://localhost -UseBasicParsing

You should receive a valid HTTP response from IIS.

Create or edit the IIS site bindings in the IIS Manager GUI, or use PowerShell if you manage bindings as code. For the reverse proxy rule, save this rewrite configuration in the site’s web.config:

<configuration>
  <system.webServer>
    <rewrite>
      <rules>
        <rule name="RAGProxy" stopProcessing="true">
          <match url="(.*)" />
          <action type="Rewrite" url="http://127.0.0.1:8000/{R:1}" />
        </rule>
      </rules>
    </rewrite>
  </system.webServer>
</configuration>

After saving the file, recycle the site or app pool. Then test the local FastAPI endpoint directly before you expose it through IIS:

Invoke-WebRequest http://127.0.0.1:8000/health -UseBasicParsing

You should get a JSON response that includes status: ok.

11) Open the firewall safely

Do not expose the Python service port to the internet. Only open the public web ports you actually need.

On the VPS as the non-root sudo user, allow IIS traffic through Windows Firewall:

New-NetFirewallRule -DisplayName "IIS HTTP" -Direction Inbound -Protocol TCP -LocalPort 80 -Action Allow
New-NetFirewallRule -DisplayName "IIS HTTPS" -Direction Inbound -Protocol TCP -LocalPort 443 -Action Allow

If you are still testing locally and have not installed TLS yet, keep only port 80 open for the moment. Once you confirm the site is working, add HTTPS and then remove any temporary exposure you no longer need.

Check the current rules:

Get-NetFirewallRule -DisplayName "IIS *" | Format-Table -AutoSize

12) Add TLS with a real certificate

For production, terminate HTTPS at IIS. You can use a trusted certificate from your certificate workflow or a publicly trusted CA. Bind it in IIS after the certificate is imported into the Local Computer certificate store.

After you install the certificate, assign it to the HTTPS site binding in IIS Manager. Then test with PowerShell from the server:

Invoke-WebRequest https://localhost -UseBasicParsing

From your local computer, test the public endpoint as well. Replace the hostname with your own domain that points to the server.

curl -I https://server.example.com

A successful response should show 200 or a similar valid status and a certificate that matches your domain.

If you are planning DNS and certificate work at the same time, our DNS, SSL, and email guide is a useful companion before cutover day.

13) Smoke test the full RAG path

Now test the app through the public web path, not just on localhost. First hit the health endpoint.

curl https://server.example.com/health

You should see a JSON response with status set to ok. Then try the answer endpoint with a small JSON payload.

curl -X POST https://server.example.com/ask -H "Content-Type: application/json" -d '{"question":"What documents are indexed?"}'

You should receive a JSON object with your query text and a matches array. If the array is empty, the service is running but your retrieval table has no loaded data yet.

Load one sample document to confirm the search path works:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "INSERT INTO documents (source, chunk, embedding) VALUES ('kb-sample', 'This is a sample chunk for testing.', ARRAY[0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0]::vector);"

After inserting a real embedding in your production pipeline, rerun the same endpoint test and confirm you receive ranked results.

14) Verify reboot persistence

A production deployment is not finished until it survives a reboot. Reboot the server during a maintenance window.

Restart-Computer

Reconnect after the reboot, then check the core services again:

Get-Service postgresql-x64-17, rag-api, W3SVC

You should see PostgreSQL, your API service, and World Wide Web Publishing Service in the Running state.

Test the API again from the server:

Invoke-WebRequest http://127.0.0.1:8000/health -UseBasicParsing

Then test from your local computer against the public hostname one more time. A clean response after reboot confirms that the database, service wrapper, IIS, and firewall rules are all persistent.

Troubleshooting the most likely failures

1. PostgreSQL will not accept connections

Run this on the server:

Get-Service postgresql-x64-17

If the service is stopped, start it with Start-Service postgresql-x64-17. If it is running but the app still fails, check pg_hba.conf and confirm the password is correct in .env.

2. The API starts locally but IIS returns 502

Check the FastAPI process and port:

Get-NetTCPConnection -LocalPort 8000

If nothing is listening, inspect the service logs or run the startup script manually:

C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe -ExecutionPolicy Bypass -File C:\opt\rag-app\start-api.ps1

If it fails, the error message usually points to a Python package issue or a bad database string.

3. HTTPS fails but HTTP works

Check the IIS binding and certificate store:

Get-ChildItem Cert:\LocalMachine\My

If the certificate is missing, import it again and rebind the site in IIS. If the cert exists but the domain does not match, renew or replace it before going live.

4. Search returns no useful matches

Inspect the table row count:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "SELECT count(*) FROM documents;"

If the count is zero, your ingestion pipeline has not loaded data yet. Import chunks and embeddings before you test the answer endpoint again.

5. Windows Firewall blocks external access

Verify the rule exists and is enabled:

Get-NetFirewallRule -DisplayName "IIS HTTPS" | Format-List *

If it is disabled or missing, recreate it with New-NetFirewallRule and test again from your client.

Rollback and recovery

Keep a rollback path ready before every schema change or major app release. The safest quick rollback is usually the last working app directory plus a database backup.

Create a logical backup of PostgreSQL before you change the schema:

"C:\Program Files\PostgreSQL\17\bin\pg_dump.exe" -U postgres -d ragdb -f C:\opt\rag-app\data\ragdb-backup.sql

Restore it if you need to revert:

"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -f C:\opt\rag-app\data\ragdb-backup.sql

For the application itself, keep the previous working copy in a separate folder, such as C:\opt\rag-app.prev. If the new build breaks, stop the service, restore the previous directory, and start the service again. That is a straightforward rollback that support teams can walk through quickly during an incident.

Hostperl customers who need a private chatbot or document assistant can run this stack on a suitable Hostperl VPS for smaller deployments or a dedicated server when database growth and memory pressure become part of daily operations.

If you want a deployment that stays supportable after launch, keep the retrieval database local, test your backups, and avoid exposing the app port directly to the internet.

FAQ

Can I use SQLite instead of PostgreSQL with pgvector?
Not for this tutorial. PostgreSQL with pgvector gives you a cleaner path for similarity search, backups, and service management on Windows Server.

Do I need IIS?
No, but IIS is the Windows-native way to terminate HTTPS and forward requests to the local API on this platform.

Can this run on Windows Server Core?
Yes, with more manual work. This guide assumes a standard Windows Server install with IIS management available.

What should I monitor first?
Track PostgreSQL service state, free disk space, API response time, and the size of the documents table. Those four signals usually catch trouble early.

How do I scale this later?
Move the app to a larger Hostperl dedicated server, keep PostgreSQL local or on a separate database host, and add a queue for ingestion jobs before you expand the corpus.