Windows Server RAG Hosting with pgvector and IIS

What this tutorial builds
Windows Server RAG hosting works well when you need a private chatbot or internal knowledge assistant on a Windows stack, but want the retrieval layer in PostgreSQL with pgvector. In this guide, you will set up Windows Server 2022 or Windows Server 2025, install PostgreSQL with vector search, publish a FastAPI app on IIS as a reverse proxy, and verify the full request path from browser to database.
This tutorial assumes a fresh server. You will start at first login, create a non-administrator account, tighten the firewall, deploy the app, and test recovery before production. If you run Hostperl VPS or Hostperl dedicated server infrastructure, this is the kind of controlled rollout our support teams expect before a cutover.
Hostperl VPS is a practical choice for smaller private RAG deployments, while a dedicated server makes more sense once your document corpus, worker queue, and database need more memory or steadier IOPS.
Architecture and prerequisites
The stack in this tutorial stays deliberately simple:
- Windows Server 2022 or Windows Server 2025
- PostgreSQL 17 on Windows
- pgvector extension for similarity search
- Python 3.12 for the FastAPI service
- IIS with URL Rewrite and Application Request Routing acting as a reverse proxy
This split keeps the LLM or embedding model out of the web tier. It also gives you a cleaner rollback path if the API needs a restart or the database needs a restore. For a broader deployment pattern, compare this with our RAG application hosting on Ubuntu Server with PostgreSQL and our RAG vector database hosting guide.
1) Connect to the server
On your local computer, open PowerShell or Terminal and connect with the default documentation IP shown below:
ssh root@203.0.113.10Replace 203.0.113.10 with the real public IP assigned to your Hostperl server. If your access model uses a non-admin account first, connect with that account instead, but keep the same server IP.
After you log in, open a second terminal on your computer and keep the original root or administrator session open until the new account is verified.
2) Detect the operating system and update Windows
On the VPS as root, confirm the platform and build level before you change anything else. On Windows Server, use PowerShell rather than Linux commands.
Get-ComputerInfo | Select-Object WindowsProductName, WindowsVersion, OsHardwareAbstractionLayer, OsBuildNumberThis confirms whether you are on Windows Server 2022 or 2025. You should see the product name and build number in the output.
Install-WindowsUpdate -MicrosoftUpdate -AcceptAll -AutoRebootThis command works if the PSWindowsUpdate module is already present. If it is not, install it first with:
Set-ExecutionPolicy RemoteSigned -Scope Process -Force; Install-Module PSWindowsUpdate -Force; Import-Module PSWindowsUpdateThen rerun the update command. Reboot if Windows asks for it, and reconnect before continuing. A current patch level matters here because IIS, TLS, and PowerShell modules are easier to support when the OS is fully updated.
3) Create a non-administrator deployment account
Do not run the service day to day from the built-in administrator. Use a dedicated local account for the app and a separate admin account for maintenance.
On the VPS as root, create a local user called deploy and add it to the Administrators group:
New-LocalUser -Name deploy -Password (Read-Host -AsSecureString "Enter a strong password for deploy")That command prompts you for a password without echoing it. Then add the account to the admin group:
Add-LocalGroupMember -Group Administrators -Member deployOpen a second remote session and test the account before you close the original one.
On your local computer, open a new PowerShell session and connect as deploy using your normal remote access method, such as RDP or your provider console. Then verify group membership on the server:
whoami /groupsYou should see the Administrators group listed. Keep the original root or admin session open until this succeeds.
4) Prepare directories and a working layout
On the VPS as the non-root sudo user, create a predictable layout for the RAG service, database assets, and logs. In Windows Server, that means normal NTFS paths rather than Linux-style directories.
New-Item -ItemType Directory -Force -Path C:\opt\rag-app, C:\opt\rag-app\logs, C:\opt\rag-app\data, C:\opt\rag-app\modelsThis creates a clean application tree under C:\opt\rag-app. You will place the FastAPI app, a local configuration file, and any cached document artifacts there.
Check the paths exist:
Get-ChildItem C:\opt\rag-appYou should see the four directories you just created.
5) Install PostgreSQL and pgvector
For Windows Server RAG hosting, PostgreSQL stores your document chunks and embeddings. pgvector adds similarity search without forcing you into a separate vector service.
On the VPS as the non-root sudo user, install PostgreSQL 17 from the official Windows installer or your standard software deployment method. After installation, verify the service and version:
Get-Service postgresql* | Format-Table -AutoSizeYou should see the PostgreSQL service in the Running state once setup completes.
"C:\Program Files\PostgreSQL\17\bin\psql.exe" --versionThis confirms the client tools are installed. Adjust the path if your PostgreSQL version differs.
Next, open postgresql.conf. The file path is usually under the PostgreSQL data directory, for example C:\Program Files\PostgreSQL\17\data\postgresql.conf. Add or confirm these settings:
listen_addresses = 'localhost'
shared_buffers = 1GB
max_connections = 100These values keep the database local to the machine, which is the safer default for a private RAG app. Save the file, then open pg_hba.conf and make sure local trust is not used for production. A safer minimum is:
# TYPE DATABASE USER ADDRESS METHOD
local all all scram-sha-256
host all all 127.0.0.1/32 scram-sha-256Restart PostgreSQL after editing the files:
Restart-Service postgresql-x64-17Now create the database and extension.
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d postgres -c "CREATE DATABASE ragdb;"Then enable pgvector in that database:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE EXTENSION IF NOT EXISTS vector;"Confirm the extension is present:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "\dx"You should see vector in the extension list. For a deeper database performance path, our PostgreSQL tuning article covers maintenance tradeoffs that matter once your chunk table grows.
6) Create the schema for documents and embeddings
Now create tables for document chunks and metadata. Keep them small and explicit.
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE TABLE documents (id bigserial PRIMARY KEY, source text NOT NULL, chunk text NOT NULL, embedding vector(1536) NOT NULL, created_at timestamptz NOT NULL DEFAULT now());"That example uses a 1536-dimensional embedding, which matches many common embedding models. If your model uses a different size, change the column definition before you load data.
Add an index for similarity search:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "CREATE INDEX documents_embedding_idx ON documents USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);"Check the table and index exist:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "\d documents"If you see the column list and the ivfflat index, the retrieval layer is ready.
7) Install Python and build the RAG API
On the VPS as the non-root sudo user, install Python 3.12 and the runtime packages you need. Use the Windows installer if Python is not already present, then confirm it works:
python --versionYou should see Python 3.12.x or a compatible version. Next, create a virtual environment and install the app dependencies.
cd C:\opt\rag-apppython -m venv .venvC:\opt\rag-app\.venv\Scripts\pip.exe install fastapi uvicorn psycopg[binary] pgvector python-dotenvThat command installs the web framework, ASGI server, PostgreSQL driver, pgvector integration, and environment file support.
Create the application file:
notepad C:\opt\rag-app\app.pyPaste the following content, then save and close Notepad:
from fastapi import FastAPI
from pydantic import BaseModel
import os
import psycopg
from pgvector.psycopg import register_vector
app = FastAPI()
DB_DSN = os.getenv("DB_DSN", "dbname=ragdb user=postgres password=CHANGE_ME host=127.0.0.1 port=5432")
class Query(BaseModel):
question: str
@app.get("/health")
def health():
return {"status": "ok"}
@app.post("/ask")
def ask(payload: Query):
with psycopg.connect(DB_DSN) as conn:
register_vector(conn)
with conn.cursor() as cur:
cur.execute(
"SELECT source, chunk FROM documents ORDER BY embedding <=> %s::vector LIMIT 3",
([0.0] * 1536,)
)
rows = cur.fetchall()
return {"question": payload.question, "matches": [{"source": r[0], "chunk": r[1]} for r in rows]}
This is a minimal retrieval service. In production, you would replace the zero vector placeholder with an actual embedding generated by your model pipeline.
Test the syntax before you go further:
C:\opt\rag-app\.venv\Scripts\python.exe -m py_compile C:\opt\rag-app\app.pyIf the command returns no output, the file compiles cleanly.
8) Add environment settings and secure the file
Create a local environment file that stores the database connection string outside the source code.
notepad C:\opt\rag-app\.envUse this content, then save and close the editor:
DB_DSN=dbname=ragdb user=postgres password=CHANGE_ME host=127.0.0.1 port=5432Replace CHANGE_ME with the real database password you set for PostgreSQL. Then restrict access to the app directory so only the service account and administrators can read it.
icacls C:\opt\rag-app /inheritance:ricacls C:\opt\rag-app /grant deploy:(OI)(CI)F Administrators:(OI)(CI)FThose ACLs remove inherited permissions and grant full control only to deploy and Administrators. That matters because the database password should not be readable by casual users or shared service accounts.
9) Create the Windows service for the API
You need the RAG API to start on boot and restart on failure. One straightforward production approach on Windows Server is to run the FastAPI app under a service wrapper.
On the VPS as the non-root sudo user, create a startup script:
notepad C:\opt\rag-app\start-api.ps1Paste this content and save it:
Set-Location C:\opt\rag-app
$env:PYTHONUTF8 = "1"
.\.venv\Scripts\uvicorn.exe app:app --host 127.0.0.1 --port 8000Check the script file exists:
Get-Content C:\opt\rag-app\start-api.ps1Now create a scheduled task or service wrapper. If you use NSSM, the service can be installed like this:
nssm install rag-api C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exeSet the arguments to:
-ExecutionPolicy Bypass -File C:\opt\rag-app\start-api.ps1Then start the service:
Start-Service rag-apiVerify it is running:
Get-Service rag-apiIf you prefer Task Scheduler instead of NSSM, keep the same script and schedule it at startup. The important part is unchanged: the API must start automatically after a reboot.
10) Configure IIS as a reverse proxy
IIS will accept external HTTPS traffic and forward it to the local FastAPI service on port 8000. That keeps the Python app off the public interface.
On the VPS as the non-root sudo user, install the IIS role and management components if they are not already present:
Install-WindowsFeature Web-Server, Web-Http-Redirect, Web-Filtering, Web-Request-Monitor, Web-Mgmt-ToolsThen add URL Rewrite and Application Request Routing. After installation, confirm IIS responds locally:
Invoke-WebRequest http://localhost -UseBasicParsingYou should receive a valid HTTP response from IIS.
Create or edit the IIS site bindings in the IIS Manager GUI, or use PowerShell if you manage bindings as code. For the reverse proxy rule, save this rewrite configuration in the site’s web.config:
<configuration>
<system.webServer>
<rewrite>
<rules>
<rule name="RAGProxy" stopProcessing="true">
<match url="(.*)" />
<action type="Rewrite" url="http://127.0.0.1:8000/{R:1}" />
</rule>
</rules>
</rewrite>
</system.webServer>
</configuration>After saving the file, recycle the site or app pool. Then test the local FastAPI endpoint directly before you expose it through IIS:
Invoke-WebRequest http://127.0.0.1:8000/health -UseBasicParsingYou should get a JSON response that includes status: ok.
11) Open the firewall safely
Do not expose the Python service port to the internet. Only open the public web ports you actually need.
On the VPS as the non-root sudo user, allow IIS traffic through Windows Firewall:
New-NetFirewallRule -DisplayName "IIS HTTP" -Direction Inbound -Protocol TCP -LocalPort 80 -Action AllowNew-NetFirewallRule -DisplayName "IIS HTTPS" -Direction Inbound -Protocol TCP -LocalPort 443 -Action AllowIf you are still testing locally and have not installed TLS yet, keep only port 80 open for the moment. Once you confirm the site is working, add HTTPS and then remove any temporary exposure you no longer need.
Check the current rules:
Get-NetFirewallRule -DisplayName "IIS *" | Format-Table -AutoSize12) Add TLS with a real certificate
For production, terminate HTTPS at IIS. You can use a trusted certificate from your certificate workflow or a publicly trusted CA. Bind it in IIS after the certificate is imported into the Local Computer certificate store.
After you install the certificate, assign it to the HTTPS site binding in IIS Manager. Then test with PowerShell from the server:
Invoke-WebRequest https://localhost -UseBasicParsingFrom your local computer, test the public endpoint as well. Replace the hostname with your own domain that points to the server.
curl -I https://server.example.comA successful response should show 200 or a similar valid status and a certificate that matches your domain.
If you are planning DNS and certificate work at the same time, our DNS, SSL, and email guide is a useful companion before cutover day.
13) Smoke test the full RAG path
Now test the app through the public web path, not just on localhost. First hit the health endpoint.
curl https://server.example.com/healthYou should see a JSON response with status set to ok. Then try the answer endpoint with a small JSON payload.
curl -X POST https://server.example.com/ask -H "Content-Type: application/json" -d '{"question":"What documents are indexed?"}'You should receive a JSON object with your query text and a matches array. If the array is empty, the service is running but your retrieval table has no loaded data yet.
Load one sample document to confirm the search path works:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "INSERT INTO documents (source, chunk, embedding) VALUES ('kb-sample', 'This is a sample chunk for testing.', ARRAY[0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0]::vector);"After inserting a real embedding in your production pipeline, rerun the same endpoint test and confirm you receive ranked results.
14) Verify reboot persistence
A production deployment is not finished until it survives a reboot. Reboot the server during a maintenance window.
Restart-ComputerReconnect after the reboot, then check the core services again:
Get-Service postgresql-x64-17, rag-api, W3SVCYou should see PostgreSQL, your API service, and World Wide Web Publishing Service in the Running state.
Test the API again from the server:
Invoke-WebRequest http://127.0.0.1:8000/health -UseBasicParsingThen test from your local computer against the public hostname one more time. A clean response after reboot confirms that the database, service wrapper, IIS, and firewall rules are all persistent.
Troubleshooting the most likely failures
1. PostgreSQL will not accept connections
Run this on the server:
Get-Service postgresql-x64-17If the service is stopped, start it with Start-Service postgresql-x64-17. If it is running but the app still fails, check pg_hba.conf and confirm the password is correct in .env.
2. The API starts locally but IIS returns 502
Check the FastAPI process and port:
Get-NetTCPConnection -LocalPort 8000If nothing is listening, inspect the service logs or run the startup script manually:
C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe -ExecutionPolicy Bypass -File C:\opt\rag-app\start-api.ps1If it fails, the error message usually points to a Python package issue or a bad database string.
3. HTTPS fails but HTTP works
Check the IIS binding and certificate store:
Get-ChildItem Cert:\LocalMachine\MyIf the certificate is missing, import it again and rebind the site in IIS. If the cert exists but the domain does not match, renew or replace it before going live.
4. Search returns no useful matches
Inspect the table row count:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -c "SELECT count(*) FROM documents;"If the count is zero, your ingestion pipeline has not loaded data yet. Import chunks and embeddings before you test the answer endpoint again.
5. Windows Firewall blocks external access
Verify the rule exists and is enabled:
Get-NetFirewallRule -DisplayName "IIS HTTPS" | Format-List *If it is disabled or missing, recreate it with New-NetFirewallRule and test again from your client.
Rollback and recovery
Keep a rollback path ready before every schema change or major app release. The safest quick rollback is usually the last working app directory plus a database backup.
Create a logical backup of PostgreSQL before you change the schema:
"C:\Program Files\PostgreSQL\17\bin\pg_dump.exe" -U postgres -d ragdb -f C:\opt\rag-app\data\ragdb-backup.sqlRestore it if you need to revert:
"C:\Program Files\PostgreSQL\17\bin\psql.exe" -U postgres -d ragdb -f C:\opt\rag-app\data\ragdb-backup.sqlFor the application itself, keep the previous working copy in a separate folder, such as C:\opt\rag-app.prev. If the new build breaks, stop the service, restore the previous directory, and start the service again. That is a straightforward rollback that support teams can walk through quickly during an incident.
Hostperl customers who need a private chatbot or document assistant can run this stack on a suitable Hostperl VPS for smaller deployments or a dedicated server when database growth and memory pressure become part of daily operations.
If you want a deployment that stays supportable after launch, keep the retrieval database local, test your backups, and avoid exposing the app port directly to the internet.
FAQ
Can I use SQLite instead of PostgreSQL with pgvector?
Not for this tutorial. PostgreSQL with pgvector gives you a cleaner path for similarity search, backups, and service management on Windows Server.
Do I need IIS?
No, but IIS is the Windows-native way to terminate HTTPS and forward requests to the local API on this platform.
Can this run on Windows Server Core?
Yes, with more manual work. This guide assumes a standard Windows Server install with IIS management available.
What should I monitor first?
Track PostgreSQL service state, free disk space, API response time, and the size of the documents table. Those four signals usually catch trouble early.
How do I scale this later?
Move the app to a larger Hostperl dedicated server, keep PostgreSQL local or on a separate database host, and add a queue for ingestion jobs before you expand the corpus.
