Enterprise Deployment

Self-Hosted & Local
AI Code Reviews

Run PRInspector on AI models hosted on your own hardware with Ollama. Code diffs go only to your model server — never to OpenAI, Anthropic, or any other third-party AI provider — which makes it easier to meet your own security and compliance requirements.

Self-Hosted Architecture

Source Node
GitHub (via the Diffnix GitHub App)
Agent Node
Diffnix review worker (your server)
Inference Node
Ollama model server (your GPUs)

Hardware Guide

Approximate requirements for 4-bit quantized models. Any model available in Ollama can be used.

ProfileBest ForSuggested HardwareExample Models
StarterGetting started and everyday pull requestsGPU with 8GB+ VRAM, or Apple Silicon with 16GB+ unified memoryqwen2.5-coder:7b (default), phi3:mini
BalancedLarger or more complex diffsGPU with 16GB+ VRAM, or Apple Silicon with 32GB+ unified memoryqwen2.5-coder:14b
Maximum qualityDeepest reviews on critical codebasesGPU with 24GB+ VRAM (e.g. RTX 4090), or Apple Silicon with 64GB+ unified memoryqwen2.5-coder:32b

How Self-Hosting Works

Self-hosted deployments are set up together with our team. Here's what's involved.

STEP 01

Run Ollama on Your Hardware

Install Ollama on a GPU server inside your network and pull the models PRInspector should use. Keep port 11434 reachable only from the Diffnix worker.

GPU server
# On a GPU server inside your network
ollama serve
ollama pull qwen2.5-coder:7b
ollama pull phi3:mini
STEP 02

Point Diffnix at Your Model Server

Set the worker's environment variables to your Ollama address and choose a model for each review tier. Nothing to add to your repositories.

Worker environment
# Diffnix worker environment (.env)
OLLAMA_HOST=http://10.0.0.45:11434      # your Ollama server
OLLAMA_MODEL_STANDARD=qwen2.5-coder:7b  # main review + custom rules
OLLAMA_MODEL_TACTICAL=phi3:mini         # fast baseline checks
OLLAMA_CONCURRENCY=3                    # parallel model requests
OLLAMA_TIMEOUT=120000                   # per-request timeout (ms)
STEP 03

Start Diffnix and Connect GitHub

Start the Diffnix services with Docker Compose and connect your GitHub organization through a GitHub App pointed at your server. Diffnix still reaches GitHub to receive pull request events and post review comments.

Diffnix server
# On your Diffnix server
docker compose up -d --build

# The worker checks Ollama on startup and
# pulls any configured model that is missing