Self-Hosted & Local
AI Code Reviews
Run PRInspector on AI models hosted on your own hardware with Ollama. Code diffs go only to your model server — never to OpenAI, Anthropic, or any other third-party AI provider — which makes it easier to meet your own security and compliance requirements.
Self-Hosted Architecture
Hardware Guide
Approximate requirements for 4-bit quantized models. Any model available in Ollama can be used.
| Profile | Best For | Suggested Hardware | Example Models |
|---|---|---|---|
| Starter | Getting started and everyday pull requests | GPU with 8GB+ VRAM, or Apple Silicon with 16GB+ unified memory | qwen2.5-coder:7b (default), phi3:mini |
| Balanced | Larger or more complex diffs | GPU with 16GB+ VRAM, or Apple Silicon with 32GB+ unified memory | qwen2.5-coder:14b |
| Maximum quality | Deepest reviews on critical codebases | GPU with 24GB+ VRAM (e.g. RTX 4090), or Apple Silicon with 64GB+ unified memory | qwen2.5-coder:32b |
How Self-Hosting Works
Self-hosted deployments are set up together with our team. Here's what's involved.
Run Ollama on Your Hardware
Install Ollama on a GPU server inside your network and pull the models PRInspector should use. Keep port 11434 reachable only from the Diffnix worker.
# On a GPU server inside your network ollama serve ollama pull qwen2.5-coder:7b ollama pull phi3:mini
Point Diffnix at Your Model Server
Set the worker's environment variables to your Ollama address and choose a model for each review tier. Nothing to add to your repositories.
# Diffnix worker environment (.env) OLLAMA_HOST=http://10.0.0.45:11434 # your Ollama server OLLAMA_MODEL_STANDARD=qwen2.5-coder:7b # main review + custom rules OLLAMA_MODEL_TACTICAL=phi3:mini # fast baseline checks OLLAMA_CONCURRENCY=3 # parallel model requests OLLAMA_TIMEOUT=120000 # per-request timeout (ms)
Start Diffnix and Connect GitHub
Start the Diffnix services with Docker Compose and connect your GitHub organization through a GitHub App pointed at your server. Diffnix still reaches GitHub to receive pull request events and post review comments.
# On your Diffnix server docker compose up -d --build # The worker checks Ollama on startup and # pulls any configured model that is missing