
Inside Diffnix PRInspector: How AI Actually Reviews Your Code (No Black Box)
Most AI code review tools are opaque by design. You submit a diff, comments appear, and you are supposed to trust the output. This post opens PRInspector completely every pipeline stage, every architectural decision, every tradeoff. so you understand exactly what is happening to your code and why.
Introduction
There is a specific frustration that comes from using a tool you cannot reason about. The AI review comment looks plausible. You are not sure if it is correct. You cannot check the reasoning. You end up either trusting it blindly or ignoring it blindly, and neither of those is useful behavior for a senior engineer evaluating findings on production-bound code.
This is not a blog post about what PRInspector does. It is about exactly how it works, from the moment a PR is opened to the moment a comment appears on a specific line number. The data models, the reasoning stages, the context decisions, the output calibration. All of it, without the marketing layer.
If you are evaluating AI code review tools, this post tells you what questions to ask every vendor. If you are already using PRInspector, this post makes the output more actionable.
Why "Just Send It to GPT" Fails
Before getting into how PRInspector works, it helps to understand why the obvious approach breaks down. The obvious approach: take the diff text, append "review this," send to an LLM API, post the response as a PR comment.
Three specific things break.
Treating the diff as flat text. A unified diff is a structured artifact. The @@ hunk headers encode file position. Context lines anchor changes in the surrounding code. Added and removed lines together describe a transformation. A model treating this as flat text loses that structure entirely. It cannot distinguish "this was added" from "this provides context for what was added."
Operating without codebase context. The changed lines are almost never enough to evaluate correctness. What are the imports? What calls this function? What interface is being implemented? None of that is in a typical diff. A model reasoning without that context is guessing, and it produces comments that sound authoritative about things it cannot actually know.
Generic, unfocused reasoning. A single prompt asking "what is wrong with this code?" produces comments at whatever granularity the model defaults to: sometimes style, sometimes real bugs, sometimes imagined issues with unchanged context lines. There is no principled way to calibrate specificity when the task is not scoped.
PRInspector is built specifically against these three failure modes. Each stage of the pipeline addresses one of them.
Stage 1: Structured Diff Parsing
PRInspector does not receive a diff as a string. It parses the raw diff payload from your Git provider into a typed data model before any analysis runs.
typescript
interface ParsedHunk {
file: string;
language: LanguageTag;
originalStartLine: number;
modifiedStartLine: number;
removedLines: CodeLine[];
addedLines: CodeLine[];
contextLines: CodeLine[];
hunkIndex: number;
}
interface ParsedDiff {
prTitle: string;
prDescription: string;
baseBranch: string;
headBranch: string;
changedFiles: ParsedHunk[];
totalAdditions: number;
totalDeletions: number;
}Language detection runs on file extension first, with a content-based fallback for polyglot files — SQL strings embedded in Python, JavaScript in HTML templates, and similar patterns. The originalStartLine and modifiedStartLine values from the @@ hunk header are preserved as metadata through every subsequent stage. Every finding produced later in the pipeline maps back to a precise line number in the modified file.
That line precision matters more than it sounds. A review comment posted to the wrong line number signals that the tool does not understand what it is looking at. Structural parsing guarantees the positional mapping is exact before any reasoning begins.
Stage 2: Context Enrichment
This is the stage that separates PRInspector from tools that only see the diff. For each parsed hunk, context enrichment fetches supplemental information from the repository that was not in the diff.
python
def enrich_hunk(hunk: ParsedHunk, repo: RepoClient) -> EnrichedHunk:
# Fetch the full file at the base commit
file_content = repo.fetch_file(hunk.file, ref=hunk.baseSha)
# Extract all imports and module-level definitions
imports = extract_imports(file_content, hunk.language)
# Find where the changed function is called from
callers = find_callers(hunk.addedLines, repo, hunk.language)
# Resolve any interface or base class the changed code implements
interfaces = resolve_interfaces(hunk.addedLines, file_content)
return EnrichedHunk(
hunk=hunk,
imports=imports,
callers=callers[:MAX_CALLER_CONTEXT], # bounded to prevent token overflow
interface_definitions=interfaces,
)The MAX_CALLER_CONTEXT bound is not arbitrary. Fetching every call site for a widely-used utility function would overflow the model's context window and dilute the relevant signal. The enrichment step uses a relevance scoring pass to select the callers most likely to reveal correctness issues for the specific change. A function called in 60 places sends the 4-6 most structurally relevant call sites.
💡 Tip: Your PR description directly shapes context enrichment quality. The enrichment step uses the PR title and description as its primary statement of intent when deciding which context is relevant. A description that says "this guards against null account_id in the referral flow" directs enrichment to fetch the account model and referral callers. "Fix bug" directs nothing. The model will still do its best, but you are giving it less to work with.
Stage 3: Reasoning Prompt Construction
Enriched hunk data does not go into a single "review this code" prompt. Each hunk is evaluated across four dimensions, each in its own scoped prompt with the context appropriate to that type of reasoning.
Correctness: Does the code do what the PR description says it should? Are there logic errors, missing guards, or incorrect assumptions given the enriched calling context?
Security: Does the change introduce injection vectors, unvalidated inputs, missing authorization checks, or weakened authentication?
Performance: Does the change introduce N+1 patterns, unbounded loops, repeated database calls, or algorithmic regressions?
Maintainability: Does the change reduce readability, violate established patterns in the file, or introduce complexity without a clear reason?
Each dimension receives the context most useful for its type of reasoning. Correctness gets the caller sites and model definitions. Security gets the input paths and data flow. Performance gets the surrounding loop structures and any database call patterns visible in the enriched context.
python
CORRECTNESS_PROMPT_TEMPLATE = """
You are reviewing a {language} code change for CORRECTNESS ONLY.
PR Description: {pr_description}
Changed Function: {function_name}
File: {file_path}
CHANGED CODE (unified diff):
{diff_hunk}
CALLING CONTEXT:
{caller_excerpts}
IMPORTS IN THIS FILE:
{imports}
Evaluate whether the implementation matches the stated intent.
Identify logic errors, missing guards, or incorrect assumptions
given the calling context above.
Do NOT comment on style, security, or performance in this evaluation.
For each issue found:
- Line number in the modified file
- Severity: critical | major | minor
- Explanation of the problem
- Suggested fix (concrete, not abstract)
Respond in JSON only. If no issues, return {"issues": []}.
"""📌 Insight: Dimension decomposition does two things at once. It reduces hallucinations by constraining the model's reasoning scope, and it produces findings specific enough to act on. "This guard on line 22 does not cover the null case for account_id" is actionable. "There might be some issues with error handling" is not.
Stage 4: Cross-Hunk Reasoning
Individual hunk analysis is necessary but not sufficient. Real bugs often live in the relationship between changed files rather than in any single change.
After all individual hunk analyses complete, PRInspector runs a cross-hunk reasoning pass that looks for three patterns:
Symbol propagation gaps. A type or function is renamed in one file but not updated across the other files changed in the same PR.
Missing paired changes. An API method signature changes but the test file does not reflect the new signature.
Intra-PR state conflicts. Two hunks in the same PR modify shared state in ways that interact adversarially when the PR is taken as a whole.
typescript
function crossHunkReasoning(hunks: AnalyzedHunk[]): CrossHunkIssue[] {
const symbolChanges = extractSymbolChanges(hunks);
const missingPropagations = findMissingPropagations(symbolChanges, hunks);
const sharedStateConflicts = detectSharedStateConflicts(hunks);
return [...missingPropagations, ...sharedStateConflicts];
}This stage catches the category of bugs that look fine in every individual file but create problems when the full PR is treated as a unit.
Stage 5: Output Formatting and Confidence Scoring
Raw LLM output does not ship directly to the developer. Each finding goes through a post-processing layer that:
Assigns a confidence score based on how much supporting context was available for that finding
Filters out findings below the configured threshold (default: 0.65)
Maps findings to precise line numbers using the GitHub or GitLab comment API format
Groups related findings to prevent comment spam on the same issue
The default behavior surfaces the top five high-confidence findings per PR. Everything below threshold is available in the detail view but does not appear as inline comments unless the team has configured a lower threshold for their risk profile.
This mirrors how a good reviewer operates: lead with the most important observations, do not file every potential issue as a blocking comment.

Real-World Use Case: The Double-Commit Bug
A fintech team submitted a 340-line PR refactoring their payment processing module across 8 files. Reviewers focused on the architectural changes and approved after 35 minutes. The refactor looked sound. One hunk contained a subtle implementation error.
The developer correctly added a transaction context manager to wrap database operations. But a conn.commit() call from the previous implementation was left in place after the transaction block closed.
python
# Before the refactor
def finalize_order(order_id: str, items: list):
conn = get_db()
conn.execute("UPDATE orders SET status='processing' WHERE id=?", [order_id])
conn.commit()
for item in items:
conn.execute("UPDATE inventory SET qty=qty-1 WHERE sku=?", [item.sku])
conn.commit()
# After the refactor — the version that was approved
def finalize_order(order_id: str, items: list):
conn = get_db()
with conn.transaction():
conn.execute("UPDATE orders SET status='processing' WHERE id=?", [order_id])
for item in items:
conn.execute("UPDATE inventory SET qty=qty-1 WHERE sku=?", [item.sku])
conn.commit() # orphaned line from previous implementationThe refactor looked correct to reviewers: the transaction wrapper was present, the loop was inside it, and the structure was clean. What was invisible without knowing the connection manager implementation: conn.transaction() commits on clean exit. The orphaned conn.commit() on the last line calls commit on an already-closed transaction, raising TransactionAlreadyClosedError on PostgreSQL backends.
The test suite passed because it ran against SQLite, which handles already-committed transactions permissively.
PRInspector produced this finding:
Line 47 [Correctness / Critical]
conn.commit()is called explicitly after thewith conn.transaction()block exits. Based on the transaction context manager implementation indatabase/connection.py(lines 88-94), this manager commits on clean exit. This double-commit raisesTransactionAlreadyClosedErroron PostgreSQL. Callers inorder_processor.py(lines 112, 187, 203) do not catch this exception. The orphaned commit on line 47 should be removed.
The finding required three pieces of cross-file context that were not in the diff: the connection manager behavior, the caller locations, and the exception handling at each call site. All three were fetched by the context enrichment step.
For a broader look at why logic bugs like this survive normal code review processes, that post covers the structural mechanisms in detail.

Advanced Tips for Getting More From PRInspector
Tune confidence thresholds per repository, not globally
The default 0.65 threshold is a starting point. A payments service should run at 0.45 to surface more findings including uncertain ones. A marketing site can run at 0.80. The teams that report "too much noise" almost universally have not calibrated thresholds for their specific risk profile. Thirty minutes of threshold tuning in week one eliminates most noise complaints.
yaml
# .diffnix/config.yml
pr_inspector:
confidence_threshold: 0.45
max_comments_per_pr: 8
dimensions:
security:
confidence_threshold: 0.35 # lower: prefer false positives for security
maintainability:
enabled: false # disable for velocity-focused sprintsUse the reasoning trace before dismissing a finding
Every PRInspector finding includes a reasoning trace showing which context excerpts contributed to it. When a finding seems wrong, check the trace before dismissing it. Most apparent false positives are findings grounded in outdated context — the code the model fetched has changed since the last enrichment. The trace tells you in 30 seconds whether to act on the finding or ignore it.
Start with the cross-hunk view on large PRs
For PRs over 200 lines across multiple files, the cross-hunk analysis is often more valuable than the individual hunk findings. It surfaces relationship-level issues that file-by-file analysis cannot see. If you are reviewing a refactor or a feature that touches multiple modules, read the cross-hunk tab before looking at individual file findings.
⚠️ Warning: PRInspector's context enrichment is bounded by the PR scope. A correctness issue caused by code entirely outside the PR — a caller in a separate service, a contract defined in a package the PR does not touch — cannot be detected. AI review supplements integration testing and observability. It does not replace them.
Conclusion
The reason most AI code review tools feel unreliable is not the model quality. It is the architecture around the model. Sending a diff to a chat endpoint and asking for a review produces chat-quality output. Building a pipeline that parses structure, enriches context, reasons in scope, and gates output on confidence produces reviewer-quality output.
That is the difference between a tool that engineers ignore after one week and one they trust enough to act on.
Diffnix is a private, AI-powered code intelligence platform that understands your code — not just scans it. If you want to understand the full landscape of what AI review can and cannot do, the complete AI code review guide covers it end to end. And if you want to see how PRInspector handles security-specific findings, how Diffnix catches vulnerabilities at the PR stage covers that dimension in detail.
Try PRInspector on your next PR. Setup takes under five minutes.