Filed at 3,000 m, lands at 300 m.AI and agents

Your AI code reviewer still needs a linter

A linter is a smoke alarm and an LLM is a building inspector. Keep linting as the deterministic gate (security rules on every touched file, everything else on changed lines) and let the model review intent and suggest repairs that go back through the same gate.

In this descent, 2 stops

Once an AI reviewer is commenting on every merge request, the static analysis stage starts to look like a redundant line item. It isn’t. A good AI reviewer makes your linter more important, because the linter is what frees the reviewer to look at the things that matter.

Only deterministic checks get a veto

Only your static analyser gets to block a merge on its own. Your large language model (LLM) reviewer runs afterwards with the analyser’s results in hand, and its findings are advice, with one exception: when it raises a security concern, a human decides. Collapse the two into one AI step and you get a review that is slower, more expensive and harder to trust than the one it replaced.

The analogy I keep coming back to is a smoke alarm and a building inspector. The smoke alarm is cheap and not very bright. It can’t tell burnt toast from a kitchen fire, but it never forgets to check. The inspector can tell you the deck won’t survive the next storm, but bills by the hour and occasionally has an off day. Nobody cancels their smoke alarms because they booked an inspection.

A gate has to give the same answer twice. When a build fails, the developer needs to see which rule fired, on which line, and reproduce it on their own machine. A linter does exactly that, whether it’s PMD for Apex, ESLint for Lightning Web Components or Salesforce Code Analyzer running both. An LLM won’t, even with the temperature turned down, and a model upgrade can change what it flags without anyone touching a config file. Your ruleset has a Git history. The model’s judgement lives on someone else’s server.

Then there’s attention. Every extra instruction in a review prompt competes with the others. Ask the model to check class suffixes and hardcoded IDs alongside the business logic, and it will be admirably thorough about class suffixes.

Linting isn’t only for code, either. In a Salesforce org a lot of the system lives in configuration: objects, fields, flows, permission sets. The same deterministic checks apply there, and the humble ones earn their keep: a field with no description or help text, a flow nobody bothered to explain. It looks like housekeeping until someone has to work out what Status_2__c is for. These days that someone is often an AI coding agent, reading the metadata for context.

The review process I work within covers six areas: conformance to our design patterns, framework naming, domain naming, static rule violations, behaviour (intent, bugs, regressions and performance) and, arguably the most important, security. Split by owner, it looks like this.

CheckExamplesScopeOwnerWhy
Pattern conformanceSalesforce Object Query Language (SOQL) kept in selector classes, modules called from trigger handlersDiff lines + contextBothStructure is rule-checkable, “right pattern for this problem” is not
Framework naming conventionsClass suffixes such as Selector, Module, Action and ServiceDiff linesLinterPure pattern matching, no judgement required
Domain naming conventionsNames that match the business domain or capability the code belongs toDiff linesLLMA rule can match a pattern, but it can’t tell which domain a class belongs to
Rule violationsMissing descriptions, queries or data manipulation language (DML) inside loopsDiff linesLinterDeterministic, cheap and repeatable
BehaviourRecursive operations that burn through limits, repurposing a field shared across domainsDiff lines + contextLLMNeeds the story, the surrounding code and judgement
SecuritySOQL injection and hardcoded credentials (rules), a new @AuraEnabled method exposing records to guest users (the model)Any touched fileBoth, then a humanRules catch the known patterns every time, the model reasons about exposure no rule can describe, and a human owns the risk

Security is the row where two layers still aren’t enough. A linter catches the known shapes, like a query that ignores the running user’s permissions or a hardcoded credential, and catches them every time. What it can’t see is a perfectly compliant method that hands every customer’s address to a guest user because the story said to show the address.

The LLM can reason about that kind of exposure, but it’s probabilistic, and security is exactly where “usually spots it” isn’t good enough. So security findings from the model aren’t advice you scroll past. They go to a human, who either fixes the issue or signs off on why it’s fine. Anything touching access, sharing or credentials gets that human review regardless of what the tools said.

A merge request passes through the smoke alarm first: security rules on every touched file and quality rules on changed lines only, with any blocking violation stopping it. If it passes, the findings go to the building inspector, the LLM review, which checks behaviour, pattern fit, domain naming and data exposure. The review is advisory except on security, where findings need human sign-off before merge.
How a merge request gets reviewed.

Nothing merges without a human, least of all anything the model flags as a security concern.

This split has costs. It means two systems to maintain, and a ruleset without an owner rots quietly. It also matters less for small teams: if three developers share a modest codebase, one careful human reading the lint output is plenty. The split pays off when volume makes reviewer attention the scarce resource, which in a large org is roughly always.

Lint the diff for quality, the whole file for security

Scope linting by risk. Gate quality rules on the lines a merge request changes, run security rules across every file it touches, and leave the rest of the legacy alone until someone opens it. A full-repo gate sounds rigorous and in practice gets switched off. This stop is mostly about the smoke alarm (the linter), with the building inspector (the LLM reviewer) turning up near the end.

Run a sensible lint ruleset across a monorepo of 10,000-plus files and 100-plus developers and you get a violation count nobody wants to read. Much of it is older than some of the team. No product owner will fund a sprint to fix lint. The report becomes wallpaper, the job gets marked “allow failure”, and eventually someone removes it during a pipeline tidy-up.

A smoke alarm that goes off every day doesn’t get anyone out of the building. It gets the battery taken out.

Diff-scoped linting turns the gate into a ratchet. New code meets the standard from its first commit, and old code improves as people touch it. The count only moves one way.

There are two ways to scope a diff. Scoping to changed files lints every file the branch touches, so a one-line fix can drag in dozens of old violations. Scoping to changed lines only reports violations on lines the branch added or modified, which is fairer. The catch is whole-method rules like length or cyclomatic complexity: a new line can tip the method over, but the violation is reported at the declaration, which nobody changed.

So I split them by risk. Quality rules only report on lines the change added or modified, so nobody gets blocked over someone else’s old formatting. Security rules run across every touched file in full, and any violation blocks, including one that’s been sitting there for years. That feels unfair the first time it happens to you, and it’s also the point: touch the file and you own its exposure.

The reason is reachability. Add @AuraEnabled to a method, and a query a few lines below it, untouched for years, is suddenly reachable from the browser by anyone with access to that class. The diff shows one new annotation, while the risk sits in the lines around it.

None of this depends on which linter you run. PMD, ESLint and Salesforce Code Analyzer all accept a list of files, so point yours at the changed ones instead of the whole workspace, and do the same for your flow and metadata checks.

Two things sit outside this scoping. Dependency and secret scanning run across the whole project as separate status checks, because a vulnerable library is live whether or not this sprint touched it. And greenfield projects can skip the scoping altogether: with no legacy to protect, lint the lot.

Scope decides where a rule looks. Severity decides whether it blocks, and getting severity wrong does more damage than a missing rule. Every rule in the ruleset should earn its place, and if nobody can say what it prevents, it’s noise with a config file.

I learned this the slow way. When I first became a tech lead, I weighted cosmetic violations the same as functional ones, and every merge request had to be spotless before it went in. The result was a frustrated team and features waiting on formatting debates. The code was a canvas and I was curating the gallery.

Over time I started picking my battles. Consistency still matters, because it makes code faster to read and review across a big team. But a cosmetic finding changes nothing at runtime, so fixing it gets weighed against timelines and budgets like any other cost. A useful test: if this violation shipped, would anything behave differently?

TierExamplesWhat happens
BlockingSecurity rules, limit-consuming operations inside loops, hardcoded IDs, tests without asserts, framework class naming, missing descriptionsFails the merge request
AdvisoryComplexity thresholds, long parameter lists, debug statements, queries without a filterComment on the merge request, developer’s call
CosmeticDeclaration order, variable and parameter naming, one declaration per lineA formatter or auto-repair fixes it, never blocks

The ruleset I work with actually has two levels, required and recommended, and both the advisory and cosmetic tiers sit in recommended. It’s still worth knowing which is which, because only the cosmetic ones should go to a fixer without a second look.

The line isn’t only functional versus cosmetic, either. Some rules look cosmetic but are load-bearing, because other rules depend on them. Framework naming, like the Selector suffix, sits in the blocking tier for exactly that reason. A rule such as “no SOQL outside selector classes” only works if every selector is reliably named. Let that naming drift and every rule built on top of it drifts too.

Better still, take layout out of review entirely. When an agent writes a good share of the code, blocking a developer on brace placement is arguing with a tool. An opinionated formatter like Prettier (natively for Lightning Web Components, via a community plugin for Apex) rewrites whitespace, brace placement and line wrapping on save or in a pre-commit hook. The argument about where the brace goes gets settled once, in a config file, and never again in a merge request. It won’t rename a variable, so a thin layer of cosmetic lint rules stays, but most of that tier disappears.

The analyser’s output then goes into the LLM review as context, with a plain instruction: these are already flagged, don’t repeat them, look for what a rule can’t see. For changes touching sharing, access or credentials, give it the surrounding code as well as the diff. The model is also handy for triaging likely false positives and drafting suppression comments, but it doesn’t get to clear the gate.

Auto-repair is the obvious next step: if the model can spot the problem, let it fix it. The speed case is real. In GitHub’s 2024 public beta, developers committed Copilot Autofix fixes for pull request alerts in a median of 28 minutes, against 1.5 hours by hand (GitHub). That’s vendor data about a vendor product, but the direction is hard to argue with.

The case for supervision is just as real. GitHub’s own documentation warns that generated fixes can change what the program does, only partly fix the problem, or introduce a new vulnerability (GitHub Docs). A 2026 study of 319 LLM-generated patches for 64 Java vulnerabilities found only 24.8% were fully correct, while 51.4% failed on both security and functionality (Al-Maamari, 2026). A pilot for another 2026 study, with eight participants patching Python code, found that people out of their depth on security trusted the assistant’s patch with little verification (Helpful or Harmful?, 2026).

So the model can repair, inside a fence. Where a deterministic fixer exists, like ESLint’s --fix or a code formatter, use that first: same input, same fix. Cosmetic and advisory findings on changed lines are fair game for the model, as a separate commit. That commit goes back through the same lint gate and test run, then gets human review like any other change. Missing descriptions are the exception worth making: they block, but the model can draft one from how the field is used, and a human only has to confirm it’s true.

Security findings get a suggested fix in a comment, never an applied one. And the model never adds a suppression, because silencing a rule is a decision, not a repair, and a human approves every one. The inspector can sketch the fix. Someone with their name on the building signs it off.

The payoff comes when the two layers start to compound. When the LLM reviewer flags the same thing across several merge requests, stop paying tokens to find it and write it as a rule. In an org that uses the selector pattern, the classic candidate is SOQL living outside a selector class. A custom PMD rule can catch that by walking the syntax tree.

A rule like that has limits worth knowing. It checks where a query lives, not whether it’s the right query, so a selector that quietly drops half the records the story needed passes cleanly. That’s still the inspector’s job, and the inspector now has more attention for it because the alarm handled the rest.

Before your next sprint planning, export the last month of LLM review comments and find the one that turns up most. Write it as a lint rule and add it to the gate. If it can’t be expressed as a rule, that’s useful too: you’ve found where the line between smoke alarm and inspector sits in your codebase.

Nathan Avatar

More about Nathan

Where next

A consistent coding framework is what makes 100 developers behave like one team

A framework is a core set of classes, the conventions around them, and someone enforcing both. Its product is that every team's code has the same shape, so anyone can review anyone's work and the next platform change is three edits, not every class. That only holds if someone owns it.

10,000 m to 300 m. Architecture, 10 minute read.