heydeer Sign in
Back to HeyDeer

Guides / Code review

How to evaluate AI code review on your own pull requests

Choose an AI code reviewer by running it on changes your team understands. Record which findings deserve a fix, which claims are wrong, and how much time the team spends checking them. A comment count alone cannot tell you whether a reviewer helped.

HeyDeer team ·

Download the evaluation worksheet (CSV)

Here is a trial process you can use with HeyDeer or another review tool.

Pick changes that represent your work

Start with a small set of recent pull requests: a routine feature, a bug fix, a refactor, and a change to a sensitive path such as authorization or billing. Include changes that your team believes are correct. Those help reveal unsupported warnings.

Ten PRs can be a manageable starting point for a pilot. It is not enough to establish a universal accuracy ranking. Expand the sample before making claims beyond your own team’s experience.

For every PR, save the base and head commit hashes. If you compare tools, give them the same revision. A review of the original change and a review after fixes answer different questions.

Record the context and settings

Keep the repository instructions, review depth, included files, and available context alongside the results. Record the date and product configuration too. If one reviewer receives an explanation of an important business rule, give the other reviewer equivalent information where its configuration allows it. Document any difference you cannot remove.

Evaluate the workflow you would use after the trial. If your normal process includes a follow-up review after a fix, include that run in the time and cost records.

Judge each claim against the code

Have a developer familiar with the change classify each distinct finding. When a claim is disputed, record the evidence and ask another developer to check it.

How to classify review findings
ClassificationWhat it means
Valid and actionableThe issue exists, and the team would change the code because of it.
Valid, deferredThe issue exists, but fixing it is outside this change or not currently worth the cost.
False positiveThe claimed issue does not exist under the relevant code and requirements.
DuplicateAnother finding already describes the same underlying issue.
UnresolvedThe available evidence is insufficient to decide.

A comment that the team ignores is not automatically false. It may be correct but low priority, already known, or hard to understand. Keep those reasons separate so that the configuration change addresses the actual problem.

Count repeated descriptions of one bug as one distinct issue. Retain the duplicate count separately to measure how much extra reading the reviewer creates.

Measure the time needed to use the feedback

For each PR, record the tool’s elapsed review time, the developer’s time checking findings, and the actual charge. Keep code-fixing time separate: fixing a useful discovery is different from spending ten minutes disproving a warning.

Open the CSV worksheet in your spreadsheet app. Use one row per tool and review run; keep follow-up runs on separate rows. Record costs in the same currency when comparing tools.

What to record in the evaluation worksheet
RecordInclude
PR and revisionPR URL, base commit, head commit, and review date.
ConfigurationTool and version, review depth, instructions, files, and available context.
FindingsCounts for each classification above, with evidence for disputed claims.
Time and costElapsed review minutes, developer triage minutes, actual charge, and currency.

Report the counts before reducing everything to a score. If you calculate precision, the share of assessed findings judged correct, define which findings enter its denominator. Keep unresolved findings visible. If no findings were reported, precision is undefined rather than 100%.

Recall, the share of real defects a reviewer finds, needs a known set of defects. Reviewing your ordinary PRs does not establish that every bug has been found. You can add a separate exercise using known historical defects, but label those results separately and withhold the later fix from the reviewer.

Choose a rollout threshold before expanding

Decide what would make the trial useful for your team. That might be finding previously missed issues in a critical path while keeping triage time acceptable. A threshold should reflect your repository and reviewers, not a number borrowed from another company’s benchmark.

In HeyDeer, start with manual reviews on one repository. Use the same worksheet to check whether a change to instructions or feedback scope improves the results. Once the team trusts the feedback, expand automatic reviews gradually. The rollout guide describes those settings, and reducing noise maps common problems to specific adjustments.

Our internal evaluation explains how we assessed reviewer contributions during development. It does not independently rate the completed HeyDeer workflow or establish its recall. Use it as context for the evaluation method, then test the product against your own changes.

Compare the current plans with your observed trial charges. The credit guide explains what each review uses and where to see its final cost.

Try the method on your own pull requests

Explore an example review, or start a HeyDeer trial and record your findings in the worksheet. The quickstart walks you through connecting a repository and running your first review.