Skip to content
Putting technology to work.
Insights to guide decisions and action.

Search articles

Adding AI code review to client development: the decisions come before the tool

Table of contents · 4 items

You had AI read your pull requests, but it made so many comments that nobody read them anymore. The review comments filled up with machine output, and human feedback got buried. A few weeks after adoption it comes to this, and the tool is quietly switched off. Trying AI code review can go this way.

OpenCodeReview, which Alibaba has released as open source, tries to avoid this "too noisy" problem not through smarter AI but through how the work is divided up. Looking at how it works also shows what you need to decide when bringing it into development for clients.

Much of the work is not left to AI

According to its GitHub repository and README, OpenCodeReview is a command-line tool written in Go and released under the Apache-2.0 license. After about two years of use inside Alibaba, it was made public in May 2026, and its GitHub stars passed 40,000 in a little over four months (as of September 24, 2026).

What sets it apart is that it splits the review process into deterministic processing and AI processing. Which files to review, how to bundle them, which rules apply, which line of code each comment should be attached to: these tasks, which must not go wrong, are handled by ordinary program logic, and the part that reads the code and makes judgments is left to AI. Built-in checks cover null pointer dereferences, thread safety, XSS, SQL injection and more.

You choose the model yourself, and it is said to support the Anthropic, OpenAI Chat Completions, OpenAI Responses and AWS Bedrock connection protocols. In the benchmark published by its developer (200 pull requests in 10 languages drawn from 50 public repositories), it is described as scoring higher than Claude Code on precision and F1 when both use the same model, while using roughly one-ninth of the tokens. On the other hand, its recall, the share of real defects it catches, is lower than Claude Code's, and the developer describes this as a design that leans toward reducing false positives.

We think that handing the positioning of comments to a mechanism separate from the AI matters a great deal in practice. Even a few comments pointing at lines that do not exist are enough to undermine trust in the entire review.

In client development, first decide where the code goes

So far, this has been about the tool. Once you bring it into your process for client development, accuracy is no longer the first issue, because the code in your hands belongs to your clients.

Decision ItemWhat to confirm specifically
Where the code is sentWhich model provider receives it, and how storage and training are handled
Contractual permissionWhether the NDA includes a clause on using third-party services
Explaining it to the clientWhether to tell the client in advance that you use it, and how to note it in deliverables
Locus of ResponsibilityWhether defects the AI missed are handled the same way as in conventional review

Being able to choose the model provider means that the choice becomes a contract issue in itself. According to the README, it sends changed files to the configured LLM, and the agent also reads entire files and searches the codebase. What gets sent is not necessarily limited to the diff. Before sending client code to an external API, confirm which clause of the contract permits it. If you adopt the tool while leaving this vague, you gain convenience at the cost of a situation you cannot explain. We also cover who handles AI-based checks and who confirms the results at the end in how to structure reviews of AI-generated code.

It is not a tool for reducing human review

The other thing to decide is how to treat the AI's comments. You cannot hand machine-generated comments to a client as they are and present the work as "reviewed." Responsibility for what you deliver to clients does not shift when you change tools.

A realistic setup is to place AI before human review. AI catches obvious defects and rule violations first, and people spend their time on design judgments and on whether the business specifications make sense. With this division of labor, the total amount of review does not go down, but people can put their time into the places that need judgment. We summarize what human reviewers should point out in the two types of code review comments.

The published benchmark figures may not apply directly to your company either. They cover 200 cases in 10 languages, and both the language mix and the code size differ from your own projects. If you adopt it, the reliable approach is to run a few dozen of your own past pull requests that contain no client code, and decide after reading the comments it actually produces and what it misses.

How to proceed with a trial

Pick one internal project and start with a repository that contains no client code. After looking at the volume and quality of the comments, expand to projects where the contractual checks are done. If you do it the other way around, you end up stopping after people have seen how convenient it is, and that causes the most friction.

On September 24, 2026, we checked the GitHub repository information, the README and the bundled skill description (SKILL.md). The benchmark figures are the developer's own claims, and our editorial team has not reproduced them. We have also not installed or run OpenCodeReview or verified the accuracy of its comments. This article organizes the issues to consider in an adoption decision and how to proceed.

For bringing AI into your development process or designing how you organize reviews, please consult GleamHub.

Share this articleXFacebook
Kakeru Suzuki

Fascinated by the possibilities of technology, has had a deep interest in programming and digital art since student days

Turn this article's theme into your company's next step

Concrete steps forward for your organization.

We organize your desired architecture, legacy systems, and operational requirements to formulate your next steps toward execution.

  • Desired architecture
  • Integration with existing environments
  • Operational requirements
Consult on development & operations initiatives

You can consult with us from the initial conceptual stage. Details from this article will be carried over to the inquiry form.

Receive the latest articles by email