PERSONAL RESEARCH / Agentic coding

Review the change.
Verify the result.

Verification, automated review and approval answer different questions. My current working synthesis for reviewing agent-generated changes.

Article date
Base notes dated
Sources checked
DENIZ WETZ / Personal research synthesisALL RESEARCH

01RESEARCH NOTE

Verification needs something to check

A coding agent needs observable acceptance criteria: tests, a build, expected outputs or a running interface. Anthropic recommends giving Claude checks it can execute and using the results to guide corrections. This creates a feedback loop with evidence about the outcome. [01]

My interpretation is that passing checks provides evidence within their scope. A successful build says little about authorization rules; a browser test covers the journey it exercises. The agent checking its own work is also different from another reviewer examining the change.

02RESEARCH NOTE

Automated review needs an explicit boundary

Claude Code Review examines changes in repository context and reports findings. Its managed review does not approve or block a pull request; the check concludes neutrally. Findings require a separately configured gate to stop a merge. Claude can also initiate its local /code-review skill itself. [02]

Anthropic’s SDLC playbook combines automated findings and correction loops with code-owner approval enforced through branch protection. This is a documented workflow; its controls need to be deliberately configured in a project. [03]

03RESEARCH NOTE

Human review has different meanings

GitHub recommends using Copilot review alongside careful human review and checking its suggested changes. [04]

OpenAI describes a different experiment: agents review and sometimes merge their own changes, while people define goals and validate outcomes. The report limits generalization to environments with comparable tooling and investment. It is an engineering practice report with a specific scope. [05]

These sources describe different products, environments and responsibilities. They do not support a single blanket answer about manually reading every line.

04RESEARCH NOTE

Read recommendations in their scope

BSI and ANSSI’s 2024 guidance emphasizes developers checking and understanding generated code, with automated tests and security checks as additional layers. [06]

NIST’s SSDF 1.1 takes a broader approach: PW.7.1 lets organizations choose human review, tool analysis, or both; PW.7.2 calls for carrying out that choice, recording findings and triaging them. These different scopes do not establish that any LLM reviewer is sufficient or that every line always needs manual review. [07]

05RESEARCH NOTE

My working synthesis

For my own research, I favour small, reviewable changes, executable checks and a separate review perspective. I want review to concentrate on intent, architecture, permissions, data boundaries and assumptions that tests may omit.

I distinguish a review comment from permission to merge or deploy. Retry limits and escalation rules keep correction loops bounded; production changes need a clear owner and an explicit acceptance decision.

What evidence justifies accepting this change, and what remains uncertain? More reviewers can expand coverage, but their number alone does not establish independent judgment or complete correctness.

REFERENCES / 07

Sources.

This article builds on dated personal research notes. The source check covers the cited claims and links, rather than an exhaustive literature review. Document dates and any access limits appear below.

  1. Anthropic — Best practices for Claude Code (opens in a new tab)

    DocumentationAccessed

    Give Claude a way to verify its work. Continuously maintained documentation.

  2. Anthropic — Claude Code Review (opens in a new tab)

    DocumentationAccessed

    Managed review check behavior; local /code-review skill and automatic invocation.

  3. Anthropic — The AI-native SDLC playbook (opens in a new tab)

    Engineering articlePublished Accessed

    Concrete PR workflow; code-owner approval and branch protection.

  4. GitHub — Responsible use of Copilot agents (opens in a new tab)

    DocumentationAccessed

    Limitations; Copilot review complements human review.

  5. OpenAI — Harness engineering: leveraging Codex in an agent-first world (opens in a new tab)

    Engineering articlePublished Accessed

    Agent review and merge workflow; limits to transferability and open questions.

  6. BSI / ANSSI — AI Coding Assistants (opens in a new tab)

    GuidancePublished Accessed

    Original document: §3.3, p. 9. Document content checked in the base research notes on 28 September 2026. ANSSI announcement: 4 October 2024. Publisher page checked on 11 October 2026; PDF recheck unavailable.

  7. NIST — Secure Software Development Framework (SSDF), Version 1.1 (opens in a new tab)

    GuidancePublished Accessed

    Final publication; PW.7.1 and PW.7.2, printed p. 14 (PDF p. 23).