Kousigan A

"Looks good" isn't evidence: how to verify AI-generated code

AI-generated code is not done when it compiles. Check the requirement, the diff, tests and static checks, and get a review from outside the model that wrote it.

Last verified against official docs on

When a coding agent says “done”, what you have is a candidate. To verify AI-generated code, check it against the requirement and read the diff. Run the tests, lint, type checks and the build. Then get a review from something other than the model that wrote it. Code generation is not completion. Verification is completion.

A successful build is one piece of evidence, not proof the requirement was met. I learned that from AI-assisted work.

This guide matches each claim to the evidence behind it and puts the checks in order. It also has two prompts to copy, and what six coding agents offer to run checks for you. The tool behaviour comes from each tool’s official docs, last verified on 4 October 2026; none of it was tested for this post.

Part 2 of this series has a short verify step as the last stage of its workflow. This part goes deeper: how do you know the code is actually correct?

A passing build can hide a missing requirement

The risky workflow is a short one:

Prompt → AI writes code → app runs → "looks good" → done

An application running is not proof that the change is correct. A build can pass while:

  • a requirement is missing
  • an edge case is broken
  • authorisation is wrong
  • existing behaviour has regressed
  • the architecture is violated
  • the tests don’t cover the new path

GitHub’s tutorial on reviewing AI-generated code says the same thing in one line: “Be skeptical of code that “looks right” but doesn’t match your intent.”

Match every claim to its evidence

A claim about AI-generated code is only as good as the check behind it. Each claim has a better kind of evidence than the agent’s word:

Claim Better evidence
Code compiles Build
Types are valid Type checker
Style follows rules Linter
Existing behaviour works Tests
New requirement works New or updated tests
No accidental changes Git diff
Architecture is respected Review
Security-sensitive change is safe Security review
Requirement is actually satisfied Acceptance criteria

Each row catches something different. A clean build tells you nothing about the requirement, and a green test suite tells you nothing about a file the agent changed by accident.

The model that wrote the code can’t be its only judge

Asking the same AI that wrote the code whether it works isn’t meaningful. It writes the code, you ask, and it says “Yes, the implementation looks good”. That’s one signal, counted twice. I’d never use the same signal for generation and verification.

Signal Counts as evidence?
The AI says it works No
The AI says tests pass, and you haven’t seen them run No
The app opens No
The build passes Partly
The tests pass Partly
Tests, static checks, the diff and a requirement review Yes

The chain I’d want is generation, then independent checks that give different signals, then a human review. When an AI does review, give it a fresh start. Anthropic’s best practices explain why: a reviewer in a fresh sub-agent context “sees only the diff and the criteria you give it, not the reasoning that produced the change”. Part 1 covers setting up a reviewer sub-agent.

The checks, from requirement to evidence

The checks for AI-generated code follow the same order as Prompt 1 further down.

Re-read the requirement first

Turn the requirement into acceptance criteria, then hold the code against them:

Requirement → acceptance criteria → implementation → verification

Then ask “Did we actually solve the original problem?” rather than “Did the AI finish the task?”

Read every changed file, then the diff

Every changed file and the final diff answer one question: what exactly changed? Look for:

  • unrelated modifications
  • accidental refactoring
  • unnecessary dependencies
  • changed configuration
  • removed validation
  • security-sensitive changes
  • generated files
  • unexpected API changes

The diff is your audit trail.

Ask whether the tests prove the behaviour

Swap “did you write tests?” for a better question: do the tests prove the behaviour we care about? Good tests for a change cover the happy path, edge cases, failure cases, regression cases and the business rules that matter.

A passing test suite can still miss an untested requirement. Also check that the tests changed for the right reasons. GitHub’s tutorial puts it bluntly: “Watch for tests that are deleted or skipped, instead of fixed.”

Run lint, type checks and the build

Lint, type checks, the build and static analysis give the same answer every time. Don’t make the AI reason about things your tooling can verify deterministically. Let the tools catch what humans miss.

Try the edge cases and failure paths

Edge cases and failure paths are where plausible code breaks. Anthropic’s best practices list “the trust-then-verify gap” as a common failure. In their words: “Claude produces a plausible-looking implementation that doesn’t handle edge cases.” The hostile review prompt below is built to find them.

Look for security and dependency risks

Check every new dependency. GitHub’s tutorial asks you to “Check if suggested dependencies exist and are actively maintained”, and warns about “hallucinated or suspicious packages (such as packages that don’t actually exist)”. The same page lists “hallucinated APIs, ignored constraints, or incorrect logic” as things to look for. A security-sensitive change gets a security review.

Record the evidence

Write down what was checked, the command or evidence used, the result, and anything still unverified. Prompt 1 below asks the agent for exactly that, so you get the evidence in writing.

Verification hooks and review commands in six coding agents

Most coding agents can run verification checks for you through hooks, so nobody has to remember. Part 3 explains what a hook is. Here is what each tool offers for verification:

Agent After each edit When it tries to finish Built-in review Command approval
Claude Code (hooks, commands, permissions) A PostToolUse hook runs after a tool call succeeds, for example a linter after edits; exit code 2 shows its error output to Claude A Stop hook that exits with code 2 “Prevents Claude from stopping”; an agent hook can check the tests first (experimental) /code-review (alias /review) checks the current diff for correctness bugs; /security-review checks it for vulnerabilities In Manual mode, Bash commands need approval, except a built-in set of read-only commands; allow, ask and deny rules change that. In auto mode, the built-in starting mode in the terminal and VS Code from v2.1.283, a classifier reviews most actions instead of you
OpenAI Codex (hooks, code review, permissions) A PostToolUse hook runs after Bash, file edits (apply_patch) and MCP calls; “block” replaces the tool result with your feedback A Stop hook that returns “block” makes Codex continue, with your reason as the next prompt /review reads uncommitted changes, a branch diff or a commit, and reports findings “without changing your working tree” “Ask for approval” runs routine local commands in the workspace, and asks before going online or outside it
Gemini CLI (hooks, reference, commands, settings) An AfterTool hook (“run tests” is a listed use) can add text to the tool result An AfterAgent hook can reject the final response and force a retry, with your reason as the new prompt No local review command in the commands reference; /setup-github sets up GitHub Actions to review pull requests The default mode prompts for approval on each tool call; auto_edit approves edits only
Cursor (hooks, Agent Review, Bugbot, Run Modes) An afterFileEdit hook fires after the agent edits a file, “useful for formatters” A stop hook can send a followup_message as the next user message to keep iterating (five loops by default) /agent-review reviews local changes and can run automatically; Bugbot reviews pull requests Allowlisted actions run without approval; Auto-review sends other calls to a classifier that “is not a security boundary”
Windsurf, now documented as Devin Desktop (hooks, terminal) A post_write_code hook runs after Cascade writes a file (“Run linters, formatters, or tests”); on exit code 2 Cascade sees the error No documented hook that sends Cascade back; post_cascade_response runs asynchronously, for logging Not covered in the docs checked Four auto-execution levels, from Disabled (every command needs approval) to Turbo (all run except the deny list)
GitHub Copilot (hooks, code review, CLI) A postToolUse hook (cloud agent and CLI) runs after each tool completes successfully and can add context for the model An agentStop hook that returns “block” forces another turn, with your reason as the prompt Copilot code review on pull requests and in IDEs; by default its reviews don’t count toward required approvals The CLI asks the first time it needs a tool that can change or run files; --deny-tool beats the allow options

Each row comes from that tool’s official docs, last verified on 4 October 2026; none was tested. Copilot’s agent hooks in VS Code are in Preview.

All six can run a hook after an edit or a tool call. Five document a hook that sends the agent back for another pass when it tries to finish; the Windsurf docs don’t. Four have a built-in AI review of your changes.

Claude Code is the worked example here. Its hooks guide shows an agent hook on the Stop event that checks the tests before Claude may stop. It goes in .claude/settings.json:

.claude/settings.json
{
  "hooks": {
    "Stop": [
      {
        "hooks": [
          {
            "type": "agent",
            "prompt": "Verify that all unit tests pass. Run the test suite and check the results. $ARGUMENTS",
            "timeout": 120
          }
        ]
      }
    ]
  }
}

The same page says agent hooks are experimental, and to prefer command hooks for production workflows.

What the agent can check, and what stays with you

AI and tooling are good at the checks that have a clear answer. Humans own the decisions that need judgement:

AI and tooling can verify Humans should verify
Tests Product intent
Lint and type checks Business decisions
The build Architectural trade-offs
Static analysis Acceptable risk
Dependency checks UX decisions
Diff inspection Security-critical decisions
Repetitive review and test generation Whether the solution solves the user’s problem

GitHub says the same about its own reviewer. Its code review docs say: “Copilot is not guaranteed to spot all problems or issues in a pull request. Sometimes it will make mistakes. Always validate Copilot’s feedback carefully. Supplement Copilot’s feedback with a human review.” AI accelerates engineering, but it doesn’t remove engineering judgement.

Two prompts to copy

Neither prompt relies on a tool-specific feature. Prompt 1 goes further than Part 1’s implement and verify prompt and Part 2’s follow-up: as well as each result, it asks for the command or evidence behind every check.

Prompt 1: verify the work

Paste it after the agent finishes coding, before you accept the change.

Prompt 1: verify the work
The implementation is complete.

Do not assume the task is done.

Verify the work independently.

1. Re-read the original requirement.
2. Identify the acceptance criteria.
3. Review every changed file.
4. Inspect the final diff.
5. Run relevant tests.
6. Run lint and type checks.
7. Run the build where applicable.
8. Check edge cases and failure paths.
9. Look for unintended behaviour changes.
10. Look for security or dependency risks.

For each check, report:
- what you checked
- the command or evidence used
- the result
- anything still unverified

If something fails, investigate the root cause and fix it.
Do not commit, push, deploy, run destructive commands or make external
changes unless I explicitly authorise it.
If two reasonable attempts fail, stop and report the failure rather than
repeatedly guessing.

Do not claim completion without verification evidence.

Prompt 2: hostile review

Use it for a second opinion. Run it in a fresh session or a sub-agent if your tool has them, so the reviewer sees the diff and not the reasoning behind it.

Prompt 2: hostile review
Change to review: <the diff, branch or commit>
Requirement: <paste the requirement or acceptance criteria>

Assume this implementation contains a serious bug.

Try to prove it wrong.

Look for:
- edge cases
- incorrect assumptions
- security issues
- race conditions
- error handling problems
- regressions
- missing tests
- violations of existing architecture

Do not rewrite the code yet.
Do not commit, push, deploy, run destructive commands or make external
changes.

Report findings with evidence.

A reviewer told to find gaps “will usually report some, even when the work is sound”, say Anthropic’s best practices. Act only on findings that affect correctness or the requirement.

A checklist for AI-generated code: the definition of done

The full checklist, a definition of done for AI-written code, ready to copy:

Definition of done
AI CODING: DEFINITION OF DONE

□ Requirement understood
□ Relevant existing code inspected
□ Implementation completed
□ Tests added/updated
□ Tests passed
□ Lint passed
□ Type checks passed
□ Build passed
□ Final diff reviewed
□ Unintended changes checked
□ Edge cases considered
□ Security implications considered
□ Acceptance criteria verified
□ Remaining uncertainty documented
□ Human review completed where necessary

If you find the same verification work happening on every task, that’s a workflow worth packaging. Part 5 of this series is about when repeated work like that deserves a skill, and how to write one.

Quick answers

How do you verify AI-generated code?

Verifying AI-generated code means checking it against the original requirement, reading the final diff, and running tests, lint, type checks and the build. Then get a review from someone or something other than the model that wrote it, and keep a record of what is still unverified.

Can an AI coding agent review its own code?

An AI coding agent can review code, but the same session that wrote it is biased towards its own reasoning. A reviewer in a fresh session or sub-agent, which sees only the diff and the criteria, gives a more independent signal.

Is a passing build enough to trust AI-generated code?

A passing build only shows that the code compiles. AI-generated code can build cleanly while a requirement is missing, an edge case breaks, authorisation is wrong or the tests skip the new path.

What should a human check in AI-generated code?

A human should own the decisions tools cannot make. That means product intent, business rules, architectural trade-offs, acceptable risk, user experience (UX) and security-critical choices, and whether the change solves the user’s actual problem.

Can a coding agent run tests automatically after each edit?

Claude Code, OpenAI Codex, Gemini CLI, Cursor, Windsurf and GitHub Copilot all document hooks that run a command after the agent edits a file or calls a tool. A test or lint command there runs every time, not only when the agent remembers.

More on AI coding is in the AI category.

Cartoon of Kousigan making a heart with his hands

Kousigan A

Build with AI. Ship with care.

Software engineer building real products with AI assistants. I write down what works: prompts, rules, skills and agents, and how to ship safely.