For most of the history of software, the hard part was writing the code. You designed a solution, typed it out, and reviewing it was a smaller, later step. With Claude Code, Codex and OpenCode in the loop, that order has flipped. An agent can produce a working change in minutes. Deciding whether that change is correct, complete and safe to merge is now where most of your time and judgment goes.
That shift changes what a good workflow looks like. If you still organise your day around writing, you end up with long agent sessions, half-read diffs and a vague feeling that something was skipped. If you organise it around review, you get smaller units of work, clearer evidence and far fewer surprises a week later. This article is a practical guide to reviewing AI generated code in that second way.
Coding moved from writing to reviewing
When a person writes code, review catches mistakes in something the author understood line by line. When an agent writes code, the author is fast, confident and occasionally wrong in ways that look plausible. The failure modes are different:
- Silent omissions. You ask for five changes in one message and get four. The summary says "done".
- Scope drift. The agent fixes the bug you asked about and also refactors a helper nobody asked it to touch.
- Confident explanations. The final message describes what the agent intended, which is not always what the diff contains.
- Context bleed. In a long session, instructions from an earlier request quietly influence a later one.
None of these are solved by reading faster. They are solved by changing the unit you review.
Why one long chat is the wrong unit of review
A long agent conversation feels efficient. You keep adding requests and the agent keeps working. The problem shows up at review time. The conversation now contains several goals, partial attempts, reverted ideas and tool output. To answer "is request number four finished?" you have to scroll through all of it and reconstruct what happened.
Long sessions also make every later request carry the earlier context with it. That context can help, but it can also mislead: a constraint that applied to the first task gets applied to the fourth. When something goes wrong, it is hard to tell which request caused it.
The fix is simple to state: one request, one task, one review.
Make the task the unit of review
One request per session
Give each request its own session with only the context it needs: the repository, the project instructions and the request itself. When it finishes, you review one request and one answer. If you need a change, you continue that same conversation with a follow-up, so the agent keeps the context of that task and nothing else.
This does not mean fresh sessions are always cheaper. Caching and repeated setup can change the numbers. What you get reliably is focus: a small, readable history and a clean diff tied to a single goal.
Write a definition of done into the request
Agents stop when they believe the task is complete. Tell them what complete means:
- the files or area they may change
- the behaviour that must work afterwards
- the check that proves it, such as a test command or a manual step
- what they must not touch
A request like "Add CSV export to the invoices list. Only change files under src/invoices/. Add a test that exports with a date filter applied. Run npm test -- invoices and paste the result." gives you something concrete to verify.
Ask for evidence, not reassurance
The final message of an agent is a claim. Ask it to end with evidence you can check: the test command it ran and its output, the list of files it changed, and anything it could not do. "I could not run the migration locally" is far more useful than "everything works".
A review checklist for agent output
For each finished task, go through the same short list:
- Does the diff match the request? Read the file list first. Anything outside the stated scope needs a reason.
- Is every part of the request done? Compare the request line by line with the diff, not with the summary.
- Do the checks actually pass? Re-run the test command yourself for anything that matters.
- Did it add or remove dependencies? Look at the lockfile and package manifest.
- Are there new secrets, logs or debug output? Agents sometimes leave print statements or hardcoded values.
- Is error handling real? Look for swallowed exceptions and broad catches added to make a test pass.
- Would you understand this in three months? If not, ask the agent to simplify or explain it in a comment.
Then decide: accept, send a follow-up in the same conversation, or discard the task and try a clearer request.
Size the review to the risk
Not every change deserves the same depth. A copy fix and an authentication change should not get the same treatment.
| Change | Review depth |
|---|---|
| Copy, styling, docs | Read the diff, glance at the result |
| Isolated feature behind a flag | Diff, tests, one manual run |
| Shared utilities, data models | Diff line by line, tests, check every caller |
| Auth, payments, migrations, security | Plan reviewed before coding, full diff review, manual verification, second opinion |
For the riskiest changes, review the plan before any code exists. A plan is cheap to change. A merged migration is not. Our guide to Claude Code plan mode shows how to get a plan you can actually review.
Keep parallel work reviewable
Running several agents at once is where review discipline pays off most. Two agents writing to the same checkout will eventually collide: one edits a file the other is testing, or a migration runs in the middle of another task. Separate git worktrees give each task its own working directory and branch, so each one ends with its own diff that you can review on its own. See how to run multiple Claude Code sessions without file conflicts for the setup.
Parallelism also changes the bottleneck. With three or four agents running, finished work arrives faster than you can read it. Keep a simple rule: a task is not done when the agent stops, it is done when you have reviewed it. Anything finished but unreviewed is work in progress.
Where VibeiDE fits
VibeiDE is a desktop app built around this workflow. It coordinates the Codex, Claude Code and OpenCode CLIs you already installed, using your own provider accounts. Each project has its own task queue, each task keeps its saved provider conversation so a follow-up continues it, and finished tasks wait in Ready for review, which is an unread completion state rather than a code approval. Changes shows the file diffs for the project or a selected worktree, the composer can give every task its own worktree and branch, and Plan council has one provider draft a plan and the other review it before you decide to implement. If you want to try this way of working, the quick start takes about three minutes, and there is a 14-day trial without a card.
Takeaways
- Treat review as the main job and writing as the delegated one.
- Make one request one task, with its own session and its own diff.
- Put a definition of done and a verification step into every request.
- Trust evidence over summaries, and re-run the checks that matter.
- Match review depth to risk, and review plans before code for risky changes.
- Keep parallel tasks in separate worktrees, and count unreviewed work as unfinished.
For more on running several agents without losing track, see running AI coding agents in parallel and how VibeiDE compares with a wall of separate terminals.



