Skip to content
AI agents & MCP

AI Coding Agents: How Teams Use Them Without Losing Control

By Sebastiaan Jansen · · 17 min read

AI coding agents read a repository, propose a plan and edit code toward a ticket you have already defined. They don't replace the engineer who owns the outcome. Teams that stay in control treat the agent as a fast pair programmer with a narrow brief, its own branch and a human review before anything merges.

The agent does the mechanical work. A person decides the goal, the boundaries and whether the result is fit to ship. Most of the practice comes down to keeping that split.

What AI coding agents do

An AI coding agent is a model that can use tools: it reads files, searches the repo, runs tests, edits code and opens a pull request. A chat window that only suggests a snippet is not an agent. The difference is the loop. The agent looks at the result of each tool call, picks the next step, and keeps going until it hits a stop you set.

Good stops are concrete:

  • The ticket's acceptance checks pass, or the agent reports which ones it couldn't meet.
  • The diff stays inside the paths named in the task.
  • The tests you named have been run, with the output attached to the pull request.
  • A human has to approve before merge, and the agent never pushes to the default branch.

Start with a failing test, a rename the suite can verify, or a plan that lists files and waits for approval. "Make onboarding faster" has no file list and no check, so the agent will wander and then write a confident summary anyway.

Agents, copilots and scripts

A copilot completes the line you're typing. A script repeats fixed steps. An agent works out its own steps from a goal, which is why the scope has to be written down before the run starts. Agentic coding is that loop running on a real branch: plan, edit, run, revise, stop, and hand the diff to a reviewer. Claude Code, Codex, Cursor and similar tools differ in how they read the repo and how tightly you can constrain a run.

What one loop looks like

A bounded run is a short sequence you can narrate afterward.

  1. The agent reads the ticket and the files named on it. If a named file is missing, it stops and says so.
  2. It writes a plan in a few lines: which function changes, which test is added, which commands it will run. If the ticket says the plan needs a comment first, it waits.
  3. It edits only those files, on the feature branch.
  4. It runs the command from the ticket and pastes the exit code. A non-zero exit is a result. It is not a cue to try different commands until one passes.
  5. It opens the pull request with the ticket link, the plan and the log. Then it stops, even if it can see a nearby cleanup.

People skip the stop at step 5. A green test feels like permission to keep going, and "while you're in there, the import order in that folder is messy" is how a three-line fix turns into a review nobody wants to read. If the nearby cleanup is real, it gets its own ticket and its own run.

Common mistakes in the first week

Teams that have just turned agents on tend to repeat the same four misses.

The prompt is the ticket. The brief lives in a chat nobody else can see, and the tracker still says "checkout bug". The reviewer opens a diff with no statement of what "done" was. Put the brief on the ticket, then tell the agent to read that and nothing else.

The stop is a vibe. "Looks good" and "make it clean" are not checks. Name the test command and the paths. If you can't name them, the task is still a conversation, and a conversation belongs with a copilot.

The agent marks its own homework. It deletes an assertion, skips a test, or widens a type to any so the compiler goes quiet. Read the test diff before the production diff. A new test that cannot fail is not a test.

Two runs share a branch. Both started from the same commit, neither re-read, and the second run rewrites the first run's file. Use one task per branch, and make a write fail when the file changed underneath it.

None of these are model problems. They show up on strong models and weak ones, because the constraints are missing.

How teams use AI agents for software development

AI agents for software development are most useful on work that is already specified. The ticket exists, the repo is the source of truth, and someone on the team can reject the result without it turning into an argument about taste.

Step Person Agent
Name the outcome and the non-goals Engineer or tech lead Nothing
Point at files, tests and constraints Engineer, on the ticket Reads them
Draft a plan Reviews it Writes it from the repo
Edit on a branch Nothing Does the edits
Run the named tests Reads the log Runs them
Review the diff and merge Engineer Nothing

The agent's column has more entries, but control sits in the person's column. Skip one of the person's rows and the agent fills the gap with a guess.

Where they help on a normal week

  • A first fix next to a failing test. Paste a stack trace and the suspect commit, ask for a failing test first and then a fix, and review both.
  • Mechanical refactors with a suite behind them: renames, import updates, a type that has to change in twenty call sites. The suite decides whether it worked.
  • Boilerplate that follows a local pattern, such as a new route that has to match three existing ones. Name the file that sets the pattern.
  • PR descriptions. Once the diff exists, ask for a summary a reviewer can scan: what changed, what wasn't run, and what needs a person.

There's a fifth case that looks like a feature but is really transcription. A support note says the empty-state copy on the billing page is wrong, and the correct sentence is already in the ticket. The agent changes that string in the component the ticket names and leaves the page styling alone. The work is small, the check is a snapshot or a glance, and a person still merges it. It works for the same reason the refactor does: the decision was made before the run.

Where they should not run alone

  • Unclear product behavior. If two engineers would build it differently, the agent shouldn't be the one choosing.
  • Sensitive surfaces like auth, payments, permissions, cryptography and migrations that rewrite customer data. An agent can draft the change, but the owner of that surface merges it.
  • Cross-cutting design. "Split this service" touches deploys, data and on-call. That needs a design review first, and maybe an agent on one slice afterward.
  • Anything you can't revert: production writes, force-pushes, secret rotation, customer emails. These aren't coding tasks and don't belong on an agent's tool list.

Copilot, Cursor, Claude Code, Codex, Aider and hosted cloud agents all speed up the inner loop, but none of them know which ticket matters this week unless you tell them. More bad runs come from a missing ticket than from the choice of model.

How to tell a ticket is ready

Read the ticket once and try to predict the diff. If you can name the files and the test command without opening the repo, it's ready for a run. If you catch yourself writing "and related files" or "use your judgment", it isn't. Cut it down until the prediction is boring.

A ready ticket has five lines and no narrative:

  • Outcome, one sentence, in the present tense. "A null tax_region uses the store default."
  • Non-goals. "No rate changes. No UI."
  • Allow list and deny list of paths.
  • The command that defines done, including the expected exit.
  • The branch name and the reviewer.

The stack trace, the suspect commit and a snippet of the wrong output can sit under those lines. They're evidence, and they don't replace the five lines. An agent that gets only the evidence will pick a plausible cause and fix that. Sometimes it will be right, but you won't be able to tell from the pull request, because the pull request will match whatever cause the agent chose.

A mistake that looks like success

The suite is green and the ticket's sentence is satisfied, but the agent got there by changing a fixture. The checkout test used to build a cart with a null region and expect a 500. After the run it builds a cart with a region already set, so the assertion never exercises the bug. The command exited 0, and the allow list was respected, because the fixture lives next to the test.

The review step for this is mechanical. Open the test and check that the input still matches the bug report: if the report says null, the test input is null. Then check that reverting the production change makes the test fail. If it still passes, the test isn't holding the fix in place, and the pull request goes back.

That's also why "run the tests" is a weak stop on its own. The stop is "run these tests, and the new one fails on the parent commit." Ask for that parent-commit run in the pull request body. It's one extra command, and it catches the fixture edit.

Rules that keep a run inside the repo

Control here means a few checks you can apply after a run, not a policy page nobody opens.

One task gets one branch. The agent starts from the default branch and pushes only to a feature branch. Protected branches reject direct pushes, so a prompt can't merge anything on its own.

Paths work as an allow list. Name the directories on the ticket, and if the diff touches anything else, stop. That keeps a real bug from hiding inside a 40-file "cleanup" pull request.

Writes are version-checked. The agent writes a file only if it still matches the version it read, so a second run (or a person) can't silently overwrite that work. On a conflict, the agent re-reads or stops.

Tools come from an explicit list. Allow test commands and git; refuse production credentials, package publishes and deploys. An MCP server is how Claude Code or Codex gets that fixed list. A catalogue you can read beats a prompt that says "use any tool you need."

The plan is a checkpoint. Once a change goes past a few files, the agent writes the plan and waits, and approval is a comment on the ticket.

Secrets stay out of the prompt. Strip tokens from tickets and logs. If a tool needs a credential, the tool holds it.

Every run leaves a trail: branch, ticket link, commands, exit codes and files touched. A reviewer shouldn't have to rebuild the session from scrollback.

Falrow keeps these rules next to the ticket. Claude Code and Codex plan from the git repository over a fixed MCP tool list, and writes are version-checked. On another host, the same rules can live in a repo instruction file and in CI. Having the checks matters more than which product enforces them.

A worked example: one error, one pull request

This walk-through is illustrative. The figures show the shape of the task, not a measured incident.

Ticket. "Checkout returns 500 when tax_region is null, starting after yesterday's rate-table deploy. Expected: treat a missing region as the store default, and add a test. Do not change tax rates or the checkout UI."

Scope on the ticket.

  • Allow: api/checkout/tax.ts, api/checkout/tax.test.ts
  • Deny: migrations, UI, rate tables, anything under billing/
  • Done when: the new test fails on the old code, passes on the branch, and npm test -- api/checkout exits 0
  • Branch: fix/tax-region-null

Plan, before any edit. Read tax.ts and the test. Find the call that assumes tax_region is a string. Add a test where null resolves through getStoreDefaultRegion(). Guard against the null without touching the rate math. Run npm test -- api/checkout and stop until it exits 0.

A reviewer comments: "Use the existing helper. Do not add a new default." The agent revises that line of the plan, then makes the edits.

The diff. One null guard, one test, and a pull request that quotes the ticket and includes the test output. If billing/invoice.ts shows up in the diff, reject the run even though the tests passed.

The person. They read the guard, confirm the helper and merge. The agent never sees production. The path stays short because the ticket was short. Runs get riskier as tickets get vaguer, and model strength has less to do with it.

Where this run usually goes wrong

Three variants show up the first time a team tries this ticket.

The agent doesn't find the helper on its first search, so it adds its own getStoreDefaultRegion() that returns a hardcoded region. The new test passes against the hardcode. The comment on the plan exists to prevent exactly this, which is why the plan is a checkpoint and not a formality. If the run skipped the checkpoint, the reviewer sends it back with the same sentence: use the existing helper, delete the new one.

The agent "fixes" the 500 by catching every exception in the checkout handler and returning a default tax of zero. The new test passes. The rate math is untouched, so the deny list looks clean. But the behavior is wrong, because a real tax failure is now a silent zero. The ticket said "treat a missing region as the store default", not "swallow errors". The reviewer reads the catch, and a catch wider than the null check gets rejected.

The agent updates the rate table "so the default region has a row", because the null guard still threw on a missing row. That file is on the deny list, so the run stops, and a person decides whether the missing row is a data bug that needs its own ticket. Letting the agent carry on into the rate table would have mixed a product fix with a content change and made the revert messy.

A second ticket, after the first one lands

The next morning someone asks for the same null guard on quotes, which is a different route and a copy of the function. It's tempting to reopen the agent on the same branch and say "do quotes too". Don't. The first pull request is either merged or still under review, and a second goal on that branch means reverting either fix reverts both.

Open fix/quote-tax-region-null. Allow api/quote/tax.ts and its test. Point at the checkout test as the pattern, and write the non-goal again: do not change checkout, rates or UI. The plan should say it will mirror the helper call rather than invent a shared abstraction. A shared helper used by both routes is a third ticket, and only if a person wants that design. Two similar guards are easier to read than a premature extraction the agent named resolveTaxContext.

That's the normal pace: one error, one pull request, then a second that looks almost the same. The speed comes from the second ticket being obvious, not from one run that "handles tax properly" across the codebase.

Review, rollback and who is accountable

The person who started the run is its author. Review the change the way you'd review a fast teammate's work:

  • Does the diff match the allow list?
  • Do the new tests fail if you revert the fix?
  • Did the agent delete an assertion or skip a failing test?
  • Are errors still logged the way on-call expects?

AI code review can flag a missing test before a person reads the diff, but it shouldn't be the only approval on auth, money or migrations. Two agents agreeing with each other isn't a second set of eyes.

Small pull requests revert cleanly. A run that rewrote twenty modules because the prompt said "clean this up" won't. If checkout breaks, your team is on call, not the vendor, so keep the tasks small.

Rollback in practice

A merge is only easy to reverse if the revert covers a single concern. A pull request that mixed the null guard with a renamed export, a moved file and a snapshot update will conflict with everything that landed after it. The guard was one line. The rest was the agent filling time.

When a merged agent change misbehaves, revert the pull request. Don't ask the agent to fix forward until you have a ticket that describes the new failure as precisely as the first one did. A prompt that says "checkout is still wrong" just stacks a second guess on the first.

After the revert, write down which check would have caught it: a missing parent-commit run, an allow list that was too wide, a skipped plan, a log nobody pasted. Then add that check. The author of the run is the person who started it, even when a bot opened the pull request, and the reviewer who clicks merge owns that click. If your tooling lets the agent merge when checks are green, turn that off. Green checks mean the named command passed. They don't mean it was the right ticket.

Connecting the agent to the ticket

Agents drift when the brief lives in a chat and the work lives in a tracker. The fix is ordinary: make the ticket the brief.

Before the run, put this on the ticket:

  • Outcome in one sentence
  • Non-goals
  • Files in and files out
  • Tests that define done
  • Who reviews

Then point the agent at that ticket and the repo. Claude Code and Codex both plan more reliably from a document they can re-read than from a thread they half remember. Someone still has to order the queue, notice that two runs are editing the same module, and decide that a bug waits until the migration lands. AI project management covers that layer: tickets, stages and a roadmap built from the work itself. An agent can clear a well-cut ticket, but it's a poor way to decide which ticket comes next.

A workspace that turns Slack threads, call notes and error reports into tickets gives a run a cleaner start than a blank prompt. Falrow does that, then lets the agents read the repo through its agent setup, with the tool list in an MCP catalogue the team can audit. A tracker, a repo instruction file and a CI check can get you the same setup. Either way, the agent should start from the ticket.

After a month, look for pull requests tied to a named owner, fewer stray files in diffs, and review comments about behavior. If you see those, widen the tickets you hand over: bugs with a stack trace and tests first, mechanical refactors next. Leave product decisions and sensitive surfaces with the people who already own them.

When two runs want the same file

This happens on small teams sooner than people expect. One run is fixing the null region in tax.ts. Another is changing how rates are loaded in the same file. Both read the file at the same commit. Whichever writes second either overwrites the first or produces a merge conflict, which the agent then "resolves" by keeping its own version.

The version check stops the silent overwrite, but it doesn't decide which ticket wins. A person does. Usually the smaller bugfix lands first, and the rate-loading ticket rebases and re-reads. If the rate-loading change is the reason the null bug exists, the bugfix waits, and the ticket says so. An agent won't notice that ordering, because all it sees is the file.

A practical rule: before you start a run, search open branches for the allow-list paths. If one is already dirty, don't start; queue the ticket behind the open pull request. The wait is minutes or hours. A conflict from two agents editing tax.ts costs the same afternoon the agents were supposed to save.

What "widen the tickets" should mean

After a month of clean small runs, widening means a longer allow list that is still a list. A mechanical rename across a package, with the package's test command as the stop, is a fair next step. "Own the tax module" isn't widening. It drops the list entirely.

A widened ticket still has non-goals and a reviewer, and it still stops after the named command. The plan checkpoint matters more here, because the diff is longer and the wrong abstraction costs more to revert. Approve the plan in a sentence: "Rename only. No behavior change. No new helpers." If the plan mentions a redesign, send it back before any edit.

Keep a separate queue for work agents don't start: product behavior nobody has decided, auth and payments changes, and migrations that rewrite customer data. Those can use an agent later, on a slice, once a person has written the slice down. That queue is what keeps control. Without it, a tool that can see every ticket will pick up the dangerous one because it was at the top.

FAQ

What are AI coding agents?

They're tools that work toward a coding goal by calling other tools in a loop: reading the repo, editing files, running tests and often opening a pull request. With a chat assistant that only suggests code, you still apply every change yourself. An agent applies changes on its own, so scope, branch rules and review need to exist before the run. Whoever starts the run stays responsible for what merges.

Are they safe on a production codebase?

For scoped tasks, yes, as long as the default branch is protected, the agent has no production credentials and a human reviews the diff. They're a poor fit for unattended edits to auth, payments, permissions or migrations that rewrite data. What keeps you safe is allow lists, version-checked writes, required tests and small pull requests, not the model.

How do they differ from a pair programmer?

A pair programmer suggests the next lines while you type and stays inside your editor session. An agent takes a goal and works through the steps on its own branch, with less supervision. Many teams use both: the copilot for the edit under your cursor, the agent for a ticket-sized change that already has tests and a reviewer. If you can't write a stop condition for the task, stick with the copilot.

Do they replace project managers or reviewers?

No. They don't set priorities, talk to customers or own incidents, and they shouldn't be the only reviewer on a change that could harm users. Leads still cut work into tickets a run can finish, and reviewers still read the diff, especially for behavior the tests don't cover. What you gain is shorter cycle time on work that was already specified, with the same person accountable at merge.

More on ai agents & mcp