Skip to content
AI agents & MCP

Codex and GitHub: From Issue to Pull Request With Guardrails

By Sebastiaan Jansen · · 7 min read

A codex github workflow takes a GitHub issue, gives a coding agent a branch and a written task, and gets back a pull request a person can review. The agent reads the repository, edits only what the issue allows, and stops at tests and review. Whether that pull request is safe to ship depends on permissions, checks, and merge rules.

How a codex github workflow moves an issue

The unit of work is the pull request. A chat disappears when you close it. A pull request keeps a branch, a diff, a description, checks, and reviewers. GitHub's pull request model was built for that kind of review, and an agent should use it the way a teammate would.

  1. Someone writes an issue with the outcome, the files in scope, and the checks that define done.
  2. The agent branches from the default branch and never commits to main.
  3. It reads the named code and the tests next to it, makes the edit, runs the named checks, and pushes.
  4. The pull request says what changed, what stayed the same, and which checks ran.
  5. A human reads the diff. Required reviewers and status checks still apply, and a person merges.

The loop breaks on a vague wish ("make billing better") or on an agent that can push to a protected branch. Use codex github for a bug with a failing test, a small API change with a written contract, or a migration that already has a plan. Leave wide redesigns to a person. Agents are fast at local edits and bad at deciding whether the product should change.

Set up an openai codex github connection with least access

An openai codex github connection needs only enough access to clone, branch, and open a pull request. Extra rights won't make the diff any better.

Give a machine user or a GitHub App read access, write on feature branches, and permission to open pull requests. Leave out org admin and any right to bypass branch protection. On the default branch, require a pull request and status checks, and dismiss stale reviews when new commits land. If a person can't push to main, the agent shouldn't either. Keep CODEOWNERS on so billing/ and auth/ wait for the same owners as before.

Scope, secrets, and the base branch

Put the boundary in the issue: the paths the agent may edit, the behavior those edits must produce, and what is out of scope (drive-by refactors, dependency bumps, files it must not touch). Done means the commands that must pass, plus the manual check a reviewer will repeat.

Name the base branch, because main and release/2.4 are different tasks. Give the agent CI's test setup or documented fixtures, never production credentials or a deploy token. For a third-party test, a sandbox key is enough. Agents quote context back, so rotate any token that ends up in an issue or pull request body.

Where GitHub Copilot and Codex take different jobs

People comparing github copilot codex setups are choosing between an editor assistant and an agent that opens its own pull request. The two can share a repository.

Use the editor when a person is already in the file. You accept or reject each suggestion, run the tests, and open the pull request yourself. That suits exploration and code you are still designing.

Use the issue-to-PR agent when the task is already specified and you want a draft with CI attached. Here the guardrails are branch protection, required checks, owners, and a written scope. A green check means those tests passed. It doesn't tell you they were the right tests, or that the diff matches the issue.

What the issue must contain

Treat the issue as the spec. A new teammate should be able to implement it without a meeting.

  • Outcome: what is true after the change, in one short paragraph.
  • Current behavior: an error string, a request example, or a failing test name.
  • Constraints: APIs you must not break, flags to respect, libraries to keep.
  • Files: the paths in play, and a do-not-edit list.
  • Checks: exact commands, such as npm test -- billing/retry and npm run lint.
  • Out of scope: nearby ideas the agent should leave alone.

A pasted transcript is not a spec. If the work crosses several services, split it. AI project management works the same way outside GitHub: each ticket needs an outcome, a file boundary, and a definition of done.

Guardrails on every agent pull request

Put the gates in GitHub settings and the issue template. An agent can ignore a prompt, but it can't get past branch protection.

Gate What it stops Who owns it
Protected default branch Direct pushes, including from the agent Repository admin
Required status checks A merge that never ran tests CI config
Reviewers and CODEOWNERS Silent edits in auth, billing, or infra Code owners
File list in the issue Unrelated refactors in the same PR Issue author
Secret scanning Tokens committed inside a fix Org security settings
Human merge A green check treated as approval The reviewer

Name the commands that ran and the ones that didn't. "npm test -- billing/retry, end-to-end tests skipped" tells a reviewer more than "tests passed." On a large pull request, read the tests themselves. A test that only asserts a mock returned true hasn't proved anything about the behavior.

AI code review keeps the same split: a tool can flag risk, and a person decides whether this is the change the issue asked for. The agent may push fixes in response to comments, but it may not resolve a thread it disagrees with. Any commit after approval restarts review.

Worked example: a failed invoice retry

On a provider timeout the invoice stays pending. A later retry may complete it once, and a timeout must never set succeeded. Right now billing/retry.ts treats any response without status: "failed" as success, so a timeout can send a receipt for a charge that doesn't exist. The issue limits edits to that file and billing/retry.test.ts, which must cover "timeout leaves the invoice pending." Pricing, receipts, migrations, timing, and queues stay untouched. The checks are npm test -- billing/retry and npm run lint, on fix/invoice-retry-timeout from main.

The agent adds the test, watches it fail, then changes the code so only an explicit success counts as succeeded. The pull request notes that the receipt sender and retry interval are unchanged. But the test uses a fresh invoice id, so it misses a second attempt on an invoice that is already pending. The reviewer asks for a timeout followed by one charge. The agent pushes that commit, review restarts, the billing/ owner approves, and a person merges.

How planning and tool access fit

A GitHub issue makes a good handoff when the work already lives there. A Slack thread, a call note, or a production error with no spec makes a poor one. An agent that reads the repo, the issue, and the team's rules writes a smaller diff than one working from a pasted paragraph. Claude Code for teams plans from shared context, then lets a person reject the write.

Keep tool access as narrow as repo access. An MCP server can expose issues, docs, and repo actions, so turn on only what the task needs. Falrow lists those tools on its MCP tool catalogue. Its agents feature lets Claude Code and Codex plan from the git repository under team rules, with version-checked writes.

Record the revision the agent read so a later push can refuse a file that has moved. Git still won't catch two clean edits that contradict each other. Once branches are waiting, reviewer time becomes the limit.

Falrow is in private beta for teams that want tickets, cycles, and forecasts in one self-hosted SQLite database. Access and pricing are on request at /access. For the agent product setup, see OpenAI Codex for software teams.

FAQ

Does Codex open the GitHub pull request itself?

Yes, if its GitHub App or machine user can push a branch and open a pull request. Branch protection, required checks, and required reviewers still apply, and a person merges after those gates pass.

Is GitHub Copilot the same product as Codex?

No. Copilot suggests edits while a person is in the file, and that person usually opens the pull request. An agent takes a written issue and returns a branch. Use the editor for exploration, and the agent when the ticket already has a file list and test commands.

What should a GitHub issue contain before an agent starts?

The outcome, the current behavior, files in and out of scope, the base branch, and the exact commands that must pass. Add a failing test or a concrete error when you have one. Leave out transcripts and broad goals.

Should an agent merge its own pull request?

No. It may open pull requests, but it may not bypass branch protection or approve its own code. A person merges once the diff is accepted and the checks are green, and new commits dismiss the old approval.

More on ai agents & mcp