Skip to main content
Some codebases punish a quick edit. The change compiles, the obvious tests pass, and a week later you discover the new code was pasted on top of assumptions it never understood. Get It Right is built around the idea that sometimes you have to try and fail to understand a codebase, so it makes the attempt, checks it mechanically, has a separate agent judge it, and tries again with what was learned.

How a run goes

Each pass through the loop has the same shape:
  1. Implement. An agent makes the change. This is the only step that writes code.
  2. Check. Your lint, test and build commands run in parallel. Any lane you did not configure is skipped.
  3. Review. A separate reviewer agent, which has read-only tools plus a shell, evaluates the result. If a check failed it skips straight back to implementing.
  4. Decide. The reviewer returns a verdict: pass ends the loop; continue goes back to implement with the feedback; refactor first runs an agent that undoes the problematic work (it does not reimplement), then starts a fresh implement pass; stuck pauses for you.
The checks have the last word on success. A reviewer that says pass over a failing build is overridden to continue. The reverse is not true: a reviewer can still fail a green build, because it can see what exit codes cannot, such as a page that renders blank or a handler nobody wired up. stuck is for when the review could not be done, such as a missing tool or an app that will not start, not for hard problems. It escalates to you, and your answer goes to the reviewer rather than to the implementer, so a finished build does not get reimplemented because a reviewer was blocked.
A Get It Right run with a passing reviewer verdict

The reviewer's verdict on an iteration. Example data.

Inputs worth setting

context_bridge controls memory across iterations. With feedback_only, each implementation starts fresh but sees the reviewer’s accumulated feedback and the earlier summaries. full keeps the whole conversation for both. none starts every pass cold and relies on variation between attempts.

Reading a run

Open the chat and each iteration appears as a step with its three check results and the reviewer’s grade, strategy and feedback. The reviewer’s feedback is the only thing the next attempt gets beyond your original request, so it is worth reading: it tells you what the agent believes is wrong.

Cost

Every iteration is at least one implement run and one review run on a flagship model by default, up to five times. Use max_retries to cap it, a cheaper model for mechanical phases, and review_enabled: false when the checks alone are enough.