Back to Blog

Human-in-the-Loop: The Verification Seam That Makes Agentic Development Actually Ship

Human-in-the-Loop: The Verification Seam That Makes Agentic Development Actually Ship

Everyone selling AI dev tooling in 2026 has the same demo: prompt, watch the agent write a feature, merge. The demo always works. The day-two reality - the part nobody films - is the agent confidently shipping the wrong feature, or shipping the right one while quietly breaking three other pages. By the time anyone notices, the agent has moved on, the diff has scrolled off the screen, and someone is rolling back on a Friday afternoon.

The fix is not a smarter agent. The fix is a verification seam: one place in the loop where a human renders a verdict, and the rest of the cycle - generation, replay, ticketing, the next agent run - arranges itself around that verdict. Six stages. One human click. The loop closes on itself. This post is how we run that loop at Lastest, why it works, and how we have it dialled in tightly enough that we are starting to put the orchestrator itself on autopilot.

Wireframe of the six-stage agentic dev cycle. Tests cover the major areas, an AI subagent makes an improvement, Lastest replays and surfaces a diff, a human renders one of three verdicts, the verdict creates a ticket, and the agent picks the ticket up - closing the loop. The human seam in the centre is highlighted as the only stage with irreplaceable human judgment.
Six stages, one human seam. Everything outside the seam is either commodity infrastructure or zero-marginal-cost replay.

The six-stage agentic development cycle

The shape of the loop matters more than any single tool inside it. We have iterated on this exact diagram for about eighteen months and it has stabilised into six stages, in this order:

  1. Tests cover the major areas. Before an agent touches anything, the product has a baseline of Playwright specs plus visual baselines for every user-facing surface that matters: cart, checkout, auth, settings, billing, the dashboard, the marketing pages a sales call lands on. Not every pixel - the major areas. This is the floor.
  2. An AI dev makes an improvement to an area. A subagent picks up a ticket and edits one area. Cursor, Claude Code, Codex, whichever - doesn't matter. The constraint is that one subagent runs against one ticket against one branch.
  3. Lastest replays and surfaces what changed. The whole spec suite runs against the new branch. Three diff engines (pixel, structural, perceptual) compare current against the approved baseline and against the uploaded design or spec. Anything not green lands in a sorted review queue.
  4. A human renders the verdict. The reviewer opens the queue, looks at the baseline-vs-current-vs-diff triple, and clicks one of three buttons: pass, miss, or regression. That is the entire role.
  5. The verdict files the next ticket. Pass closes the original ticket. Miss reopens it as an improvement (the build doesn't match the intent). Regression spawns a bugfix ticket (something that was working is now broken). The ticket body is autogenerated from the diff - the human never types a triage description.
  6. The agent picks the ticket up. The orchestrator hands the new ticket back to stage 02. The loop continues. Approved baselines from stage 04 become the new floor at stage 01.

This is the entire architecture. There are no other stages and there is no other place where the human is needed. The reviewer at stage 04 is the only human in the picture, and their entire job collapses to picking one of three lanes.

Why the human seam belongs at stage 04, not stage 02 or stage 03

Every team building one of these loops eventually argues about where the human goes. Almost every team puts the human in the wrong place on the first try. There are three intuitive options and only one survives contact with a real workload.

Option A · Human at generation (stage 02)

The human writes the prompt, watches the agent type, intervenes when it goes off the rails. This is what most pair-programming demos show. It does not scale - you have made the agent a transcription tool and capped its throughput at one human's attention span. By Wednesday afternoon, the human is exhausted and reading line-by-line through a 600-line refactor. The agent's leverage is gone.

Option B · Human at replay (stage 03)

The human writes the assertions, the agent generates the code, the replay runs the assertions. This is the world before AI agents and it is exactly the world that produced the 40% test-maintenance tax. Every time the agent expands the surface, somebody has to extend the assertion library. The agent generates features at 10x; the human writes assertions at 1x; the assertion library becomes the bottleneck within two sprints.

Option C · Human at the verdict (stage 04)

The agent does the work, the suite replays itself, the diff engines do the comparison, and the human looks at the result. One screen. One of three buttons. No prompt writing. No assertion writing. No code reading unless they want to. This is the only seam where the cost of human attention scales sub-linearly with the surface area of the app, because the diff engines have already done the filtering before the human gets there.

The pattern only works if the diff is good enough that the human can trust the filter. Which means: visual diff against the approved baseline (regression-safety), visual diff against the uploaded design or spec (intent-conformance), three engines disagreeing on purpose so cosmetic noise can be told apart from structural drift. We covered why this matters in Validation Is the Final Frontier of AI Development - the short version is that the verdict cannot live anywhere else in the loop, because nowhere else has enough signal for the human to be cheap.

The subagent topology: one orchestrator, four lanes, one verdict

Inside stages 01-03 and 05-06, the work is split across four specialised subagents. They never talk to each other directly. They report to an orchestrator that reads tickets, spawns lanes, and waits for the verdict to come back from the human at stage 04.

Diagram of the subagent fan-out. An orchestrator at the top spawns four lanes: a test author (subagent A), a feature implementer (subagent B), a validator running Lastest (subagent C), and a ticket scribe (subagent D). All four converge at a central Lastest review surface where the human renders the verdict. A dashed feedback arrow returns the verdict to the orchestrator to seed the next iteration.
Four lanes. Each subagent has one tool and one output. The orchestrator never reviews - that is the human's job.

Subagent A · Test author

Crawls the app, drafts Playwright specs for the major areas, captures the first visual baseline of each one. Runs once at the start of a project, then incrementally when a brand-new area appears (a new page, a new modal, a new flow). Its tool is Playwright codegen plus the Lastest spec recorder. Its output is one spec file and one baseline per area. It does not edit code; it only adds coverage.

Subagent B · Implementer

The one most teams already have: a Cursor or Claude Code session that reads the ticket, opens a branch, makes the change, commits. The thing to notice is that B never decides whether its own work is good. It only ever produces a branch. The verdict on whether the branch lands is rendered three stages downstream.

Subagent C · Validator

Replays the entire spec suite against B's branch, runs three diff engines against every baseline, and sorts the resulting flags by severity. Pixel-only deltas under 1% drop to the bottom. Structural shifts (DOM changes, layout reflow) bubble up. Perceptual ship-blockers (contrast collapses, missing-icon class shifts) sit at the top. The sorted queue is what the human sees at stage 04.

Subagent D · Ticket scribe

Reads the verdict from stage 04, looks up the diff and the suspect commit range, and writes the next ticket. The ticket body is templated - the scribe doesn't write prose, it fills slots. Verdict = pass means "close, update baseline, link to merge." Verdict = miss means "reopen, attach intent diff, assignee = B." Verdict = regression means "new ticket, attach baseline diff, suspect commit, severity, assignee = B." All three ticket shapes are deterministic.

The fan-out matters because it isolates failure modes. If the implementer subagent goes off the rails, only stage 02 is affected - the validator still runs, the human still sees the bad diff, the verdict still fires. If the validator misbehaves (a false positive on a font swap, say), the human catches it in one click and moves on. No subagent owns the verdict, which means no subagent can corrupt the loop.

The three verdicts and what they produce

Almost every other piece of the system flows from the fact that the human has exactly three buttons. Two-button systems (approve / reject) lose the distinction between "you didn't build what I asked" and "you broke something." Four-button systems (approve / minor / major / reject) ask the human to do triage they cannot do at the speed the loop demands. Three is the right number because it maps one-to-one onto the three states the next ticket can be in.

Three-column diagram showing the verdicts. Pass produces a close-ticket with shipped baselines. Miss produces an improvement ticket because the build does not match the design. Regression produces a bugfix ticket with a suspect commit and severity. Each column shows the diff signal that triggers it, the templated next ticket, and how the agent picks the work up.
Three verdicts, three deterministic ticket templates. The human picks the lane; the scribe writes the ticket; the orchestrator routes the next iteration.

Pass · close the ticket

All diff engines are green, or the only deltas are inside the region the ticket said would change. The reviewer accepts. The branch merges. The new screenshots become the next baseline. Lastest signs the approval with the reviewer's identity and a timestamp, so three months from now the audit trail can answer "who approved this and when." The orchestrator moves on to the next ticket on the queue.

Miss · improvement ticket

This is the verdict that two-button systems lose. The build runs. The build doesn't regress anything. But the build is not what was asked for - the agent re-interpreted the design, or shipped a slightly different layout, or used the wrong colour token. The intent diff (current vs uploaded design) is red even though the baseline diff (current vs previous baseline) is green. The reviewer clicks miss, the scribe reopens the ticket as an improvement with the gap quoted out of the spec, and subagent B re-prompts itself with the new context. The branch stays open.

Regression · bugfix ticket

Something that used to work is now broken. The baseline diff is red somewhere the ticket did not say it would change. The reviewer clicks regression, the scribe spawns a new bugfix ticket linked to the suspect commit, the merge is blocked, and the orchestrator routes the bugfix back to subagent B with the diff attached. The original ticket stays open behind it. The bugfix gets the same six-stage treatment as everything else.

The reason this collapses to three buttons is that there is no fourth verdict that the loop cannot already handle. "Approve with comments" is just miss with a longer note - the scribe captures the note and routes it as an improvement. "Approve but please also…" is two tickets: pass on the first, new ticket for the second. Two tickets is cheaper than one ambiguous verdict.

What a day in the loop actually looks like

We have been running this loop on our own product for about a quarter. On a typical day, the orchestrator runs between five and ten iterations, the human spends fifteen to thirty minutes total at the review surface, and the bulk of the day is silent agent work that the reviewer never sees because nothing got flagged.

Horizontal timeline of one workday from 09:00 to 17:00 showing eight loop iterations. Most of the day is silent agent work in grey. Short coloured bars mark the moments the human reviews a diff - five pass, two miss, one regression. Total human attention across the day is twenty-two minutes. The day shipped six merges.
A single workday: 8 iterations, 6 merged, 22 minutes of human attention, 0 silent regressions. The seam costs about three minutes per merge.

The shape of the day matters more than the raw number. The human's attention arrives in punctuated bursts - two to four minutes per iteration, scattered across the day. It does not block on agent work, because the agent is running in the background. It does not block agent work, because the agent has already moved on to the next ticket the moment the verdict fires. There is no all-hands review meeting, no end-of-day diff pile, no Friday-afternoon rollback because the agent shipped something on Wednesday that nobody noticed.

The day in the diagram had one regression caught at 10:41 - subagent B introduced a contrast collapse on the checkout CTA. The seam caught it the moment validator C ran, the bugfix ticket was on the board by 10:42, B was already drafting the fix by 10:43, and the recovery merged before lunch. Without the seam, the same regression ships, lands in production overnight, gets noticed by a customer in the morning, and costs a sprint to triage.

The throughput numbers: why "no seam" is more expensive than the seam

The most common objection to a human-in-the-loop seam is that it slows the loop down. It does, in one specific sense: a setup with no validation at all merges PRs faster than a setup that asks a human to look at a diff. The argument unravels the moment you measure anything else.

Grouped bar chart comparing three setups across four metrics over a fourteen-day window. The no-validation setup ships 48 PRs per week but produces 9 silent regressions and 5 rollbacks. The LLM-as-judge setup ships 44 PRs but still produces 6 silent regressions and 3 rollbacks. The Lastest human-in-the-loop setup ships 41 PRs, with 0 silent regressions, 0 rollbacks, and an average of 4 minutes of human review per PR.
Setup C ships 85% as fast as the no-seam setup, with zero silent regressions and zero rollbacks. The seam tax is four minutes per PR; the cost it avoids is a week per incident.

Three setups, fourteen days, two internal apps, 133 merged PRs in total. Setup A merges fastest because no human is in the way - and produces nine silent regressions in two weeks, five of which required a rollback. Setup B (LLM-as-judge) ships almost as fast as A and still leaks six silent regressions, because another model is not an audit trail. Setup C, the human-in-the-loop seam, ships 85% of A's PRs and produces zero silent regressions and zero rollbacks.

The seam costs four human minutes per PR. The thing it avoids costs roughly one engineer-week per silent regression (triage, hotfix, communication, sometimes a public postmortem). At anything more than one regression per quarter, the seam pays for itself an order of magnitude over. At the rates Setup A actually produces them, it pays for itself thirty-eight times over inside two weeks.

Where this strategy sits among the alternatives

The seam is not a competitor to type checks or unit tests or scripted E2E. It is the only strategy that lands in the upper-right of the confidence-versus-coverage matrix, because everything else trades one axis for the other.

Two-by-two matrix plotting validation strategies. Type checks and unit tests sit low coverage but high trust. Scripted E2E and pixel-only visual regression sit mid coverage mid trust. LLM-as-judge sits high coverage but low trust. Only the Lastest human-in-the-loop seam - visual diff against two references plus three diff engines plus a one-click human verdict and an audit trail - lands in the top-right quadrant.
Every strategy is useful. Only one of them clears both axes alone.

You still want type checks and unit tests - they catch a specific bug class at zero marginal cost. You still want scripted E2E for flows that need procedural assertions ("user can complete a purchase end-to-end"). You probably want an LLM triage layer in front of the human, so the queue gets pre-sorted. None of them are sufficient on their own. The seam is what makes them composable: each strategy catches what it catches, the visual diff catches everything else, and the human renders one verdict across the combined signal.

How to bootstrap the loop in your own project

The whole architecture sounds expensive to build, and it can be, but the practical bootstrap is small. We did it ourselves on the Lastest codebase in about a week of part-time work. The steps map one-to-one onto the six stages.

  1. Stage 01: pick five areas, write five specs. Cart, checkout, auth, settings, the highest-traffic landing page. Generate Playwright specs with the codegen, record the first baseline for each in Lastest. You now have a floor. Skip the long tail - the major areas are 80% of the regression-catch value.
  2. Stage 02: route one Cursor/Claude/Codex session per ticket. Treat the agent as a non-reviewing implementer. One ticket per branch. If you have a Linear or GitHub project, the orchestrator can be a single shell script that pulls the top ticket off the queue.
  3. Stage 03: point Lastest at the branch on every push. The free tier of lastest on GitHub covers this end-to-end - pixel, structural and perceptual diffs out of the box. Wire it into the existing CI; the run takes about three minutes for a hundred specs.
  4. Stage 04: assign one reviewer. The seam needs a single person on the hook for the verdict. Pair-of-eyes review is fine but not required. The reviewer reads the queue twice a day and clicks one of three buttons.
  5. Stage 05: turn on the three ticket templates. Lastest can emit a webhook on each verdict with the diff link and suspect commit attached. Wire it to your ticket tracker so close / improvement / bugfix all land as templated tickets, not free-text comments.
  6. Stage 06: let the orchestrator pull the next ticket. The simplest orchestrator is a cron job. The next-simplest is a daemon that listens to the ticket-tracker webhook. The fanciest is the orchestrator we run, but you do not need it on day one.

The loop becomes self-improving once it is closed. After about a sprint of running, the diff queue gets shorter (the agent learns from being repeatedly rejected on the same class of regression), the baselines stabilise (fewer cosmetic deltas because the agent stops touching things it does not need to), and the reviewer's daily attention budget drops from thirty minutes to fifteen. The loop pays compounding interest. The seam is not a tax - it is the bearing the whole machine runs on.

The next move: putting the orchestrator on autopilot

We mentioned this in the intro and we will return to it now: once the loop is closed and stable, the orchestrator itself can be made autonomous. The agent does not need a human to pull the next ticket; it only needs a human at the verdict. We are running an internal experiment right now where the orchestrator runs overnight, the implementer ships against the queue, the validator replays, and any iteration with a non-green diff parks itself at the top of the review queue for the morning. The human walks in, drinks coffee, renders six verdicts in fifteen minutes, and the day's worth of work merges.

This is the part where the seam becomes the entire shape of agentic development. Generation runs forever for free. Replay runs forever for free. The verdict is the only thing that costs anything, and it costs - in our measurements - about three minutes per merge. Everything that is not the verdict has already been industrialised. The teams that win the next year of agentic shipping are the ones that have built the cleanest seam, not the ones with the smartest model.

Recap: the six stages, one more time

If you remember nothing else from this article, remember the shape of the loop:

  1. Cover the major areas with tests and visual baselines.
  2. Let a subagent improve one area.
  3. Replay through Lastest; let the diff engines sort the queue.
  4. Render a verdict: close, improvement, or bugfix.
  5. The verdict files the next ticket, deterministically.
  6. The agent picks the ticket up; the loop continues.

Five of those six stages do not need a human. The sixth one - the verdict - is the seam that makes the other five trustworthy. Lastest is the surface where the seam happens. It is open source on GitHub, free to self-host, and runs in CI in about three minutes per hundred specs. If you have an agent shipping features right now and nobody is reviewing the diffs, you do not need a smarter agent. You need the seam.