Back to Blog

Testing AI-Written Code: The New Edge-Case Problem

Testing AI-Written Code: The New Edge-Case Problem

An AI agent just shipped a PR. Every check is green - types, units, integration, the LLM reviewer that left two thumbs up. The agent has already moved on. Should you trust it?

The honest 2026 answer is: not yet, and not because the model is bad. The model is good at the things you can see. It is systematically bad at the things you can’t - boundaries, null paths, race windows, the places where the contract you mocked has drifted out from under you. None of those show up as red checks. They show up as Friday incidents.

84% of developers use AI tools. Only 33% trust the output. We unpacked that adoption gap in AI Test Generation: Hype vs. Reality. That post asked what happens when AI writes the tests. This one asks what happens when AI writes the production code - a much bigger surface and a much louder failure mode.

What AI-written code gets right

AI implementers are excellent at the parts of code where most of the training data lives. Syntax is correct. Imports resolve. Patterns are idiomatic. Happy path runs on the first try about 88% of the time in our triage. CRUD, validation, simple forms, glue code - solved.

This is the layer surface-level testing exercises. Types pass because the AI’s annotations match its own code. Unit tests pass because the AI wrote them against the inputs the AI thought of. The first three layers of the pyramid go green almost by construction - the AI’s model of the world is internally consistent.

What it gets wrong, systematically

Past the happy path, the pass rate falls off a cliff.

Grid wireframe of edge-case categories versus pass rates. The happy-path row sits at 96% pass. Boundary conditions, null and empty state, race conditions, integration drift, and business-rule edges fall in a stepped cliff to 22% for integration drift. The visual contrast between the green happy-path row and the red lower rows is the headline.
The same model that passes the happy path 96% of the time passes integration-drift tests 22% of the time. The cliff is at row three and the cliff is the whole story.

Five categories show up over and over:

  • Boundary conditions. Off-by-one on the last page of pagination. Fence-post errors on a slider that goes 0–100 inclusive but is consumed by a function that expects 0–99. The AI tested the middle of the range and called it good.
  • Null and empty state. The first-time-user empty inbox renders undefined. A brand-new account throws on the settings page because the role lookup returned null. These are every user’s first session - the AI’s mental model of “a user” is the seeded test fixture.
  • Race conditions. Double-click submits twice. Optimistic UI shows the new state while the request is in flight and loses. The AI tested it once, sequentially, on a fast machine; production users live in the race window.
  • Integration drift. The backend renamed customer_id to account_id three weeks ago. The mock still says customer_id. Unit tests green, integration tests green against the mock, first real API call in staging throws. Tricentis and the QA Trends 2026 reports call this out specifically: AI-generated code passes surface tests and breaks where the contract has moved under the mock.
  • Business-rule edges. Discount applies twice because nobody fed the loyalty-tier override into the context window. Tax calculated before shipping when the spec says after.

These are not bugs in the model. These are bugs that any system writing code without external observation produces. The fix is not a better implementer; it is a better observer.

Why asking AI to test its own code doesn’t work

The intuitive response is “let the AI write more tests.” This is the trap. The same model that missed the edge case in the implementation will miss it again in the test - a second draw from the same biased distribution, the agreement mistaken for “validation.”

Two-panel wireframe. Left panel shows a closed loop: the model writes the code, the model writes the tests, the model grades itself, ship. Right panel shows the verdict seam: the suite plus visual diff replays feed a queue, and a human renders the verdict that closes the loop from outside.
Left: the closed loop - model writes code, model writes tests, model grades itself, ship. Right: the verdict seam - suite plus visual diff replays and a human renders the verdict. The human is the only ground-truth signal in the picture.

The industry is not transparent enough about this, so let us be: AI cannot write decent assertions for its own code. A test is only useful to the extent it knows something the code being tested does not. When the test author and the code author are the same model with the same context window, the test knows exactly what the code knows - no more. It will assert the things the code already does. It will not assert the things the code forgot. This is by definition.

LLM-as-judge does not save you either. Another model is a correlated draw from a similar distribution. We covered the numbers in Human-in-the-Loop - LLM-as-judge leaks six silent regressions per fourteen-day window; the human-in-the-loop seam leaks zero. The difference is not intelligence. It is that the human is outside the loop.

The new assertion stack

Types and unit tests are cheap and catch a specific bug class at zero marginal cost - keep them. The fix is to recognise they are necessary and no longer sufficient, and add the layers above.

Six-layer stack wireframe, bottom-up: types, unit tests, integration tests, end-to-end tests, visual regression, human verdict. The bottom three layers are shaded as inside the closed loop where the AI can grade itself. The top three layers are highlighted as external observers, with the human verdict on top as the only layer that knows intent.
Six layers, bottom-up. The first three are inside the closed loop - the AI can grade itself there. Layers four and five are external observers. The verdict on top is the only layer that knows intent.

The stack is additive. Types catch type errors. Units catch function-in-isolation regressions. Integration tests catch contract drift - if the mocks are honest, which is the thing the AI cannot self-verify. E2E catches happy-path flow breaks. Visual regression catches what the user sees: undefined rendered to the DOM, a missing empty state, a CTA that collapses to invisible. The human verdict catches the last thing nothing else can: “this is not what I asked for.”

Each layer catches a bug class the layer below is structurally blind to. The verdict closes the final gap - the build can be visually correct and be the wrong feature, and only a human knows the difference.

The velocity gap nothing else closes

The reason this stack matters now is that the agent ships features at 10x while the assertion library moves at 1x. The moment you have a competent AI implementer, hand-written assertions become the pipeline’s rate limiter.

Line chart wireframe across twelve sprints. The AI throughput line climbs steeply to 10x. The hand-written assertion line creeps up to 2.6x. The shaded gap between the two curves represents shipped edge cases. A third line for visual diff plus verdict tracks alongside AI throughput, showing the seam scales with surface area rather than headcount.
Twelve sprints. AI throughput climbs to 10x. Hand-written assertions creep to 2.6x. The shaded gap is where every shipped edge case lives. The visual diff plus verdict line scales with surface area, not headcount.

Every other assertion layer scales with engineer-hours. The diff engines scale with surface area; the human verdict scales with the number of unresolved diffs, which the diff engines have already filtered. One reviewer renders verdicts on ten merges in fifteen minutes when most diffs are green. That is the only economics that keeps up with an agent shipping eight branches a day.

The human verdict still does what nothing else does

There is no AI-only path to agent-velocity shipping without silent regressions. The architecture of the seam is laid out in Human-in-the-Loop. The load-bearing claim: the human is the only entity in the loop not generated by the model, and therefore the only one that can render an external verdict.

The agent writes the code; the suite runs; the visual diff sorts the queue; the human picks pass, miss, or regression. Edge cases caught at the visual layer (because the user sees them) or the verdict layer (because the human knows what was asked for). Nothing about this requires a smarter model - it requires the seam in place.

What to do on Monday

If an AI agent is shipping branches right now and nobody is reviewing the diffs, you do not need a smarter agent. You need the assertion stack to extend past the closed-loop zone:

  • Keep types, units, integration. Floor, not ceiling. Don’t let the AI grade itself on them and call that “tested.”
  • Add visual baselines for every user-facing surface the agent can touch.
  • Wire one reviewer into the loop - three buttons, two short review windows a day.
  • Green diff: merge. Red and intended: accept the new baseline. Red and not intended: you just caught an edge-case regression the AI shipped silently. The seam paid for itself.

Lastest is the open-source surface where this happens - Docker, ~3 minutes per hundred specs in CI, free to self-host. If you are shipping AI-written code into production and the only thing checking it is more AI, try Lastest on GitHub. The seam is the part of the loop you cannot get from a model.