P07 min readUpdated

When more automation hides a broken user journey

Pipelines get greener while users still struggle. Why automated tests pass over a broken journey, and how I keep functional truth ahead of technical coverage.

AI draftHuman oraclePASS / FAILobservables only

Automation is not the enemy. Losing the plot is. Teams celebrate a rising coverage figure and a pipeline that has been green for weeks while a checkout, an onboarding flow or an approval path still fights the person it was built for. The checks are real and they pass. They are answering a question nobody outside Engineering asked.

Test automation hides a broken journey once the number of checks becomes the measure of quality. The more of them there are, the more convincing the green looks, and the harder it gets for anyone to say out loud that the product does not work.

Why automated tests pass while the journey is broken

The tests are rarely at fault. They look where they were built to look.

What the suite doesWhat the user meets
Tests each service on its own, with its neighbours mockedThe handover between them: a booking that is saved but never reaches the back office
Uses seeded, tidy data: a new account, one item, a valid cardAn account with years of history, an expired card and a half-finished order
Finds a field by its identifier and fills itHas to find the field, understand its label and get past the keyboard covering the next button
Asserts that no error was shownNo error, and no confirmation email either
Was updated to match the new behaviour when it failedThe new behaviour was the defect
Counts the lines of code that ranWhether the outcome was right is not counted

None of those rows is bad practice. Isolated tests are fast, seeded data is repeatable, and updating a check after a deliberate change is correct. The trouble is cumulative. Each one moves the suite a step further from what a person does, and the dashboard reports the total as if the distance were zero.

Start from the journeys a user must complete

I start with the business question: what must a user complete for us to have earned this release? The answer is a short list of journeys in the business’s own words. Sign up and reach the first useful screen. Book and pay. Submit a request and have it approved. If Product cannot produce that list quickly, that is the first finding.

Then I ask which of those journeys we have actually exercised, and I mean three things by that.

  • By a human. Someone who reads the screen, hesitates where a user would and notices the thing no assertion was written for.
  • With real data. Not production customers, but data shaped like theirs: accounts with history, awkward names and addresses, a payment method that needs a second step, a record migrated from the old system.
  • On a surface that matters. The phone most users hold, in the hand and not on a device farm alone, the browser the approvers use, the inbox where the confirmation lands. If the journey crosses from app to email to back office, so does the tester.

The result is journey evidence: for each journey, who walked it, on which build, on what, and what they saw.

Journey evidence beside the dashboard (illustrative)
  • Pipeline on build 2.14.0-rc1: 1,480 automated checks, all passing.
  • Journey: a new customer signs up, verifies their email and makes a first booking.
  • Walked by hand on rc1: iPhone, Android phone and Chrome on a laptop, each with a new account.
  • Found: the verification email takes four minutes to arrive. Its link opens the web sign-in page, not the app the customer signed up in.
  • Found: that page asks for a password the customer has not set yet.
  • Automated checks in this area: the sign-up request returns 201 and the verification endpoint accepts a valid token. Both are correct. Neither follows the email.
  • Status: a new customer on mobile cannot complete the journey. Raised with Product before the readiness note is written.

Where automated checks belong

Technical checks still run. I often speed them up with AI so they do not consume the week, on the terms set out in AI in QA without fake confidence: the assistant drafts and I own the expected result. They sit behind the mission, not in front of it.

The order matters. First get evidence that the journey works for a person. Then automate the parts you now understand, so the next release cannot quietly undo them. A check written after someone has walked the journey has an oracle that a person has seen with their own eyes. A check written beforehand, from the story alone, encodes a guess. The same order decides what goes into the regression core: journeys first, then the checks that guard them.

Serve the user. Protect the business decision. Automate what remains.

The smell test: green dashboard, hesitant Product

A useful smell test: if your dashboard is green and Product still hesitates, you are measuring the wrong thing. That hesitation is data. It usually means someone has tried the feature, or watched a customer try it, and what they saw does not match the report.

I do not answer it with more numbers. I ask what they are worried about, in their words, and it is nearly always a journey: they are not sure a new customer can get through. Then we walk it together on the release candidate. Either it works and the worry is retired with evidence, or it does not and we have found what the suite could not see. Afterwards the readiness note leads with the journeys and puts the pipeline figures underneath as support.

The objection: manual testing does not scale

Correct, and I am not asking it to. The human pass I am describing is a handful of journeys per release, not a manual regression of the whole product. It can stay that small because automation carries the rest. What does not scale is finding out from customers.

The opposite objection comes from testers who have been burned: the suite is noise, so stop investing in it. I do not agree with that either. A suite that guards understood behaviour frees the hours that journey work needs. What has to change is what gets reported upward, so that a count of checks stops standing in for the state of the product.

A third reaction is quieter: Engineering feels accused. It helps to say plainly that the suite did its job. It was asked whether the code behaves as written, and it does. Nobody asked it whether a new customer can get from the email to a booking.

With journey evidence in front, the dashboard goes back to what it is good at: an alarm for regressions in behaviour the team already understands. Release conversations get shorter, because Product is reading about things it recognises. The coverage figure can keep rising, and nobody mistakes it for the answer.

Questions I get asked

Why do automated tests pass when the feature is broken for users?

Because they check what they were written to check. Most suites test services in isolation with mocked neighbours and tidy seeded data, and assert on things a developer anticipated. Users meet the joins between systems, awkward real-world data and steps outside the application such as emails. A journey can fail in any of those places while every individual check still passes.

Is high test coverage a sign of good quality?

It shows how much of the code ran during testing, which is useful for spotting areas with no tests at all. It does not show that the right outcomes were asserted or that a person can complete a journey. Treat coverage as a measure of the suite, and use evidence from walked journeys to describe the state of the product.

Should you test a journey manually before automating it?

For a new or changed journey, yes. Walking it first with realistic data on the surfaces users have shows what correct looks like and where it breaks. Checks written afterwards guard behaviour someone has actually seen. Checks written first, from the story alone, tend to encode the author’s assumptions and then pass whether or not the journey works.

Filed under API checks, automation and AI in QA. Terms: Functional testing.