P17 min readUpdated

AI in QA without fake confidence

I use AI to speed up API and performance work. I do not let a model decide that a release is safe. Where it helps in QA, and where a fluent answer gets mistaken for evidence.

AI draftHuman oraclePASS / FAILobservables only

The dangerous moment with AI in QA is a quiet one. Someone pastes a user story into an assistant, gets back thirty tidy test cases and a paragraph saying the feature is well covered, and attaches both to the ticket. Nothing in that output has touched the build. It reads like evidence, so it gets treated as evidence.

I use AI in my own work. On recent engagements I have used Cursor and ChatGPT to speed up REST API automation in Bruno and performance checks in k6, and it has taken a lot of typing out of the week. It has not changed who decides whether a check passed.

Where AI helps in software testing

AI is useful in QA when the task is repetitive scaffolding and the result can be checked against the system. Four jobs fit that description in my work.

  • Drafting API collections. Requests, headers, environment variables and first-pass assertions for Bruno or Postman. The wiring is tedious to type and quick to check by running it.
  • Shaping k6 scenarios. The structure of the script, the stages and the thresholds, from a description of the journey. The endpoints and the numbers come from me and from Engineering.
  • Summarising logs. A first read of a long log or trace that tells me where to look. Then I look.
  • Proposing edge cases. Inputs and states I might not have thought of. Each one is a suggestion until someone has tried it on the build.

What these share is the direction of trust. The assistant produces something I can run, read or try, and the system under test tells me whether it was right. Once that direction reverses, and the assistant is telling me about the system, I am holding an opinion and not a result.

The rule: the assistant drafts, a human owns the oracle

An oracle is how you know a result is correct: the expected outcome and your reason for believing it. In functional testing that reason is the acceptance criteria, a third party’s contract, a business rule or a conversation with Product. It is never the tool that typed the check.

Pass or fail still comes from observables: the status code and body of a response, a timing, a row in the database, an email that arrived, a release gate everyone agreed. An assistant can suggest an assertion. Whether it is the right assertion is a question about the business, and the assistant was not in the refinement session.

TaskWhat the assistant draftsWhat turns it into evidence
API checkRequests, variables and a first assertionI rewrite the assertion from the contract and watch it pass and fail on a real build
k6 scenarioScript structure and load shapeThresholds agreed with Engineering, a named environment, results read by a person
Log summaryWhere the errors seem to clusterI open the lines it points to before repeating the claim
Edge casesA list of candidatesEach is executed on the build or recorded as not tried
Test cases from a storySteps and expected results that restate the storyProduct confirms what correct means, and a person runs them

What fake confidence looks like

The practical risk is invented certainty: pretty reports that never touched the system under test. I watch for it in my own work first. A generated assertion that passes because it checks what the response happened to contain. Test cases derived from a story that restate the story, gaps included. An edge-case list attached to a ticket as though listing were testing. A log summary quoted in a readiness note by someone who never opened the log.

In each case fluency has replaced contact with the product. We are used to judging a colleague’s understanding by how clearly they write, and that habit does not transfer. An assistant’s output is fluent whether it is right or wrong, so the polish tells you nothing. The same failure exists without AI, in pipelines that stay green over a broken journey. AI makes it cheaper to produce.

Keeping AI out of the release decision

I keep AI away from silent authority. No model output becomes a gate by itself. A summary can tell me where to look and a draft can save me the typing, but the line in the readiness note that says a journey was verified has a person behind it who ran it on a named build.

If confidence is low, or the domain is policy-sensitive, a person reviews before anything is acted on. That covers money, personal data, access rights and anything a regulator takes an interest in. It is the same design I expect from any product that mixes automation with judgement: the automated part handles the volume, a named person makes the call, and the record shows which was which.

Two working rules sit beside that. Client data, credentials and production payloads stay out of the prompt unless the client has approved the tool for them. And whatever an assistant drafted is reviewed like any other change before it joins the pack, by someone who can say what each check is for.

Use AI like a junior pair that types fast and occasionally hallucinates. Pair it. Do not promote it to release manager.

The objection: it is right most of the time

Often it is, and that is the difficulty. A reviewer who has read twenty correct drafts reads the twenty-first less carefully, and nothing in the text marks the one that is wrong. So I size the review to the consequence and not to my recent luck. A drafted request to a search endpoint gets a read and a run. An assertion on a payment amount is rewritten from the contract every time.

The other thing I hear, usually from whoever holds the budget, is that AI could write the regression suite. It can certainly type one. The decision about which journeys must hold every release still belongs to the people accountable for them, and a large suite generated without that decision is a maintenance bill with a green badge on it.

Some testers take the opposite position and refuse to use it at all. That costs them the typing time and buys no safety, because the safety was never in the typing.

The practical win is cycle time: less of the week on wiring boilerplate, more on the paths that actually break users. It holds only while evidence stays evidence. When a release goes wrong, nobody can stand behind “the assistant said it was covered”. They can stand behind a named person, a build number and the result that person saw.

Questions I get asked

How do you use AI in software testing safely?

Use it for work you can check against the system: drafting API requests, the structure of performance scripts, a first read of logs, candidate edge cases. Verify each against a real build before relying on it. Keep expected results owned by a person, keep client data out of tools the client has not approved, and never let generated output act as a release gate without a named reviewer.

Can AI write test cases from user stories?

It can draft them, and the draft will look complete. The weak point is the expected results, which restate the story along with any gaps or wrong assumptions in it. Treat the output as a starting list. Product confirms what correct looks like, a tester adds what the story left out, and nothing counts as tested until it has been run on a build.

Should AI-generated summaries go into a release readiness report?

Only as a pointer to evidence. A generated summary of logs or test results can help a tester find what matters, but the statement that reaches decision-makers should come from a person who has checked the underlying data on a named build and environment. If nobody opened the source, the summary is an opinion and the report should not present it as a result.

Filed under API checks, automation and AI in QA.