⏱ 9 min read · 1,719 words

Four Things AI Still Can’t Do in Software Testing

Four Things AI Still Can’t Do in Software Testing

August 17, 2026
⏱ 3 min read · 446 words

Four Things AI Still Can’t Do

Every conversation about AI and testing eventually lands on the same question: what’s actually left for the human?

It’s a fair question. AI now writes regression suites, drafts test data, and produces a first-pass exploratory charter faster than most testers can open their laptop. If you’re measuring the job by output volume, the human side of that ledger is shrinking fast.

But volume was never the hard part of testing. Judgement was. And judgement is where AI still runs out of road.


Curiosity

A good tester notices that the error message changed tone halfway through the flow. They wonder why. They pull the thread. AI doesn’t wonder — it pattern-matches. When the pattern doesn’t exist yet, AI has nothing to anchor to.


Risk judgement

Which failures matter most for this release, for this audience, at this stage of the programme? That judgement requires context about the business, the users, the politics, and the history of the system. AI doesn’t have any of that unless someone gives it to the model carefully — and even then, it can’t weigh it the way someone who’s sat in those rooms can.


Stakeholder communication

When something breaks and a programme director needs to understand whether to hold the release, they need a human who can read the room, explain the risk in plain language, and make a recommendation they’ll stand behind. That is not a transcript of test results. It is a conversation.


Challenge under pressure

A delivery team under deadline pressure will rationalise almost anything. An independent tester’s job is to hold the line — to say “this risk is real, and we shouldn’t ship until we understand it.” That requires confidence, independence, and accountability. None of which are tokens in a prompt.


Why this list isn’t shrinking

Every one of these four requires something AI doesn’t have: a stake in the outcome, a memory of what’s happened before, and the standing to say something unwelcome to someone senior. You can prompt a model to simulate scepticism. You can’t give it a reputation to protect or a Monday-morning meeting to sit through when the release it approved goes wrong.

That’s not a knock on the technology. It’s a description of where the actual value of a tester has always lived — underneath the test cases, in the judgement calls nobody put in the ticket.

The testers who treat this list as their job description, not their fallback, are the ones AI makes more valuable, not less.


Resync is New Zealand’s independent QA consultancy. We test what others build.