A compliance coordinator at a regional insurer had been producing the quarterly regulatory summary for six months, and it had become one of the more reliable things she did. Internal audit notes, the latest regulatory guidance, departmental updates: her structured workflow pulled them together, and the results were consistently strong enough that her director praised the clarity and the legal team used the summaries as reference material ahead of board meetings. In the seventh month, one paragraph described a new reporting requirement as though it were already in force. It was not. The requirement had been proposed and was still under consultation, a distinction that mattered a great deal to the people reading the summary and mattered not at all to the model that produced the sentence, which had simply drawn on source material that blurred exactly that line.
Nobody in that situation asked whether the AI was at fault. That question doesn't get asked, because it doesn't lead anywhere useful. The summary went out under her name, was read as her work, and would be defended, if it needed defending, by her.
That instinct, to treat the tool as having contributed something separate from what the person is accountable for, tends to grow stronger exactly when the tool's contribution grows larger. When AI produces a full first draft that only needs minor editing, the sense of authorship weakens in a way that light editing of a colleague's work never quite does. You didn't write it. You refined it. And refinement feels like a smaller claim on the result than creation does.
The organisation on the receiving end doesn't share that feeling. A senior associate at a consulting firm once put an AI-generated competitor analysis in front of a client after making what she considered minor adjustments. The client's chief strategy officer challenged one paragraph, which characterised a rival's product launch as defensive, and the client team had direct market intelligence suggesting exactly the opposite: an offensive move aimed at capturing a new segment. She could not defend the characterisation, because she had neither originated it nor independently checked it. The analysis had been plausible. It had also been wrong, and the plausibility was doing all the work that verification should have done.
What she took from the recovery, the revised analysis, the follow-up call, the internal debrief, wasn't embarrassment so much as a change in how she read her own review stage. She had been checking the document as though someone else had written it, when the moment she put it in front of the client, she became its author. That holds whether the document is two pages or forty slides, whether the audience is internal or external, whether the stakes are routine or the kind that follow you afterwards. Once it becomes a daily operating stance rather than an abstract principle, review stops being a check for whether something looks right and becomes a check for whether you're prepared to defend it against someone who knows the subject better than you do.
"Human in the loop" is the phrase most workplaces reach for when they want to sound careful about this without deciding anything. It describes a position in a process, not a standard the person in that position is actually meeting. Someone can be technically in the loop, glancing at a summary before it goes out, nodding at a recommendation before forwarding it, and still be offering nothing that would count as review if the content were ever challenged.
The more useful question is narrower and less comfortable: could you defend what you're sending? Not does this look right, which fluent AI output is specifically good at passing, but could you stand behind every claim in the document if someone who understood the subject better than you pushed back on it. That question forces a different kind of attention. It asks you to locate the specific statements carrying risk, not just to confirm that the piece reads well as a whole, and it exposes the gap between skimming for coherence and having actually formed a view about whether each claim is true.
Answering it honestly doesn't require distrust of the tool or a return to writing everything from nothing. Most AI-assisted work will pass the question without difficulty, because most of it is routine, low-stakes and easy to verify quickly. The value of asking it is that it makes visible the small number of statements in any piece of work where you genuinely haven't checked, only accepted, and those are the ones that create the kind of consequence a compliance coordinator or a senior associate ends up explaining afterwards.
Consequence should set the level of review, and by default it rarely does. A communications director at a public health charity built two safeguards into her AI-assisted process rather than relying on remembering to be careful. External-facing text never went out without a substantive human edit: not a skim, not a formatting pass, but a real engagement with the content that adjusted phrasing and checked that the tone matched the audience. Alongside that, anything touching patient data, funding figures, regulatory language or public health claims required a second reviewer, someone who didn't need to understand the AI workflow, only to read the output with fresh eyes and a clear brief about what to watch for.
Over eighteen months, that second reviewer caught three issues the communications director had missed on her own: a funding figure rounded in a way that could mislead a donor, a health claim that overstated how settled a preliminary finding actually was, a patient reference that needed further anonymising. None of the three would have caused a crisis on its own. Each one, reaching its intended audience, would have required a correction carrying real reputational cost. The safeguard didn't prove itself with a dramatic save. It proved itself the ordinary way, by quietly catching small problems before they became larger ones, often enough that her organisation extended the same second-review requirement to every piece of grant reporting produced with AI support.
What made both safeguards durable rather than performative was that neither demanded heroics. A rule that costs ten minutes of structured attention survives contact with a busy week. A rule that depends on someone remembering to be unusually vigilant does not. The useful exercise, for anyone building their own version of this, is to think about the errors they are actually likely to make rather than the ones that would be most dramatic if they happened. Quiet distortions, confident overstatements, missing qualifiers: these are worth designing against, because they are common enough as a category, even if unpredictable in timing, to justify a permanent check rather than an occasional one.
None of this settles what a court or a regulator would decide about liability in any particular jurisdiction or profession, and that question sits well outside what a workflow habit can answer. What it does settle is the more immediate, more common one: whether the person sending the work has a process they could point to, and a clear sense of what they actually own.
Every professional who uses AI seriously enough will eventually watch something they sent forward turn out to be wrong. What differs is not whether that happens but what it reveals. For some it will expose an approach that mistook fluency for verification. For others, it will confirm a process built well enough that the error is the exception rather than the pattern it might otherwise have become. The tool was never holding the responsibility in the first place. It was always waiting with you, whether or not you had noticed.
Anthony Velland
The distinction between generating and owning a piece of work, and the practical review habits that protect it, is developed in full in AI Without Guesswork.
Ad · Amazon affiliate link.