A human resources director had been using AI well for months. It drafted internal updates, structured meeting agendas, and turned her rough notes into briefing documents her director actually praised. When a sensitive restructuring came up and she needed to tell a group of employees their roles were changing, she did what had worked every time before. She gave the tool clear context, named the audience, outlined the key messages, and asked for a draft.
What came back was competent. The structure was logical. The language was professional. Every message she had asked for was there. And it was wrong in a way no amount of editing could fix. The sentences were correct without being kind. The tone was clear without carrying any sense that a person, addressing other people about something that would unsettle them, had written it. She spent over an hour trying to soften it, and each change she made improved one sentence while making the paragraph around it feel more artificial than before.
The problem was not her input. She had briefed the task well. The problem was that she had asked AI to do a piece of work that was never suited to it in the first place, and no revision was going to change that, because the fault sat underneath the wording, not inside it.
Most people who use AI regularly at work eventually settle into one of two habits. They either try to route everything through it, on the theory that more use signals more competence, or they pull back from it whenever a task feels important, treating caution as the safer default. Neither habit is really a decision. Both are a way of avoiding the actual question, which is not how often AI gets used but where it belongs.
Some work is structured. It has a definable shape, a clear input, an expected form of output, and a reasonably objective way to check whether the result is good. A first-pass summary, a formatting pass, a routine report built from familiar inputs: these have edges you can see, which means a tool can operate inside those edges without much risk of drifting off course. Other work does not have that shape at all. It depends on context nobody has written down, on the particular history between the people involved, on a read of the room that changes the correct answer from one situation to the next. A restructuring announcement is not really a writing task. It is a judgement about what a specific group of anxious people need to hear, delivered in language that has to feel like it came from someone who understands what they are about to go through. No amount of context in a prompt substitutes for actually being the person accountable for how that lands.
This is really the same distinction that runs through good delegation generally, long before AI entered the picture. You hand off the parts of a task that are procedural and keep the parts that depend on judgement nobody else has the standing to exercise. AI simply makes the temptation to skip that distinction easier to give in to, because the output always arrives looking finished, whether or not the task was ever appropriate to hand over.
Five things tend to tell you which side of that line a task sits on, and they are worth asking about deliberately rather than relying on a feeling that something seems too important to risk.
How structured is the task, and how would you actually know if the result was good? If you can describe both the shape of a good answer and a reasonably objective way to check for it, a tool has something solid to work against. If the only real test is whether it feels right to the people who will receive it, that test cannot be delegated.
How much of the task depends on context that lives in your head rather than on the page? Office politics, a colleague's history with a particular client, the fact that a phrase landed badly the last time someone used it: none of that travels into a prompt, however carefully it is written, because you would have to already know it mattered before you thought to mention it.
How relational is the work? Anything that exists primarily to manage how one person feels about what another person is telling them sits close to the human-first end by default, almost regardless of how well it can otherwise be briefed.
How reviewable is the eventual output? Some drafts can be checked properly in a few minutes because the standard for correctness is visible on the page. Others cannot be reviewed in any way that would catch what actually matters, because what matters is a tone or a nuance a checklist has no way to register.
And who is actually answerable for the result? The closer a piece of work sits to a decision, a relationship, or a consequence you will personally have to stand behind, the stronger the case for keeping your own hand in it from the start, whatever assistance you use once the foundation is set.
None of these questions demands a firm yes or no. They are there to move a task, honestly, toward one end of a spectrum or the other, and most real work sits somewhere in between rather than cleanly at either extreme.
The restructuring communication did eventually work, once the director changed how she used the tool rather than whether she used it. She wrote the first draft herself: the framing, the specific words she wanted said, the tone she knew the room needed. Only after that did AI enter the process, and its role changed completely. It checked her structure for gaps, flagged a paragraph that read more harshly than she had probably intended, and suggested a clearer way to sequence one section without touching the sentences that carried the actual message. The result kept the two things the AI-first version could never hold onto at the same time: her judgement about what these particular people needed to hear, and a second pass that caught what a tired first draft, written under pressure, is prone to miss.
That is usually what the right answer to this looks like in practice. It is rarely a straight choice between using AI for a task and refusing it outright. A single piece of work commonly has several stages, and different stages can sit in different places on that same spectrum. The framing and the words that carry emotional weight stay with the person accountable for them. The structural check, the consistency pass, the second pair of eyes on phrasing: those can reasonably move to a tool, once the part that actually needed a human has already been decided.
This is not a fixed map of safe and unsafe categories, and treating it as one would misread what is actually going on. A task that sits firmly on the human-first side in one organisation, one relationship, or one moment can sit differently somewhere else, and the same task can move across that line as circumstances change. What stays constant is the discipline of actually asking the question before defaulting to habit, rather than letting either enthusiasm or caution answer it for you.
Judgement, in the end, is what this comes down to. Someone who reaches for AI on every task without asking whether it belongs there has not adopted the tool more fully than someone who thinks carefully about where to place it. They have simply stopped making a decision that was theirs to make. And someone who avoids the tool entirely on anything that matters is making the opposite mistake, treating caution as a substitute for judgement rather than an input into it. The stronger position sits between those two habits, in the deliberate act of deciding, task by task and sometimes stage by stage, where a tool genuinely helps and where the work needs a person who is actually there.
Anthony Velland
AI Without Guesswork sets out this kind of placement decision as one of the core disciplines behind reliable, everyday AI use, not an occasional exception to it.
Ad · Amazon affiliate link.