What to Do When AI Gives You a Bad Answer


What to Do When AI Gives You a Bad Answer

The summary comes back thin. Or the structure has drifted somewhere you didn't ask it to go. Or the tone is wrong in a way you can't quite name but instantly recognise. Whatever the shape of it, the reaction is almost always the same: type something like "make this better" or "try again, more detail this time," and send it back.

That instinct is understandable. It's also not a diagnosis. It's a second guess dressed up as a correction, and the tool has no more information to work with than it did the first time. Sometimes the retry helps a little. Often it just fails in a different place, and after two or three rounds of this, most people either give up on the task, accept something mediocre, or start again from nothing.

The good news is that most weak outputs trace back to a small, identifiable set of causes, and almost all of them sit on your side of the exchange rather than the model's. Something in what you provided, how you framed the request, or where in your workflow you brought the tool in created the conditions for a poor result. That's not a criticism. It's useful, because it means there's somewhere specific to look before you touch the wording again.

There are three questions worth asking, in this order, before you send anything back for another pass.

Did the input clearly state what the output needed to accomplish, for whom, and at what level of detail? If not, the fix is to add that information, not to bolt on a vague instruction like "improve this." Vague corrections feel productive without actually telling the tool what was missing, so the next attempt is just a different guess rather than a better one.

Did the request actually contain more than one task? A single instruction that asks for a summary and a set of recommendations, or an analysis and a strategic framing, is really two jobs wearing one label. Splitting them and running them in sequence, letting the first output feed the second, usually resolves more than another round of rephrasing ever will.

Is this the right point in the process for AI to be doing this particular piece of work? Sometimes the input was fine and the task was well defined, and the problem is that the work itself needed a person's judgement before anything else touched it.

Most weak answers fall into one of those three buckets, and the clue to which one is usually visible in the shape of the failure itself.

When the output is vague because the brief was

The most common cause, by some distance, is missing context or an unstated purpose. A request that asks for a summary of three weeks' progress across two workstreams, with plenty of raw notes attached but no signal about who is reading it or what they need to decide, tends to come back as a wall of text that touches everything and emphasises nothing. That isn't the tool struggling. It's the tool doing exactly what an unframed request invites: treating every point in the notes as equally important, because nothing told it otherwise.

The diagnostic clue is in how the weakness reads. Output that feels generic or unfocused usually means the task's purpose was never spelled out. Output that's in the right territory but at the wrong level of detail usually means depth or audience was left undefined. Output with the wrong tone usually means nobody told the model who would actually be reading it.

A financial analyst at an asset management firm learned to spot this pattern over about six weeks of paying closer attention to her own failures. When she started, she estimated that roughly one output in four needed significant rework, and most of her fixes were the unfocused kind: more detail here, a different phrasing there, occasionally just "try again with more emphasis on risk factors." The results were inconsistent, which is what you'd expect from corrections that weren't aimed at anything in particular. Once she began tracking the pattern instead of just reacting to it, she noticed within a few weeks that her single most frequent problem was missing audience specification. She was giving the tool strong analytical material but rarely stating who would read the output or what decision it needed to support. Adding one sentence about audience and decision context to her inputs changed her first-pass results noticeably. The corrections that remained became smaller, adjustments of emphasis rather than full rebuilds, and her estimate of outputs needing serious rework fell to roughly one in ten.

What changed wasn't her prompting technique. It was that she stopped guessing at fixes and started checking a specific thing first.

When the request was actually two requests

The second pattern shows up differently. Rather than a flat, generic result, you get something uneven: parts of it land well, other parts feel thin or slightly off-topic, as though two different jobs had been squeezed into one answer. That's usually because they were. A request that asks for a summary and next steps in the same breath is asking for accuracy and compression on one hand, and interpretation and judgement on the other, and when a single pass has to deliver both, each one tends to weaken the other.

A marketing manager at a technology company ran into this during a product launch. She'd been asking for one output that combined competitive positioning, messaging recommendations and a draft of customer-facing copy, and every version came back with positioning that felt generic, messaging that quietly contradicted the positioning, and copy that could have belonged to a different product. Three rounds of revision ate most of a morning before she separated the request into its actual components: positioning first, reviewed and corrected; then messaging, built on the approved positioning; then copy, built on the approved messaging. The final draft needed one light edit instead of another rewrite, and the whole sequence, done properly, took less time than the single combined attempt she'd already abandoned.

The fix here isn't more detail. It's separation, handling each distinct job on its own terms and letting the output of one inform the input of the next.

When the task should never have gone to AI first

The third category is the one people are most likely to miss, because it doesn't look like a prompting problem at all. Sometimes the input is genuinely good and the task is genuinely singular, and the output is still wrong in a way no amount of revision seems to fix. That's usually a sign the work was handed to the tool at the wrong point in the process.

A human resources director experienced this during a sensitive restructuring. She'd used AI well for months on internal updates and briefing documents, so when she needed to write to employees whose roles were changing, she followed her usual approach: clear context, defined audience, key messages outlined, and a draft requested. What came back was structurally sound and professionally worded, and it was also wrong. The tone lacked the human weight the situation needed. The sentences were correct without being compassionate. She spent over an hour revising it, and each adjustment improved one part while making another feel slightly more artificial, because the problem wasn't in the sentences. It was in the order of operations. The message needed a person's judgement about tone and sensitivity to come first, with AI brought in afterwards to check structure and flag gaps, not the other way round. Once she reversed that sequence, the revised version took less time to finish than the original had taken to fix, and it read like something a person had actually written, because it was.

These three causes, missing context, bundled tasks and placement error, account for most of the recoverable failures you'll run into. They don't account for all of them. Occasionally a tool will simply handle a particular kind of request inconsistently, or a combination of instructions will create signals that don't resolve cleanly, and no amount of tracing will explain why. Those cases are rarer than they feel in the moment, though, and they're usually identifiable by elimination: if the input is clear, the task is singular, and the placement is right, and the output is still off, that's probably just one of those edge cases, and a straightforward retry is a reasonable response.

The habit worth building isn't a new skill so much as a change in direction. Instead of building forward from task to input to output, you work backward from the weak output to the input to the task, looking for the point where the chain actually weakened. It takes less time than a single unfocused retry, and it tends to teach you something that a retry never does: which of your own habits are producing the failure in the first place. Over time, the diagnosis starts happening before you even send the request, and the failures it was built to catch simply become less frequent.

Anthony Velland

Interested in going further?

The three-question diagnostic in this piece is one part of the fuller working method for structuring reliable AI use laid out in AI Without Guesswork.

Ad · Amazon affiliate link.