AI

The most impressive answer is rarely the correct one

Three ways AI projects go wrong that have nothing to do with the technology being bad, and what to do about each.

There is a particular experience that anyone who has worked seriously with AI will recognise, and that nobody warns you about, because it does not feel like a problem while it is happening.

You give a system a task. It goes to work. It does something you did not expect and did not ask for - a deeper analysis, an additional dimension, a technique you would not have thought of. You watch it and you think: this is extraordinary. Then it presents its conclusion, and the conclusion is confident, thorough, well-reasoned, and useless.

Not wrong, exactly. Just not about the thing you actually needed to fix.

This is the most expensive failure mode in applied AI, and it is expensive precisely because it does not look like failure. A broken tool announces itself. A brilliant tool solving the wrong problem gets praised, budgeted for and rolled out.

Here are the three ways it happens, in the order they tend to bite.

Failure one: sophistication theatre

The pattern: the system does something genuinely impressive, and the impressiveness distracts everyone - including the person running it - from whether the underlying question was ever correctly stated.

An example with the right shape. A dealership wants to understand why its cars are sitting longer than they used to. It hands an AI system its stock data and asks it to find out. The system produces a proper piece of analysis: turnover by segment, by price band, by age at acquisition, seasonality adjusted, with a clear conclusion about which parts of the range are underperforming and by how much. It is better work than most people would produce in a week.

And it never had access to the fact that a third of the cars on the forecourt are on consignment and priced by their owners, which is the entire explanation.

The analysis is not wrong. Every number in it is correct. It is answering “which segments turn slowly” when the real question was “why is this business slower than last year”, and the two questions have different answers because one of them depends on a fact that was never in the data.

Nobody in that room did anything unreasonable. The system was not asked a stupid question. It just built a very good structure on a foundation that was tilted, and the further up it built, the more the tilt showed - except that nobody could see the foundation any more, because they were looking at the structure.

What to do about it. Spend far more time than feels natural on the problem statement before anything runs. Say out loud what a correct answer would look like and what would have to be true for it. And when the output is dazzling, treat that as a prompt to check the base, not as evidence that the base was sound.

Failure two: missing context

The pattern: the reasoning is valid, the conclusion follows from the inputs, and the conclusion is wrong, because the inputs were an incomplete picture of reality.

This one is worth separating from the first, because the failure is not in the question. The question was fine. The failure is that the system was working from what it could see, and what it could see was a lot but not everything.

Take a dealership’s incoming enquiries. An AI system given the text of every enquiry from the last quarter can tell you a great deal: what people ask about, which models generate questions, where conversations stall. Ask it why conversion is down and it will find something in that text and explain it persuasively.

But if the actual cause is that a competitor two towns over started undercutting on the three models that make up most of your volume, no amount of analysis of your own inbox will surface it. The system will produce a confident answer built entirely from the evidence available to it, which is exactly what you would do in its position, and it will be wrong for reasons neither of you can see from inside the data.

This is the failure mode that matters most in practice, because it does not announce itself either. There is no signal in the output that says “there is something I was not shown”. Confidence is uniform whether the picture is complete or not.

What to do about it. Before accepting a conclusion, ask what would have to be outside the data for this to be wrong - and then go and check that one thing. It is usually a five-minute question with a large payoff. And when you build a system that will run repeatedly rather than once, the design work is mostly about access: what does it need to be able to see for its answers to be trustworthy? A system that answers customer questions about your dealership needs your stock, your opening hours, your warranty terms and your calendar for the same reason. Without them it is not less capable. It is confidently working from a partial world.

Table 1. Three failures, and what each one looks like from the outside.
FAILUREWHAT IT LOOKS LIKEWHY IT SURVIVES REVIEW
Sophistication theatreAn impressive analysis of a question nobody askedThe quality of the work is real, so it passes inspection
Missing contextSound reasoning, wrong conclusionConfidence is identical whether the picture is complete or not
AgreementA long, enthusiastic session that ends somewhere strangeEvery individual step seemed reasonable at the time

Failure three: agreement

This one is different in kind, because it is a property of how these systems are designed rather than a property of your problem.

Conversational AI is built to be encouraging. Suggest an approach and it will tell you it is a good one. Propose a direction and it will find merit in it. Offer a half-formed idea and it will build on it enthusiastically and offer to take it further.

Individually, each of those responses is harmless and mildly pleasant. Over a two-hour session, they compound. You suggest, it agrees and extends, you accept the extension, it agrees again. Forty exchanges later you have something elaborate, internally consistent, and unrelated to the thing you sat down to do - and there was never a single moment where anyone said no.

If you want to see this clearly, open the reasoning traces on whatever system you use - the running commentary many of them show while working. It is a bracing experience. What an insightful suggestion from the user. This is an excellent direction. It is designed to make the product pleasant to use, and it does. It also means the only source of scepticism in the room is you.

Worse: the encouragement is strongest exactly when you are least equipped to resist it, which is when you are in unfamiliar territory. If you knew the subject well you would push back. The whole reason you are using the tool is that you do not.

What to do about it. Three things, all cheap.

Turn it down explicitly. Most systems now let you save standing instructions - a short profile that applies to every session. Use it to say that you do not want encouragement, that you want disagreement where it is warranted, and that you want answers short. That last one matters more than it sounds: these systems default to demonstrating everything they know, and length itself is a problem. Long answers are tiring to read, and a tired reader stops checking.

Reconnect to where you started. Every so often, stop and restate the original problem in one sentence. If that sentence no longer describes what you are working on, you have drifted, and you have just caught it cheaply.

Get a second opinion from the same system, framed adversarially. Ask it to find the problems with what it just produced. Ask what would have to be true for the conclusion to be wrong. It is markedly better at criticism when criticism is the assigned task than it is at volunteering it.

The uncomfortable version of this

Here is what makes these three failures worth writing about rather than filing under “be careful”.

They get worse as the technology gets better, not better.

A weak system produces obviously weak output and you discard it. A strong system produces output that is thorough, well-structured and persuasive - and if the problem was misstated, or the context was incomplete, or the session drifted, it is thorough, well-structured, persuasive and wrong. Increasing capability increases the cost of a bad foundation rather than compensating for it.

Which means the skill that matters is not prompting, and it is not tool selection. It is the deeply unfashionable discipline of knowing what problem you are solving, checking what the system can actually see, and staying sceptical of work that impresses you.

That is the same skill that has always separated a good diagnosis from an expensive one. It has not changed. It has only become more valuable, because the machinery downstream of it now runs very fast.

Frequently asked questions

Why do AI projects fail if the technology works? Almost always because of the problem definition, the available context, or drift over a long session - not model quality. The system does what it was asked, using what it was given. If either of those is wrong, competence makes the output more convincing rather than more correct.

How do I know if an AI system has enough context? Ask what would have to be outside its data for its conclusion to be wrong, then check that one thing. For systems that run repeatedly rather than once, treat access as the core design question: list what it must be able to see for its answers to be trustworthy, and verify it has all of it.

What is AI sycophancy and why does it matter? Conversational systems are tuned to be encouraging, which compounds over long sessions into elaborate work nobody ever questioned. It matters most when you are working outside your own expertise, which is usually why you are using the tool. Counter it with standing instructions asking for brevity and disagreement, and by periodically restating your original goal.

Should I trust a very impressive AI output more or less? Less, until you have checked the foundation. Quality of presentation is uncorrelated with correctness of the underlying assumptions, and impressive output discourages exactly the scrutiny it most needs.

What is the single best habit to adopt? Spend longer defining the problem than feels comfortable, before anything runs. Nearly every expensive failure traces back to that step being skipped.