The Visibility Trap: Measuring AI Answers Beyond the Ranking with Stuart Bruce, Purposeful Relations
Stuart Bruce, co-CEO of Purposeful Relations, argues AI is not just a tool. It is a stakeholder describing you to everyone else...
Watch videoEveryone now knows that most AI projects fail. Most of any type of project fails after all. The 80% figure from RAND and the 95% figure from MIT have been in enough LinkedIn posts that you’ve already heard the standard advice: start with the business problem, sort the data out first, get an executive sponsor.
None of that helps you on the Tuesday morning when a project has clearly failed and three people disagree about why.
I went back through the primary sources to see if the conventional wisdom misses anything and I think it does. This first part covers understanding what actually broke. Part two gets into the argument that follows, which is a different problem with a different literature behind it.
The most-cited finding from RAND’s 2024 study is that 84% of interviewees named a leadership failure as the primary cause of AI projects going wrong. Wrong problems specified up front leads to priorities switched mid-flight, and timelines in weeks for work that takes months.
Read the methodology section and RAND flags that most of the 65 interviewees were non-managerial engineers not business executives, and the authors write plainly that the results may therefore skew towards identifying leadership failures. They asked builders why things broke, and the builders said management broke them.
The same bias is probably sitting in your post-mortem meeting. Engineers locate the failure upstream in the brief. Leadership locates it downstream in execution. Both accounts are sincere, both are structurally biased, and the one that wins is usually the one held by whoever chairs the meeting.
Collect both but don’t weight them. And if you are the person who commissioned the external review, notice which population you sampled before you accept its conclusion.
The first is whether the project needed AI at all.
RAND’s interviewees described being told to apply machine learning to datasets with a handful of dominant patterns that a few if-then rules would have captured. Those projects can succeed on their own terms and still be failures, because they should never have been commissioned in the first place. The question is hard to raise after the fact because it indicts the initial decision rather than the delivery. Which is precisely why somebody has to be asking it.
The second is whether leadership was expecting a deterministic system.
Every model carries some randomness. Executives who expected repeatability get variable output, lose confidence in the product, and then, RAND found, lose confidence in the data science team as well. The team gets blamed for a property of the technology that nobody explained to anyone, or if they did, optimism and complexity made it fall on deaf ears. If this is what happened, rebuilding won’t fix it, because the problem is an expectation rather than a model.
The generic version of this finding is “data quality”, which tells you nothing and appears in every vendor deck.
Thirty of RAND’s 50 industry interviewees raised data issues, and the recurring pattern is that organisations have plenty of historical data collected to satisfy compliance or logging requirements, but not analysis. One interviewee described it well: companies think they have great data because they get a slew of weekly reports, but do not realise the data may not meet its new purpose.
The example in the report is an e-commerce site that logged which links users clicked, but not what else was on screen at the time, or what search brought them there. The volume is there but it lacks the context that would make it trainable for AI.
There is a staffing version of the same problem, and it is the one mid-market firms should worry about most. RAND’s interviewees described data engineering as low-prestige work, one calling data engineers the plumbers of data science. Turnover is high, and departing engineers take critical undocumented knowledge with them of which datasets are reliable and how their meaning has shifted over time. Rediscovering that costs months, and that time can kill projects whose sponsors have already started asking when they will see something.
Three questions for your next AI post-mortem, before anyone gets to the technical account:
Who did we ask, and what would they be predisposed to say?
Did this problem require a model, or did somebody want one?
Was our data collected for the purpose we then used it for?
Answer those honestly and you will usually know what broke. Unfortunately, that is the easy bit. The answers don’t tell you what to do next, and the argument about that is the one that damages organisations. We’ll get into it next week.
Richard
Stuart Bruce, co-CEO of Purposeful Relations, argues AI is not just a tool. It is a stakeholder describing you to everyone else...
Watch video
Stuart Bruce is co-founder and co-CEO of Purposeful Relations, which advises communications teams on AI,...
Read more
Everyone now knows that most AI projects fail. Most of any type of project fails after all. The 80% figure from...
Read moreGet ahead with the most actionable insights, playbooks and real-world AI use cases you can adopt right now, in your inbox every week