×

The Gap Between AI Spending and AI Value

Abstract illustration of heavy AI spending versus thin realized value

Two numbers about enterprise AI seem to contradict each other. McKinsey says 72% of organizations use generative AI in at least one function. Kearney says roughly 46% of AI proof-of-concepts get scrapped before they ever launch. So almost everyone is using AI somewhere. But almost half of what gets built never actually ships. I don't think these two facts disagree. They are describing the same gap, just from two different angles.

The Gap Isn't Proof the Technology Doesn't Work

It's tempting to read a scrap rate close to 50% as proof that AI is overhyped. I don't think that's the right read. I've done a lot of architecture reviews, and PoCs almost never fail because the model can't answer the question. They fail for a different reason. It was never really just: can this model give a good answer on a clean, curated sample? Frontier models have been able to do that for a while now. What a demo never tests is something else entirely. Can the system handle real traffic? Can it handle messy, real-world data? Can it survive the audit and compliance checks that only show up once something is headed to production? Those are the questions that actually kill projects.

Proving a model can answer questions on a clean sample isn't useless. But it proves something nobody seriously doubted in the first place. It doesn't fund the harder work: making that capability reliable at scale.

What a PoC Actually Measures Versus What Production Requires

A typical PoC measures answer quality on a small set of examples, often hand-picked to look good. Production asks completely different questions. What happens when there's heavy load? What happens when the data is messy? Who owns the system at 2am when it breaks? What does the audit trail look like six months later? None of that shows up in a demo. And none of it gets funded either, because the budget was scoped around one goal: "prove the model works." Nobody budgeted for the real goal, which is "build something people can depend on."

That's what the 46% scrap rate is really telling us. The underlying AI capability didn't fail. The project was just scoped to answer the easy question. It never had a budget, or even a plan, for the hard one.

What Actually Closes the Gap

The teams that actually get AI into lasting production don't have a better model. Frontier model access is close to a commodity now. What sets them apart is the unglamorous layer built around the model. That means real evaluation against labeled ground truth. It means pipelines that can tolerate messy input. And it means someone who actually owns failures, so the project doesn't turn into an orphaned pilot the moment the demo goes well.

None of that is a model upgrade. It's the same discipline that turns any prototype into a real product. AI is just new enough that a lot of teams are still skipping that step.

The Real Conclusion

High adoption and a high scrap rate aren't actually contradictory. Together, they describe an industry that is still funding the demo phase. It calls that demo "production." Then it acts surprised when the demo can't survive conditions it was never built for. The spending gap closes the same way it always has. Not with a bigger model. But by turning "this is possible" into a system people can actually depend on.

Building something with AI?

I design and ship production systems across ML, deep learning, GenAI, LLMs, and RAG. Happy to talk through what you're working on.

Book a Free Call