The hardest part of AI due diligence is that the demo always works. Whether the company has built something genuinely defensible or wrapped a thin prompt around a frontier model someone else trained, the demo on the screen looks roughly the same: a clean interface, an impressive output, a confident founder. The surface tells you almost nothing. Diligence is the discipline of getting underneath it — and in AI, what is underneath ranges from a deep, compounding technical moat to a feature that any competent team could rebuild in a fortnight.

I have run and reviewed this kind of diligence across a long career on both sides of the table — as an operator building products, as an investor including as CIO at the Nielsen Innovate Fund, and now through GenovateAI's AI Due-Diligence Memo, where the deliverable is a code-level, independent verdict on an AI asset. The patterns that distinguish real depth from a thin wrapper are consistent enough that they can be written down. This piece does that.

The single question that organises all AI diligence: if the frontier model providers shipped this capability natively tomorrow, what would be left of this business?

— What you are actually assessing

Eight layers, not one demo.

A serious AI diligence does not assess “the AI.” It assesses eight distinct layers, each of which can be strong or hollow independently of the others:

01
Product reality
Does the thing in the deck match the thing in production? How much of the demo is real versus staged, and what is the gap between the best-case output and the median one?
02
AI depth
What is genuinely theirs? Proprietary models, fine-tuning, a meaningful orchestration layer, evaluation infrastructure — versus a system prompt over a third-party API.
03
Data moat
Is there a proprietary, compounding data asset that competitors cannot easily replicate — or is the “data advantage” just public data and customer inputs anyone could gather?
04
Engineering quality
Code-level: architecture, test coverage, evaluation harnesses, the ability to ship improvements safely. The difference between a team that can compound and one that has plateaued.
05
Security & compliance
Data handling, model governance, and exposure under regimes like the EU AI Act — especially for higher-risk use cases.
06
Inference-cost margin
The unit economics that demos hide. What does it cost to serve one user at scale, and does the business have positive gross margin once token costs are honest?
07
Roadmap defensibility
Is the moat widening or narrowing? Does the roadmap depend on capabilities the model providers are likely to absorb, or on assets that get harder to copy over time?
08
Red flags
The patterns below — each individually survivable, but in combination a signal that the AI story is thinner than the pitch.
— The red flags

Patterns that signal a thin wrapper.

No single flag is fatal — plenty of good companies trip one. The signal is in the cluster. When several of these appear together, the probability that you are looking at a thin wrapper dressed as a defensible AI business rises sharply.

  • The demo is the only artifact. There is a polished demo and a pitch, but no evaluation harness, no error analysis, no honest account of where the system fails. A team that has built something real can tell you precisely where it breaks. A team that cannot has not looked.
  • “Proprietary model” that is a system prompt. The IP, on inspection, is a prompt template and some glue code over a frontier API. That can still be a real business — but it is not a model moat, and a deck that calls it one is signalling either confusion or spin.
  • A data moat made of other people’s data. The “unique dataset” turns out to be publicly scrapable, licensed non-exclusively, or simply the customers’ own inputs that the customers could take elsewhere. No compounding, no exclusivity, no moat.
  • Accuracy claims with no denominator. “95% accurate” on what test set, against what baseline, measured how? Headline metrics with no methodology behind them are marketing, and the absence of a real evaluation set is itself the finding.
  • Inference economics that only work at the demo. The unit cost to serve a real user at scale has never been calculated, or is quietly negative. A business that has not modelled its token costs at volume does not yet know whether it has a business.
  • A roadmap that is one model release from obsolescence. The core value proposition is a capability the frontier providers are visibly moving toward shipping natively. If the answer to “what survives if the providers do this themselves?” is “not much,” the moat is borrowed time.
  • Founders who deflect code-level questions. Technical diligence requests are met with NDAs, “it’s complicated,” or redirection to the demo. Real teams are usually relieved to talk to someone who understands the stack. Persistent deflection is a tell.
  • No evaluation or monitoring in production. The system ships outputs but has no harness measuring quality over time and no drift detection. This is the difference between a model that compounds and one that silently degrades — and its absence is a maturity flag.

The demo always works. Diligence is the discipline of finding out what is underneath it.

— The verdict

Thin wrapper is not the same as bad bet.

An important nuance, because it is where careless diligence goes wrong in the other direction: “thin wrapper” is a description of technical depth, not a verdict on the investment. Some of the better businesses in this cycle are thin wrappers in the technical sense — their defensibility lives in distribution, a specific workflow, a regulated niche, switching costs, or brand, not in proprietary models. The error is not investing in a wrapper. The error is investing in a wrapper while believing you are buying a model moat, and pricing the deal accordingly.

So the job of AI diligence is not to issue a pass/fail on technical depth. It is to locate the defensibility honestly — to say plainly “the AI here is shallow, but the distribution and switching costs are real, and that is what you are paying for” or “the AI is genuinely deep, but the unit economics do not yet close.” A good memo gives the investor the real shape of the asset, with the red flags named and the moat located where it actually is. That is the deliverable: not a demo reaction, but a defensible, code-level verdict an investment committee can act on.

— See the artifactThe AI Due-Diligence Memo — structure, verdict layers & sample View →

The frontier models will keep getting better, faster, and cheaper, and that movement will keep dissolving the moats of companies whose only advantage was access to a capability that is becoming a commodity. The discipline that protects capital in that environment is not a sharper demo sense. It is the willingness to get underneath the demo, ask the eight-layer questions, read the red flags as a cluster, and locate the defensibility where it actually lives — or report, honestly, that there isn’t any.

An independent verdict on an AI asset

Thirty minutes with Liron to scope a code-level AI Due-Diligence Memo on a target — an independent read on product reality, AI depth, data moat, unit economics, and red flags before you commit capital.

Book an AI Decision Call