The vertical AI success stories are real, and they invite an obvious conclusion: stop buying general-purpose tools, go narrow, embed AI in one specific workflow. It is a tidier lesson than the evidence actually supports — and the gap between the two is where a lot of money gets misallocated.

Where measured value shows up — and how unevenly

Start with the strongest single piece of evidence, because it is genuinely strong. A field experiment across more than five thousand customer-support agents (about 5,172, in the peer-reviewed version) found a workflow-embedded generative AI assistant produced an average productivity gain of roughly 15% — but the gain was heavily concentrated among less-experienced workers, while the most experienced staff saw little speed gain and, on some measures, small declines in quality.

Two things about that study matter for the "go narrow" argument. First, the value didn't come from narrowness as such; it came from a task with large, exploitable variation in human skill, and the mechanism the researchers identify is the diffusion of tacit knowledge from the best workers to the rest. Second — and decisively — the study had no general-purpose comparison arm. It simply cannot tell you whether narrow beats broad, because it never tested broad. Anyone citing it for that conclusion, as I nearly did, is over-reading it.

The vertical picture elsewhere is similarly uneven. In pharmaceutical discovery, AI-originated molecules have reportedly cleared Phase I at high rates, with Phase II success reverting to roughly the historical average — but that finding comes with an explicit small-sample caveat and from an industry-interested source, so treat it as suggestive, not settled. In clinical documentation, ambient AI scribes show real but modest and product-dependent time savings, with at least one study finding no significant effect.

Narrow and embedded, in other words, is not a guarantee. It is not even clearly the lever.

The counter-fact — and what it actually shows

Here is the result that should stop the "go narrow" reflex, handled carefully this time.

In a controlled study of diagnostic reasoning, a general-purpose, off-the-shelf large language model — as horizontal as it gets — significantly outperformed physicians working alone. When physicians were given that same model, one study found only a small, non-significant improvement over their baseline; an independent replication, however, found a clear and significant improvement (from roughly 66% to 74%), attributing the difference to how much clinicians actually used the tool.

Be precise about what this does and doesn't establish. These were vignette exercises, not redesigned clinical workflows — so they cannot prove that "redesign is the decisive variable," and I won't claim they do. What they do show is twofold: a horizontal model can be more than capable enough on a specialist task, and whether users benefit turns on how well the tool is actually taken up, not on how narrow the model is. The capability was there in every arm. What varied was use.

What actually distinguishes the deployments that pay

Strip out narrowness and two things remain, neither about scope.

The first is genuine uptake and integration into how the work is done — the replication's own explanation for its result, and consistent with the macroeconomic evidence that transformative technologies pay off only alongside complementary organisational investment, which lags the technology by years. A tool that sits beside the work, unused or half-used, produces the non-significant result; a tool people actually reach for produces the significant one.

The second is measurement honesty. This is more my judgment than a citable finding, so I'll label it as such: in my experience the deployments that realise value are the ones that defined, up front, what outcome they expected to move and against what baseline — and the ones that didn't are the ones now telling stories about "efficiency" they can't quantify. Notice this discipline is scope-agnostic. It doesn't favour narrow or broad. It favours legible.

The decision, reframed

So the portfolio question is not "narrow or broad?" It is: for this use case, will the tool actually be used the way it needs to be, and can we measure whether it worked?

Where the answer is yes, the returns are real and can be large. Where it's no, a narrow deployment fails for the same reason a broad one would — nobody uses it well and nobody can prove it did anything — and you'll have paid to learn something you could have known in advance.

The uncomfortable version, and the honest one: your AI may well be capable enough already for the task in front of you. What usually hasn't changed is whether anyone actually uses it, and whether you'll be able to tell.