← All posts

Essay · August 2026

Vertical AI vs Generic AI: Why Narrow Wins, and Why It Isn't the Model

Vertical AI does not win because its model is better. Harvey, the legal AI platform valued at $11 billion in March 2026, runs frontier models from Anthropic, Google DeepMind and OpenAI — the same models anyone can call through an API. What a vertical player owns is a domain benchmark that exists nowhere else, the workflow surface where the work already happens, the compliance envelope around the data, and the humans who configure it inside the customer. Those four layers are buildable at any company size; the model layer is not.

In March 2026 Harvey raised $200 million at an $11 billion valuation. It sells AI to lawyers: the majority of the AmLaw 100, more than 500 in-house legal teams, and by its own count 100,000 lawyers across 1,300 organizations in 60 countries. Every one of those lawyers already had a general-purpose assistant on their phone, and most of them had it for free.

The obvious explanation is that Harvey built a model that understands law better than the general models do. It didn't. Harvey runs Anthropic, Google DeepMind and OpenAI models — the same frontier models available to anyone with an API key and a credit card. The company published a post about it in March under the title "Why Harvey is Multi-Model by Design."

So the narrow player is winning while running the generalist's engine. Which decides whether "go vertical" is advice you can act on at your size, or just a story about venture rounds you'll never raise.

The moat everybody assumes, and the one that's actually there

Read Harvey's own architecture post and there is no claim about a proprietary legal model anywhere in it. The three stated reasons for running multiple providers are quality, reliability and choice. Quality, because no single model is best at everything, so the research team routes each task to whichever model tests best on it. Reliability, because when one provider hits capacity limits or an outage, work reroutes without the lawyer noticing — and because model availability differs by region, which matters when your customer is in Australia and their data has to stay there. Choice, because Harvey requires every provider to meet the same terms: zero data retention, no human review of customer data, no training on customer inputs.

"GPT-5.4, Now Live in Harvey," March 5. "Opus 4.8, Now Live in Harvey," May 28. A company that ships a supplier's release as its own product launch is telling you exactly where its value isn't.

The picture isn't uniform. Ambience Healthcare, which sells ambient documentation AI to hospitals, describes its platform as powered by proprietary reasoning models purpose-built for healthcare. So vertical AI is not one architecture: some build models, some rent them. Where they converge is everything wrapped around the model, and that wrapping is the part a general assistant can't reach.

Four layers a general assistant can't get to

Layer

What it looks like in practice

Why a general assistant can't have it

Domain evaluation

Harvey runs BigLaw Bench and expert preference testing through an in-house Applied Legal Research team

The lab has no access to your expert judges or your definition of "correct"

Workflow surface

Harvey inside iManage, Box, Ironclad, Docusign, Microsoft 365; Ambience inside Epic, Oracle Cerner, athenahealth

A chat window makes the user come to it; the work stays where it already lives

Compliance envelope

Harvey's ISO 42001 certification, localized data processing across US, EU and Australia

Governance is per-industry and per-jurisdiction, not a product feature

Deployment humans

Harvey's embedded legal engineering teams; more than 25,000 custom agents built by customers on the platform as of March 2026

Self-serve software has no one inside your building

Ambience's framing of the second row is sharper than anything a vendor deck usually admits: the platform adapts to the specific care setting and specialty without requiring workflow redesign or staff retraining. Generic tools push that burden onto the user, then call the result an adoption problem.

The number only a narrow system can produce

On 19 August 2026 Ardent Health announced an enterprise rollout of Ambience across its network: 30 acute care hospitals, roughly 280 sites of care, more than 1,800 affiliated providers in six states. The rollout followed a pilot that ran across 17 specialties and 7 languages, with more than 140,000 patient encounters documented.

The reported pilot results, as published by Ambience:

  • 45% decrease in documentation time, from Epic User Action Log data
  • 5 hours per week saved per clinician
  • 90% encounter usage rate among pilot providers
  • 70% of pilot clinicians reported reduced cognitive load
  • 100% of pilot clinicians said it improved their job satisfaction

Sort that list by where each number came from. The last two are clinicians answering a questionnaire about how they feel, which is worth something and is not evidence. The 45% came out of Epic's audit log, the same system the hospital already uses to bill, to staff, and to defend itself when an auditor asks what happened during a shift.

Narrow scope is what makes attribution possible, and that has almost nothing to do with intelligence. "Documentation time in Epic" was a number before the AI showed up and will still be a number after it's switched off, so the change is measurable by subtraction. "Productivity gain from a general assistant" is not a number that exists in any system anywhere, which is why every board discussion about it collapses into competing anecdotes.

Veeva: the slow version of the same bet

Veeva Systems sells industry cloud software to life sciences: clinical, regulatory, safety, commercial. In November 2015 it reported quarterly revenue of $106.9 million and more than 375 customers. On 3 June 2026 it reported quarterly revenue of $882.9 million, up 16% year over year, more than 1,500 customers, full-year guidance above $3.6 billion, and a stated goal of a $6 billion revenue run rate by 2030.

CEO Peter Gassner described the current phase as "moving from an industry-specific application company to an industry-specific application and AI agent company."

Veeva did not win life sciences with AI. It spent more than a decade accumulating an industry data model, a regulated-content footprint and a customer base, and it is now monetizing that with agents layered on top. Vault CRM added 27 customers in that single quarter and passed 150 live. The agents are new. The moat isn't.

The vertical asset compounds slowly, and AI cashes it in. That turns the planning question inside out: not "which AI should we buy this year," but "what narrow asset are we accumulating that an agent will be able to cash in three years."

What this looks like without a billion dollars

You are not going to out-model a frontier lab, and you were never in that race. Three of the four layers above cost nothing but discipline, and all four are available at fifteen people.

Start by picking a workflow that already has a number. It has to run daily or close to it, and its metric has to live in a system you own: CRM, helpdesk, phone system, accounting. If nobody in the room can state the current number, you don't have a project yet, you have a demo. Our ROI method starts at the same place.

Then write the evaluation before you write the agent, because this is the step everyone skips and it's the one that does the work. Pull 30 to 50 real cases from last month. Have your best person mark what a correct outcome looks like on each one, in their own words, including the two or three where they had to think about it. That file is your BigLaw Bench. It costs an afternoon, it never needs a vendor, and it is the only instrument that will tell you whether next quarter's model upgrade helped or quietly broke something — which, given how much results drift between runs, is not a question you can answer by reading release notes. Teams that skip this don't find out they were wrong; they just keep arguing about vibes for two quarters.

Put the agent where the work already happens: inside the CRM record, in the ticket, on the call. Every step a person has to take toward the AI is adoption you're spending down. And own the boring layer — logs, permissions, escalation, and a defined answer to what the agent does when it refuses or isn't sure.

None of that is a technology decision. It's four choices about scope, measurement, placement and accountability, the same four Harvey and Ambience made, at a scale where you can make them in an afternoon.

When narrow doesn't win

The pattern gets oversold, so here's where it breaks:

  • Low frequency work. Building an evaluation costs roughly the same whether the workflow runs 400 times a month or four. At four, it never pays back.
  • No system of record. If the work lives in people's heads and a WhatsApp thread, there is no before-number, and the project ends in a debate nobody can win.
  • Rules with no internal owner. Domain-specific means somebody has to be personally accountable for what "correct" means and keep that definition current. Without that person you get a narrow tool drifting on stale assumptions, which is worse than a general one.

There's also the genuine long tail. When every task is different, a general assistant is simply the better answer, and this was never either/or anyway — Harvey's customers still use Copilot, Ardent's clinicians still search the web.

The question to take into your next planning meeting

Not which AI is smartest. Name one thing your team does every day inside a system that already counts it, then say out loud what the current number is. If that sentence comes out easily, you've found the project. If it doesn't, the honest next step is instrumentation, not AI, and that one is cheap — which is the part nobody wants to hear in a meeting about AI.

And if it doesn't, the person to send this to is whoever owns that system. The argument is with them, not with the AI budget.

The free two-minute AI audit is built to run exactly that check with you, and it will tell you when the answer is that you don't need an agent yet. When a workflow does clear the bar, Grow2.ai builds the agent against a KPI agreed in the contract and ships the pilot in 14 days; if the KPI isn't met, you don't pay.

Frequently asked questions

What is vertical AI?

Vertical AI is a product built for one industry or one professional workflow, rather than for general use. Harvey builds for legal work, Ambience Healthcare for clinical documentation, Veeva for life sciences. The defining trait is not the model but the scope: a vertical product is judged against a metric that already exists inside the customer's own systems.

Does vertical AI use its own model?

Not necessarily, and often not. Harvey runs frontier models from Anthropic, Google DeepMind and OpenAI, and published its reasoning in "Why Harvey is Multi-Model by Design" in March 2026. Ambience Healthcare, by contrast, says it uses proprietary reasoning models built for healthcare. Both approaches exist; neither is what creates the advantage.

Is vertical AI better than ChatGPT or Claude?

Not at general reasoning — vertical products usually run those same models underneath. It performs better on a specific job because it sits inside the system where that job happens, is evaluated against domain-specific test cases, and carries the compliance terms the industry requires. For open-ended or one-off tasks, a general assistant is the better tool.

Can a small business build vertical AI?

Yes, for its own operations. You cannot compete on the model layer, but the four layers that matter — a domain evaluation set, placement inside the working system, a compliance and permissions envelope, and someone accountable for the definition of correct — are all achievable by a small team. The evaluation set is typically 30 to 50 real cases from the past month.

How do I know if a workflow is narrow enough?

Take a support queue as the example: it qualifies if tickets arrive every day and the helpdesk already reports first-response time. Frequency plus an existing metric in a system you control are the two tests. Where no one can state today's number, the honest first step is measurement rather than automation.

Why does it matter where a reported AI result came from?

Because survey results and system telemetry are not the same class of evidence. In the Ardent Health pilot, the 45% reduction in documentation time was drawn from Epic User Action Log data — the hospital's own audit trail — while the reported improvements in cognitive load and job satisfaction came from clinicians answering a questionnaire. Both were published in the same announcement by Ambience, the vendor. When you evaluate any AI result, including your own, sort the numbers by origin before you act on them.

How much does it cost to start with an AI agent?

Grow2.ai runs a fixed-price 14-day pilot against a KPI agreed in the contract; if the KPI isn't met, you don't pay. Current pricing is listed on [grow2.ai](https://grow2.ai/en/). The free two-minute AI audit comes before any of that and often concludes that a given workflow isn't ready yet.

AI agents for business — 2–3 emails a month

Breakdowns, cases and tools already working inside companies.

No spam. Unsubscribe in one click.