In March 2026 Harvey raised $200 million at an $11 billion valuation. It sells AI to lawyers: the majority of the AmLaw 100, more than 500 in-house legal teams, and by its own count 100,000 lawyers across 1,300 organizations in 60 countries. Every one of those lawyers already had a general-purpose assistant on their phone, and most of them had it for free.
The obvious explanation is that Harvey built a model that understands law better than the general models do. It didn't. Harvey runs Anthropic, Google DeepMind and OpenAI models — the same frontier models available to anyone with an API key and a credit card. The company published a post about it in March under the title "Why Harvey is Multi-Model by Design."
So the narrow player is winning while running the generalist's engine. Which decides whether "go vertical" is advice you can act on at your size, or just a story about venture rounds you'll never raise.
The moat everybody assumes, and the one that's actually there
Read Harvey's own architecture post and there is no claim about a proprietary legal model anywhere in it. The three stated reasons for running multiple providers are quality, reliability and choice. Quality, because no single model is best at everything, so the research team routes each task to whichever model tests best on it. Reliability, because when one provider hits capacity limits or an outage, work reroutes without the lawyer noticing — and because model availability differs by region, which matters when your customer is in Australia and their data has to stay there. Choice, because Harvey requires every provider to meet the same terms: zero data retention, no human review of customer data, no training on customer inputs.
"GPT-5.4, Now Live in Harvey," March 5. "Opus 4.8, Now Live in Harvey," May 28. A company that ships a supplier's release as its own product launch is telling you exactly where its value isn't.
The picture isn't uniform. Ambience Healthcare, which sells ambient documentation AI to hospitals, describes its platform as powered by proprietary reasoning models purpose-built for healthcare. So vertical AI is not one architecture: some build models, some rent them. Where they converge is everything wrapped around the model, and that wrapping is the part a general assistant can't reach.
Four layers a general assistant can't get to
Layer | What it looks like in practice | Why a general assistant can't have it |
|---|---|---|
Domain evaluation | Harvey runs BigLaw Bench and expert preference testing through an in-house Applied Legal Research team | The lab has no access to your expert judges or your definition of "correct" |
Workflow surface | Harvey inside iManage, Box, Ironclad, Docusign, Microsoft 365; Ambience inside Epic, Oracle Cerner, athenahealth | A chat window makes the user come to it; the work stays where it already lives |
Compliance envelope | Harvey's ISO 42001 certification, localized data processing across US, EU and Australia | Governance is per-industry and per-jurisdiction, not a product feature |
Deployment humans | Harvey's embedded legal engineering teams; more than 25,000 custom agents built by customers on the platform as of March 2026 | Self-serve software has no one inside your building |
Ambience's framing of the second row is sharper than anything a vendor deck usually admits: the platform adapts to the specific care setting and specialty without requiring workflow redesign or staff retraining. Generic tools push that burden onto the user, then call the result an adoption problem.
The number only a narrow system can produce
On 19 August 2026 Ardent Health announced an enterprise rollout of Ambience across its network: 30 acute care hospitals, roughly 280 sites of care, more than 1,800 affiliated providers in six states. The rollout followed a pilot that ran across 17 specialties and 7 languages, with more than 140,000 patient encounters documented.
The reported pilot results, as published by Ambience:
- 45% decrease in documentation time, from Epic User Action Log data
- 5 hours per week saved per clinician
- 90% encounter usage rate among pilot providers
- 70% of pilot clinicians reported reduced cognitive load
- 100% of pilot clinicians said it improved their job satisfaction
Sort that list by where each number came from. The last two are clinicians answering a questionnaire about how they feel, which is worth something and is not evidence. The 45% came out of Epic's audit log, the same system the hospital already uses to bill, to staff, and to defend itself when an auditor asks what happened during a shift.
Narrow scope is what makes attribution possible, and that has almost nothing to do with intelligence. "Documentation time in Epic" was a number before the AI showed up and will still be a number after it's switched off, so the change is measurable by subtraction. "Productivity gain from a general assistant" is not a number that exists in any system anywhere, which is why every board discussion about it collapses into competing anecdotes.
Veeva: the slow version of the same bet
Veeva Systems sells industry cloud software to life sciences: clinical, regulatory, safety, commercial. In November 2015 it reported quarterly revenue of $106.9 million and more than 375 customers. On 3 June 2026 it reported quarterly revenue of $882.9 million, up 16% year over year, more than 1,500 customers, full-year guidance above $3.6 billion, and a stated goal of a $6 billion revenue run rate by 2030.
CEO Peter Gassner described the current phase as "moving from an industry-specific application company to an industry-specific application and AI agent company."
Veeva did not win life sciences with AI. It spent more than a decade accumulating an industry data model, a regulated-content footprint and a customer base, and it is now monetizing that with agents layered on top. Vault CRM added 27 customers in that single quarter and passed 150 live. The agents are new. The moat isn't.
The vertical asset compounds slowly, and AI cashes it in. That turns the planning question inside out: not "which AI should we buy this year," but "what narrow asset are we accumulating that an agent will be able to cash in three years."
What this looks like without a billion dollars
You are not going to out-model a frontier lab, and you were never in that race. Three of the four layers above cost nothing but discipline, and all four are available at fifteen people.
Start by picking a workflow that already has a number. It has to run daily or close to it, and its metric has to live in a system you own: CRM, helpdesk, phone system, accounting. If nobody in the room can state the current number, you don't have a project yet, you have a demo. Our ROI method starts at the same place.
Then write the evaluation before you write the agent, because this is the step everyone skips and it's the one that does the work. Pull 30 to 50 real cases from last month. Have your best person mark what a correct outcome looks like on each one, in their own words, including the two or three where they had to think about it. That file is your BigLaw Bench. It costs an afternoon, it never needs a vendor, and it is the only instrument that will tell you whether next quarter's model upgrade helped or quietly broke something — which, given how much results drift between runs, is not a question you can answer by reading release notes. Teams that skip this don't find out they were wrong; they just keep arguing about vibes for two quarters.
Put the agent where the work already happens: inside the CRM record, in the ticket, on the call. Every step a person has to take toward the AI is adoption you're spending down. And own the boring layer — logs, permissions, escalation, and a defined answer to what the agent does when it refuses or isn't sure.
None of that is a technology decision. It's four choices about scope, measurement, placement and accountability, the same four Harvey and Ambience made, at a scale where you can make them in an afternoon.
When narrow doesn't win
The pattern gets oversold, so here's where it breaks:
- Low frequency work. Building an evaluation costs roughly the same whether the workflow runs 400 times a month or four. At four, it never pays back.
- No system of record. If the work lives in people's heads and a WhatsApp thread, there is no before-number, and the project ends in a debate nobody can win.
- Rules with no internal owner. Domain-specific means somebody has to be personally accountable for what "correct" means and keep that definition current. Without that person you get a narrow tool drifting on stale assumptions, which is worse than a general one.
There's also the genuine long tail. When every task is different, a general assistant is simply the better answer, and this was never either/or anyway — Harvey's customers still use Copilot, Ardent's clinicians still search the web.
The question to take into your next planning meeting
Not which AI is smartest. Name one thing your team does every day inside a system that already counts it, then say out loud what the current number is. If that sentence comes out easily, you've found the project. If it doesn't, the honest next step is instrumentation, not AI, and that one is cheap — which is the part nobody wants to hear in a meeting about AI.
And if it doesn't, the person to send this to is whoever owns that system. The argument is with them, not with the AI budget.
The free two-minute AI audit is built to run exactly that check with you, and it will tell you when the answer is that you don't need an agent yet. When a workflow does clear the bar, Grow2.ai builds the agent against a KPI agreed in the contract and ships the pilot in 14 days; if the KPI isn't met, you don't pay.