AI in Healthcare

The economics of clinical AI: unit economics that actually work

The teams building durable clinical AI companies share a discipline about margin structure that many peers lack.

Priya RaghavanDecember 8, 20254 min read

Inference cost is not incidental

Founders often model inference cost as trivial. At clinical scale, it is not. A single specialty deployment can generate millions of model calls per month. This directly compresses gross margin.

Design for margin from day one

The teams with healthy long-term margins made two choices early: they picked model architectures they could afford to run, and they built caching and retrieval layers that dramatically reduce redundant inference.

Pricing

Per-seat pricing rarely covers inference. Per-transaction or per-patient models tend to align better with value delivered and cost incurred.

The contract structure

Include the right to renegotiate pricing when usage exceeds forecasts by a defined margin. Enterprise buyers will accept this if introduced early.

The token trap

Teams that price per seat while their underlying model calls scale with document volume eventually discover the mismatch the hard way, usually during a renewal conversation with a health system that has quietly tripled its usage. The fix is not a smarter prompt; it is a pricing unit that moves with the same variable that drives cost.

We have watched founders redesign their entire billing schema mid-contract because the original unit, a flat monthly fee, bore no relationship to compute consumed. Customers noticed the disconnect before the vendor did, because their own usage dashboards made the asymmetry visible.

A hypothetical worth studying

Consider a documentation assistant that charges per clinician seat but calls a large model on every note, every revision, and every summarization pass. At modest adoption the economics look fine. At full rollout across a hospital's ambulatory network, inference cost can silently overtake the seat fee, and margin quietly goes negative without any line item flagging it.

The companies that avoid this outcome instrument cost per transaction from the earliest pilot, long before it matters financially, so the warning shows up in a spreadsheet rather than in a board meeting.

The counterargument

Some argue that obsessing over unit margin too early trades growth for discipline that competitors without the same rigor will simply out-fundraise. There is truth to this in categories where land-grab dynamics genuinely reward speed over efficiency, and a founder who slows down to build a perfect pricing model can lose the account entirely.

The resolution most experienced operators land on is not to ignore growth, but to build the cost telemetry in parallel with the land-grab, so that when the market matures and pricing pressure arrives, the company already knows exactly where its margin lives.

What we tell portfolio companies

Model cost per unit of clinical work, not per user, from the first commercial contract onward, and revisit it every time model providers change their pricing tiers, which happens more often than most teams budget for. Treat model provider concentration as a margin risk, not just a technical dependency, and negotiate contract terms accordingly.

The vendor concentration risk

Companies that built their entire cost model around a single model provider's published rate card learned, often mid-contract, that those rates are not stable commitments but starting points subject to renegotiation whenever the provider's own economics shift. A pricing plan that assumes today's inference cost persists for three years is really a bet on someone else's roadmap, not a business plan.

The more disciplined teams now negotiate multi-provider flexibility into their architecture even before it is cheaper to do so, treating the ability to switch models as insurance against a cost shock rather than a nice-to-have engineering exercise.

Margin as a fundraising signal

Investors evaluating clinical AI companies now ask for gross margin broken out by customer cohort, not just in aggregate, because a blended figure can hide a large enterprise account that is actually unprofitable once true inference cost is allocated to it. Founders who can produce that breakdown on request signal a level of financial maturity that increasingly separates fundable companies from merely fast-growing ones.