Clinical foundation models: a builder's map for 2025
Which layers of the stack are consolidating, which are still open, and where the real defensibility lives.

The stack
The clinical AI stack has clarified into four layers: foundation models, clinical adapters, workflow surfaces, and evaluation infrastructure. Different layers have different rates of consolidation.
Where the moats are
Foundation models are consolidating around a handful of providers. Adapters are fragmenting by specialty. Workflow surfaces are becoming the primary battleground. Evaluation is the layer most under-invested relative to its long-term importance.
Advice for founders
Do not build a foundation model unless you have a proprietary data advantage nobody else can match. Do build the adapter, the workflow, or the evaluation layer for a specific clinical domain.
The regulatory overlay
Every layer intersects with regulation eventually. Founders who ignore this until product-market fit tend to hit a wall shortly after.
The data assembly problem beneath the model
Most of the differentiated value in a clinical foundation model does not come from novel model architecture, which is broadly shared knowledge across the field, but from the process of assembling, cleaning, and labeling clinical data at a scale and quality that a general-purpose model provider cannot easily replicate. That process requires clinical expertise embedded directly in the data pipeline, not just applied afterward to review outputs.
This is why partnerships with health systems for data access have become one of the most contested and valuable relationships in the category. A founder with exclusive or preferential access to a well-curated, longitudinal dataset from a specific clinical domain has a durable advantage that a larger, better-capitalized competitor cannot simply purchase.
Advice we give founders in this layer
Pick a narrow clinical domain where you can plausibly become the best-informed builder in the world, rather than trying to build a general clinical model that competes directly with the largest AI labs on breadth. The labs will win the breadth competition; a focused team can still win depth in a domain the labs have not prioritized.
Instrument your product from day one to capture structured feedback on model outputs from real clinical use, because that feedback loop, more than any initial training data advantage, is what compounds into a defensible moat over multiple product iterations. Founders who treat deployment purely as distribution, rather than as an ongoing data-generation engine, give up their best long-term advantage.
Where consolidation is already visible
The infrastructure layer — model hosting, fine-tuning tooling, and evaluation frameworks tailored to clinical use — is consolidating quickly around a small number of well-funded providers, and founders should generally not compete there. Building a proprietary version of infrastructure that a handful of large platforms already offer well is rarely a good use of a small team's engineering time.
The application layer closest to a specific clinical workflow remains genuinely open, and it is where most of the interesting founder opportunity sits today. The gap between a generic model capability and a workflow that a clinician will actually trust and adopt is still large, and closing that gap requires deep domain judgment that infrastructure providers are not positioned to supply themselves.
Regulation will shape the roadmap, not just the launch
A clinical foundation model that is continuously updated raises a genuinely unresolved regulatory question: does each meaningful retraining constitute a new device requiring reassessment, or can a company operate under a predetermined change control plan that anticipates ongoing learning. The answer differs by use case and jurisdiction, and it is not yet fully settled anywhere.
Founders should build their model update process assuming regulators will eventually require detailed documentation of what changed between versions and why, even in jurisdictions where that requirement is not yet explicit. Retrofitting that discipline after a model has been in clinical use for a year is far more disruptive than building it in from the first deployed version.



