Foundation models meet the clinic: what the first year of deployment actually taught us
Twelve months of hospital pilots have revealed a clear pattern — general-purpose models plateau fast, and verticalized systems win the long tail.

The honeymoon is over
Health systems that rushed to pilot general-purpose foundation models in 2023 have quietly begun to sunset most of those deployments. The demos were compelling. The workflows were not. A model that summarizes a chart brilliantly in a boardroom often fails on the third patient of a Monday clinic — the one with three comorbidities, two conflicting notes, and a medication list that hasn't been reconciled in nine months.
The clinicians we spoke with across eight academic centers describe the same arc: initial enthusiasm, an accuracy cliff at the edges of the training distribution, and a slow return to structured tooling with narrower scope.
Where verticalization compounds
The teams that are still growing usage are the ones that specialized. Ambient scribes tuned to specific specialties. Radiology copilots that only handle chest CT. Pre-authorization agents restricted to a single payer's rule set. In each case, the surface area shrinks and the model's failure modes become legible enough to trust.
This is not a story about smaller models. It is a story about smaller problems. The winning teams pair a large backbone with a tight retrieval layer over the institution's own data, then evaluate weekly against a fixed clinical rubric.
The evaluation gap
Almost every serious deployment we studied has built its own eval harness — because none of the public benchmarks reflect the messiness of real EHR data. Founders who ship an evaluation product alongside their model are winning enterprise trust faster than those who ship raw capability.
Expect procurement to standardize on this in 2025. Health systems will ask, in writing, for a clinical evaluation methodology before signing.
What we're funding next
We are actively looking for teams building the workflow substrate underneath these models: consent capture, provenance, model routing, and human-in-the-loop review at the point of care. The application layer will fragment. The plumbing will consolidate.

