Rent or Own? The Hospital AI Question of the Year
Dear readers,
Every hospital board we speak with eventually asks the same question. Not "which model is the smartest," but: what will this cost us at scale, and who stays in control?
Until now, the honest answer was uncomfortable. Strong opinions everywhere, verifiable numbers almost nowhere. The debate between renting intelligence per token and running open-weight models yourself has been driven by vendor marketing on one side and ideology on the other.
So we did the unglamorous work. Today we publish a white paper: Should Your Hospital Rent AI by the Token, or Run It Itself? It synthesizes 30 sources, the core of them peer-reviewed, with every number dated, sourced, and graded by evidence class. Where a number comes from a practitioner blog, the paper says so. Where the evidence does not exist, the paper says that too, in seven explicit gaps.
Five findings that change the conversation:
1. The floor is falling while the ceiling rises. The price of a fixed capability level collapsed more than 280x between November 2022 and October 2024 (Stanford AI Index 2025). Meanwhile, running the absolute frontier becomes 3x to 18x more expensive per year (MIT, 2025). Cheap gets cheaper. The cutting edge gets pricier.

Figure 1: Fixed-capability prices collapsed 280x in 23 months, while frontier running costs keep rising.
2. Break-even is real, and it is tiered. Small open models on consumer hardware pay back in weeks to three months. Mid-size deployments need months. Large clusters take 4 to 69 months against premium APIs, and fail entirely against value-priced APIs (Carnegie Mellon, 2025, across 54 scenarios). Scale alone is not a business case.

Figure 2: Months to break-even by deployment tier. Large clusters fail against value-priced APIs.
3. Token volume arrives faster than boards expect. One heavy clinician workflow already consumes a modeled 3.5 million tokens per day, roughly $450 to $650 per month at premium API prices. A mid-size hospital crosses practitioner break-even thresholds sooner than its budget cycle suggests.
4. The answer is not cloud versus local. It is a tiered architecture, we call it proximity AI. This is the finding we care most about, because it matches what we build. High-volume routine work runs on-device or on-premise. Big-context orchestration runs on privately hosted open models, at $0.10 to $1.20 per million tokens, with zero-retention and European hosting options. Frontier APIs handle the hardest reasoning tail. Compute belongs as close as possible to where the work happens, in time and space. We call that principle Proximity AI, and the economics literature now supports it independently.
5. Cost is not the only ledger. Across the EU, employee representatives hold consent-level rights over monitoring-capable systems (Germany, the Netherlands, Austria, Italy), with consultation duties in France, Spain, and Sweden. A German labor court has already ruled that introducing AI speech recognition in a hospital was a change of operations requiring a negotiated social plan. Deployment architecture determines labor-relations exposure. Certifications do not dissolve it.
Who should read it: hospital CIOs, CMIOs, CFOs, and the people who advise them. It is written for the boardroom, not the lab. Jargon is defined once, assumptions are stated, and a 90-day plan is included.
And yes, hospitals are already doing this. A German university hospital running a compliant on-premise platform reports 44% of its active users saving at least 30 minutes per week (JMIR AI, 2026).
The question is no longer whether open-weight models are good enough for hospital workflows. The question is whether your cost model has caught up with the evidence.
Read the white paper here
We are happy to walk your team through the break-even model with your own numbers. That is where the paper stops being theory.
Bart