Break-even analysis: when on-prem GPUs beat reserved cloud
Cloud GPU rental at $1.50-3/hr looks expensive on paper. Owning an H100 looks expensive on the invoice. The crossover depends on utilization more than anything else — here's the math, with caveats.
A used H100 costs roughly $25,000 in 2026. Renting an H100 from Lambda Labs or Runpod costs roughly $2/hour. Naive math says owning pays off after 12,500 hours = ~1.4 years of 24/7 usage. But the real math has a lot more in it. This post walks through what we actually use on engagement decisions.
The straightforward case
For a fully-saturated, 24/7 production workload, the break-even between on-prem and rented cloud is roughly:
months_to_break_even = capex / (rental_rate × 24 × 30)
= $25,000 / ($2 × 720)
= 17 months
After 17 months, the on-prem H100 is paying for itself relative to the rental. Over a 3-year life, it saves ~$25K vs. continuous rental.
This is the headline number. It's also wrong in most real engagements because the assumptions don't hold.
What changes the math
Utilization rate. Most workloads don't saturate 24/7. If your usage is 8 hours/day weekdays only, the on-prem GPU sits idle 75% of the time. The break-even stretches to 17 / 0.25 = 68 months — 5.7 years. By then the H100 is two generations behind and worth less than a 4090.
Power and cooling. ~700W under load × $0.15/kWh × 24 × 30 = $76/month if running 24/7. Less if idle. Data center rack and cooling adds another 20-30%.
Operations overhead. Replacement parts, monitoring, dealing with the inevitable failure at 3 AM. Conservatively 0.05 FTE @ $150/hr = $1,300/month.
Depreciation risk. A new GPU generation lands every 18-24 months. Buying an H100 in 2026 means it's a B200 era by 2027 and a backup-machine candidate by 2029. Resale value at 3 years: maybe 30% of capex.
Idle cloud savings. Modern providers like Lambda Labs charge only when the instance is running. If you can spin down at night, weekends, and during low-traffic periods, your effective rental hours drop.
A realistic comparison
Let's price out three real scenarios over 3 years.
Scenario A: 24/7 production workload at full utilization
On-prem:
- Capex: $25,000 (one H100 + workstation)
- Power: $76/mo × 36 = $2,736
- Ops: $1,300/mo × 36 = $46,800
- Total over 3 years: $74,536
- Resale at year 3: -$7,500 (30% recovery)
- Net: $67,036
Cloud (Lambda Labs):
- Rental: $2/hr × 24 × 30 × 36 = $51,840
- Ops: $500/mo × 36 = $18,000 (less, since hardware is theirs)
- Total: $69,840
On-prem wins by ~$2,800. Marginal. Could go either way depending on actual electricity, ops overhead, and resale.
Scenario B: Business-hours workload (8h/day weekdays)
On-prem (same capex, same ops, lower power):
- Capex: $25,000
- Power: $40/mo (running fewer hours) × 36 = $1,440
- Ops: $1,300/mo × 36 = $46,800
- Total: $73,240 minus resale = $65,740
Cloud (only paying when running, ~2,000 hr/yr):
- Rental: $2/hr × 2,000 × 3 = $12,000
- Ops: $500/mo × 36 = $18,000
- Total: $30,000
Cloud wins by ~$35K. Significant. When utilization is partial, the cloud's pay-per-use model dominates.
Scenario C: Burst workload (most of the year quiet, 2 months at saturation)
On-prem:
- Capex: $25,000
- Power: $30/mo (mostly idle) × 36 = $1,080
- Ops: $1,300/mo × 36 = $46,800
- Total: $72,880
Cloud (rent only during burst, ~2 months × 720 hr × 2/hr × 3 years):
- Rental: $4,320
- Ops: $500/mo × 36 = $18,000
- Total: $22,320
Cloud wins by $50K. When workload is mostly quiet with occasional bursts, on-prem is a terrible match.
The pattern
Above ~70% sustained utilization, on-prem wins. Below ~30%, cloud wins. Between 30-70%, it's a toss-up that depends on specific assumptions, and the qualitative factors usually break the tie.
Real engagements rarely sustain 70% utilization. Most production workloads have diurnal patterns (lower at night), weekly patterns (lower on weekends), and seasonal patterns. Effective utilization for "always-on" services is often 30-50%.
The hybrid model
The cleanest answer for most engagements is hybrid:
Steady-state on-prem. Buy enough GPU to cover the baseline load — the constant ~30% of traffic that's reliably there.
Burst to cloud. Spin up cloud instances during peak periods. Spin them down when the peak passes.
Calibration / training on cloud. One-off heavy compute (custom imatrix calibration, fine-tuning runs) goes to cloud spot instances at $0.50-1/hour. No reason to own infrastructure for jobs that run once a quarter.
The hybrid model gets the cost benefits of on-prem for the baseline plus the elasticity of cloud for everything else. Operational complexity goes up, but it's manageable with the right control plane.
Don't forget the GPU generation cycle
H100 in 2026 is no longer the flagship. B200 has been shipping for over a year. By 2027 you'll have B300 or whatever Nvidia names the next refresh. By 2028 H100 is yesterday's silicon.
This is the durable argument for cloud: someone else holds the depreciation risk. They write off the GPU, you rent the current-generation hardware at competitive rates. Three-year-old on-prem hardware is a productivity tax — slower than the latest, hard to resell, hard to replace.
If your workload doesn't need the latest GPU (a 7B model on a 4090 runs fine and probably will for years), the depreciation argument is weaker. If you need flagship hardware, renting is structurally better positioned.
What we recommend
A rough heuristic from a half-dozen engagements:
| Workload shape | Recommendation |
|---|---|
| 24/7 saturated production | On-prem if utilization is real; otherwise hybrid |
| 8-hour business day | Cloud |
| Burst / seasonal | Cloud |
| Internal tool, small team | 4090 workstation under someone's desk |
| Customer-facing, multi-tenant | Cloud, scaled up and down |
| Compliance-bound (data residency) | On-prem regardless of cost |
| Calibration / training only | Cloud spot, every time |
Caveats to the cost numbers
Every number in this post is indicative. Real engagement numbers vary based on:
- Specific GPU model and resale market
- Local electricity rates
- Cloud provider, region, and any negotiated discount
- Reserved vs. on-demand pricing
- Whether the workload is genuinely 24/7 or has hidden idle time
- Headcount and pay scales for ops support
We've seen engagements where on-prem broke even in 8 months and others where it never did over 3 years. The math is honest but the inputs vary a lot. Always run the calculation against your actual inputs before committing capex.