TL;DR
- Question: is self-hosting an LLM cheaper than a managed API? Not by default. It depends on a break-even you can calculate.
- Not about cost: data privacy, fine-tuning and latency control can decide it before the math does.
- The formula:
best-case cost=monthly price÷ tokens per month.real cost=best-case cost÷busy fraction. Self-hosting wins whenreal costis below theAPI price. - Made-up example: a $3,000 per month GPU against a $2 per 1M tokens API needs to be busy more than 57% of the time.
- Hidden costs: redundancy, egress, storage, supporting resources and people all raise the break-even.
When self-hosting is the right call
- Data privacy and residency. Regulated industries, or data that can’t be sent to a third-party API, often point to self-hosting. Private cloud endpoints can also meet some of these rules.
- Fine-tuning and custom models. If you need a custom fine-tuned model, running it yourself gives you full control. Some managed providers offer fine-tuning too, so check before you rule them out.
- Latency and concurrency control. Strict latency guarantees may need dedicated GPUs instead of shared capacity.
- Cost. At high, steady volume, self-hosting can be cheaper than a managed API. It only pays off past a break-even point, and there is no universal number for it, so you have to calculate your own.
Sources: General Compute and Azumo list these reasons. Lyceum states there is no general crossover figure.
The cost is not only the GPU
Before comparing with an API, you need to know what one million tokens really costs you on your own GPU. I’ll build it in two steps: the GPU cost first, then the idle cost.
The numbers below are made up to keep the math simple. Replace them with your own.
- GPU: an H200 that costs $3,000 per month
- Speed: 1,000 tokens per second
- Month: 730 hours
Step 1: the GPU cost
tokens per hour=tokens per second× 3,600- 1,000 × 3,600 = 3.6M tokens per hour
tokens per month(if the GPU never stops) =tokens per hour× 730- 3.6M × 730 = 2,628M tokens per month
best-case cost=monthly price÷ millions oftokens per month- $3,000 ÷ 2,628 = $1.14 per 1M tokens
This is the best case. It assumes the GPU produces tokens every second you pay for.
Step 2: add the idle cost
A rented GPU costs the same when it’s busy and when it’s idle. A reserved GPU costs the same at ten percent load as at ninety (Lyceum). So the real cost is the best-case cost divided by the fraction of time the GPU is busy.
real cost=best-case cost÷busy fraction
| GPU busy | Calculation | Cost per 1M tokens | Money wasted on idle time |
|---|---|---|---|
| 100% | 1.14 ÷ 1.00 | $1.14 | $0 |
| 50% | 1.14 ÷ 0.50 | $2.28 | $1,500 |
| 20% | 1.14 ÷ 0.20 | $5.71 | $2,400 |
idle money wasted = monthly price × (1 − busy fraction). At 20% busy, $3,000 × 0.80 = $2,400.
The same GPU, the same price, and the cost per token is five times higher. Only the usage changed.
In the 20% case, you are wasting $2,400 of your $3,000 every month on idle time.
The break-even formula
Now compare your real cost with an API. Self-hosting wins when your real cost per 1M tokens is below the API price.
I’ll keep the same made-up GPU and add a made-up API price of $2 per 1M tokens.
Step 3: find the break-even
- Self-hosting wins when:
best-case cost÷busy fraction<API price- 1.14 ÷
busy fraction< 2
- 1.14 ÷
- Rearrange it:
busy fraction>best-case cost÷API price- 1.14 ÷ 2 = 0.57
- The GPU must be busy more than 57% of the time.
Check it with the table from before. At 50% busy, the cost is $2.28, which is above $2, so the API wins. At 80% busy, the cost is $1.43, so self-hosting wins.
The same answer in tokens
break-even volume = monthly price ÷ API price
- $3,000 ÷ $2 = 1,500M tokens per month
Check: 57% of the 2,628M tokens the GPU can make in a month is about 1,500M.
One check before you trust it
The break-even volume must fit inside what the GPU can make. Here 1,500M fits inside 2,628M, so one GPU is enough. If the break-even were above 2,628M, one GPU could never reach it, and you would need more GPUs.
Also compare at the same token mix. APIs usually charge different prices for input and output tokens, so use one blended price that matches your real traffic (Lyceum).
Hidden costs
The GPU is not the only fixed cost, and each item below raises your break-even.
Redundancy
A failover node means paying for a second GPU that mostly waits, which roughly doubles your GPU cost in the simple case (General Compute).
Egress
Providers charge for data leaving the cloud past a free allowance (AWS), which is small for text but matters for audio and video.
Storage
Model weights and snapshots are billed per GB per month (AWS EBS), which is rarely the biggest line but grows with every model version you keep.
Supporting resources
Load balancers, monitoring and networking add 20 to 40 percent on top of the GPU cost according to Azumo, which is one source’s assumption, not a rule.
Maintenance and people
General Compute estimates 2 to 4 weeks of senior engineer time for setup and 5 to 10 hours a week to maintain it, and this is often the biggest cost because it depends on your team.
Back to the formula
Replace monthly price with total monthly cost (GPUs, failover, supporting resources, storage, egress and people), and the break-even volume goes up.
Conclusion
Self-hosting is the right call when privacy, fine-tuning or latency requires it, or when your volume is high and steady enough to pass the break-even.
Know your numbers: measure your own throughput, your busy fraction and your total monthly cost, then plug them into the formulas above.
If the numbers say managed, you can start there and move to self-hosting once real usage justifies it, or run a hybrid with managed for bursts and a smaller self-hosted base (General Compute).
All figures in the examples are made up, and most sources here are vendor blogs, so verify before you rely on them.