blog/guides·Aug 20, 2026·4 min·by the eroq team
Self-hosting uncensored models vs an uncensored AI API — the real math
GPU rental, utilization curves and ops time — an honest break-even analysis of running your own uncensored stack versus flat per-call credits.
"Why pay per call when a rented 4090 is $300 a month?" is the most reasonable objection to an uncensored AI API — and sometimes it's right. Here is the actual math, including the parts that don't fit in a tweet.
The sticker math
A rented consumer GPU (4090-class) runs roughly $0.35–0.50/hour — call it $250–370/month always-on. An A100-class card for larger models runs $1.10–1.80/hour — $800–1,300/month. A quantized 13B roleplay model on the 4090 serves a real product's chat traffic; a 70B wants the A100 or a multi-GPU rig.
Against flat credits (a completion is 1–3 credits ≈ 1–3¢, full table), the naive break-even on chat is:
| Monthly completions | API cost (RP mini) | Self-host (4090) |
|---|---|---|
| 100,000 | ~$830 | ~$300 + ops |
| 30,000 | ~$250 | ~$300 + ops |
| 10,000 | ~$83 | ~$300 + ops |
Somewhere around 30–40k completions a month, the GPU line crosses under the API line. Case closed for self-hosting at scale? Only if the next four sections don't apply to you.
What the sticker math hides
Utilization is the whole game. The GPU bills 24/7; your users arrive in evening peaks. Real products see 15–35% utilization, which doubles or triples your effective per-call cost. The API's flat price is someone else's batching — you pay for output, not idle silicon. Failed generations even refund themselves.
Ops is a salary, not a footnote. Model updates, CUDA driver roulette, OOM crashes at your Saturday peak, quantization regressions after every fine-tune update. Budget real engineering-hours a month for a production inference stack — at contractor rates, that's your API bill right there before serving a single token.
Tuning is the invisible cost. Raw uncensored checkpoints repeat, drift and break the fourth wall under roleplay pressure — the failure modes are specific and fixing them (sampling profiles, anti-repetition, register control) is exactly the work you were hoping to skip.
Chat was the easy modality. The moment your product wants images, video, voice or transcription, self-hosting multiplies: separate models, separate VRAM, separate pipelines — plus storage and a CDN for outputs. One API key covering all ten is the argument that usually ends the debate for small teams.
When self-hosting genuinely wins
Playing it straight, self-hosting is the right call when all four hold:
- Steady, high, predictable volume (≥50k completions/month without evening cliffs);
- One modality, one model you've already validated;
- An engineer who wants to own inference (data-locality or compliance requirements count double);
- Tolerance for a weekend of downtime while something recompiles.
That's a real profile — some of the best products in the space run this way. It just isn't most products, and it's almost never products at the start, when iteration speed is worth more than margin.
If privacy is what's pushing you toward your own GPU, read the provider's retention terms before the GPU quotes. Here, eroq's own models — RP+ and RP mini included — retain nothing upstream, and Private mode (a Studio toggle, or private: true on the image and video endpoints) keeps no file and no library entry and replaces the prompt with stars in the usage ledger. Outside Private mode, prompts are kept in the usage ledger for billing and abuse handling; if your compliance requirements rule even that out, that's condition 3 above, and self-hosting is the honest answer.
The hybrid that actually works
The pattern we see succeed: prototype and launch on the API (a key, 50 free credits, five-minute quickstart), instrument your real traffic, and revisit the math at your genuine volume — flat public prices make the API side of the spreadsheet a five-minute job. If chat volume alone crosses the line, move that workload to your own GPU and keep images, video and voice on the API. Statelessness makes the swap boring: it's one base URL per workload.
FAQ
Is self-hosting an uncensored model cheaper than an API?
At volume, on chat alone, eventually yes — the naive crossover sits around 30–40k completions a month against 1-credit RP mini calls. Then apply the corrections: 15–35% real utilization doubles or triples your effective per-call cost, and ops time is a salary line. Products that model those two honestly usually find the crossover much further out than the sticker math suggests.
What GPU do I need to run an uncensored roleplay model?
A quantized 13B on a 4090-class card serves a real product's chat traffic at roughly $250–370 a month always-on. A 70B wants an A100-class card or a multi-GPU rig, in the $800–1,300 range. Those are rental prices as of September 2026 — check current rates before you build a spreadsheet on them.
What does eroq cost at 100,000 completions a month?
100,000 credits on RP mini, which is about $833 at the $100-for-12,000 pack rate, or roughly $1,000 at the entry pack. Plan credits roll over forever, so a subscription that overshoots one month is not wasted — it banks.
Can I run both, and move workloads later?
Yes, and it is the pattern that usually works. Prototype on the API, instrument your real traffic, then move only the workload that crosses the line — almost always chat — onto your own GPU while images, video, voice and transcription stay on one key. The API is stateless, so the swap is one base URL per workload rather than a migration.
Infrastructure decisions deserve receipts, not vibes — run both columns on your own numbers.
Make this with the models behind the post — start with 50 free credits , or browse every engine and its price .