Pricing
Per GPU, per month. Published, so you can tell whether you can afford us.
Enterprise buyers do not mind talking to sales. They mind not being able to work out the order of magnitude before they do. The bands are below, along with the arithmetic for deciding whether this pays for itself.
Pricing bands
1–8 GPUs
$180
per GPU, per month
Below four GPUs there is usually not enough contention to be worth managing. We will tell you that on the call.
9–32 GPUs
$140
per GPU, per month
The band most first deployments land in. Typically two to five teams sharing a fleet.
33–128 GPUs
$110
per GPU, per month
Multiple sites or zones, compliance labels in active use, quota that changes monthly.
129+ GPUs
Talk to us
and we will publish the band
The curve keeps going down. At this size the conversation is about placement policy and support terms rather than unit price.
Annual commitment
15% off the monthly rate. No multi-year lock-in on offer, because a three-year contract with a company this young is not a good trade for you.
Trial
30 days on up to 8 GPUs. No card, no commitment, and no automatic conversion to a paid plan at the end of it.
Do the arithmetic first
Whether this pays for itself.
We would rather you worked this out on the page than discovered it in month three.
a 16 × H100 fleet, three teams
- Fleet
- 16 × H100 80GB
- Band
- 9–32 GPUs, $140 per GPU per month
- Atlas, monthly term
- $2,240 / month
- Atlas, annual term
- $1,904 / month · $22,848 / year
- Hardware it manages
- roughly $400k–450k of GPUs at list
- As a share of hardware cost
- about 5–6% per year
- Against one platform engineer
- roughly 2–3 weeks of a fully-loaded senior salary, per year
The test that actually matters
The pricing model
Why it is not priced per token.
Hosted providers charge per token because tokens are what they sell — they own the GPU, and the token is the unit of the thing you are renting. That model does not transfer to hardware you already bought.
If we charged per token, you would be paying twice for the same silicon: once to NVIDIA, and again to us every time you used it. Worse, it would put our incentive on the wrong side. We would want your token volume to rise. You want your existing fleet to go further, which usually means fewer tokens, better cache hit rates, and smaller models where they suffice.
Capacity pricing keeps those aligned. We are paid for the GPUs under management, so the only way we grow inside an account is by being useful enough that you put more of your fleet on it — which you will not do unless it is working.
It also means your bill is predictable. Finance can forecast a per-GPU line. Nobody has to explain a spike caused by a misbehaving retry loop.
Included at every band
There is no feature gating by tier. SSO is not an enterprise upsell; it is table stakes for the only audience we sell to.
- Every model in the catalogue, and any open-weight model you bring
- Unlimited teams, keys, deployments and environments
- Unlimited console users — we do not charge per seat
- SSO via SAML or OIDC, and SCIM provisioning
- Audit log with SIEM export
- Quota, reservations and admission policy
- Per-team metrics, chargeback data and capacity reporting
- Engine version pinning and rollback per deployment
What we do not charge for
And, more usefully, why not.
- Tokens
- You already own the GPU. Charging per token would be charging you twice for the same silicon.
- Data egress
- There is barely any. Telemetry is a few kilobytes a minute per node.
- Seats
- Per-seat pricing makes platform teams ration access to their own tooling, which is the opposite of the point.
- Models
- No per-model fee and no premium tier for larger models. A 70B deployment costs the same as an 8B one because it consumes the GPUs you already declared.
- Support
- Included at every band. A paid support tier is a tax on having a problem.
Pricing questions
- How much does Atlas cost?
- It is priced per managed GPU per month, in published bands: $180 at 1–8 GPUs, $140 at 9–32, $110 at 33–128, and a negotiated rate above that. There is no per-token charge, no egress charge and no per-seat charge. A 16-GPU fleet is $2,240 a month on a monthly term, or about $1,904 a month on an annual one.
- Why not price per token like everyone else?
- Per-token pricing exists because hosted providers own the hardware and the tokens are what they are selling you. You own the hardware. Charging you per token would mean charging you a second time for a GPU you already bought, and it would give us an incentive to want your token volume to go up rather than your fleet to go further. Capacity pricing puts our incentive on the same side as yours.
- Do we pay for idle GPUs?
- You pay for GPUs under management, whether or not they are busy. If a chunk of your fleet is genuinely idle, take it out of management and the bill goes down — and our capacity reporting will be the thing that tells you it is idle in the first place.
- Is there a free tier?
- No, but there is a 30-day trial on up to 8 GPUs with no card and no commitment, and the capacity calculator on this site is free and ungated forever because it is useful whether or not you ever talk to us.
- At what point does self-hosting beat a hosted API?
- Roughly when your hosted API spend passes the fully-loaded cost of the hardware that would replace it, which for most teams is somewhere in the low tens of thousands of dollars a month. Below that, a hosted API is usually the right answer and we will say so. The reasons to self-host below that line are data residency and latency, not cost.
Prices are in USD and exclude VAT and sales tax. They are the published bands as of this page's last update; a quote is fixed for the term of the contract. Work out what your fleet can serve before deciding whether the per-GPU figure is worth it.
Get a quote with your actual fleet in it
Tell us the GPU count and we will send a figure, not a discovery call. If your fleet is too small for this to make sense, we will say that instead.
Priya Raghavan replies within one working day. No sequence, no SDR, no calendar tennis.