Frontier seats, direct API spend, or both: every buyer arrives with a bill. Whether AI is worth buying is a question you've already answered. What our product puts in front of you instead: given what each source costs and what constraints each workload carries, where should it run?
Any workload can be served from three places, and they price in different units:
The frontier closed models have no open-weight substitute at the top end, and their list prices reflect it. A share of your work will need them; the everyday majority runs well on the tier below.
The tier of open-weight models that handles the bulk of everyday work sells an order of magnitude or more below frontier list. That gap is capturable today, through a governed gateway, with no hardware.
Owned or VPC capacity is a fixed cost that amortizes. Whether it beats the APIs depends on one variable more than any other: how much of the time the machine is actually serving.
The spread inside "open" is real, and it's why curation is a product decision. The strongest open models price near frontier rates; the workhorse tier below them runs far cheaper. Deciding which model earns which workload is the difference between a price gap on a slide and one on your invoice.
The cost of an owned GPU-hour is the cost of running it plus the fixed cost of owning it, spread over the hours it actually serves. Because the fixed cost is recovered only over busy hours, unit cost rises steeply as utilization falls.
Independent research measured that curve directly: effective cost spanning $0.21 to $15.25 per million output tokens on identical hardware, a penalty of 2.5 to 24 times across the range of typical enterprise load.
Two consequences follow:
The crossover exists in every deployment; where it falls depends on inputs specific to yours. The shape of the answer, though, is consistent:
| Workload profile | Cheapest source | Why |
|---|---|---|
| Low or bursty volume | Governed third-party API | Pooled operator utilization no single tenant matches. |
| High, steady volume | Owned or VPC capacity | Fixed cost amortizes across sustained hours. |
| Sustained volume under a sovereignty constraint | Owned or VPC capacity | Compliant alternatives price well above commodity rates. |
| Dual use: training alongside inference | Owned capacity | One asset displaces two rental bills. |
| Peak headroom above any baseline | Governed third-party API | Burst without sizing hardware for peak. |
A breakeven percentage quoted to the decimal, before anyone has measured your demand shape, is a modelling artifact dressed as a fact. We hold a working cost model, re-computable from its inputs, and we'll walk through it with you in diligence. But the number that decides your posture is yours:
We're paid for governance, the application catalog, and operations, and the fee is charged in every posture. A customer who runs entirely on third-party APIs pays it; so does one who buys a cluster.
Our income stays flat whether you buy hardware or stay on APIs, so when we tell you a workload should move, or stay put, the advice has no thumb on the scale.
The commercial structure →Move everyday seat work from frontier list prices to governed open-weight APIs under your own keys. No hardware, no procurement cycle, immediate difference on the bill.
The work that genuinely needs a frontier model keeps it, through the same gateway, under allowances. Nobody argues about it, because the audit trail shows which work that is.
When a workload's sustained volume justifies fixed capacity, on-prem or VPC, your own numbers make the case. If the case stays weak, you simply keep the capex-free postures, with the full product intact.
Thirty minutes with your actual spend, and we'll show you what the posture mix would look like, and what we'd wait to measure before recommending anything else.