Home / Economics
Economics

You already pay an AI bill. The real question is where each workload should run.

Frontier seats, direct API spend, or both: every buyer arrives with a bill. Whether AI is worth buying is a question you've already answered. What our product puts in front of you instead: given what each source costs and what constraints each workload carries, where should it run?

Three sources, three cost shapes.

Any workload can be served from three places, and they price in different units:

Frontier APIs

Top capability, top price

The frontier closed models have no open-weight substitute at the top end, and their list prices reflect it. A share of your work will need them; the everyday majority runs well on the tier below.

Open-weight APIs

The immediate price gap

The tier of open-weight models that handles the bulk of everyday work sells an order of magnitude or more below frontier list. That gap is capturable today, through a governed gateway, with no hardware.

Your hardware

Priced per GPU-hour

Owned or VPC capacity is a fixed cost that amortizes. Whether it beats the APIs depends on one variable more than any other: how much of the time the machine is actually serving.

The spread inside "open" is real, and it's why curation is a product decision. The strongest open models price near frontier rates; the workhorse tier below them runs far cheaper. Deciding which model earns which workload is the difference between a price gap on a slide and one on your invoice.

The one variable that dominates

Utilization decides who wins.

The cost of an owned GPU-hour is the cost of running it plus the fixed cost of owning it, spread over the hours it actually serves. Because the fixed cost is recovered only over busy hours, unit cost rises steeply as utilization falls.

Independent research measured that curve directly: effective cost spanning $0.21 to $15.25 per million output tokens on identical hardware, a penalty of 2.5 to 24 times across the range of typical enterprise load.

Two consequences follow:

  • An operator pooling demand across many customers runs at utilization no single tenant can match. That's why hosted APIs undercut a lightly used private cluster.
  • A buyer whose demand is genuinely sustained inverts the advantage: the fixed cost spreads across enough hours to fall below what any provider charges to cover the same asset plus margin.
lightly used private cluster fixed cost spread over few busy hours sustained enterprise demand ownership beats any provider's price utilization → (share of hours actually serving) effective $ / million tokens measured range on identical hardware: 2.5–24×
Utilization follows your demand shape. A business-hours workforce caps it; training or fine-tuning filling the idle hours lifts it.

Which posture wins, by workload profile.

The crossover exists in every deployment; where it falls depends on inputs specific to yours. The shape of the answer, though, is consistent:

Workload profileCheapest sourceWhy
Low or bursty volumeGoverned third-party APIPooled operator utilization no single tenant matches.
High, steady volumeOwned or VPC capacityFixed cost amortizes across sustained hours.
Sustained volume under a sovereignty constraintOwned or VPC capacityCompliant alternatives price well above commodity rates.
Dual use: training alongside inferenceOwned capacityOne asset displaces two rental bills.
Peak headroom above any baselineGoverned third-party APIBurst without sizing hardware for peak.
Your numbers decide

Your breakeven comes from your own meters.

A breakeven percentage quoted to the decimal, before anyone has measured your demand shape, is a modelling artifact dressed as a fact. We hold a working cost model, re-computable from its inputs, and we'll walk through it with you in diligence. But the number that decides your posture is yours:

  • The same metering that enforces your budgets measures your real demand shape.
  • The posture recommendation is computed from that usage data.
  • Until the data says otherwise, the cheap posture is where you stay.
Incentives

Our fee is the same in every posture.

We're paid for governance, the application catalog, and operations, and the fee is charged in every posture. A customer who runs entirely on third-party APIs pays it; so does one who buys a cluster.

Our income stays flat whether you buy hardware or stay on APIs, so when we tell you a workload should move, or stay put, the advice has no thumb on the scale.

The commercial structure →

What this looks like in practice.

Capture the gap now

Move everyday seat work from frontier list prices to governed open-weight APIs under your own keys. No hardware, no procurement cycle, immediate difference on the bill.

Keep frontier where it earns it

The work that genuinely needs a frontier model keeps it, through the same gateway, under allowances. Nobody argues about it, because the audit trail shows which work that is.

Let your meters call the crossover

When a workload's sustained volume justifies fixed capacity, on-prem or VPC, your own numbers make the case. If the case stays weak, you simply keep the capex-free postures, with the full product intact.

Bring last month's AI bill.

Thirty minutes with your actual spend, and we'll show you what the posture mix would look like, and what we'd wait to measure before recommending anything else.

Talk to us