Home / Deployment postures
Deployment Postures

One product, three places models can run.

Two layers deploy under different rules. The platform (console, gateway, hosted applications, member memory) always runs under your control, on your hardware or in your own cloud tenancy, never at a model provider. The models are the layer that moves: each workload runs wherever its economics and constraints point.

The three postures, side by side.

PostureWhere models runCost shapePrimary draw
On-premises Customer-owned certified hardware, in your building. Fixed cost that amortizes; wins at high sustained volume. Strongest sovereignty; ownership economics.
Your cloud (VPC) A cloud tenancy in your name: your existing one, or one we procure and stand up for you during onboarding. Rented compute, no capex; data stays in your tenancy. Control without a data center, and without needing prior cloud maturity.
Governed third-party API Provider endpoints, frontier or open-model hosts, under your own keys. Pure per-token; cheapest at low or bursty volume. Immediate start; frontier access; commodity open-model prices.

Every posture runs through the same governance. Every request flows through your own gateway (identity, metering, budgets, the content-free audit) before any model sees it, wherever that model runs. The data-boundary label states plainly when a workload's prompts leave your perimeter, and to whom.

The posture mix is set per application, per group.

One group's chat can run on a local model while another group's runs against a frontier provider under your key with an allowance, falling back to the local model when it's spent. That per-request governance is the same gateway in every posture, which is what makes the mix administrable by a generalist instead of a platform team.

Start light

Platform on a small, GPU-free footprint in your own tenancy, every model reached through governed APIs. Capture the open-model price gap immediately, with no hardware procurement and no data-center visit.

Add capacity when your numbers say so

The metering that enforces your budgets also measures your real demand shape. When a workload's sustained volume justifies fixed capacity, the recommendation comes from your own usage data.

Move without a project

Moving a workload between postures changes where its model runs and nothing else. The same gateway, policies, applications, and member memory carry over, so the move becomes an economic decision you take when the numbers support it.

Isolation posture is a configurable ladder.

Independent of where models run, you choose how connected the deployment is. Each rung is a real operating mode with its tradeoffs stated up front, so you can hold exactly the line your constraints require.

Rung 1 · strictest

Air-gapped

No connectivity. For classified, government, and contractually air-gapped environments. Fewest capabilities, strongest guarantee.

for mandates that require it
Rung 2 · common case

Controlled

Internet-connected, data flow governed. Documents and prompts never leave; only explicitly permitted, logged outbound calls do: user-invoked web search, docs lookups, whitelisted APIs, third-party models under your key.

most deployments land here
Rung 3 · most open

Connected

Fully connected, for cost-driven buyers with no residency mandate. Still governed: the gateway, budgets, and audit apply to every call regardless.

lightest to operate
Hardware

Certified, prebuilt, and yours.

For the on-premises posture we deploy on prebuilt enterprise AI systems: the Dell, HPE, and Lenovo AI factory lines built on NVIDIA AI Enterprise. Never custom-assembled hardware.

  • Your capex, your depreciation, your balance-sheet benefit. The hardware is yours outright.
  • A small set of blessed configurations. We deploy known-good configurations of the certified lines, which is what keeps fleet operations dependable.
  • Transparent bill of materials. Hardware at vendor pricing and the NVIDIA AI Enterprise entitlement appear on your bill of materials as what they are. No bundle fog.
Start anywhere

Start where you are, even with no cloud at all.

For a buyer segment defined by not having a platform team, requiring an existing, well-run cloud tenancy would be a strange thing to ask. So we don't.

If you have no cloud presence at all, we procure and stand up a tenancy in your name as a fast onboarding step, with ownership landing where it belongs: your tenancy, your keys, your data. The platform layer needs only modest, GPU-free compute, so it imposes no hardware decision of its own.

More in the FAQ →

Posture-mapping is part of the product.

Bring your constraints (residency, budget, existing spend, appetite for hardware) and we'll map your workloads to postures on your numbers.

Talk to us