Two layers deploy under different rules. The platform (console, gateway, hosted applications, member memory) always runs under your control, on your hardware or in your own cloud tenancy, never at a model provider. The models are the layer that moves: each workload runs wherever its economics and constraints point.
| Posture | Where models run | Cost shape | Primary draw |
|---|---|---|---|
| On-premises | Customer-owned certified hardware, in your building. | Fixed cost that amortizes; wins at high sustained volume. | Strongest sovereignty; ownership economics. |
| Your cloud (VPC) | A cloud tenancy in your name: your existing one, or one we procure and stand up for you during onboarding. | Rented compute, no capex; data stays in your tenancy. | Control without a data center, and without needing prior cloud maturity. |
| Governed third-party API | Provider endpoints, frontier or open-model hosts, under your own keys. | Pure per-token; cheapest at low or bursty volume. | Immediate start; frontier access; commodity open-model prices. |
Every posture runs through the same governance. Every request flows through your own gateway (identity, metering, budgets, the content-free audit) before any model sees it, wherever that model runs. The data-boundary label states plainly when a workload's prompts leave your perimeter, and to whom.
One group's chat can run on a local model while another group's runs against a frontier provider under your key with an allowance, falling back to the local model when it's spent. That per-request governance is the same gateway in every posture, which is what makes the mix administrable by a generalist instead of a platform team.
Platform on a small, GPU-free footprint in your own tenancy, every model reached through governed APIs. Capture the open-model price gap immediately, with no hardware procurement and no data-center visit.
The metering that enforces your budgets also measures your real demand shape. When a workload's sustained volume justifies fixed capacity, the recommendation comes from your own usage data.
Moving a workload between postures changes where its model runs and nothing else. The same gateway, policies, applications, and member memory carry over, so the move becomes an economic decision you take when the numbers support it.
Independent of where models run, you choose how connected the deployment is. Each rung is a real operating mode with its tradeoffs stated up front, so you can hold exactly the line your constraints require.
No connectivity. For classified, government, and contractually air-gapped environments. Fewest capabilities, strongest guarantee.
for mandates that require itInternet-connected, data flow governed. Documents and prompts never leave; only explicitly permitted, logged outbound calls do: user-invoked web search, docs lookups, whitelisted APIs, third-party models under your key.
most deployments land hereFully connected, for cost-driven buyers with no residency mandate. Still governed: the gateway, budgets, and audit apply to every call regardless.
lightest to operateFor the on-premises posture we deploy on prebuilt enterprise AI systems: the Dell, HPE, and Lenovo AI factory lines built on NVIDIA AI Enterprise. Never custom-assembled hardware.
For a buyer segment defined by not having a platform team, requiring an existing, well-run cloud tenancy would be a strange thing to ask. So we don't.
If you have no cloud presence at all, we procure and stand up a tenancy in your name as a fast onboarding step, with ownership landing where it belongs: your tenancy, your keys, your data. The platform layer needs only modest, GPU-free compute, so it imposes no hardware decision of its own.
More in the FAQ →Bring your constraints (residency, budget, existing spend, appetite for hardware) and we'll map your workloads to postures on your numbers.