The questions buyers actually ask, including the awkward ones. If yours isn't here, ask it directly: plain answers are a product principle.
No. The separation is architectural: we operate the control plane (configuration, health, versions, metering) and are structurally unable to read the data plane (prompts, responses, documents, member memory, model weights). Telemetry that leaves your deployment carries control-plane fields only.
The same constraint contractually binds our subcontractors: any vendor access is scoped to the control and operations plane, never to your data.
For every model call: identity, application, model, decision, timestamp, and token counts. Message content is never recorded, by construction. That makes the trail complete enough for an auditor ("who used what, when, how much") and structurally incapable of becoming a surveillance log of what your employees wrote.
Inside your perimeter: on your hardware or in your own cloud tenancy, never as a SaaS we host. It holds your keys, your access rules, and your audit trail. Updates are pulled from a mirror inside the boundary on a schedule we gate; nothing is pushed in from outside, by any vendor, including us.
Yes. Isolation is a configurable ladder: air-gapped (no connectivity), controlled (the common case: data stays in, only explicitly permitted and logged calls go out), or connected. Stricter postures carry fewer conveniences and cost more to operate, and we price that honestly.
No, and the whole product is built on that premise. The console is designed for the IT administrator who runs Google Workspace, Okta, or the M365 admin center, and it's measured against the concrete jobs that person must complete unaided: give a team an AI application, handle "the AI is slow," handle "we need model Y," offboard a person and prove it, pass an audit. Kubernetes stays on our side of the line.
Yes, through the governed third-party posture: frontier providers plug in under your own keys, behind your own gateway, with per-person allowances and the same content-free audit as everything else. When an allowance is spent, the app falls back automatically to a local model. Frontier access becomes something you govern.
A curated menu with honest tiers. Hot models are pinned in GPU memory and answer instantly. Available models load on demand, with a first-call wait stated in the console. Everything else runs as an external model under your own provider keys, through the same gateway, with its capacity taken from the provider's own rate limit. Models appear under their real names, and the menu is kept current as the frontier moves, through the managed update pipeline.
The tiering is honest physics: GPU memory is the real ceiling on what can answer instantly, and we'd rather disclose that than let you discover it mid-rollout.
Two different questions get two different measurements. Whether a model fits is a question about GPU capacity, answered in units the console shows you. Whether it will hold up for everyone pointed at it is a question about measured throughput, and GPU count alone can't answer it: across model classes, the same hardware spans a roughly forty-fold range in tokens per second, so any tool converting GPU count into "people served" is guessing.
Capacity starts as a labelled industry-benchmark estimate, gets measured on your own hardware, and reads out in human terms: "about 40 people comfortably," with the basis shown. The console also pairs what you've committed through grants with the peak actually observed, so you can see over-commitment before your users feel it.
No, by design. Builder surfaces are aimed at a technical persona most of our buyers don't employ, so we put that engineering into finished apps instead. If you need something beyond the shelf, customer-funded builds are one of the ways the catalog grows.
Some catalog applications are proven open-source products we've adopted and integrated: licenses inventoried, notices preserved, branding intact where a license requires it. Two things matter for you:
Possibly none. Two of the three postures carry no hardware at all: your workloads can run in your own cloud tenancy or behind governed third-party APIs. If and when your sustained volume justifies owned capacity, we deploy on prebuilt certified enterprise AI systems (the Dell, HPE, and Lenovo AI factory lines), purchased by you, on your balance sheet, from a small set of blessed configurations we operate fleet-wide.
The hardware vendor's field service, working as our subcontractor under back-to-back commitments. You hold one SLA, with us; we hold our vendors to response times that let us honor it. We manage Dell, Red Hat, and NVIDIA behind the scenes, and their access stops at the control plane.
No. The platform layer needs only modest, GPU-free compute. If you have no cloud presence, we procure and stand up a tenancy in your name as part of onboarding, with ownership landing with you: your tenancy, your keys, your data. A missing cloud footprint is an onboarding step for us, and ownership still lands with you.
Every update, platform, models, and applications, is validated against real workloads before your fleet takes it, then delivered through a signed release path and pulled from inside your perimeter on a schedule we gate. Drift between the running system and its declared configuration is detected and automatically reverted. Your team stays out of it unless you've asked to gate specific windows.
You keep what's yours, which is most of it: the hardware is on your balance sheet, the cloud tenancy is in your name, the keys are yours, and your data never left your perimeter to begin with. Adopted open-source applications remain licensed to run. What you lose is us: the governance layer, the catalog updates, and the operations. Exit is a conversation we'll have with you before you sign.
A base deployment subscription (standard apps + governance + operations, with a per-seat governance license), per-app add-on subscriptions priced individually on the shelf, a one-time fixed-scope onboarding fee, and pass-throughs (hardware, the NVIDIA AI Enterprise entitlement, third-party model usage under your keys) visible at their real rates. The fee is the same in every posture, and we take no per-token margin. Details on the pricing page.
If you have a platform team that wants to own model serving, GPU scheduling, upgrade gating, egress policy, and 24/7 on-call, that's a legitimate path, and the hardware vendors will happily sell it to you. Most of our customers don't have that team and never will. What they'd have to hire, we operate, and because we run a whole fleet, our operating cost per deployment keeps falling while a single self-hosted cluster pays its full cost alone.
Early, and we'd rather you hear it from us. The governance core, gateway, and console are built and fault-tested, and we're deploying with a small number of design partners now. What early buys you: direct access to the team that builds the product, catalog priorities steered by your needs, and design-partner terms. If you need a vendor with a decade of logos, we're not that yet, and we won't pretend otherwise.
Plain answers, from the people building it.