Call a deployment
Open Admin → GPUs for your organization and choose How to call endpoint on a running vLLM deployment. The panel gives you the deployment-specific URL and a copyable request.stream: true. The model field is optional. If supplied, AI
Reserve pins it to the repository actually deployed, so a request cannot use
one deployment URL to target another model. On streaming requests AI Reserve
sets stream_options.include_usage on your behalf (overriding an explicit
false) so the final chunk always carries token counts for usage attribution.
The path contains the deployment’s internal ID. Copy the complete URL from the GPUs page rather than constructing it
from the HuggingFace endpoint name.
Authentication and access
- Send an active AI Reserve API key belonging to the same organization as the deployment.
- User-scoped and team-scoped keys can call their organization’s deployment; normal key status, IP allowlist, organization IP blocklist, geographic, and API rate-limit controls still apply.
- Spend-based gates do not apply to this route. GPU capacity is prepaid hourly from the organization wallet and inference requests add no per-token spend, so the checks that meter token spend — organization wallet balance, per-user and per-team budget caps, and the Cognition Token balance gate — are intentionally not evaluated here. A principal over its budget cap can still call a running deployment. To restrict who can drive a deployment, pause or revoke keys, scope the key’s IP allowlist, or use the organization IP blocklist; to stop the spending itself, pause the deployment.
- A key from another organization receives the same
404as an unknown deployment. The API never confirms that another organization’s deployment exists. - Paused, failed, deleted, provisioning, externally drifted, and legacy HuggingFace-toolkit deployments cannot serve this OpenAI-compatible route.
HF_TOKEN, raw HuggingFace URL, and deployment
management APIs are never exposed to the caller.
Billing and usage
GPU capacity is prepaid hourly when the deployment starts and while it remains running. Inference requests are not charged by token on top of that hourly rate. Requests still record model, API key/user/team attribution, token counts when the vLLM response supplies them, latency, and success/error status at$0 request cost.
Running endpoints bill continuously, including idle time. Pause a deployment
to stop future hourly charges; the unused remainder of the current prepaid
hour is not refunded.
Supported protocol
The managed gateway currently promises:POST /v1/gpu/<deployment-id>/chat/completions- OpenAI-compatible JSON responses
- OpenAI-compatible SSE streaming
- vLLM text-generation deployments only
/v1/models, completions, embeddings, and generic
HuggingFace-toolkit task routes are not part of this managed customer API.