Skip to content
Quick Start - Concrete AI

Quick Start - Concrete AI

Use Concrete AI to run models through secure, OpenAI-compatible inference endpoints. This guide shows how to give your application the access it needs to call a Dedicated Inference deployment or a public model on the on-demand endpoint. AI API Keys are managed with exo ai api-key and are zone-scoped, so pass --zone on every call.

Prerequisites

You need:

First Steps via CLI

The following steps create and manage an AI API Key. This is the credential your application presents when it calls an inference endpoint. First decide which deployments or public models the key may access, then create the key and use its value as a bearer token in your requests.

Scopes

A key holds two independent access lists:

  • models: names of public models served on the on-demand endpoint
  • deployments: IDs of Dedicated Inference deployments in your organization

An empty list denies access on that side, and a key with both lists empty can authenticate but reaches nothing.

Deployments in error or deleted state cannot be added to a key; any other state is accepted, so you can scope a key to a deployment that is still being provisioned. Models must be public models.

Create the Key

Create a key that can access one deployment:

exo ai api-key create my-first-key \
  --deployment <deployment-id> \
  --zone <zone>

The output contains the plaintext key:

ID:     <ai-api-key-id>
Name:   my-first-key
Value:  xYz...

Warning

Store the value somewhere safe right away. It is only shown at creation and cannot be recovered later.

Parameters:

  • NAME: 1 to 50 characters. Letters, digits, spaces and _'()- are accepted, but the name must start and end with a letter or digit.
  • --model: public model names
  • --deployment: deployment UUIDs
  • --all-models: grant access to all public models
  • --all-deployments: grant access to all deployments

To grant access to every public model instead, pass --all-models. To cover every Dedicated Inference deployment, pass --all-deployments. Both flags together give the key access to everything:

exo ai api-key create my-first-key \
  --all-models \
  --all-deployments \
  --zone <zone>

A key restricted to two deployments:

exo ai api-key create app-prod \
  --deployment 32d6938e-c2aa-43a9-8291-87652c91de91,a780ef42-ccde-4a21-b618-acd84f86070a \
  --zone <zone>

Call the Endpoint

Use the key as a bearer token against your deployment URL:

curl -X POST "https://<deployment-id>.inference.<zone>.exoscale-cloud.com/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <key-value>" \
  -d '{
    "model": "swiss-ai/Apertus-v1.5-70B",
    "messages": [{"role": "user", "content": "Write a joke."}],
    "max_tokens": 300
  }'

For a dedicated deployment:

curl "https://<deployment-id>.inference.<zone>.exoscale-cloud.com/v1/models" \
  -H "Authorization: Bearer <key-value>"

For a public model on the shared endpoint, the URL is https://inference.<zone>.exoscale-cloud.com/v1 and model is the public model name:

curl -X POST "https://inference.<zone>.exoscale-cloud.com/v1/chat/completions" \
  -H "Authorization: Bearer <key-value>" \
  -H "Content-Type: application/json" \
  -d '{"model": "<public-model-name>", "messages": [{"role": "user", "content": "Hello"}]}'

A 403 means the key is valid but its scope does not cover the target. A 401 means the key is unknown or revoked. On the on-demand endpoint, 429 means the organization’s consumption quota is exhausted.

Manage Keys

exo ai api-key list --zone <zone>

Returns every key in the organization, including revoked ones still in their 30-day retention window.

exo ai api-key show <key-id> --zone <zone>

Shows the key’s name, scope, and timestamps. The plaintext value is never shown outside the creation output.

exo ai api-key update <key-id> --all-models --zone <zone>

Flags you omit stay unchanged. To cut access to all public models, pass --all-models false. A revoked key cannot be updated.

Clean Up

Revoke the key when you no longer need it:

exo ai api-key revoke <key-id> --zone <zone>

Revocation returns an operation and completes asynchronously. Once propagated, the key gets a 401 on inference endpoints. The key entry stays visible in api-key list with revoked-at set and is deleted after 30 days.

First Steps via Portal

In the Exoscale Portal, select Concrete AI to create API Keys.

Create an AI API Key

Under AI API Keys click Add, add the Name of the API keys, select the inference end point and after clicking Add again the Secret will be shown. Copy the Secret and please be aware that this value is only available at the moment of creation and will not be retrievable after leaving the page.

Next Steps

Last updated on