Overview - Concrete AI
Concrete AI is Exoscale’s AI infrastructure suite for building and operating production AI workloads on sovereign European infrastructure. It brings together high-performance GPUs for workloads you operate yourself, Dedicated Inference for models served on dedicated GPUs, On-Demand Inference for pay-as-you-go public models, and managed vector databases for retrieval-augmented generation and semantic search.
Architecture
Concrete AI is organised in layers, allowing you to choose between operating the infrastructure yourself and using managed AI services. At the infrastructure layer, GPU instances provide dedicated NVIDIA GPUs for workloads where you manage the model, runtime, and application. At the managed inference layer, Dedicated Inference turns a model into an OpenAI-compatible endpoint backed by dedicated GPU resources. On-demand inference provides a shared endpoint for public models.
For Dedicated Inference, a model is imported and stored securely in Exoscale Object Storage before it is deployed. A deployment runs one or more replicas of that model on dedicated GPUs and exposes its own endpoint. Increasing the replica count adds capacity for concurrent requests; the endpoint URL remains the same.
Applications send requests to either a dedicated deployment endpoint or the shared on-demand endpoint. AI API Keys authenticate those requests and define which deployments and public models the calling application may access. On each request, the inference router checks the bearer token against its key set, verifies the key’s scope covers the target model or deployment, and applies the organization’s consumption quota for public models. Requests to dedicated deployments are not metered.
Vector databases can sit alongside GPU Servers or the inference layer when an application needs retrieval-augmented generation or semantic search. The application retrieves relevant data from its managed PostgreSQL or OpenSearch vector database, then includes that context in its model request.
Terminology
Understanding the key concepts used in Concrete AI will help you deploy models efficiently.
- AI API Key
- An AI API Key is a bearer credential that authenticates inference requests on Exoscale AI endpoints. One key can cover several Dedicated Inference deployments, several public models on the shared on-demand endpoint, or both. The plaintext value is returned once, at creation.
- API Key Scope
- The set of resources a key can access, defined by two lists:
modelsfor public models on the on-demand endpoint anddeploymentsfor Dedicated Inference deployments. - Revoked API Key
- A key that no longer authenticates requests. Revoked keys stay listed for 30 days, then are deleted.
- Deployment
- A deployment represents a running inference endpoint backed by dedicated GPU resources. Each deployment exposes an OpenAI-compatible API endpoint and runs on your own GPUs at
<deployment-id>.inference.<zone>.exoscale-cloud.com. - Model
- A model is an AI artifact (for example from Hugging Face) that is imported into Exoscale and stored securely in Exoscale Object Storage for deployment.
- Replica
- A replica is a running copy of a deployment used to scale inference capacity and handle concurrent requests.
- GPU Count
- Defines how many GPUs are assigned to a single model instance, enabling large models to run across multiple GPUs.
- Dedicated GPU
- A physical GPU allocated exclusively to a deployment rather than shared with other customers. This helps keep performance more predictable because a deployment is not competing with unrelated workloads for the same GPU.
Features
- AI API Key Scoped Access
- Each key carries a
modelslist and adeploymentslist. Use--all-modelsor--all-deploymentsto grant access to every public model or every deployment in the organization, or pass explicit names and IDs. An empty list grants no access on that side. - AI API Key Shown Once
- The plaintext Secret for the key is returned only in the
create-ai-api-keyresponse, in thevaluefield. It cannot be retrieved later. - Dedicated GPU Performance
- Each deployment runs on dedicated NVIDIA GPUs, ensuring consistent performance without resource contention.
- Bring Your Own Model
- Import public, gated, or private Hugging Face models, including models owned by your Hugging Face organization.
- OpenAI-Compatible API
- Integrate deployed models directly into applications using a standard OpenAI-compatible API
- Flexible Scaling
- Scale deployments by changing replicas. Scale to zero to stop GPU billing while keeping the endpoint URL. AI API Keys are managed independently. The endpoint will not answer requests while scaled to zero, but you can scale it back up later.
- Sovereign & Secure
- All deployments run in European data centers with strong isolation between organizations, supporting GDPR and data sovereignty requirements.
- Transparent Pricing
- Billing is based on per-second GPU usage and standard Object Storage costs, with no token-based pricing.