Run federated queries in Managed ClickHouse®
With federated queries in Managed ClickHouse®, you can read and pull data from S3-compatible object storage, Azure Blob Storage, any web resource accessible over HTTP, and external databases such as PostgreSQL and MySQL.
Learn more about capabilities and applications of federated queries in About querying external data in Managed ClickHouse®.
About running federated queries
Federated queries are written using specific SQL statements and can be run from CLI, for instance. To run a federated query, just send a query over an external S3-compatible object storage including relevant S3 bucket details. A properly constructed federated query returns a specific output.
Prerequisites
The prerequisites depend on the table function or table engine used in your federated query.
Access to S3 and URL sources
To run a federated query, the ClickHouse service user connecting to the cluster requires grants to the S3 and/or URL sources. The main service user is granted access to the sources by default, and new users can be allowed to use the sources with the following query:
GRANT CREATE TEMPORARY TABLE, S3, URL ON *.* TO usernameThe CREATE TEMPORARY TABLE grant is required for both sources. Append WITH GRANT OPTION
to let the user pass the privileges on to others:
GRANT CREATE TEMPORARY TABLE, S3, URL ON *.* TO username WITH GRANT OPTIONAzure Blob Storage access keys
To run federated queries using the azureBlobStorage table function or the
AzureBlobStorage table engine, get your Azure Blob Storage keys using one of the
following tools:
From the portal menu, select Storage accounts, go to your account, and click Security + Networking > Access keys. View and copy your account access keys and connection strings.
Supported sources
Managed ClickHouse can query S3-compatible object storage, Azure Blob Storage, HTTP
resources through the url() table function, and other databases such as PostgreSQL and
MySQL. Both table functions and persistent tables backed by the matching table engine
(S3, URL, AzureBlobStorage, PostgreSQL, MySQL) are available.
Note
Reading from an external source also requires the matching privilege, for example
GRANT S3 or GRANT URL, in addition to CREATE TEMPORARY TABLE. See
Manage users and roles.
Run a federated query
See some examples of running federated queries to read and pull data from external sources.
Query using the azureBlobStorage table function
Pass the storage account connection parameters explicitly in the query. Before you start, fulfill relevant prerequisites, if any.
SELECT and azureBlobStorage
SELECT *
FROM azureBlobStorage(
'DefaultEndpointsProtocol=https;AccountName=ident;AccountKey=secret;EndpointSuffix=core.windows.net',
'ownerresource',
'all_stock_data.csv',
'CSV',
'auto',
'Ticker String, Low Float64, High Float64'
)
LIMIT 5INSERT and azureBlobStorage
INSERT INTO FUNCTION
azureBlobStorage(
'DefaultEndpointsProtocol=https;AccountName=ident;AccountKey=secret;EndpointSuffix=core.windows.net',
'ownerresource',
'test_funcwrite.csv',
'CSV',
'auto',
'key UInt64, data String'
)
VALUES (1, 'column2-value');Query an external database
Query a PostgreSQL database, such as Managed PostgreSQL,
directly with the postgresql() table function. The user needs the POSTGRES privilege in
addition to CREATE TEMPORARY TABLE:
SELECT *
FROM postgresql(
'PG_HOST:PG_PORT',
'defaultdb',
'customers',
'PG_USER',
'PG_PASSWORD'
)
LIMIT 5The mysql() table function works the same way for MySQL sources, with the MYSQL
privilege. For dimensions you join constantly, create a persistent table with the
PostgreSQL or MySQL table engine instead, so the connection details live in one place.
Query using the s3 table function
Before you start, fulfill relevant prerequisites, if any.
SELECT and s3
SQL SELECT statements using the S3 and URL functions are able to query public resources using the URL of the resource. For instance, let’s explore the network connectivity measurement data provided by the Open Observatory of Network Interference (OONI).
WITH ooni_data_sample AS
(
SELECT *
FROM s3('https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/clickhouse_export/csv/fastpath_202308.csv.zstd')
LIMIT 100000
)
SELECT
probe_cc AS probe_country_code,
test_name,
countIf(anomaly = 't') AS total_anomalies
FROM ooni_data_sample
GROUP BY
probe_country_code,
test_name
HAVING total_anomalies > 10
ORDER BY total_anomalies DESC
LIMIT 50INSERT and s3
When executing an INSERT statement into the S3 function, the rows are appended to the corresponding object if the table structure matches:
INSERT INTO FUNCTION
s3('https://bucket-name.s3.region-name.amazonaws.com/dataset-name/landing/raw-data.csv', 'CSVWithNames')
VALUES (1, 'column2-value');Query a private S3 bucket
Before you start, fulfill relevant prerequisites, if any.
Private buckets can be accessed by providing the access token and secret as function parameters.
SELECT *
FROM s3(
'https://private-bucket.s3.eu-west-3.amazonaws.com/dataset-prefix/partition-name.csv',
'some_aws_access_key_id',
'some_aws_secret_access_key'
)Depending on the format, the schema can be automatically detected. If it isn’t, you may also provide the column types as function parameters.
SELECT *
FROM s3(
'https://private-bucket.s3.eu-west-3.amazonaws.com/orders-dataset/partition-name.csv',
'access_token',
'secret_token',
'CSVWithNames',
'order_id UInt64, quantity Decimal(18, 9), order_datetime DateTime'
)Query using the s3Cluster table function
Before you start, fulfill relevant prerequisites, if any.
The s3Cluster function allows all cluster nodes to participate in the
query execution. Using default for the cluster name parameter, we can
compute the same aggregations as above as follows:
WITH ooni_clustered_data_sample AS
(
SELECT *
FROM s3Cluster('default', 'https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/clickhouse_export/csv/fastpath_202308.csv.zstd')
LIMIT 100000
)
SELECT
probe_cc AS probe_country_code,
test_name,
countIf(anomaly = 't') AS total_anomalies
FROM ooni_clustered_data_sample
GROUP BY
probe_country_code,
test_name
HAVING total_anomalies > 10
ORDER BY total_anomalies DESC
LIMIT 50Query using the url table function
Before you start, fulfill relevant prerequisites, if any.
SELECT and url
Let’s query the Growth Projections and Complexity Rankings dataset, courtesy of the Atlas of Economic Complexity project.
WITH economic_complexity_ranking AS
(
SELECT *
FROM url('https://dataverse.harvard.edu/api/access/datafile/7259657?format=tab', 'TSV')
)
SELECT
replace(code, '"', '') AS `ISO country code`,
growth_proj AS `Forecasted annualized rate of growth`,
toInt32(replace(sitc_eci_rank, '"', '')) AS `Economic Complexity Index ranking`
FROM economic_complexity_ranking
WHERE year = 2021
ORDER BY `Economic Complexity Index ranking` ASC
LIMIT 20INSERT and url
With the URL function, INSERT statements generate a POST request, which
can be used to interact with APIs having public endpoints. For instance,
if your application has a ingest-csv endpoint accepting CSV data, you
can insert a row using the following statement:
INSERT INTO FUNCTION
url('https://app-name.company-name.cloud/api/ingest-csv', 'CSVWithNames')
VALUES (1, 'column2-value');Query a virtual table
Before you start, fulfill relevant prerequisites, if any.
Instead of specifying the URL of the resource in every query, it’s possible to create a virtual table using the URL table engine. This can be achieved by running a DDL CREATE statement similar to the following:
CREATE TABLE trips_export_endpoint_table
(
`trip_id` UInt32,
`vendor_id` UInt32,
`pickup_datetime` DateTime,
`dropoff_datetime` DateTime,
`trip_distance` Float64,
`fare_amount` Float32
)
ENGINE = URL('https://app-name.company-name.cloud/api/trip-csv-export', CSV)Once the table is defined, SELECT and INSERT statements execute GET and POST requests to the URL respectively:
SELECT
toDate(pickup_datetime) AS pickup_date,
median(fare_amount) AS median_fare_amount,
max(fare_amount) AS max_fare_amount
FROM trips_export_endpoint_table
GROUP BY pickup_dateINSERT INTO trips_export_endpoint_table
VALUES (8765, 10, now() - INTERVAL 15 MINUTE, now(), 50, 20)