Config file reference
This page describes the different configuration options for Polars On-Prem. The config file is a
standard TOML file with different sections. Any of the configuration can be overridden using
environment variables in the following format: PC_CUBLET__section_name__key.
Example configuration files can be found at Example Configurations.
See the sidebar for extensive documentation on important components and their configuration together.
Top-level configuration
| Key | Type | Description |
|---|---|---|
cluster_id |
string | Logical ID for the cluster; workers and scheduler that share this ID will form a single cluster. e.g. prod-eu-1; must be unique among all clusters. |
instance_id |
string | Unique ID for this node within the cluster, used for addressing and leader selection. e.g. scheduler, worker_0; must be unique per cluster. |
license |
string | Absolute path to the Polars On-Prem license file required to start the process. Shorthand for license.on_prem_enterprise.license_path.e.g. /etc/polars/license.json. See License. |
memory_limit |
integer | Hard memory budget (in bytes) for all components in this node; enforced via cgroups when delegated. Only a plain integer is accepted here; quantity strings such as 10Gi are not. See Resource limits.e.g. 1073741824 (1 GiB), 10737418240 (10 GiB). |
cpu_reserved |
integer or string | CPU reserved for this node, in fractional cores. Nodes with a higher reserved CPU count are given priority when the scheduler assigns tasks, and the value is used for the resource percentage calculations in the observatory. Accepts a plain number, or a string with a decimal SI (m, k, M, G, …), binary SI (Ki, Mi, Gi, …) or scientific suffix.e.g. 8, "7.5", "500m". |
[license] section
The license key can be specified either as a string (shorthand for an enterprise license file
path, see Top-level configuration) or as one of the sections below. See
License for details.
| Key | Type | Description |
|---|---|---|
license.on_prem_enterprise |
object | Offline licensing using a license file. |
license.on_prem_enterprise.license_path |
path | Absolute path to the Polars On-Prem license file. e.g. /etc/polars/license.json. |
license.on_prem |
object | Online licensing against the Polars control plane. Configure the fields below on the leader node. |
license.on_prem.cert_dir |
path | Directory used to store the licensing certificates on the leader node. e.g. /etc/polars/certs. |
license.on_prem.workspace_id |
string | Workspace ID (UUID) used to authenticate against the control plane. |
license.on_prem.client_id |
string | Client ID used to authenticate against the control plane. |
license.on_prem.client_secret |
string | Client secret used to authenticate against the control plane. |
[scheduler] section
| Key | Type | Description |
|---|---|---|
enabled |
boolean | Whether the scheduler component runs in this process.true for the leader node, false on pure workers. |
allow_local_sinks |
boolean | Whether workers are allowed to write to a shared/local disk visible to the scheduler.false for fully remote/storage-only setups, true if you have a shared filesystem. |
allow_local_scans |
boolean | Whether queries are allowed to read data from the local filesystem (e.g. file:///path/to/data).false for fully remote/storage-only setups, true if you have a shared filesystem. |
deny_anonymous_users |
boolean | Whether to reject queries submitted without a username. Setting this to true requires all queries to be sent with a username; false (the default) allows anonymous queries. See Anonymous Users. |
n_workers |
integer | Expected number of workers in this cluster; scheduler waits for the latter to be online before running queries. e.g. 4. |
default_partitions_per_worker |
integer | Default number of partitions per worker when not specified in the query settings. Defaults to 1. |
anonymous_result_location |
object | Destination for results of queries that do not have an explicit sink. Supports a local mounted path (must be reachable on the exact same path and allow_local_sinks enabled) or an object store on S3, Google Cloud Storage, or Azure Blob Storage. All options must be network reachable by scheduler, workers, and client.e.g. /mnt/storage/polars/results. See Anonymous Results.e.g. s3://bucket/path/to/key |
anonymous_result_location.local |
object | Object used for local disk-backed anonymous results. |
anonymous_result_location.local.path |
path | Local path where anonymous results are stored. e.g. /mnt/storage/polars/results. |
anonymous_result_location.s3 |
object | Object used for S3-backed anonymous results. |
anonymous_result_location.s3.url |
string | S3 bucket url. e.g. s3://bucket/path/to/key. |
anonymous_result_location.s3.aws_endpoint_url |
string | Storage option configuration, see scan_parquet(). |
anonymous_result_location.s3.aws_region |
string | Storage option configuration. e.g. eu-east-1 |
anonymous_result_location.s3.aws_access_key_id |
string | Storage option configuration. |
anonymous_result_location.s3.aws_secret_access_key |
string | Storage option configuration. |
anonymous_result_location.s3.external_endpoint |
string | Endpoint to use when generating the presigned URLs handed out to clients. Set this when clients reach the bucket at a different address than the cluster does; it takes priority over aws_endpoint_url for presigning only.e.g. https://minio.example.com:9000. |
anonymous_result_location.s3.presign_duration |
string | How long the presigned URLs generated for results stay valid. Either an ISO 8601 duration format or a jiff friendly duration format (see https://docs.rs/jiff/0.2.18/jiff/fmt/friendly/). Defaults to 8h.e.g. 1 hour.e.g. PT1H. |
anonymous_result_location.s3.allow_deletes |
boolean | Whether clients may delete anonymous results from this location. Defaults to true for S3. |
anonymous_result_location.gcs |
object | Object used for Google Cloud Storage-backed anonymous results. |
anonymous_result_location.gcs.url |
string | GCS bucket url. e.g. gs://bucket/path/to/key. |
anonymous_result_location.gcs.google_service_account_path |
string | Storage option configuration, see scan_parquet().e.g. /etc/polars/gcs-service-account.json. |
anonymous_result_location.gcs.presign_duration |
string | How long the presigned URLs generated for results stay valid. Same formats as the S3 variant above. Defaults to 8h. |
anonymous_result_location.gcs.allow_deletes |
boolean | Whether clients may delete anonymous results from this location. Defaults to false. |
anonymous_result_location.abs |
object | Object used for Azure Blob Storage-backed anonymous results. |
anonymous_result_location.abs.url |
string | Azure container url. e.g. az://container/path/to/key. |
anonymous_result_location.abs.azure_storage_account_name |
string | Storage option configuration, see scan_parquet(). |
anonymous_result_location.abs.azure_storage_account_key |
string | Storage option configuration. |
anonymous_result_location.abs.presign_duration |
string | How long the presigned URLs generated for results stay valid. Same formats as the S3 variant above. Defaults to 8h. |
anonymous_result_location.abs.allow_deletes |
boolean | Whether clients may delete anonymous results from this location. Defaults to false. |
client_service |
object | Object used for configuring the bind address of the client service. This is the service used by the polars-cloud Python client. Defaults to 0.0.0.0:5051. |
client_service.bind_addr |
string | Bind address for the client service. e.g. 0.0.0.0:5051. |
client_service.bind_addr.ip |
string | IP address for the client service bind address. e.g. 192.168.1.1. |
client_service.bind_addr.port |
integer | Port for the client service bind address. e.g. 5051. |
client_service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-1. |
worker_service |
object | Object used for configuring the bind address of the worker service. This is an internal service used by the workers. Defaults to 0.0.0.0:5050. |
worker_service.bind_addr |
string | Bind address for the worker service. e.g. 0.0.0.0:5050. |
worker_service.bind_addr.ip |
string | IP address for the worker service bind address. e.g. 192.168.1.1. |
worker_service.bind_addr.port |
integer | Port for the worker service bind address. e.g. 5050. |
worker_service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
checkpoint |
object | Enable checkpointing for queries. This requires that the worker has checkpoint_location configured with S3 storage. See worker.checkpoint_location. |
checkpoint.enabled |
boolean | Whether checkpointing runs on this cluster. Must be set to true to enable checkpointing. |
checkpoint.period |
string | Period at which checkpoints will be created. Once the period has passed after a stage has completed, a checkpoint is created. Either an ISO 8601 duration format or a jiff friendly duration format (see https://docs.rs/jiff/0.2.18/jiff/fmt/friendly/) e.g. 5 secs.e.g. PT5S. |
[worker] section
| Key | Type | Description |
|---|---|---|
enabled |
boolean | Whether the worker component runs in this process.true on worker nodes, false on the dedicated scheduler. |
heartbeat_period |
string | Interval for worker heartbeats towards the scheduler, used for liveness and load reporting. Either an ISO 8601 duration format or a jiff friendly duration format (see https://docs.rs/jiff/0.2.18/jiff/fmt/friendly/) e.g. 5 secs.e.g. PT5S. |
shuffle_location |
object | Object used for shuffle data storage. Supports worker-local disk, a shared filesystem, or an object store on S3, Google Cloud Storage, or Azure Blob Storage. If omitted, a temporary directory is created for this process and discarded on shutdown, so set this explicitly for anything but a throwaway cluster. See Shuffle data. |
shuffle_location.local |
object | Object used for local disk-backed shuffle data storage. |
shuffle_location.local.path |
path | Local path where shuffle/intermediate data is stored; fast local SSD is recommended. e.g. /mnt/storage/polars/shuffle. |
shuffle_location.shared_filesystem |
object | Object used for shared filesystem-backed shuffle data storage. |
shuffle_location.shared_filesystem.path |
path | Shared filesystem path where shuffle/intermediate data is stored. Must be accessible by all workers on the same path. e.g. /mnt/storage/polars/shuffle. |
shuffle_location.s3 |
object | Object used for S3-backed shuffle data storage. |
shuffle_location.s3.url |
path | Destination for shuffle/intermediate data. e.g. s3://bucket/path/to/key. |
shuffle_location.s3.aws_endpoint_url |
string | Storage option configuration, see scan_parquet(). |
shuffle_location.s3.aws_region |
string | Storage option configuration. e.g. eu-east-1 |
shuffle_location.s3.aws_access_key_id |
string | Storage option configuration. |
shuffle_location.s3.aws_secret_access_key |
string | Storage option configuration. |
shuffle_location.gcs |
object | Object used for Google Cloud Storage-backed shuffle data storage. |
shuffle_location.gcs.url |
path | Destination for shuffle/intermediate data. e.g. gs://bucket/path/to/key. |
shuffle_location.gcs.google_service_account_path |
string | Storage option configuration, see scan_parquet().e.g. /etc/polars/gcs-service-account.json. |
shuffle_location.abs |
object | Object used for Azure Blob Storage-backed shuffle data storage. |
shuffle_location.abs.url |
path | Destination for shuffle/intermediate data. e.g. az://container/path/to/key. |
shuffle_location.abs.azure_storage_account_name |
string | Storage option configuration, see scan_parquet(). |
shuffle_location.abs.azure_storage_account_key |
string | Storage option configuration. |
checkpoint_location |
object | Location where checkpoints for queries are stored. Required when the scheduler has [scheduler.checkpoint] enabled. Supports a shared filesystem or an object store on S3, Google Cloud Storage, or Azure Blob Storage. Unlike shuffle_location there is no local variant, since checkpoints must survive the loss of a single worker. See Checkpointing. |
checkpoint_location.shared_filesystem.path |
path | Shared filesystem path where checkpoint data is stored. Must be accessible by all workers on the same path. e.g. /mnt/storage/polars/checkpoints. |
checkpoint_location.s3 |
object | Object used for S3-backed checkpoint storage. |
checkpoint_location.s3.url |
path | Destination for checkpoint data. e.g. s3://bucket/path/to/key. |
checkpoint_location.s3.aws_endpoint_url |
string | Storage option configuration, see scan_parquet(). |
checkpoint_location.s3.aws_region |
string | Storage option configuration. e.g. eu-east-1 |
checkpoint_location.s3.aws_access_key_id |
string | Storage option configuration. |
checkpoint_location.s3.aws_secret_access_key |
string | Storage option configuration. |
checkpoint_location.gcs |
object | Object used for Google Cloud Storage-backed checkpoint storage. |
checkpoint_location.gcs.url |
path | Destination for checkpoint data. e.g. gs://bucket/path/to/key. |
checkpoint_location.gcs.google_service_account_path |
string | Storage option configuration, see scan_parquet().e.g. /etc/polars/gcs-service-account.json. |
checkpoint_location.abs |
object | Object used for Azure Blob Storage-backed checkpoint storage. |
checkpoint_location.abs.url |
path | Destination for checkpoint data. e.g. az://container/path/to/key. |
checkpoint_location.abs.azure_storage_account_name |
string | Storage option configuration, see scan_parquet(). |
checkpoint_location.abs.azure_storage_account_key |
string | Storage option configuration. |
task_service |
object | Object used for configuring the bind address of the task service. This is an internal service in the worker for receiving tasks from the scheduler. Defaults to 0.0.0.0:5052. |
task_service.bind_addr |
string | Bind address for the task service. e.g. 0.0.0.0:5052. |
task_service.bind_addr.ip |
string | IP address for the task service bind address. e.g. 192.168.1.1. |
task_service.bind_addr.port |
integer | Port for the task service bind address. e.g. 5052. |
task_service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
task_service.public_addr |
string | Address at which this service is reachable by the scheduler. Defaults to the bind address if not set. This field is required when the bind address is 0.0.0.0.e.g. 192.168.1.1. |
task_service.public_addr.ip |
string | IP address for the task service public address. e.g. 192.168.1.2. |
task_service.public_addr.port |
integer | Port for the task service public address. e.g. 5052. |
task_service.public_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
shuffle_service |
object | Object used for configuring the bind address of the shuffle service. This is an internal service in the worker for serving shuffle data to the other workers. Defaults to 0.0.0.0:5053. |
shuffle_service.bind_addr |
string | Bind address for the shuffle service. e.g. 0.0.0.0:5053. |
shuffle_service.bind_addr.ip |
string | IP address for the shuffle service bind address. e.g. 192.168.1.1. |
shuffle_service.bind_addr.port |
integer | Port for the shuffle service bind address. e.g. 5053. |
shuffle_service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
shuffle_service.public_addr |
string | Address at which this service is reachable by the other workers. Defaults to the bind address if not set. This field is required when the bind address is 0.0.0.0.e.g. 192.168.1.1. |
shuffle_service.public_addr.ip |
string | IP address for the shuffle service public address. e.g. 192.168.1.2. |
shuffle_service.public_addr.port |
integer | Port for the shuffle service public address. e.g. 5053. |
shuffle_service.public_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
env_vars |
object | Environment variable overrides used by Polars or user code, as a table of key/value string pairs. e.g. { POLARS_MAX_THREADS = "8" }. |
extras.hdfs.enabled |
boolean | Enable HDFS support. |
extras.pyiceberg.enabled |
boolean | Enable PyIceberg support. |
[observatory] section
| Key | Type | Description |
|---|---|---|
enabled |
boolean | Enable sending/receiving profiling data so clients can call result.await_profile().true on both scheduler and workers if you want profiles on queries; false to disable. |
max_metrics_bytes_total |
integer | How many bytes all the worker host metrics will consume in total. The full amount is allocated up front to keep memory usage predictable, and the amount available per node is max_metrics_bytes_total / n_workers. If a system-wide memory limit is specified then this is added to the share that the scheduler takes. For every worker, about 50 bytes of metrics are stored per second. |
database_path |
string | Location to use for storing profiling data. An SQLite database file will be created here, or if a file already exists it will be opened. If this points to a directory, a file in that directory will be created. Polars On-Prem will automatically add the cluster_id to this file name to ensure uniqueness within the directory. |
service |
object | Object used for configuring the bind address of the observatory service. This is an internal service in the scheduler for receiving profiling data from all nodes. Defaults to 0.0.0.0:5049. |
service.bind_addr |
string | Bind address for the observatory service. e.g. 0.0.0.0:5049. |
service.bind_addr.ip |
string | IP address for the observatory service bind address. e.g. 192.168.1.1. |
service.bind_addr.port |
integer | Port for the observatory service bind address. e.g. 5049. |
service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
rest_api.enabled |
boolean | By default enabled for exposing the observatory REST API. This is a public service for accessing the profiling data and host metrics data through a web interface. |
rest_api.service |
object | Object used for configuring the bind address of the observatory REST API service. Defaults to 0.0.0.0:3001. |
rest_api.service.bind_addr |
string | Bind address for the observatory REST API service. e.g. 0.0.0.0:3001. |
rest_api.service.bind_addr.ip |
string | IP address for the observatory REST API service bind address. e.g. 192.168.1.1. |
rest_api.service.bind_addr.port |
integer | Port for the observatory REST API service bind address. e.g. 3001. |
rest_api.service.bind_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
[monitoring] section
| Key | Type | Description |
|---|---|---|
enabled |
boolean | Enable sending/receiving monitoring data to the observatory service. If enabled, it will use the address specified in observatory_service.public_addr. |
host_metrics |
object | Object used for configuring the host metrics exporter. |
host_metrics.enabled |
boolean | Enable/disable exporting host metrics from this node |
host_metrics.disk_usage_metrics |
boolean | Whether to export disk usage metrics for the shuffle directory. |
host_metrics.disk_io_metrics |
boolean | Whether to export disk read/write throughput metrics for the shuffle directory (enabled by default if IO accounting is enabled for the used cgroup). |
[static_leader] section
| Key | Type | Description |
|---|---|---|
leader_instance_id |
string | ID of the leader node; should match the scheduler’s instance_id.Typically scheduler to match your scheduler node. |
scheduler_service.public_addr |
string | Address at which the scheduler client service is reachable from this node. e.g. 192.168.1.1. |
scheduler_service.public_addr.ip |
string | IP address for the scheduler client service public address. e.g. 192.168.1.1. |
scheduler_service.public_addr.port |
integer | Port for the scheduler client service public address. e.g. 5051. |
scheduler_service.public_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
observatory_service.public_addr |
string | Address at which the observatory service is reachable from this node. e.g. 192.168.1.1. |
observatory_service.public_addr.ip |
string | IP address for the observatory service public address. e.g. 192.168.1.1. |
observatory_service.public_addr.port |
integer | Port for the observatory service public address. e.g. 5049. |
observatory_service.public_addr.hostname |
string | Alternative to ip, resolved once at startup.e.g. my-host-2. |
[lineage] section
| Key | Type | Description |
|---|---|---|
enabled |
bool | Enable the OpenLineage service. |
transport.http.endpoint |
string | HTTP or HTTPS endpoint of the OpenLineage collector. e.g. https://lineage_collector:5000. |