⌘K ctrl+k
Search Shortcut cmd + k | ctrl + k

Warning DuckDB ships with safe defaults. The configuration options described on this page are advanced options, so proceed with caution when changing them.

Starting with v2.0, DuckDB supports asynchronous I/O. Instead of blocking a worker thread until requested data arrives, DuckDB issues reads in the background and keeps multiple requests in flight while worker threads process data that has already arrived. This can significantly speed up queries when synchronous I/O does not saturate the available bandwidth, e.g., when reading data from object storage such as S3.

For a detailed explanation of the design and benchmark results, see the “Asynchronous I/O in DuckDB” blog post.

Supported Formats

Asynchronous I/O is currently supported for the following formats:

Other formats, such as JSON files and DuckDB's native database format, use synchronous I/O. Asynchronous I/O is used automatically for the supported formats, including when they are read through data lake formats such as DuckLake.

Thread Pools

DuckDB uses two thread pools:

  • The regular pool contains the worker threads, which perform the actual query processing (e.g., decoding, joins and aggregations). Its size is controlled by the threads setting and defaults to the number of CPU cores. Regular threads prioritize query processing but can also perform I/O tasks when idle.
  • The asynchronous pool contains threads dedicated to blocking I/O. As these threads spend most of their time waiting for responses (e.g., HTTP requests), there are more of them than CPU cores: by default, four times the number of system threads, capped at 256. Its size is controlled by the async_threads setting.

For example, to set the number of asynchronous I/O threads to 48, run:

SET async_threads = 48;

Warning The asynchronous pool increases the total number of threads per DuckDB instance. If you run many DuckDB instances in a single process, consider setting threads and async_threads explicitly.

Read-Ahead

To keep the asynchronous threads busy, DuckDB reads ahead: it schedules the reads for upcoming scan jobs (e.g., row groups in Parquet files or byte ranges in CSV files) before the worker threads need them. While a worker thread processes the current job, the asynchronous threads fetch the data for the next jobs. If the data for a job has not arrived yet, the worker thread is free to run other tasks in the meantime.

Keeping more fetch tasks in flight consumes more memory. To determine a budget and avoid out-of-memory issues, DuckDB provides the read_ahead_depth configuration option. It can have three types of values:

  • -1 (default): unlimited depth, bounded by memory.
  • N > 0: at most N jobs ahead, with no memory budget.
  • 0: read-ahead is off, each scan task schedules I/O only for its own job.

To configure it, use the SET clause, e.g.:

SET read_ahead_depth = 5;

In the default mode, the read-ahead budget is negotiated with the same memory manager that distributes memory between operators such as joins, sorts and window functions. Under high memory pressure, the read-ahead queue shrinks to a single job and the scan behaves similarly to a synchronous scan. Once memory frees up, the queue fills up again. To limit the total memory used by DuckDB, use the memory_limit setting.

Tuning

The default settings work well in most cases. On machines with high network bandwidth, you can further increase throughput by fixing the read-ahead depth and adjusting the number of asynchronous threads and the HTTP retry settings. For example, the following configuration saturated a 25 Gbit/s network on a 64-core EC2 instance reading Parquet files from S3:

SET read_ahead_depth = 64;
SET async_threads = 48;
SET http_retries = 8;
SET http_retry_wait_ms = 50;
SET http_retry_backoff = 2;

As the row group is the unit of parallelism for Parquet scans, asynchronous I/O works best if Parquet files have at least as many row groups as the number of threads. Files with only a few very large row groups cannot keep enough requests in flight to saturate the network.

Synchronizing Writes to Disk

When DuckDB persists changes to a database on the local file system, it asks the operating system to flush the written data to stable storage. The fsync_mode configuration option controls how this is done. It can have three values:

  • STANDARD (default): uses the regular sync call of the platform, i.e., fdatasync or fsync on Unix-like systems and FlushFileBuffers on Windows.
  • FULL: on macOS, uses F_FULLFSYNC, which also instructs the drive to flush its write cache and thus guarantees durability in case of a power failure. If the file system does not support F_FULLFSYNC, DuckDB falls back to the STANDARD behavior. On other platforms, FULL is equivalent to STANDARD.
  • NONE: skips the sync call. Writes are handed over to the operating system, which flushes them to disk at its own pace.

For example, to disable syncing for a performance-critical workload, run:

SET fsync_mode = 'NONE';

The option is global, i.e., it applies to the whole DuckDB instance.

Warning Setting fsync_mode to NONE may cause data loss or database corruption if the operating system crashes or the machine loses power. Only use it if the database can be recreated from other sources.

© 2026 DuckDB Foundation, Amsterdam NL
DuckDB Home Code of Conduct Trademark Use Blog