- Installation
- Documentation
- Getting Started
- Connect
- Data Import and Export
- Overview
- Data Sources
- CSV Files
- JSON Files
- Overview
- Creating JSON
- Loading JSON
- Writing JSON
- JSON Type
- JSON Functions
- Format Settings
- Installing and Loading
- SQL to / from JSON
- Caveats
- Multiple Files
- Parquet Files
- Partitioning
- Appender
- INSERT Statements
- Lakehouse Formats
- Client APIs
- Overview
- ADBC
- C
- Overview
- Startup
- Configuration
- Query
- Data Chunks
- Vectors
- Values
- Types
- Prepared Statements
- Appender
- Table Functions
- Replacement Scans
- API Reference
- C++
- CLI
- Overview
- Arguments
- Dot Commands
- Output Formats
- Editing
- Friendly CLI
- Safe Mode
- Autocomplete
- Syntax Highlighting
- Known Issues
- Go
- Java (JDBC)
- Overview
- Define Connections
- Run Queries
- Import Data
- Handle Results
- Define Functions
- Profile and Monitor
- Troubleshoot
- Node.js (Neo)
- ODBC
- Python
- Overview
- Data Ingestion
- Conversion between DuckDB and Python
- DB API
- Relational API
- Function API
- Types API
- Expression API
- Spark API
- API Reference
- Known Python Issues
- R
- Rust
- Overview
- Connect
- Import Data
- Run Queries
- Handle Results
- Write User Defined Functions
- Profile and Monitor
- Troubleshoot
- Wasm
- Tertiary Clients
- SQL
- Introduction
- Statements
- Overview
- ANALYZE
- ALTER TABLE
- ALTER VIEW
- ATTACH and DETACH
- CALL
- CHECKPOINT
- COMMENT ON
- COPY
- CREATE INDEX
- CREATE MACRO
- CREATE SCHEMA
- CREATE SECRET
- CREATE SEQUENCE
- CREATE TABLE
- CREATE VIEW
- CREATE TYPE
- DELETE
- DESCRIBE
- DROP
- EXPORT and IMPORT DATABASE
- INSERT
- LOAD / INSTALL
- MERGE INTO
- PIVOT
- Profiling
- SELECT
- SET / RESET
- SET VARIABLE
- SHOW and SHOW DATABASES
- SUMMARIZE
- Transaction Management
- UNPIVOT
- UPDATE
- USE
- VACUUM
- Query Syntax
- SELECT
- FROM and JOIN
- WHERE
- GROUP BY
- GROUPING SETS
- HAVING
- ORDER BY
- LIMIT and OFFSET
- SAMPLE
- Unnesting
- WITH
- WINDOW
- QUALIFY
- VALUES
- FILTER
- Set Operations
- Prepared Statements
- Data Types
- Overview
- Array
- Bitstring
- Blob
- Boolean
- Date
- Enum
- Geometry
- Interval
- List
- Literal Types
- Map
- NULL Values
- Numeric
- Struct
- Text
- Time
- Timestamp
- Time Zones
- Union
- Typecasting
- Variant
- Expressions
- Overview
- CASE Expression
- Casting
- Collations
- Comparisons
- IN Operator
- Logical Operators
- Star Expression
- Subqueries
- TRY
- Functions
- Overview
- Aggregate Functions
- Array Functions
- Bitstring Functions
- Blob Functions
- Date Format Functions
- Date Functions
- Date Part Functions
- Enum Functions
- Geometry Functions
- Interval Functions
- Lambda Functions
- List Functions
- Map Functions
- Nested Functions
- Numeric Functions
- Pattern Matching
- Regular Expressions
- Struct Functions
- Text Functions
- Time Functions
- Timestamp Functions
- Timestamp with Time Zone Functions
- Union Functions
- Utility Functions
- Window Functions
- Constraints
- Indexes
- Meta Queries
- DuckDB's SQL Dialect
- Overview
- Indexing
- Friendly SQL
- Keywords and Identifiers
- Order Preservation
- PostgreSQL Compatibility
- SQL Quirks
- PEG Parser
- Samples
- Configuration
- Extensions
- Overview
- Installing Extensions
- Advanced Installation Methods
- Distributing Extensions
- Versioning of Extensions
- Troubleshooting of Extensions
- Core Extensions
- Overview
- AutoComplete
- Avro
- AWS
- Azure
- Delta
- DuckLake
- Encodings
- Excel
- Full Text Search
- httpfs (HTTP and S3)
- Iceberg
- ICU
- inet
- jemalloc
- Lance
- MotherDuck
- MySQL
- ODBC
- Quack
- PostgreSQL
- Spatial
- SQLite
- TPC-DS
- TPC-H
- UI
- Unity Catalog
- Vortex
- VSS
- Quack Remote Protocol
- Guides
- Overview
- Data Viewers
- Database Integration
- File Formats
- Overview
- CSV Import
- CSV Export
- Directly Reading Files
- Directly Reading DuckDB Databases
- Excel Import
- Excel Export
- JSON Import
- JSON Export
- Parquet Import
- Parquet Export
- Querying Parquet Files
- File Access with the file: Protocol
- Meta Queries
- Describe Table
- EXPLAIN: Inspect Query Plans
- EXPLAIN ANALYZE: Profile Queries
- List Tables
- Summarize
- DuckDB Environment
- Network and Cloud Storage
- Overview
- HTTP Parquet Import
- S3 Parquet Import
- S3 Parquet Export
- S3 Iceberg Import
- S3 Express One
- GCS Import
- Cloudflare R2 Import
- DuckDB over HTTPS / S3
- Fastly Object Storage Import
- SeaweedFS Import
- Tigris Import
- ODBC
- Performance
- Overview
- Environment
- Import
- Schema
- Indexing
- Join Operations
- File Formats
- How to Tune Workloads
- My Workload Is Slow
- Out-of-Memory Issues
- Benchmarks
- Working with Huge Databases
- Python
- Installation
- Executing SQL
- Jupyter Notebooks
- marimo Notebooks
- SQL on Pandas
- Import from Pandas
- Export to Pandas
- Import from Numpy
- Export to Numpy
- SQL on Arrow
- Import from Arrow
- Export to Arrow
- Relational API on Pandas
- Multiple Python Threads
- Integration with Ibis
- Integration with Polars
- Integration with PyTorch
- Using fsspec Filesystems
- SQL Editors
- SQL Features
- AsOf Join
- Full-Text Search
- Graph Queries
- query and query_table Functions
- Merge Statement for SCD Type 2
- Timestamp Issues
- Snippets
- Creating Synthetic Data
- Dutch Railway Datasets
- Sharing Macros
- Analyzing a Git Repository
- Importing Duckbox Tables
- Copying an In-Memory Database to a File
- Troubleshooting
- Glossary of Terms
- Browsing Offline
- Operations Manual
- Overview
- DuckDB's Footprint
- Installing DuckDB
- Logging
- User Agents
- Securing DuckDB
- Non-Deterministic Behavior
- Limits
- DuckDB Docker Container
- Development
- DuckDB Repositories
- Release Cycle
- Metrics
- Profiling
- Building DuckDB
- Overview
- Build Configuration
- Building Extensions
- Android
- Linux
- macOS
- Raspberry Pi
- Windows
- Python
- R
- Troubleshooting
- Unofficial and Unsupported Platforms
- Benchmark Suite
- Testing
- Internals
- Sitemap
- Live Demo
Overview
Beyond reading a result set row by row, the Rust client can hand a query result to Apache Arrow as a stream of record batches, register an Arrow batch as a queryable table, and return results as Polars data frames. Each of these columnar result-handling options is described below.
Apache Arrow
DuckDB is columnar, and its native way to return a result to Rust in bulk is Apache Arrow. The crate re-exports the arrow crate, so no separate Arrow dependency or version alignment is required: reach the Arrow types through duckdb::arrow.
Reading a Result as Arrow
Call query_arrow() on a prepared statement to get an iterator of RecordBatch. Collecting it materializes the whole result, while iterating it reads one batch at a time:
use duckdb::{Connection, Result};
use duckdb::arrow::record_batch::RecordBatch;
use duckdb::arrow::util::pretty::print_batches;
let conn = Connection::open_in_memory()?;
let mut stmt = conn.prepare("SELECT * FROM generate_series(1, 5)")?;
let batches: Vec<RecordBatch> = stmt.query_arrow([])?.collect();
print_batches(&batches).unwrap();
get_schema() on the returned handle reports the Arrow schema DuckDB inferred for the result.
Streaming Arrow Results
query_arrow() runs the statement to completion and buffers the whole result on the client side, so a large result is held in memory even while the iterator is advanced one batch at a time. For a result that should be consumed lazily, fetching chunks only as the iterator advances, use stream_arrow(), which is otherwise identical:
let mut stmt = conn.prepare("SELECT * FROM big_table")?;
for batch in stmt.stream_arrow([])? {
// process one RecordBatch at a time
println!("{} rows", batch.num_rows());
}
DuckDB may still materialize the result internally for some statements. The streaming iterator panics if fetching or Arrow conversion fails after execution has started.
Querying an Arrow Batch
An Arrow RecordBatch produced elsewhere in a Rust program can be registered as a DuckDB table function and queried in SQL. Enable the vtab-arrow feature, register the built-in ArrowVTab table function on the connection, and pass the batch as a query parameter with arrow_recordbatch_to_query_params(). The following is adapted from the crate's arrow_vtab example:
use duckdb::{Connection, arrow::record_batch::RecordBatch};
use duckdb::vtab::arrow::{arrow_recordbatch_to_query_params, ArrowVTab};
let conn = Connection::open_in_memory()?;
conn.register_table_function::<ArrowVTab>("arrow")?;
let params = arrow_recordbatch_to_query_params(cities_batch);
let batches: Vec<RecordBatch> = conn
.prepare(
"SELECT city, population
FROM arrow(?, ?)
WHERE coastal AND population >= 500000
ORDER BY population DESC",
)?
.query_arrow(params)?
.collect();
The batch is addressed by the name given to register_table_function() (arrow here), and DuckDB filters, orders, and aggregates it like any other table. arrow_recordbatch_to_query_params() expands a batch into the two parameters the arrow(?, ?) function expects.
Polars Data Frames
With the polars feature enabled, a query result can be returned as Polars DataFrames. Call query_polars() on a prepared statement to get an iterator of data frames, one per result chunk:
use duckdb::{Connection, Result};
use polars::prelude::DataFrame;
let conn = Connection::open_in_memory()?;
let mut stmt = conn.prepare("SELECT * FROM test")?;
let dfs: Vec<DataFrame> = stmt.query_polars([])?.collect();
To combine the chunks into a single DataFrame, use accumulate_dataframes_vertical_unchecked from polars_core. The crate re-exports polars, so its types are also reachable through duckdb::polars.
Further Reading
- Run Queries — sending the queries whose results this page reads, and reading them row by row.
- Write User Defined Functions — writing table functions, of which the built-in
ArrowVTabis one. - Import Data — appending Arrow record batches into a table with the
appender-arrowfeature.