Search Shortcut cmd + k | ctrl + k
panduck

Read documents natively into the duck_block vocabulary – DOCX, ODT, EPUB, RTF, LaTeX, Org, RST, ipynb, MediaWiki and Textile – and write them back as Pandoc JSON, without linking Pandoc

Maintainer(s): teaguesterling

Installing and Loading

INSTALL panduck FROM community;
LOAD panduck;

Example

LOAD panduck;

-- Read a document straight into flattened duck_block rows
SELECT * FROM read_docx_blocks('report.docx');
SELECT * FROM read_epub_blocks('book.epub');
SELECT * FROM read_latex_blocks('paper.tex');

-- Or let the path pick the reader
SELECT * FROM panduck_read_blocks('notes.org');
SELECT panduck_format_for('paper.tex');       -- 'latex'
SELECT panduck_can_read('archive.docx');      -- true

-- Formats with no file: parse a string directly
SELECT * FROM read_mediawiki_blocks_string('== Heading ==');

-- Outline any supported document
SELECT * FROM doc_toc('report.docx');

-- Pull one section out, by heading or by a fragment of one
SELECT * FROM doc_section('report.docx', 'Methods');
SELECT * FROM doc_section('report.docx', 'Meth', match := 'contains');

-- ...or find the section by what it SAYS rather than what it is called
SELECT * FROM doc_search_sections('report.docx', 'p < 0.05');

-- Render one back out (md, html or text)
SELECT doc_render('report.docx', 'md');

-- Write back out as Pandoc JSON that a real pandoc accepts
SELECT panduck_write_pandoc_ast('doc.json',
    (SELECT list(b) FROM read_odt_blocks('doc.odt') b)
);

-- Or get the AST as a struct without touching the filesystem
SELECT panduck_blocks_to_pandoc_ast(
    (SELECT list(b) FROM read_rtf_blocks('memo.rtf') b)
);

About panduck

Panduck reads document formats natively into the duck_block vocabulary shared with duck_block_utils, markdown and webbed – so a DOCX, an EPUB and a LaTeX paper all land in the same seven-column shape and can be queried together.

It does not link Pandoc. Pandoc is a Haskell program; the only C bindings that ever existed went unmaintained in 2017, and a Cabal foreign-library request has been open upstream since 2020. Linking the GHC runtime into a dlopen'd DuckDB extension would be a bad neighbour even if it existed. Panduck is instead compatible with Pandoc's data model: every reader is hand-written, and a checked mapping table pins the correspondence to Pandoc's AST constructors.

Readers

Function Format
read_docx_blocks(path) Word OOXML
read_odt_blocks(path) OpenDocument Text
read_epub_blocks(path) EPUB 2/3
read_rtf_blocks(path) Rich Text Format
read_latex_blocks(path) LaTeX
read_org_blocks(path) Org mode
read_rst_blocks(path) reStructuredText
read_ipynb_blocks(path) Jupyter notebooks
read_mediawiki_blocks(path) MediaWiki wikitext
read_textile_blocks(path) Textile
read_pandoc_blocks(path) Pandoc JSON AST (reaches all 43 formats pandoc reads)

Every text format also has a _string variant that parses content directly. DOCX, ODT, EPUB and RTF read lists, blockquotes, tables, images and footnotes; document metadata is extracted for every format that has any.

Dispatch

Function Description
panduck_read_blocks(path) Read any supported path, reader chosen by extension
panduck_format_for(path) Which format a path resolves to
panduck_can_read(path) Whether a path is supported
panduck_supported_extensions() The registry, as a table
panduck_register_doc_reader(...) Register a reader at runtime

Document helpers

Function Description
doc_toc(path) Table of contents for any supported document
doc_section(path, heading) The blocks under one heading (match := 'contains' for a fragment)
doc_search_sections(path, pattern) The section whose CONTENT matches, not its heading
doc_container(path, id) The blocks inside one container
doc_render(path, fmt) Render a document to md, html or text

Each of these accepts expand_embedded := true, which parses a format embedded in a document – a markdown cell inside a .ipynb – so a notebook is navigable rather than a wall of raw blocks. Opt-in: it needs the markdown extension, and an unchanged call reads identically without it.

Writing Pandoc JSON

Function Description
panduck_blocks_to_pandoc_ast(blocks) duck_block list to a Pandoc AST document
panduck_blocks_to_pandoc_blocks(blocks) duck_block list to a Pandoc block array
panduck_blocks_to_pandoc_json(blocks) duck_block list to Pandoc JSON text
panduck_write_pandoc_ast(path, blocks) Write Pandoc JSON to a file
panduck_pandoc_ast_map() The constructor mapping, as a queryable table

The governing rule is that panduck may be more faithful than pandoc – richer attributes, better block types, metadata pandoc does not extract – so long as the mapping back to valid Pandoc JSON stays total. Every fixture and a set of hand-built constructs are round-tripped through a real pandoc in CI to hold that.

Scope

This is an early release. Ten native readers are implemented and tested against reference implementations – RTF, DOCX, ODT, EPUB, LaTeX, Org, RST, ipynb, MediaWiki, Textile – plus a Pandoc AST reader, with 3108 test assertions, differential validation against a real pandoc on every fixture, and eleven jobs in CI. PDF, Markdown and HTML are read by delegating to the pdf, markdown and webbed extensions rather than by a reader here.

The readers were audited over five rounds before this release, each round asking a question the previous one could not answer – does a construct come back at all, do any words go missing, do all ten readers agree on one logical document. That found 18 defects, including silent text loss in DOCX where a hyperlink's anchor text was dropped outright. All are fixed and pinned by tests. Three findings remain open and are recorded in the documentation rather than omitted.

Known divergences from pandoc are declared and reasoned in the repository's roundtrip ledger rather than hidden – including several where pandoc is the one losing information.

Added Functions

function_name function_type description comment examples
doc_container table_macro NULL NULL  
doc_render macro NULL NULL  
doc_search_sections table_macro NULL NULL  
doc_section table_macro NULL NULL  
doc_toc table_macro NULL NULL  
panduck_block_cols macro NULL NULL  
panduck_blocks_to_pandoc_ast scalar Convert a list of duck_blocks into a Pandoc AST document struct. NULL [panduck_blocks_to_pandoc_ast(blocks)]
panduck_blocks_to_pandoc_blocks scalar Convert a list of duck_blocks into a JSON array string of Pandoc block elements. NULL [panduck_blocks_to_pandoc_blocks(blocks)]
panduck_blocks_to_pandoc_json scalar Convert a list of duck_blocks into a Pandoc AST JSON string. NULL [panduck_blocks_to_pandoc_json(blocks)]
panduck_builtin_format_for scalar Return the native builtin format name for a file path or URI. NULL [panduck_builtin_format_for('doc.docx')]
panduck_can_read scalar Check if panduck can read a given file path or URI. NULL [panduck_can_read('doc.docx')]
panduck_dependencies table_macro NULL NULL  
panduck_duck_block_spec_at_least macro NULL NULL  
panduck_duck_block_type scalar Return the SQL type definition string for duck_blocks. NULL [panduck_duck_block_type()]
panduck_ensure_extension scalar Ensure that a required DuckDB extension is loaded. NULL [panduck_ensure_extension('fts')]
panduck_expand_embedded macro NULL NULL  
panduck_expand_embedded_impl macro NULL NULL  
panduck_format_for scalar Return the document format name for a file path or URI. NULL [panduck_format_for('doc.docx')]
panduck_function_exists scalar Check if a scalar or table function exists in the catalog. NULL [panduck_function_exists('read_docx_blocks')]
panduck_glob scalar Glob files matching pattern. NULL [panduck_glob('docs/*.docx')]
panduck_is_glob macro NULL NULL  
panduck_latex_tokens table Tokenize a LaTeX string and return the sequence of LaTeX tokens. NULL [SELECT * FROM panduck_latex_tokens('\textbf{hello}')]
panduck_pandoc_api_version scalar Return the Pandoc API version supported by the panduck extension as [major, minor, patch]. NULL [panduck_pandoc_api_version()]
panduck_pandoc_ast_json scalar Convert a list of duck_blocks into a Pandoc AST JSON string (alias for panduck_blocks_to_pandoc_json). NULL [panduck_pandoc_ast_json(blocks)]
panduck_pandoc_ast_map table Return the complete constructor set of Pandoc AST types mapped to DuckDB duck_block element types. NULL [SELECT * FROM panduck_pandoc_ast_map()]
panduck_pandoc_ast_to_blocks scalar Convert a Pandoc AST JSON string into a list of duck_blocks. NULL [panduck_pandoc_ast_to_blocks('{"pandoc-api-version": [1,23,1], "meta": {}, "blocks": []}')]
panduck_pdf_blocks_impl table_macro NULL NULL  
panduck_policy_format macro NULL NULL  
panduck_quote macro NULL NULL  
panduck_read_arms macro NULL NULL  
panduck_read_arms_opt macro NULL NULL  
panduck_read_blocks macro NULL NULL  
panduck_reader_enabled scalar Check if the reader for a specified format is currently enabled. NULL [panduck_reader_enabled('docx')]
panduck_reader_extension_for scalar Return the required extension name to read a file path or URI. NULL [panduck_reader_extension_for('doc.docx')]
panduck_reader_function_for scalar Return the reader table function name that handles a file path or URI. NULL [panduck_reader_function_for('doc.docx')]
panduck_reader_kind_for scalar Return the reader kind ('doc' or 'table') for a file path or URI. NULL [panduck_reader_kind_for('doc.docx')]
panduck_reader_option_for scalar Look up a reader configuration option for a given format. NULL [panduck_reader_option_for('docx', 'toc', NULL)]
panduck_reader_registry table Return the table of all registered format readers and handlers. NULL [SELECT * FROM panduck_reader_registry()]
panduck_register_doc_reader table Register a custom document block reader for a file format or extension. NULL [SELECT * FROM panduck_register_doc_reader('custom', 'read_custom', ['.custom'])]
panduck_register_table_reader table Register a custom table reader for a file format or extension. NULL [SELECT * FROM panduck_register_table_reader('custom', 'read_custom_tbl', ['.custom'])]
panduck_registry_key_for scalar Return the registry match key for a file path or URI. NULL [panduck_registry_key_for('doc.docx')]
panduck_render_format macro NULL NULL  
panduck_render_params scalar Render named parameters map into SQL argument string. NULL [panduck_render_params(MAP {'opt': 'val'})]
panduck_renumber_blocks macro NULL NULL  
panduck_resolved_format macro NULL NULL  
panduck_source_list macro NULL NULL  
panduck_supported_extensions table Return list of supported document formats and extensions in panduck. NULL [SELECT * FROM panduck_supported_extensions()]
panduck_supported_paths macro NULL NULL  
panduck_version scalar Return the version of the panduck extension. NULL [panduck_version()]
panduck_wrap_expand macro NULL NULL  
panduck_write_pandoc_ast scalar Write a list of duck_blocks out to disk as a Pandoc AST JSON file. NULL [panduck_write_pandoc_ast('output.json', blocks)]
read_docx_blocks table Read a DOCX document and return structured document blocks. NULL [SELECT * FROM read_docx_blocks('document.docx')]
read_epub_blocks table Read an EPUB document and return structured document blocks. NULL [SELECT * FROM read_epub_blocks('book.epub')]
read_ipynb_blocks table Read a Jupyter Notebook (.ipynb) file and return structured document blocks. NULL [SELECT * FROM read_ipynb_blocks('notebook.ipynb')]
read_ipynb_blocks_string table Parse a Jupyter Notebook JSON string and return structured document blocks. NULL [SELECT * FROM read_ipynb_blocks_string('{"cells": [], "metadata": {}, "nbformat": 4, "nbformat_minor": 2}')]
read_latex_blocks table Read a LaTeX document and return structured document blocks. NULL [SELECT * FROM read_latex_blocks('document.tex')]
read_latex_blocks_string table Parse a LaTeX string and return structured document blocks. NULL [SELECT * FROM read_latex_blocks_string('\section{Title}\n\textbf{Bold}')]
read_mediawiki_blocks table Read a MediaWiki wikitext document and return structured document blocks. NULL [SELECT * FROM read_mediawiki_blocks('document.wiki')]
read_mediawiki_blocks_string table Parse a MediaWiki wikitext string and return structured document blocks. NULL [SELECT * FROM read_mediawiki_blocks_string('== Section ==\n'''Bold'''')]
read_odt_blocks table Read an ODT document and return structured document blocks. NULL [SELECT * FROM read_odt_blocks('document.odt')]
read_org_blocks table Read an Emacs Org-mode document and return structured document blocks. NULL [SELECT * FROM read_org_blocks('document.org')]
read_org_blocks_string table Parse an Emacs Org-mode string and return structured document blocks. NULL [SELECT * FROM read_org_blocks_string('* Heading\nParagraph')]
read_pandoc_blocks table Read a Pandoc AST JSON file and return structured document blocks. NULL [SELECT * FROM read_pandoc_blocks('document.json')]
read_pandoc_blocks_string table Parse a Pandoc AST JSON string and return structured document blocks. NULL [SELECT * FROM read_pandoc_blocks_string('{"pandoc-api-version": [1,23,1], "meta": {}, "blocks": []}')]
read_panduck_doc table_macro NULL NULL  
read_panduck_table table_macro NULL NULL  
read_pdf_blocks table_macro NULL NULL  
read_rst_blocks table Read a reStructuredText (RST) document and return structured document blocks. NULL [SELECT * FROM read_rst_blocks('document.rst')]
read_rst_blocks_string table Parse a reStructuredText (RST) string and return structured document blocks. NULL [SELECT * FROM read_rst_blocks_string('Title\n=====\n\nParagraph')]
read_rtf_blocks table Read an RTF file and return structured document blocks. NULL [SELECT * FROM read_rtf_blocks('document.rtf')]
read_textile_blocks table Read a Textile document and return structured document blocks. NULL [SELECT * FROM read_textile_blocks('document.textile')]
read_textile_blocks_string table Parse a Textile string and return structured document blocks. NULL [SELECT * FROM read_textile_blocks_string('h1. Header\n\nbold')]

Overloaded Functions

This extension does not add any function overloads.

Added Types

This extension does not add any types.

Added Settings

name description input_type scope aliases
panduck_allow_registration Allow panduck_register_doc_reader/panduck_register_table_reader to add readers at runtime BOOLEAN GLOBAL []
panduck_disabled_readers Comma-separated denylist of reader FORMATS, applied after panduck_enabled_readers. 'code' turns off the fallback that returns a parse tree for sources no reader claimed VARCHAR GLOBAL []
panduck_enabled_readers Comma-separated allowlist of reader FORMATS ('*' for all). Applied before panduck_disabled_readers VARCHAR GLOBAL []