Read documents natively into the duck_block vocabulary – DOCX, ODT, EPUB, RTF, LaTeX, Org, RST, ipynb, MediaWiki and Textile – and write them back as Pandoc JSON, without linking Pandoc
Installing and Loading
INSTALL panduck FROM community;
LOAD panduck;
Example
LOAD panduck;
-- Read a document straight into flattened duck_block rows
SELECT * FROM read_docx_blocks('report.docx');
SELECT * FROM read_epub_blocks('book.epub');
SELECT * FROM read_latex_blocks('paper.tex');
-- Or let the path pick the reader
SELECT * FROM panduck_read_blocks('notes.org');
SELECT panduck_format_for('paper.tex'); -- 'latex'
SELECT panduck_can_read('archive.docx'); -- true
-- Formats with no file: parse a string directly
SELECT * FROM read_mediawiki_blocks_string('== Heading ==');
-- Outline any supported document
SELECT * FROM doc_toc('report.docx');
-- Pull one section out, by heading or by a fragment of one
SELECT * FROM doc_section('report.docx', 'Methods');
SELECT * FROM doc_section('report.docx', 'Meth', match := 'contains');
-- ...or find the section by what it SAYS rather than what it is called
SELECT * FROM doc_search_sections('report.docx', 'p < 0.05');
-- Render one back out (md, html or text)
SELECT doc_render('report.docx', 'md');
-- Write back out as Pandoc JSON that a real pandoc accepts
SELECT panduck_write_pandoc_ast('doc.json',
(SELECT list(b) FROM read_odt_blocks('doc.odt') b)
);
-- Or get the AST as a struct without touching the filesystem
SELECT panduck_blocks_to_pandoc_ast(
(SELECT list(b) FROM read_rtf_blocks('memo.rtf') b)
);
About panduck
Panduck reads document formats natively into the duck_block vocabulary shared
with duck_block_utils, markdown and webbed – so a DOCX, an EPUB and a
LaTeX paper all land in the same seven-column shape and can be queried together.
It does not link Pandoc. Pandoc is a Haskell program; the only C bindings that ever existed went unmaintained in 2017, and a Cabal foreign-library request has been open upstream since 2020. Linking the GHC runtime into a dlopen'd DuckDB extension would be a bad neighbour even if it existed. Panduck is instead compatible with Pandoc's data model: every reader is hand-written, and a checked mapping table pins the correspondence to Pandoc's AST constructors.
Readers
| Function | Format |
|---|---|
read_docx_blocks(path) |
Word OOXML |
read_odt_blocks(path) |
OpenDocument Text |
read_epub_blocks(path) |
EPUB 2/3 |
read_rtf_blocks(path) |
Rich Text Format |
read_latex_blocks(path) |
LaTeX |
read_org_blocks(path) |
Org mode |
read_rst_blocks(path) |
reStructuredText |
read_ipynb_blocks(path) |
Jupyter notebooks |
read_mediawiki_blocks(path) |
MediaWiki wikitext |
read_textile_blocks(path) |
Textile |
read_pandoc_blocks(path) |
Pandoc JSON AST (reaches all 43 formats pandoc reads) |
Every text format also has a _string variant that parses content directly.
DOCX, ODT, EPUB and RTF read lists, blockquotes, tables, images and footnotes;
document metadata is extracted for every format that has any.
Dispatch
| Function | Description |
|---|---|
panduck_read_blocks(path) |
Read any supported path, reader chosen by extension |
panduck_format_for(path) |
Which format a path resolves to |
panduck_can_read(path) |
Whether a path is supported |
panduck_supported_extensions() |
The registry, as a table |
panduck_register_doc_reader(...) |
Register a reader at runtime |
Document helpers
| Function | Description |
|---|---|
doc_toc(path) |
Table of contents for any supported document |
doc_section(path, heading) |
The blocks under one heading (match := 'contains' for a fragment) |
doc_search_sections(path, pattern) |
The section whose CONTENT matches, not its heading |
doc_container(path, id) |
The blocks inside one container |
doc_render(path, fmt) |
Render a document to md, html or text |
Each of these accepts expand_embedded := true, which parses a format embedded in a
document – a markdown cell inside a .ipynb – so a notebook is navigable rather
than a wall of raw blocks. Opt-in: it needs the markdown extension, and an
unchanged call reads identically without it.
Writing Pandoc JSON
| Function | Description |
|---|---|
panduck_blocks_to_pandoc_ast(blocks) |
duck_block list to a Pandoc AST document |
panduck_blocks_to_pandoc_blocks(blocks) |
duck_block list to a Pandoc block array |
panduck_blocks_to_pandoc_json(blocks) |
duck_block list to Pandoc JSON text |
panduck_write_pandoc_ast(path, blocks) |
Write Pandoc JSON to a file |
panduck_pandoc_ast_map() |
The constructor mapping, as a queryable table |
The governing rule is that panduck may be more faithful than pandoc – richer attributes, better block types, metadata pandoc does not extract – so long as the mapping back to valid Pandoc JSON stays total. Every fixture and a set of hand-built constructs are round-tripped through a real pandoc in CI to hold that.
Scope
This is an early release. Ten native readers are implemented and tested against reference implementations – RTF, DOCX, ODT, EPUB, LaTeX, Org, RST, ipynb, MediaWiki, Textile – plus a Pandoc AST reader, with 3108 test assertions, differential validation against a real pandoc on every fixture, and eleven jobs in CI. PDF, Markdown and HTML are read by delegating to the pdf, markdown and webbed extensions rather than by a reader here.
The readers were audited over five rounds before this release, each round asking a question the previous one could not answer – does a construct come back at all, do any words go missing, do all ten readers agree on one logical document. That found 18 defects, including silent text loss in DOCX where a hyperlink's anchor text was dropped outright. All are fixed and pinned by tests. Three findings remain open and are recorded in the documentation rather than omitted.
Known divergences from pandoc are declared and reasoned in the repository's roundtrip ledger rather than hidden – including several where pandoc is the one losing information.
Added Functions
| function_name | function_type | description | comment | examples |
|---|---|---|---|---|
| doc_container | table_macro | NULL | NULL | |
| doc_render | macro | NULL | NULL | |
| doc_search_sections | table_macro | NULL | NULL | |
| doc_section | table_macro | NULL | NULL | |
| doc_toc | table_macro | NULL | NULL | |
| panduck_block_cols | macro | NULL | NULL | |
| panduck_blocks_to_pandoc_ast | scalar | Convert a list of duck_blocks into a Pandoc AST document struct. | NULL | [panduck_blocks_to_pandoc_ast(blocks)] |
| panduck_blocks_to_pandoc_blocks | scalar | Convert a list of duck_blocks into a JSON array string of Pandoc block elements. | NULL | [panduck_blocks_to_pandoc_blocks(blocks)] |
| panduck_blocks_to_pandoc_json | scalar | Convert a list of duck_blocks into a Pandoc AST JSON string. | NULL | [panduck_blocks_to_pandoc_json(blocks)] |
| panduck_builtin_format_for | scalar | Return the native builtin format name for a file path or URI. | NULL | [panduck_builtin_format_for('doc.docx')] |
| panduck_can_read | scalar | Check if panduck can read a given file path or URI. | NULL | [panduck_can_read('doc.docx')] |
| panduck_dependencies | table_macro | NULL | NULL | |
| panduck_duck_block_spec_at_least | macro | NULL | NULL | |
| panduck_duck_block_type | scalar | Return the SQL type definition string for duck_blocks. | NULL | [panduck_duck_block_type()] |
| panduck_ensure_extension | scalar | Ensure that a required DuckDB extension is loaded. | NULL | [panduck_ensure_extension('fts')] |
| panduck_expand_embedded | macro | NULL | NULL | |
| panduck_expand_embedded_impl | macro | NULL | NULL | |
| panduck_format_for | scalar | Return the document format name for a file path or URI. | NULL | [panduck_format_for('doc.docx')] |
| panduck_function_exists | scalar | Check if a scalar or table function exists in the catalog. | NULL | [panduck_function_exists('read_docx_blocks')] |
| panduck_glob | scalar | Glob files matching pattern. | NULL | [panduck_glob('docs/*.docx')] |
| panduck_is_glob | macro | NULL | NULL | |
| panduck_latex_tokens | table | Tokenize a LaTeX string and return the sequence of LaTeX tokens. | NULL | [SELECT * FROM panduck_latex_tokens('\textbf{hello}')] |
| panduck_pandoc_api_version | scalar | Return the Pandoc API version supported by the panduck extension as [major, minor, patch]. | NULL | [panduck_pandoc_api_version()] |
| panduck_pandoc_ast_json | scalar | Convert a list of duck_blocks into a Pandoc AST JSON string (alias for panduck_blocks_to_pandoc_json). | NULL | [panduck_pandoc_ast_json(blocks)] |
| panduck_pandoc_ast_map | table | Return the complete constructor set of Pandoc AST types mapped to DuckDB duck_block element types. | NULL | [SELECT * FROM panduck_pandoc_ast_map()] |
| panduck_pandoc_ast_to_blocks | scalar | Convert a Pandoc AST JSON string into a list of duck_blocks. | NULL | [panduck_pandoc_ast_to_blocks('{"pandoc-api-version": [1,23,1], "meta": {}, "blocks": []}')] |
| panduck_pdf_blocks_impl | table_macro | NULL | NULL | |
| panduck_policy_format | macro | NULL | NULL | |
| panduck_quote | macro | NULL | NULL | |
| panduck_read_arms | macro | NULL | NULL | |
| panduck_read_arms_opt | macro | NULL | NULL | |
| panduck_read_blocks | macro | NULL | NULL | |
| panduck_reader_enabled | scalar | Check if the reader for a specified format is currently enabled. | NULL | [panduck_reader_enabled('docx')] |
| panduck_reader_extension_for | scalar | Return the required extension name to read a file path or URI. | NULL | [panduck_reader_extension_for('doc.docx')] |
| panduck_reader_function_for | scalar | Return the reader table function name that handles a file path or URI. | NULL | [panduck_reader_function_for('doc.docx')] |
| panduck_reader_kind_for | scalar | Return the reader kind ('doc' or 'table') for a file path or URI. | NULL | [panduck_reader_kind_for('doc.docx')] |
| panduck_reader_option_for | scalar | Look up a reader configuration option for a given format. | NULL | [panduck_reader_option_for('docx', 'toc', NULL)] |
| panduck_reader_registry | table | Return the table of all registered format readers and handlers. | NULL | [SELECT * FROM panduck_reader_registry()] |
| panduck_register_doc_reader | table | Register a custom document block reader for a file format or extension. | NULL | [SELECT * FROM panduck_register_doc_reader('custom', 'read_custom', ['.custom'])] |
| panduck_register_table_reader | table | Register a custom table reader for a file format or extension. | NULL | [SELECT * FROM panduck_register_table_reader('custom', 'read_custom_tbl', ['.custom'])] |
| panduck_registry_key_for | scalar | Return the registry match key for a file path or URI. | NULL | [panduck_registry_key_for('doc.docx')] |
| panduck_render_format | macro | NULL | NULL | |
| panduck_render_params | scalar | Render named parameters map into SQL argument string. | NULL | [panduck_render_params(MAP {'opt': 'val'})] |
| panduck_renumber_blocks | macro | NULL | NULL | |
| panduck_resolved_format | macro | NULL | NULL | |
| panduck_source_list | macro | NULL | NULL | |
| panduck_supported_extensions | table | Return list of supported document formats and extensions in panduck. | NULL | [SELECT * FROM panduck_supported_extensions()] |
| panduck_supported_paths | macro | NULL | NULL | |
| panduck_version | scalar | Return the version of the panduck extension. | NULL | [panduck_version()] |
| panduck_wrap_expand | macro | NULL | NULL | |
| panduck_write_pandoc_ast | scalar | Write a list of duck_blocks out to disk as a Pandoc AST JSON file. | NULL | [panduck_write_pandoc_ast('output.json', blocks)] |
| read_docx_blocks | table | Read a DOCX document and return structured document blocks. | NULL | [SELECT * FROM read_docx_blocks('document.docx')] |
| read_epub_blocks | table | Read an EPUB document and return structured document blocks. | NULL | [SELECT * FROM read_epub_blocks('book.epub')] |
| read_ipynb_blocks | table | Read a Jupyter Notebook (.ipynb) file and return structured document blocks. | NULL | [SELECT * FROM read_ipynb_blocks('notebook.ipynb')] |
| read_ipynb_blocks_string | table | Parse a Jupyter Notebook JSON string and return structured document blocks. | NULL | [SELECT * FROM read_ipynb_blocks_string('{"cells": [], "metadata": {}, "nbformat": 4, "nbformat_minor": 2}')] |
| read_latex_blocks | table | Read a LaTeX document and return structured document blocks. | NULL | [SELECT * FROM read_latex_blocks('document.tex')] |
| read_latex_blocks_string | table | Parse a LaTeX string and return structured document blocks. | NULL | [SELECT * FROM read_latex_blocks_string('\section{Title}\n\textbf{Bold}')] |
| read_mediawiki_blocks | table | Read a MediaWiki wikitext document and return structured document blocks. | NULL | [SELECT * FROM read_mediawiki_blocks('document.wiki')] |
| read_mediawiki_blocks_string | table | Parse a MediaWiki wikitext string and return structured document blocks. | NULL | [SELECT * FROM read_mediawiki_blocks_string('== Section ==\n'''Bold'''')] |
| read_odt_blocks | table | Read an ODT document and return structured document blocks. | NULL | [SELECT * FROM read_odt_blocks('document.odt')] |
| read_org_blocks | table | Read an Emacs Org-mode document and return structured document blocks. | NULL | [SELECT * FROM read_org_blocks('document.org')] |
| read_org_blocks_string | table | Parse an Emacs Org-mode string and return structured document blocks. | NULL | [SELECT * FROM read_org_blocks_string('* Heading\nParagraph')] |
| read_pandoc_blocks | table | Read a Pandoc AST JSON file and return structured document blocks. | NULL | [SELECT * FROM read_pandoc_blocks('document.json')] |
| read_pandoc_blocks_string | table | Parse a Pandoc AST JSON string and return structured document blocks. | NULL | [SELECT * FROM read_pandoc_blocks_string('{"pandoc-api-version": [1,23,1], "meta": {}, "blocks": []}')] |
| read_panduck_doc | table_macro | NULL | NULL | |
| read_panduck_table | table_macro | NULL | NULL | |
| read_pdf_blocks | table_macro | NULL | NULL | |
| read_rst_blocks | table | Read a reStructuredText (RST) document and return structured document blocks. | NULL | [SELECT * FROM read_rst_blocks('document.rst')] |
| read_rst_blocks_string | table | Parse a reStructuredText (RST) string and return structured document blocks. | NULL | [SELECT * FROM read_rst_blocks_string('Title\n=====\n\nParagraph')] |
| read_rtf_blocks | table | Read an RTF file and return structured document blocks. | NULL | [SELECT * FROM read_rtf_blocks('document.rtf')] |
| read_textile_blocks | table | Read a Textile document and return structured document blocks. | NULL | [SELECT * FROM read_textile_blocks('document.textile')] |
| read_textile_blocks_string | table | Parse a Textile string and return structured document blocks. | NULL | [SELECT * FROM read_textile_blocks_string('h1. Header\n\nbold')] |
Overloaded Functions
This extension does not add any function overloads.
Added Types
This extension does not add any types.
Added Settings
| name | description | input_type | scope | aliases |
|---|---|---|---|---|
| panduck_allow_registration | Allow panduck_register_doc_reader/panduck_register_table_reader to add readers at runtime | BOOLEAN | GLOBAL | [] |
| panduck_disabled_readers | Comma-separated denylist of reader FORMATS, applied after panduck_enabled_readers. 'code' turns off the fallback that returns a parse tree for sources no reader claimed | VARCHAR | GLOBAL | [] |
| panduck_enabled_readers | Comma-separated allowlist of reader FORMATS ('*' for all). Applied before panduck_disabled_readers | VARCHAR | GLOBAL | [] |