Search Shortcut cmd + k | ctrl + k

Read GORpipe .gorz and .gord files and write .gorz as native DuckDB tables

Maintainer(s): gorfather

Installing and Loading

INSTALL gorz FROM community;
LOAD gorz;

Example

-- Write a query result out as a GORZ file (rows must be in GOR order),
-- then read it back — with an optional GOR -p style range filter.
COPY (SELECT * FROM (VALUES ('chr1', 1000, 'A', 'G'), ('chr1', 2000, 'C', 'T'))
        t(chrom, pos, ref, alt) ORDER BY chrom, pos)
  TO 'variants.gorz' (FORMAT gorz);

SELECT * FROM read_gor('variants.gorz', range := 'chr1:1500-2500');
-- Equivalently, in plain SQL on the bare file name (the chrom / pos
-- predicates are pushed down into a block seek):
SELECT * FROM 'variants.gorz' WHERE chrom = 'chr1' AND pos BETWEEN 1500 AND 2500;

-- Read a partitioned .gord dictionary table (not shown here), keeping only
-- the partitions tagged sample1 or sample99 (GOR's -f). See the GORpipe paper
-- (Guðbjartsson et al., Bioinformatics 2016, doi:10.1093/bioinformatics/btw199)
-- and "Ultra-fast joint-genotyping with SparkGOR" (Guðbjartsson et al.,
-- bioRxiv 2022, doi:10.1101/2022.10.25.513331):
SELECT * FROM read_gord('example.gord', f := ['sample1', 'sample99']);

About gorz

The gorz extension reads and writes GORpipe genomic files — block-compressed .gorz and .gord dictionaries — as native DuckDB tables. Files it writes are read and seeked natively by gorpipe, and vice-versa.

Reading

  • read_gor(path) — auto-detects .gorz vs .gord by extension.
  • read_gorz(path) / read_gord(path) — force the kind.
  • A replacement scan, so a bare literal works: SELECT count(*) FROM 'variants.gord';
  • pgor … | write x.gord folders — pass the folder itself (FROM 'x.gord'); its x.gord/thedict.gord is read.
  • Projection & parallel scan — only selected columns are read; large files scan across threads.
  • WHERE chrom/pos → block seek — range predicates seek into the right block instead of scanning the whole file.
  • range := 'chrN:start-end' — GOR's -p semantics as a hard positional filter (also 'chrN', 'chrN:start-', 'chrN:pos').
  • Partition filters — GOR's -f / -ff semantics prune a .gord's file list and expose a Source column: f := ['A','B'] is an inline tag list, ff := 'tags.txt' is a tag file (one tag per line / first tab-delimited column, # lines skipped). source := 'PN' renames the exposed column.
  • Object stores — paths resolve through DuckDB's FileSystem, so s3://… works when httpfs is loaded.

Writing

COPY (SELECT chrom, pos, ref, alt FROM my_variants ORDER BY chrom, pos)
  TO 'out.gorz' (FORMAT gorz);   -- or (FORMAT gor) for plain sorted TSV

Single-threaded, standard block-zip; validates GOR order (chromosomes ascend lexicographically, positions non-decreasing) and errors out otherwise.

References

Added Functions

function_name function_type description comment examples
read_gor table NULL NULL  
read_gord table NULL NULL  
read_gorz table NULL NULL  

Overloaded Functions

This extension does not add any function overloads.

Added Types

This extension does not add any types.

Added Settings

name description input_type scope aliases
gorz_scan_chunk_bytes Byte size of each parallel .gorz full-scan range (0 = automatic: one range per thread, at least 16 MiB each) BIGINT GLOBAL []