What it does
Ingest a cohort VCF, get an embedded analytical database back. Variants, genotypes, and
samples land in a chDB store on disk; gene and ClinVar annotations live in a DuckDB file next
to it. Queries run in-process, no client/server hop. One cohort, one directory, one ~/.vcfclick/dbs/<name> path.
Ingest is two-phase: a parse-and-stage pass writes Parquet to a tempdir on the DB volume with
no engine writes; a commit pass runs only if staging succeeded. A failed re-ingest under the
same --ingest-id rolls back to the prior
state, not to empty. Concurrent same-id ingests serialize through an fcntl file lock.
DRAGEN and GATK INFO/FORMAT fields are routed into typed columns. Multi-allelic sites must
be split before ingest (bcftools norm -m -);
ingest stops with that exact command if it finds one. Cohort allele frequencies are
sparse-aware: variants absent in a cohort are counted as 0/N, not as missing.
For families, load a pedigree with db ped and report de-novo / recessive / dominant candidates with db trio; defensible de-novo uses --keep-reference so a parent's hom-reference is proven, not inferred from absence. And combine reimplements GATK3 CombineVariants — the multi-callset merge with set= provenance that GATK4 removed — verified equal to real GATK 3.8 output.
Three ways to drive it: the CLI, an optional terminal UI, and vcfclick web — a local,
localhost-only browser app (the [web] extra) with a SQL
explorer, a natural-language→SQL box, and trio / combine panels over your cohort
database. Plus an MCP server so an LLM client can write visible, auditable SQL for you.
Native desktop apps for macOS, Windows and Linux are in development.
db createdb ingestdb ingest-batchdb infodb querydb statsdb diffdb qcdb peddb triodb dumpdb pushdb pulldb listdb rmcombinemergebenchmarkannotationswebtui--ingest-id--cohort--keep-reference--min-callsets