← back to projects

vcfclick

Small VCF databases. One per cohort. Embedded ClickHouse engine for variants, embedded DuckDB for annotations, MCP natural-language layer. No daemon, no server, no SQL required for day-to-day questions.

v0.14.1 Apache-2.0 Python 3.11 – 3.14 Embedded engines

Hosted vcfclick — gauging interest

The CLI runs locally today. If you'd want a hosted workspace to keep cohorts in the cloud, share them with collaborators, and query in the browser, we'd like to hear about your data and constraints before we build it.

Request private-cohort hosting →

What it does

Ingest a cohort VCF, get an embedded analytical database back. Variants, genotypes, and samples land in a chDB store on disk; gene and ClinVar annotations live in a DuckDB file next to it. Queries run in-process, no client/server hop. One cohort, one directory, one ~/.vcfclick/dbs/<name> path.

Ingest is two-phase: a parse-and-stage pass writes Parquet to a tempdir on the DB volume with no engine writes; a commit pass runs only if staging succeeded. A failed re-ingest under the same --ingest-id rolls back to the prior state, not to empty. Concurrent same-id ingests serialize through an fcntl file lock.

DRAGEN and GATK INFO/FORMAT fields are routed into typed columns. Multi-allelic sites must be split before ingest (bcftools norm -m -); ingest stops with that exact command if it finds one. Cohort allele frequencies are sparse-aware: variants absent in a cohort are counted as 0/N, not as missing.

For families, load a pedigree with db ped and report de-novo / recessive / dominant candidates with db trio; defensible de-novo uses --keep-reference so a parent's hom-reference is proven, not inferred from absence. And combine reimplements GATK3 CombineVariants — the multi-callset merge with set= provenance that GATK4 removed — verified equal to real GATK 3.8 output.

Three ways to drive it: the CLI, an optional terminal UI, and vcfclick web — a local, localhost-only browser app (the [web] extra) with a SQL explorer, a natural-language→SQL box, and trio / combine panels over your cohort database. Plus an MCP server so an LLM client can write visible, auditable SQL for you. Native desktop apps for macOS, Windows and Linux are in development.

db createdb ingestdb ingest-batchdb infodb querydb statsdb diffdb qcdb peddb triodb dumpdb pushdb pulldb listdb rmcombinemergebenchmarkannotationswebtui--ingest-id--cohort--keep-reference--min-callsets

Quick start

pipx install vcfclick

vcfclick db create epilepsy_2026
vcfclick db ingest epilepsy_2026 cohort.vcf.gz \
    --cohort EPI-2026 --ingest-id baseline

vcfclick db info epilepsy_2026
vcfclick db diff epilepsy_2026 --cohort-a EPI-2026 --cohort-b CONTROLS

# Serve it to an MCP client (Claude Desktop, Claude Code, ...),
# using the Python that vcfclick is installed in:
VCFCLICK_DB_NAME=epilepsy_2026 python -m vcfclick_mcp.server

Examples

What each shipped command actually does, anchored in the operational behavior.

01

Ingest a single cohort VCF

Two-phase stage-then-commit pipeline. Phase 1 parses the VCF into Parquet on the DB volume with no engine writes; Phase 2 commits only if Phase 1 succeeded.

vcfclick db create epilepsy_2026
vcfclick db ingest epilepsy_2026 cohort.vcf.gz \
    --cohort EPI-2026 --ingest-id baseline
02

Re-ingest with the same ingest-id

Reusing an --ingest-id genuinely replaces the prior data for that id (rollback then re-insert). Mid-stream parse failures roll back to the prior state, not to empty.

vcfclick db ingest epilepsy_2026 cohort_v2.vcf.gz \
    --cohort EPI-2026 --ingest-id baseline --serial
# prior 'baseline' data is fully replaced; other ingest_ids untouched
03

Compare cohort allele frequencies

Per-variant allele counts and frequencies for two cohorts in the same database, ranked by the size of the difference. Sparse-aware: a variant absent from one cohort counts as 0/N there, not as missing. Single SQL pass against the embedded engine.

vcfclick db diff epilepsy_2026 \
    --cohort-a EPI-2026 --cohort-b CONTROLS --top 20
# chrom  pos  ref  alt  ac_a  an_a  af_a  ac_b  an_b  af_b  af_diff
04

Database summary

Path, size on disk and row counts per table. db stats adds how well each typed INFO/FORMAT column is populated; db query runs any read-only SQL with --format JSON, CSV or TSV for pipelines.

vcfclick db info epilepsy_2026
# size:      176.4 MB
# variants:  1843266
# samples:   412
# ingestions:3
05

MCP natural-language query

The bundled MCP server hands Claude a schema briefing that teaches the correct sparse-aware AF pattern (cohort_size CTE against samples alone, CROSS JOIN into aggregation). Point your MCP client at the Python environment vcfclick is installed in.

VCFCLICK_DB_NAME=epilepsy_2026 python -m vcfclick_mcp.server
# then in Claude:
# "Which variants are at least 10% more common in
#  EPI-2026 than in CONTROLS?"
06

Trio de-novo / recessive / dominant candidates

Load a standard PED, then report Mendelian-model candidates with slivar-style quality gates. Defensible de novo needs --keep-reference so a parent's hom-reference is stored, not guessed from absence — a no-call parent is excluded, not falsely reported.

vcfclick db ingest fam1 trio.vcf.gz --cohort trio --keep-reference
vcfclick db ped  fam1 fam1.ped
vcfclick db trio fam1 --proband CHILD --category denovo
07

Combine caller call sets (GATK3 CombineVariants)

A reimplementation of GATK3 CombineVariants — the tool GATK4 removed. Unions call sets that may share samples, annotates each record with set= provenance (Intersection / source), resolves a shared sample by input priority, and can keep only consensus sites. Verified equal to real GATK 3.8 output.

vcfclick combine gatk.vcf.gz deepvariant.vcf.gz \
    -o consensus.vcf.gz --min-callsets 2
# each record annotated set=Intersection / gatk / deepvariant
08

Concurrent same-id ingest safety

Cross-process file lock per (DB, ingest_id). Parallel pipeline workers under the same id serialize cleanly; different ids run concurrently.

# Worker A and Worker B both running:
vcfclick db ingest cohort sample_a.vcf.gz --ingest-id batch_q2 &
vcfclick db ingest cohort sample_b.vcf.gz --ingest-id batch_q2 &
# B blocks on A's lock; staging dir and import path are race-free.

Architecture

Three embedded layers: chDB (ClickHouse engine, one process) holds variants/genotypes/samples; DuckDB holds gene + ClinVar annotations next to it; FastMCP exposes a stdio MCP server that hands Claude a schema briefing teaching the correct sparse-aware AF pattern.

The ingest pipeline is cyvcf2 (htslib via Python) → typed pyarrow Parquet → chDB bulk import. The Arrow schemas are locked column-for-column against the SQL DDL by a drift-guard test; every INSERT carries an explicit column list.

Atomicity is enforced by a two-phase stage-then-commit pattern with a commit_started flag in the rollback branch, plus an fcntl file lock keyed on the (database, ingest_id) pair. Every path interpolated into chDB SQL goes through a single sql_quote_str helper that escapes both backslashes and single quotes per ClickHouse literal rules.

Apache-2.0 Python 3.11, 3.12, 3.13, 3.14 macOS + Linux
← back to projects