Benchmarking — checked against hap.py
vcfclick benchmark compares a query
VCF against a truth VCF over a reference and confident-region BED, and reports TP/FP/FN with
precision/recall/F1 in a GA4GH-shaped summary.csv. Two engines:
a normalized keyed match and a local-haplotype engine that
resolves representation-different calls — the residual that hap.py's xcmp exists for. It's a
native reimplementation, not a wrapper: no hap.py or vcfeval at runtime.
Against a fully independent caller
The hardest honest test isn't a self-benchmark — it's a call set produced by a different method. We scored the GIAB T2T-Q100 HG002 assembly dipcall (phased, assembly-based) against the v4.2.1 short-read truth on GRCh38 chr20, and ran the exact same pair through real hap.py (v0.3.12, xcmp) and through vcfclick.
~0.2%
SNP gap vs hap.py
recall and precision, on a real independent caller
~0.2%
INDEL gap vs hap.py
after closing an initial ~15-point gap
| Stratum | Metric | hap.py (xcmp) | vcfclick |
|---|---|---|---|
| SNP | recall / precision | 1.000 / 1.000 | 1.000 / 1.000 |
| INDEL | recall / precision | 1.000 / 0.999 | 0.998 / 0.998 |
GRCh38 chr20:1–6 Mb, GIAB v4.2.1 high-confidence BED. Both tools scored the identical (truth, query) pair.
What the real comparison found
Running against a genuinely independent caller — not a perturbed copy of the truth — caught two bugs that synthetic tests could not, and both are fixed:
- Phase-insensitive genotypes. The assembly caller is phased; a phased
1|0was not matching an unphased0/1, so every shared heterozygote scored as a mismatch. Genotype comparison is now phase-insensitive (a real0/1-vs-1/1error still counts). - Het-alt indels. When a caller spells a
1/2two-allele indel as two separate0/1records, the haplotype engine now reconciles it at the locus level. - Confident-region gating. A key-matched pair outside the
confident BED — where the truth set is silent and hap.py records
UNK— was being scored as a real result. Gating the matched-pair case closed the rest of the INDEL gap (0.96 → 0.998) and lined vcfclick's not-assessable counts up with hap.py's.
Honest scope: after these fixes, both SNPs and INDELs track hap.py to about two-tenths of a percent on a fully independent caller — the residual is a handful of individual variants, not a systematic gap. The full numbers, the synthetic conformance case, and the GIAB chr20/chr1 runs are documented in the repo.