Benchmarking Against Truth Sets: Why 99.5% F1 Is Measured Where Calling Is Easy

Validating a variant caller has a standard form: run it on HG002, compare against the GIAB benchmark, report F1. The number comes back around 99.5% and the pipeline is documented as accurate. Accurate where? The comparison is restricted to high-confidence regions — regions defined partly by excluding places where methods systematically disagree, which is precisely where calling is hardest. The clearest evidence: expanding the benchmark from v3.3.2 to v4.2.1 revealed eight times more false negatives in an unchanged call set. The pipeline didn't change. The measured territory did.

Sign up to read this post
Join Now
Previous
Previous

Structural Variants: Why Finding One, Placing It, and Genotyping It Are Three Different Problems

Next
Next

Batch Confounding in Public Data: Why the Correction You Apply Can Be Worse Than the Effect You Remove