Benchmarking Against Truth Sets: Why 99.5% F1 Is Measured Where Calling Is Easy
Validating a variant caller has a standard form: run it on HG002, compare against the GIAB benchmark, report F1. The number comes back around 99.5% and the pipeline is documented as accurate. Accurate where? The comparison is restricted to high-confidence regions — regions defined partly by excluding places where methods systematically disagree, which is precisely where calling is hardest. The clearest evidence: expanding the benchmark from v3.3.2 to v4.2.1 revealed eight times more false negatives in an unchanged call set. The pipeline didn't change. The measured territory did.
Sign up to read this post
Join Now

