Batch Confounding in Public Data: Why the Correction You Apply Can Be Worse Than the Effect You Remove

Public archives put molecular data from tens of thousands of patients into open hands. What you can't download is the experimental design — someone else decided which samples went on which plate, which center sequenced what, and those choices are now fixed. When batch is independent of the biology, it's a nuisance you can model. When batch correlates with the variable of interest, the two occupy the same statistical space and no correction can separate them. Worse, correction applied to a confounded design can inflate significance: in one documented case, from 11 differentially expressed probesets to over 1,000 on the same data.

Sign up to read this post
Join Now
Previous
Previous

Benchmarking Against Truth Sets: Why 99.5% F1 Is Measured Where Calling Is Easy

Next
Next

The Optimal Cutpoint Problem: Why a High-vs-Low Survival Curve Can Be Manufactured From Noise