The Optimal Cutpoint Problem: Why a High-vs-Low Survival Curve Can Be Manufactured From Noise

A two-curve Kaplan-Meier plot communicates a clinical claim in a single image. It also requires a decision the data doesn't supply: gene expression is continuous, the log-rank test compares groups, so someone chose a number and called everything above it "high." If that number was chosen by searching every threshold for the smallest p-value, the false-positive rate is around 40% rather than 5%, and a reported p = 0.002 corresponds to a genuine p = 0.05. The resulting figure is visually identical to one from a pre-specified split. Nothing in it records how many splits were tried.

Sign up to read this post
Join Now
Previous
Previous

Batch Confounding in Public Data: Why the Correction You Apply Can Be Worse Than the Effect You Remove

Next
Next

Metabolite Annotation Confidence: Why a Named Compound Is Usually a Match, Not an Identification