Diligence on a Genomics Claim: What to Ask About the Cohort, the Split, and the Endpoint
Diligence on a Genomics Claim: What to Ask About the Cohort, the Split, and the Endpoint
N is the least informative number in the sentence. Three facts that are almost never on the slide decide whether the claim is evidence or arithmetic — and every way it can go wrong points in the same direction.
The slide says the model was validated in an independent cohort of 1,240 patients and achieved an area under the curve of 0.89. It is a good sentence. You have twenty minutes and a room full of people, and you need to decide how much of the company’s story is resting on it.
The number that feels like the evidence — 1,240 — carries almost none of it. Three other facts do, and none of them is on the slide. This piece is not about a laboratory’s validation of an assay, which is a claim about a measurement checked against a comparator. It is about the claim built on top: that the measurement predicts something. Different object, different ways of going wrong.
01The cohort: count events, and count what was searched
In any setting with a binary or time-to-event outcome, the effective sample size is the number of events, not the number of people. A cohort of 1,240 patients containing 62 recurrences is, for most statistical purposes, a study of 62. And how many events are enough is a calculation rather than a rule of thumb — it depends on the number of candidate predictors, the outcome frequency, and the strength of signal being claimed, which is why the old “ten events per variable” heuristic has been superseded by explicit sample size formulas for prediction models.1 Ask for the event count before anything else. It is a one-word answer and it reframes everything after it.
Then ask how many features were considered, not how many survived. A twelve-gene signature selected from twenty thousand measured transcripts was chosen out of a space of twenty thousand, and that is the number governing how easily the selection could have been luck. The final model’s parsimony is a property of the output, not of the search.
This matters more in genomics than almost anywhere else, because association with outcome is unusually cheap here. In a study that has aged extremely well, breast cancer outcome data were used to show that any randomly chosen set of 100 or more genes had better than a 90% chance of being significantly associated with survival. Of 47 published breast cancer outcome signatures, 28 — 60% — were no better as predictors than random signatures of identical size, and 11 were worse than the median random signature. The explanation was structural rather than sloppy: much of the transcriptome tracks proliferation, and proliferation carries most of the prognostic information in that disease.2
The implication for diligence is precise. “Significantly associated with outcome” is close to free, so the null hypothesis of no association is the wrong benchmark. The right comparisons are a random feature set of the same size, and whatever the clinician already has without paying for anything.
02The split: when was it made, and how many times has it been opened?
There is really only one question here, and it is about sequence: at what point in the analysis was the held-out data separated from everything else? If normalization, batch correction, imputation, feature selection, or cutoff choice were performed across all the samples and the split came afterward, the test set participated in its own evaluation. The reported figure is then partly a measurement of the test set rather than of the model.
It is worth asking bluntly, because this failure is neither rare nor exotic. A systematic survey found leakage documented across 17 scientific fields, collectively affecting 294 papers, in some cases producing wildly overoptimistic conclusions; in one case study the authors reproduced, correcting the leakage left complex machine learning models performing no better than decades-old logistic regression.3 These are published, peer-reviewed results in fields with active methodological communities.
The word doing the most work on the slide is independent. A random split of a single cohort is internal validation. The same institution in a later time window is temporal validation. A different institution, a different era, a different referral pattern, and a different assay platform is external validation. Those are three quite different strengths of claim, and current reporting guidance for prediction models asks authors to state plainly which data were used for development and which for evaluation, precisely because the deck version of the sentence does not distinguish them.4
Then ask the follow-up almost nobody asks: how many model versions have been evaluated against that held-out set? A test set consulted once is a test set. A test set consulted after each of forty attempts is a validation set wearing the wrong name, and the last number reported from it is the maximum of forty draws.
The claim as it usually appears, and the three quantities that determine what it is worth. None of the follow-up questions is adversarial; a team that has done the work answers all of them quickly, and usually names its own weakest link first.
03The endpoint: what was measured, and was it chosen in advance?
Ask what outcome was actually recorded, how it was defined, and when it was measured. Recurrence detected on imaging is a different endpoint from biopsy-confirmed recurrence. Progression-free survival is a different endpoint from overall survival, and a composite of several events is a different endpoint again — usually a more easily reached one.
Then ask whether that endpoint was chosen before the analysis, and where the record of that choice lives. This is the question people find rudest and it is the one with the best evidence behind it. A prospective audit of every trial published over six weeks in the five highest-impact general medical journals assessed 67 trials; only 9 reported their outcomes correctly. Of 818 prespecified secondary outcomes, 55.1% were reported at all, and 365 novel outcomes appeared without any declaration that they had been added — an average of 5.4 undeclared new outcomes per trial.5 Those were registered, protocolled, regulated studies in the most heavily scrutinized journals in medicine. A model developed retrospectively on banked specimens operates under none of that discipline, and the honest version of the answer is often “we tried several and this one worked” — which is fine, as long as it is said, and as long as the next cohort is where the claim gets tested.
Two distinctions are worth forcing into the open before you leave the room. The first is prognostic versus predictive: identifying who will do badly is not the same as identifying who will respond to a particular drug, and only the second supports a companion-diagnostic story. The second is incremental value. A marker with an AUC of 0.78 on its own sounds impressive; if adding it to a standard clinical model moves that model from 0.81 to 0.82, it is a publication rather than a product. Ask for the performance of the baseline model without the marker. If nobody has computed it, that absence is itself the answer.
What good looks like
None of this is a gotcha. Every one of these questions has a legitimate answer, and small events counts, convenience cohorts, and exploratory endpoints are entirely normal at an early stage — they only become a problem when they are described as something else. A team that has done the work will answer all three in about five minutes, and will typically name its own weakest link before you find it. The tell is not a bad answer. It is a vague one, a defensive one, or the discovery that the only person who could answer it is not on the call.
The reason to spend the twenty minutes here rather than elsewhere is the asymmetry. None of these mechanisms makes a good model look worse than it is. Undercounted events, an early-leaking split, a search space nobody mentions, an endpoint chosen after the fact — all of them push the number up. That is why the questions are worth asking even when everyone in the room is honest, which they usually are. The inflation does not require anyone to intend it.
Three questions to ask
- How many outcome events were in the validation set, and how many candidate features were screened before the final model was chosen?
- At what point in the analysis was the held-out set separated — and how many model versions have been evaluated against it since?
- What endpoint was measured, was it prespecified in a record I can see, and how much does the model add over the information already available without it?
References
- Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441. doi:10.1136/bmj.m441
- Venet D, Dumont JE, Detours V. Most random gene expression signatures are significantly associated with breast cancer outcome. PLoS Comput Biol. 2011;7(10):e1002240. doi:10.1371/journal.pcbi.1002240
- Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804
- Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378
- Goldacre B, Drysdale H, Dale A, et al. COMPare: a prospective cohort study correcting and monitoring 58 misreported trials in real time. Trials. 2019;20:118. doi:10.1186/s13063-019-3173-2

