Reading a Bioinformatics Scope of Work: What’s In, What’s Out, and What Gets Billed Later
Reading a Bioinformatics Scope of Work: What’s In, What’s Out, and What Gets Billed Later
The line items describe the cheap part of the project. The expensive part is usually hiding inside a single verb — and the boundaries that end up mattering are rarely the ones anyone argues about at signature.
A scope of work is not really a description of the work. It is an agreement about who absorbs the surprises. Every boundary the document leaves unstated is a small bet on whose budget will cover something neither side has thought about yet — and because both parties are usually optimistic when they sign, the unstated boundaries tend to be the ones that decide how the project ends.
This is not about whether to build a bioinformatics team or contract one out; that decision has been made by the time the paperwork arrives. It is about the paperwork — how to read it in twenty minutes and find the three or four places where the real money lives. It applies to documents this firm writes as much as to anyone else’s.
01What’s in scope: read the nouns, not the verbs
Verbs in a scope of work are almost always unbounded. Analyze. Process. Interpret. Support. Optimize. None of them contains a quantity, which means none of them can be exceeded, which means none of them can be enforced by either side. Nouns are different, because nouns can be counted. Four are worth finding before anything else: samples, iterations, deliverable artifacts, and meetings.
Of those, iterations is the one most often missing and the one that most reliably consumes the budget. “Differential expression analysis” reads on the page as a single event. In practice the first result generates questions, and answering them means running it again — dropping an outlier, adding a covariate, restricting to a subgroup, switching to the comparison the team actually cares about now that they have seen the first one. Each of those is a full re-run plus the judgment that goes with it. A scope that says up to three analysis rounds, with a fourth available at the hourly rate tells you more about how the engagement will go than three extra paragraphs of technical detail.
The other count worth pinning down is what a sample is for billing purposes. If a specimen fails quality control, does it count against the total, get replaced at no charge, or trigger a conversation? All three are defensible. Usually none of them is written down.
02What’s out of scope: three silences
Getting the data into a state where the work can start. The economics of sequencing projects moved a long time ago. As the cost of generating data collapsed, the dominant cost shifted downstream into data management and analysis — the problem became knowledge generation rather than data generation.1 The everyday form of that shift is unglamorous: sample sheets that disagree with the manifest, phenotype columns recorded three different ways across two years, batch assignments that exist only in a departed technician’s memory. Reconciling all of it is real, skilled, slow work. It is almost never a line item, and it is the single most common trigger for the first change order. Even the NIH treats it as a distinct budgetable activity: curating data and developing supporting documentation is a named allowable cost category, explicitly separate from the routine conduct of research.2 If a funder finds the distinction worth drawing, a contract should too.
Delivering the artifact versus answering the question. A variant file, a count matrix, a summary report — these are objects. The question that prompted the project is not answered by an object; it is answered by someone deciding what the object means and being willing to say so. Those are different scopes with different prices, and the gap between them widens over time: a figure redrawn for a manuscript, a response to reviewer three, a slide for a partnering meeting, an answer to a diligence question eight months after the invoice cleared. Find out which of those the document covers. “Interpretation” on its own does not tell you.
A services section of this shape is normal and is compatible with a well-run engagement. But the two line items carrying most of the labor are verbs without counts, and the assumptions paragraph transfers three open-ended risks — data condition, storage duration, and data movement — in eleven words.
Validation, as distinct from execution. Running a pipeline and validating one for regulated use are different projects with different price tags, and the vocabulary hides it. The joint AMP/CAP recommendations are specific here: a clinical laboratory must validate its own pipeline; that validation must be appropriate to the intended clinical use, the specimen types, and the variant types of the test; and supplemental validation is required whenever a significant change is made to any component.3 So a promise to use an already-validated pipeline holds only while nothing about your assay disturbs it. If your specimen type, your panel content, or your reporting threshold sits outside the existing validation, revalidation is not an adjustment. It is a scope event, and it should be named as one before it happens rather than discovered afterward.
03What gets billed later: the items with a clock
Storage and data movement. Usually passed through at cost, usually unbounded in duration. Three sub-questions settle it: how long is data retained, who pays to move it out, and what happens on the day the contract ends. The NIH’s own rule on this is instructive, because it makes the mismatch explicit: costs for data management and sharing must be incurred during the award period, even for data and metadata that the recipient is obliged to preserve and share after the award period closes.2 The obligation outlives the budget, which means it has to be paid for in advance or it will not be paid for at all.
Reanalysis. An analysis is dated the moment it is delivered, because the knowledge base underneath it keeps moving while the dataset stands still. In a meta-analysis of rare disease sequencing, reanalysis of exome data anywhere from one month to 3.4 years after an initial negative result produced additional diagnoses in 1% to 16% of cases.4 None of that is a defect in the first analysis. It is what a fixed dataset does when the annotations, the gene–disease associations, and the classification evidence around it change. Whether revisiting the data is a fresh engagement or a retained obligation is a legitimate choice either way — but it should be a choice, made once, in writing.
Handoff and reproducibility. At the end of the engagement, do you own the outputs or the ability to regenerate them? These are not close to the same thing. When one group attempted to execute roughly 864,000 computational notebooks published on GitHub, 24.1% ran without error and only 4.0% reproduced the results their authors had recorded in them.5 That is code written by people who meant it to be shared. Code that was never meant to leave one machine does worse. If the scope does not name the environment — container image, workflow version, parameter file, reference build and annotation releases, all pinned — what you are buying is a set of numbers rather than a capability, and the difference becomes visible the first time someone asks you to run it again.
What a good scope of work actually looks like
Not longer. The best ones are often shorter than the worst, because they spend their words on counts and boundaries rather than on descriptions of standard methods any reader could have assumed. The countable things are counted, the few places where the project could plausibly double are named with a rule for what happens if they do, and nothing important is left to be settled by whoever is more reluctant to raise it.
It is worth saying plainly that this cuts in both directions. A vendor who leaves everything vague is not being generous; vagueness is a margin against work they cannot see from here, and it is a perfectly rational thing for them to want. But the same silence that permits a change order later also permits real cost to be absorbed quietly and resentfully on either side, which is how otherwise good working relationships end. The reason to ask is not suspicion. It is that both parties do better when the uncertainty is allocated deliberately rather than discovered.
Three questions to ask
- How many analysis iterations does this price cover, and what counts as a new one rather than a refinement of the last?
- Which of these am I buying — the file, the interpretation, or the defense of the interpretation to a reviewer, a regulator, or a diligence team six months from now?
- When this engagement ends, what would someone else need in order to rerun the analysis and get the same numbers — and is that in the deliverables list?
References
- Sboner A, Mu XJ, Greenbaum D, Auerbach RK, Gerstein MB. The real cost of sequencing: higher than you think! Genome Biol. 2011;12(8):125. doi:10.1186/gb-2011-12-8-125
- National Institutes of Health. Supplemental Information to the NIH Policy for Data Management and Sharing: Allowable Costs for Data Management and Sharing (NOT-OD-21-015); and Budgeting for Data Management and Sharing. Effective 25 January 2023. grants.nih.gov
- Roy S, Coldren C, Karunamurthy A, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. doi:10.1016/j.jmoldx.2017.11.003
- Chung CCY, Hue SPY, Ng NYT, et al. Meta-analysis of the diagnostic and clinical utility of exome and genome sequencing in pediatric and adult patients with rare diseases across diverse populations. Genet Med. 2023;25(9):100896. gimjournal.org
- Pimentel JF, Murta L, Braganholo V, Freire J. A Large-Scale Study About Quality and Reproducibility of Jupyter Notebooks. Proc IEEE/ACM 16th Int Conf Mining Software Repositories (MSR). 2019:507–517. leomurta.github.io

