Metabolite Annotation Confidence: Why a Named Compound Is Usually a Match, Not an Identification

Metabolite Annotation Confidence: Why a Named Compound Is Usually a Match, Not an Identification — Zetobit
Zetobit.
BIOINFORMATICS INSIGHT SERIES Metabolite Annotation Confidence Why a named compound is usually a match, not an identification THE SAME NAME IN THE SAME TABLE L-carnitine standard, same instrument, RT + MS/MS Level 1 L-carnitine library spectrum match, no standard run Level 2 L-carnitine an acylcarnitine — class, not structure Level 3 The level column is the one most often omitted. Three different claims. One identical name. ZETOBIT. Kanna Nandakumar, PhD

Bioinformatics Insight Series

Metabolite Annotation Confidence: Why a Named Compound Is Usually a Match, Not an Identification

A metabolite name in a results table can mean a compound confirmed against an authentic standard on the same instrument, or a library spectrum that happened to resemble the signal. The Metabolomics Standards Initiative created confidence levels precisely because those are different claims — and the level is the part that usually goes missing.

A metabolomics results table looks like the least ambiguous output in omics. Not a p-value, not a proportion, not a score — a chemical name. Citrate. Kynurenine. Sphingosine-1-phosphate. Compounds with structures, CAS numbers, and known positions in pathways.

What a mass spectrometer actually produces is a list of features: a mass-to-charge ratio, a retention time, an intensity, and sometimes a fragmentation spectrum. Nothing in that list is a name. The name is assigned afterward by matching the feature against reference data, and the strength of that match varies enormously between rows of the same table.

The field has a framework for expressing this. It is largely absent from the tables themselves.

Four levels, and what separates them

The Chemical Analysis Working Group of the Metabolomics Standards Initiative published a reporting scheme in 2007 that has since been widely adopted or endorsed by major journals in the field. Four levels, in descending order of evidence:

Level 1 — confidently identified compound. The feature is matched to an authentic chemical standard analyzed in the same laboratory, on the same platform, under the same conditions. Exact mass, MS/MS fragmentation, and retention time all align. This is the only level at which "identified" means what a non-specialist would assume it means.

Level 2 — putatively annotated compound. A high-quality match to a reference spectrum from a public or commercial spectral library, without an authentic standard run in parallel. The evidence points to a defined structure, but the possibility of a closely related isomer not distinguishable by the available MS/MS data cannot be entirely ruled out. This is described as the most common confidence level achieved in untargeted metabolomics.

Level 3 — putatively characterized compound class. The evidence supports a chemical family but not a specific structure: a triacylglycerol, a sulfated steroid, a dipeptide containing leucine or isoleucine. Diagnostic neutral losses or fragment ions place the compound in a class. This is commonly the highest confidence achievable for lipids and for compounds poorly represented in spectral libraries.

Level 4 — unknown molecular feature. Exact mass and isotope pattern only. Reliably detected, reliably quantified, possibly biologically important, chemically anonymous.

The distinction between Levels 1 and 2 is the one that matters most in practice and travels least well. Both produce a name. Only one of them was verified against the physical compound on the instrument that made the measurement.

Level 2 is a statement about a spectral library. Level 1 is a statement about the molecule.

Why most annotations cannot be Level 1

This is not a discipline problem. It is a supply problem.

Level 1 requires an authentic chemical standard, and the commercial availability of such standards covers a tiny fraction of the metabolome — a metabolome estimated to encompass hundreds of thousands of unique structures. For a novel or rare metabolite, a purchasable standard may simply not exist. Even where one does, running it on the same platform under matched conditions for every reported compound is a substantial cost that scales with the size of the results table.

So Level 2 is not a shortcut taken by careless labs. It is the ceiling imposed by chemistry supply chains and budgets, and it is the honest description of most untargeted results. The failure is not working at Level 2. The failure is reporting Level 2 results in language that reads as Level 1.

The isomer problem is why the distinction has teeth

If the risk of a Level 2 annotation were simply "somewhat less certain," the omission would be a stylistic matter. It isn't, because of what the uncertainty actually consists of.

Isomers — different compounds sharing a molecular formula — are ubiquitous in biology, and they often have identical exact masses with very similar, sometimes indistinguishable, MS/MS spectra. Glucose and fructose. Leucine and isoleucine. Positional isomers of lipids that differ in where a double bond sits. MS data alone frequently cannot resolve them.

This means a wrong Level 2 annotation is rarely a random wrong answer. It is usually a structurally adjacent wrong answer — a compound with the same formula, similar fragmentation, and often a completely different biological role. A metabolite reported as one member of an isomer pair, when the sample contained the other, produces a pathway analysis that is internally coherent and pointed at the wrong biology.

The same structural adjacency defeats the standard remedy from other fields. Proteomics controls annotation error with target-decoy false discovery rate estimation, generating decoys by reversing peptide sequences. In metabolomics, due to the structural diversity of small molecules and the prominence of structural isomers, creating plausible decoys is not as straightforward, and FDR estimation is not yet widely adopted. Methods have been developed — fragmentation-tree-based decoys, entropy-based approaches, and others — but they are not routine, and the field has largely operated without the error-rate control that genomics, transcriptomics, and proteomics treat as standard.

ONE FEATURE FROM THE INSTRUMENT m/z 162.1125 · RT 3.42 min · MS/MS acquired Level 1 authentic standard, same instrument · mass + MS/MS + RT a name Level 2 library spectrum match · isomer not excluded a name Level 3 compound class only · diagnostic fragment or neutral loss a name Level 4 exact mass and isotope pattern only no name Levels 1, 2 and 3 all yield something that looks like an identification in a table. Only Level 1 was verified against the physical compound on the instrument that made the measurement. Feature values shown are illustrative.

Figure 1. The MSI framework exists because a single instrument feature can be reported at four very different evidentiary standards. Three of the four produce a chemical name that reads identically in a results table. Without the level attached, a reader cannot tell which claim is being made — and the difference between Level 1 and Level 2 is the difference between a verified molecule and a plausible library match.

The levels exist. They are not consistently used.

This is the part that turns a technical nuance into a reporting problem.

A 2024 analysis put it directly: reporting qualitative MSI confidence levels is infrequently and inconsistently used by members of the metabolomics community — attributing this partly to the fact that assigning confidence scores remains a subjective process for most data reporters, and that recipients of such data lack sufficient information or tools to independently verify identifications. The same work notes a compounding problem: many reports include only chemical names, but not chemical or structure identifiers such as PubChem CIDs or InChIs, and chemical names can be highly ambiguous and misleading for data consumers.

Broader reporting compliance shows the same pattern. An evaluation of 399 public datasets across four major metabolomics repositories against MSI reporting standards found that none of the reporting standards were complied with in every publicly available study, with adherence rates ranging from 0 to 97% depending on the standard.

So the framework is a decade and a half old, endorsed by journals, and widely known — and the field's own assessment is that it is applied unevenly. The consequence is a literature in which the same word, "identified," covers claims that differ by orders of magnitude in evidentiary weight.

Parameters move the answer

One more source of variability deserves attention, because it is invisible in a results table and large.

Spectral matching depends on scoring thresholds, and those thresholds are set by the analyst or inherited from a tool's defaults. Work applying FDR estimation across 70 public metabolomics datasets found that spectral matching settings need to be adjusted for each project, and that adjusting scoring parameters and thresholds changed the number of annotations by an average of +139% relative to a default parameter set, with a range spanning −92% to +5705%.

That range is the point. Two analysts processing identical raw data with the same tool and different — both defensible — threshold choices can produce annotation lists that differ by more than an order of magnitude in length. Nothing about the sample changed.

Table 1. What a metabolite name in a results table can mean. The distinction is the whole content of the MSI framework, and it is what an unqualified name discards.
Level Evidence behind the name What it licenses
Level 1 Authentic standard on the same platform; mass, MS/MS and RT all match Pathway mapping, mechanistic claims, biomarker nomination
Level 2 High-quality library spectrum match; no in-house standard Structural hypothesis; isomer ambiguity remains open
Level 3 Compound class from diagnostic fragments or neutral losses Class-level biological context only, stated as such
Level 4 Exact mass and isotope pattern Quantitative comparison; a target for follow-up elucidation

What to change

Carry the level through to every reported compound

The level belongs in the results table as a column, not in a methods paragraph describing the pipeline in general. Different rows in the same table will legitimately sit at different levels, and collapsing them loses exactly the information a reader needs.

Report structure identifiers, not just names

InChIKeys or PubChem CIDs alongside names remove the ambiguity that plain chemical names carry and make cross-study comparison possible. This is a small change with outsized effect on data reuse — and it is one of the specific gaps the field has flagged in its own reporting.

Record the library and the matching parameters

Which spectral library, which version, which score metric, which threshold. Given that parameter choices can change annotation counts by orders of magnitude, an annotation without its matching configuration is not reproducible even by the same lab.

Reserve standards for the compounds that carry the claim

Level 1 confirmation of every annotation is not achievable. Level 1 confirmation of the handful of metabolites a paper's conclusions actually rest on usually is. Targeting authentic standards at the load-bearing compounds is the highest-value use of a limited budget.

Weight downstream interpretation by level

A pathway analysis built on Level 1 identifications carries far more weight than one built on Level 3 annotations. If the enrichment result depends on compounds annotated at Level 2 or 3, that dependency should be visible in the interpretation rather than buried.

The short version

A metabolomics results table reports chemical names, but the evidence behind those names spans from a compound verified against a physical standard to a library spectrum that resembled the signal. The MSI levels exist to make that difference visible, and the field's own assessment is that they are applied infrequently and inconsistently. Because isomers are ubiquitous, a wrong annotation is usually a structurally adjacent compound with a different biological role — a coherent answer about the wrong molecule.

The same shape, a different instrument

Readers of this series will recognize the structure. A metagenomic classifier assigns every read the closest label its database contains and has no way to say "nothing here matched well." A spectral matching pipeline assigns every feature the closest compound its library contains, with the same absence of a null option. In both cases the output is a list of names that reads as a list of findings.

Metabolomics differs in one respect that is genuinely to its credit: the field built an explicit vocabulary for the uncertainty, published it, and got journals to endorse it. The framework is not missing. It is optional in practice, and the part most often dropped is the part that distinguishes a measurement from a resemblance.

Adding a column costs nothing. Leaving it out is what turns a careful analysis into a table of names indistinguishable from certainty.

Zetobit builds and validates multi-omics pipelines, including metabolomics workflows with MSI-level tracking, structure identifiers carried to reporting, and documented spectral matching parameters. If you are standing up untargeted metabolomics or reconciling annotations across platforms and studies, we're happy to talk.

References

  1. Sumner LW, Amberg A, Barrett D, et al. Proposed minimum reporting standards for chemical analysis: Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI). Metabolomics 3:211–221 (2007). doi:10.1007/s11306-007-0082-2.
  2. Creek DJ, Dunn WB, Fiehn O, et al. Metabolite identification: are you sure? And how do your peers gauge your confidence? Metabolomics 10:350–353 (2014). doi:10.1007/s11306-014-0656-8.
  3. Spicer RA, Salek R, Steinbeck C. Compliance with minimum information guidelines in public metabolomics repositories. Scientific Data 4:170137 (2017). doi:10.1038/sdata.2017.137.
  4. Metz TO, Fiehn O, Wishart DS, et al. Introducing 'identification probability' for automated and transferable assessment of metabolite identification confidence in metabolomics and related studies. bioRxiv preprint (2024). doi:10.1101/2024.07.30.605945. Preprint — check for the peer-reviewed version before citing.
  5. Scheubert K, Hufsky F, Petras D, et al. Significance estimation for large scale metabolomics annotations by spectral matching. Nature Communications 8:1494 (2017). doi:10.1038/s41467-017-01318-5.
  6. Blaženović I, Kind T, Ji J, Fiehn O. Software tools and approaches for compound identification of LC-MS/MS data in metabolomics. Metabolites 8(2):31 (2018). doi:10.3390/metabo8020031.
  7. Salek RM, Neumann S, Schober D, et al. A decade after the metabolomics standards initiative it's time for a revision. Scientific Data 4:170138 (2017). doi:10.1038/sdata.2017.138.
  8. Metabolite identification in LC-MS metabolomics: identification principles and confidence levels. MetwareBio technical guide, accessed July 2026. Vendor-authored overview; used for the level definitions, which follow the MSI framework in reference 1.
Next
Next

Spatial Deconvolution: Why a Cell Type Map Shows You the Reference, Not the Tissue