Skip to content

The transcriptome and the proteome describe what a cell can do: which genes are read, which proteins are present. The metabolome describes what it does — the small molecules consumed, transformed and produced at the moment of sampling. Being the product of enzymatic activity rather than its program, it is often presented as the layer closest to the phenotype.

It is also the hardest to read, for a reason that has no equivalent in the other layers. This article sets out what makes the metabolome distinctive, the bottleneck that shapes the whole field, and the precautions it imposes — including when this layer is to be integrated with others.

What makes the metabolome distinctive

A transcriptome is a set of molecules of the same chemical nature, measured by a single method. A metabolome is not.

It spans considerable chemical diversity (lipids, sugars, amino acids, steroids and many other families) across a concentration range extending from millimolar to femtomolar. The consequence is direct and rarely stressed: no analytical method in existence today can detect, identify and quantify all the species present [1]. Every metabolomic study therefore observes a window, never the whole metabolome, and the choice of method determines that window.

A sampling constraint is added to this. Metabolism continues until it is stopped, so the concentrations measured may reflect handling conditions as much as the biological state of interest — a concern we meet again, in another form, in relation to sample preparation for single-cell sequencing: what happens before the measurement determines what the measurement can show.

Finally, and this point governs interpretation, the metabolome is not purely endogenous. It brings together the molecules of primary metabolism and secondary signaling metabolites, but also molecules arising from lifestyle and environmental exposure, the exposome, and those produced by the microbiome associated with the organism studied [1]. A metabolic difference between two groups of subjects may therefore reflect their diet or their gut flora as much as the disease one believes one is observing.

Targeted or untargeted: two different questions

Two approaches coexist, and conflating them leads to misplaced expectations.

The targeted approach detects and quantifies a set of known metabolites [1]. It gives reliable measurements, comparable across studies, with standards and calibration curves. Its scope is restricted by construction: you find only what you decided to look for.

The untargeted approach aims instead to detect as many compounds as possible, including unknown chemical species [1]. It permits discovery, but at the cost of an uncertainty that is the subject of the rest of this article. It also serves, and this is less often said, to compare two states without seeking to name what distinguishes them — a use we return to below.

The choice therefore follows the question: testing a hypothesis about identified metabolites, or exploring without a prior hypothesis. These two approaches differ in cost, in reliability, and in the kind of conclusion they allow.

The bottleneck: identification

Here is what truly sets metabolomics apart from the other molecular layers.

DNA, RNA and proteins are polymers built from a fixed alphabet — four bases, twenty amino acids. Identifying them amounts to reading a sequence, an operation sequencing has industrialized. Metabolites, by contrast, follow no molecular alphabet. Each is a particular chemical molecule, and identifying it requires structural elucidation — work of an altogether different nature, far longer and with a high failure rate [1].

This is no marginal difficulty: it is the central bottleneck of the field, and it is not about to be lifted.

The orders of magnitude give the measure of the problem, even though estimates differ. A reference review holds that fewer than 2 to 10% of the compounds detected in an untargeted approach can be reliably annotated, the chemical identity of most detected species remaining unknown [1]. A recent analysis of sixty-one public blood metabolomics datasets concerns a narrower set: among the one to two thousand high-confidence compounds each dataset detects, more than half remain unknown [4].

Whichever figure is taken, the conclusion does not change: in an untargeted experiment, most detected compounds have no reliable identity. A distinction that matters for what follows — some have no name at all, others carry one that is not secure, and these are not the same problems.

Why identification is the bottleneck: the other molecular layers match a finite repertoire; the metabolome requires structural elucidation.
Figure 1. Why identification is the bottleneck: the other molecular layers match a finite repertoire; the metabolome requires structural elucidation.

Confidence levels, and how they are actually used

The community responded to this problem early. Minimum reporting standards were proposed as far back as 2007, defining four confidence levels from the formally identified compound to the unidentified compound [2] — a scheme extended to five levels in 2014 by independent groups [3].

These levels answer a precise question: when a paper announces an “identified metabolite”, what is that identity based on? Not all identity evidence is equal, and the declared level indicates which kind was actually obtained.

The observation accompanying these recommendations is nevertheless severe: despite continued promotion within the community, their application remains questionable [3]. A metabolite presented as identified in a publication is not necessarily accompanied by the level that corresponds to it.

Hence a simple practical requirement, which holds as much for reading the literature as for commissioning a study: ask for the level, not just the name.

The errors that go unnoticed

Identification is not only a bottleneck. It is also a step at which critical errors can occur and pass unnoticed [3].

Two categories of error are documented. Some proposed identities are biologically implausible — the compound has no reason to be present in the organism or tissue studied. Others are incompatible with the chromatographic data, and therefore with the physicochemical properties of the proposed molecule [3]: a highly hydrophilic compound is not eluted at the retention time of a lipophilic one.

What makes these errors formidable lies in their shape. A metabolite name slots straight into a known metabolic pathway; the pathway makes sense in light of the disease under study; and the story holds together. Nothing in the data signals that the identity was wrong. It is the same structure we meet with the poorly validated antibody in chromatin analysis, with visual estimation in digital pathology, or with spectral unmixing in cytometry: a false result can be perfectly plausible.

The safeguards are known and can be planned. Check the consistency between the observed retention time and the expected physicochemical properties. Test the proposed identity against what the organism studied is liable to produce. And where a conclusion rests on a particular metabolite, confirm it by comparison with a reference compound — which costs time, and is worth doing only for the compounds the conclusion actually depends on.

The share of compounds without secure identity, according to two estimates covering different sets.
Figure 2. The share of compounds without secure identity, according to two estimates covering different sets.

What this changes for multi-omics integration

One consequence is worth anticipating when metabolomics is considered as one layer among others.

Integrating layers presupposes being able to map them onto one another: linking a change in gene expression to the corresponding protein, then to the metabolite that protein transforms. But you cannot link a change in expression to an anonymous metabolite, whereas you can link it perfectly well to a misnamed one — to the wrong gene. The first case leaves a hole in the integration, the second propagates an error through it, which is worse. If most compounds have no reliable identity, the share of the layer that is genuinely integrable shrinks accordingly.

This does not disqualify the approach, but it changes what can be expected of it. Unannotated signals are not noise: the analysis of sixty-one datasets mentioned above establishes that in-source fragments account for fewer than 10% of features, and that most abundant features show identifiable ion patterns [4]. These are therefore, for the most part, real compounds that cannot yet be named — not yet usable in mechanistic reasoning, but real. The practical consequence is to size the integrative ambition on the annotated fraction rather than on the number of features detected, a question we develop in relation to multi-omics integration.

A case: are two products equivalent?

Everything above might suggest that an untargeted approach, most of whose signals have no secure identity, allows little to be concluded. Work Inovarion contributed to shows the opposite, and the reason is instructive.

The context: growing demand for red blood cells and donor scarcity make their in vitro production a medical priority, particularly for patients with very rare blood groups. A validation question then arises — does the product obtained behave like the native one?

This is a question that reading the genome or the transcriptome alone does not settle: what we want to know is whether the cells behave in the same way, not whether they carry the same program. The metabolome answers it directly. Using a metabolomic approach, the study establishes that reticulocytes, whether native or produced in vitro, show similar metabolomic signatures, and almost identical ones after maturation into red blood cells. That result strengthens the robustness of the production protocol evaluated [5].

The methodological point is this, and it qualifies what precedes — including the integration constraint set out in the previous section: comparing does not require identifying. To find out whether two profiles resemble one another, signals measured under the same conditions may suffice, whether or not they carry a name — provided they are correctly matched from one sample to the next. It is mechanistic interpretation that requires identities — saying which pathway differs, not whether something differs.

The identification bottleneck therefore does not weigh uniformly. It is decisive for a mechanistic question, secondary for a comparison, which is why it constrains multi-omics integration — which presupposes linking names — without constraining a validation by profile comparison. Knowing which case you are in determines what should be demanded of the study.

How Inovarion can support you

Metabolomics is among the multi-omics approaches Inovarion implements, alongside genomics, transcriptomics and proteomics, and our teams have contributed to work in which it served to validate a cell product. On projects of this kind the useful questions arise early: does a targeted or untargeted approach answer the question? Is the aim to compare profiles or to establish a mechanism — since the two do not demand the same level of identification? And if the metabolic layer is to be crossed with others, what fraction will actually be usable? It is on these trade-offs, more than on the production of signals, that the value of a study is decided.

Contact an expert →

Publications

Field references

  1. Monge ME, Dodds JN, Baker ES, Edison AS, Fernández FM. Challenges in identifying the dark molecules of life. Annual Review of Analytical Chemistry, 2019;12(1):177-199. DOI
  2. Sumner LW, Amberg A, Barrett D, et al. Proposed minimum reporting standards for chemical analysis. Metabolomics, 2007;3(3):211-221.
  3. Theodoridis G, Gika H, Raftery D, Goodacre R, Plumb RS, Wilson ID. Ensuring fact-based metabolite identification in liquid chromatography–mass spectrometry-based metabolomics. Analytical Chemistry, 2023;95(8). DOI
  4. Chi Y, Mitchell JM, Zheng S, Thapa M, Li S. Assessing the metabolomics “dark matter” by a detectable khipu model. Metabolomics, 2026;22(3). DOI

Inovarion contribution

  1. Darghouth D, Giarratana MC, Oliveira L, et al. Bio-engineered and native red blood cells from cord blood exhibit the same metabolomic profile. Haematologica, 2016;101(6):e220-e222. PubMed

updated July 2026