+1(781)975-1541
support-global@metwarebio.com

Metabolomics Databases and Pathway Analysis: A Practical Guide to KEGG, HMDB, MetaboAnalyst, and Related Resources

Metabolomics results rarely arrive with one stable name, one identifier, and one unambiguous biological meaning. A compound may have several synonyms, an accurate mass may match multiple structures, and significant metabolites must still be connected to reactions, tissues, organisms, and pathways. No single resource resolves all of these tasks. KEGG, HMDB, MetaboAnalyst, and MetaboLights are not interchangeable pathway databases. KEGG organizes reactions and pathways; HMDB provides human-centered metabolite knowledge; MetaboAnalyst supports statistical and pathway analysis; and MetaboLights stores public studies and metadata. This guide explains how to select and combine these resources without overstating database matches or enrichment results.

1. Core Types of Metabolomics Databases and Analysis Resources

Metabolomics resources can be grouped by the question they answer. Metabolite-centered databases such as HMDB and LIPID MAPS organize names, structures, spectra, biological locations, and cross-database identifiers; they support annotation but do not confirm compound identity (Alseekh et al., 2021). Pathway knowledgebases such as KEGG, Reactome, and MetaCyc connect metabolites with reactions, enzymes, genes, modules, and pathways. Differences in curation, pathway granularity, and organism coverage can therefore change mapping and enrichment results (Kanehisa et al., 2025; Ragueneau et al., 2026; Caspi et al., 2020).

Analysis platforms and repositories serve different roles. MetaboAnalyst accepts processed metabolite tables, metabolite lists, and selected raw-data formats for normalization, statistical testing, enrichment, topology, network, and multi-omics analysis. MetaboLights instead archives metabolomics studies together with raw and processed data, experimental design, sample information, and other metadata. Repositories make studies discoverable and reusable; analysis platforms turn user-supplied data into statistical and biological results. Public datasets may still require format harmonization, identifier review, and metadata checks before reanalysis, but these preparation steps do not make a repository an analysis platform (Pang et al., 2024; Yurekten et al., 2024).

2. KEGG Pathway Database for Metabolomics and Multi-Omics

KEGG is most useful when a metabolite needs to be placed within a connected biochemical system. Its advantage is not merely the availability of pathway diagrams, but the linked organization of compounds, reactions, enzymes, orthologs, modules, genes, and organism-specific pathways.

Official website: https://www.kegg.jp/

2.1 How KEGG Connects Compounds, Reactions, Enzymes, and Genes

KEGG COMPOUND entries link chemical entities to REACTION records, which describe substrates and products and connect to ENZYME entries. KEGG ORTHOLOGY assigns conserved functional identifiers to genes and proteins, while MODULE and PATHWAY objects organize these elements into larger functional units and molecular networks. Together, these links provide an interpretation chain from metabolite to reaction, enzyme or gene, module, and pathway. Reference pathway maps represent generalized knowledge, whereas organism-specific maps project relevant genes and reactions onto a selected species (Kanehisa et al., 2025).

KEGG Mapper applies this structure to practical mapping. Compound, gene, and KEGG Orthology (KO) identifiers can be mapped and colored on pathway diagrams, while reconstruction functions help identify pathway or module content represented by an input list. For metabolomics, this reaction-level placement is more informative than a pathway name alone because mapped metabolites can be examined in relation to neighboring substrates, products, reactions, and enzyme nodes.

2.2 KEGG Pathway Mapping for Metabolomics Interpretation

The main contribution of KEGG to metabolomics interpretation is reaction-level context. A changed metabolite can be positioned within a local reaction neighborhood, and transcriptomic or proteomic data can be overlaid through genes, KOs, or enzymes. Agreement among metabolite abundance, enzyme-related protein abundance, and gene expression can strengthen a pathway hypothesis, whereas disagreement may point to regulation, compartmentation, timing differences, or incomplete coverage.

KEGG modules are also useful when a broad pathway is too large to interpret as a single unit. Smaller functional blocks can help distinguish a localized biosynthetic or degradation step from a general pathway label. This reduces the temptation to interpret every mapped member of a large pathway as one coordinated biological response.

2.3 KEGG Coverage, Interpretation, and Licensing Limitations

A KEGG map is a knowledge model, not a measurement of pathway activity. The presence of a reaction in the database does not show that it occurred in a particular sample, and metabolite abundance does not directly measure flux. Highly connected metabolites such as ATP, glutamate, or pyruvate may appear in many pathways and can dominate enrichment unless the matched members are inspected. Incorrect species selection or identifier conversion can also create false or missing assignments.

KEGG access and reuse conditions depend on the use case. Academic users may freely use the KEGG website, whereas non-academic use requires a commercial license. Current terms should be reviewed before service provision, automated retrieval, redistribution, or reuse in software or commercial services (KEGG, 2024).

3. HMDB for Human Metabolite Annotation and Biological Context

HMDB functions as a human-centered metabolite knowledgebase. It is most useful when a candidate compound needs standardized chemical information and human biological context rather than a pathway map alone.

Official website: https://hmdb.ca/

3.1 HMDB MetaboCards: Chemical, Spectral, and Biological Information

An HMDB MetaboCard consolidates chemical and biological information, including names, structures, formulas, exact masses, spectra, reported concentrations, tissues and biofluids, disease associations, literature, and database cross-references. HMDB 5.0 also expanded its visualization and search functions, predicted NMR and MS spectra, retention indices, and collision cross-section data (Wishart et al., 2022).

In practice, HMDB is particularly useful for checking whether alternative names refer to the same human metabolite, reviewing whether the compound has been reported in plasma, urine, cerebrospinal fluid, tissue, or another matrix, and examining possible clinical associations. It can also help reconcile identifiers before pathway analysis, because a chemically correct name may still fail to map if the analysis platform expects another identifier type.

3.2 What HMDB Cannot Confirm Without Experimental Evidence

HMDB cannot convert a database candidate into a confirmed structure. Confirmation may require an authentic standard, matched retention time, high-quality MS/MS evidence, or other orthogonal data, depending on the analytical claim. Likewise, a disease or biomarker association recorded in HMDB does not establish causality, diagnostic performance, or relevance in a new cohort. Because HMDB is human-centered, plant, microbial, and broad comparative-metabolism projects require additional organism-appropriate resources.

4. MetaboAnalyst for Statistical and Metabolomics Pathway Analysis

MetaboAnalyst is an analysis platform rather than a reference database. It combines statistical methods with pathway libraries to analyze metabolite data and support biological interpretation.

Official website: https://www.metaboanalyst.ca/

4.1 Core MetaboAnalyst Functions and Use Cases

MetaboAnalyst supports normalization and transformation, univariate and multivariate statistics, biomarker analysis, metabolite set enrichment, pathway enrichment and topology analysis, network analysis, and joint pathway analysis. The platform also integrates raw-data processing, annotation, statistical modeling, and interpretation modules for targeted and untargeted LC-MS studies (Pang et al., 2024).

The appropriate module depends on the input and research question. A concentration table can support group comparison and multivariate exploration, whereas an identified metabolite list can support over-representation analysis or metabolite set enrichment. In pathway analysis, topology metrics weight matched metabolites according to their positions in a pathway network rather than directly measuring effect size or metabolic flux. Paired gene-metabolite inputs can support joint pathway analysis. Selecting a module because it produces a familiar plot is less defensible than matching the method to the data structure.

4.2 Inputs and Parameters That Shape Pathway Analysis

MetaboAnalyst results are highly dependent on mapping and reference choices. Identifier type determines which compounds enter the analysis. Organism selection and pathway library define the biological reference framework. The background set determines the comparison baseline, while enrichment method, topology measure, and multiple-testing correction affect statistical significance and ranking. Ambiguous IDs, isomers, duplicated names, and unmatched compounds reduce coverage or assign evidence to the wrong pathways.

For targeted panels, the measured or otherwise testable metabolite universe is often a more defensible background than every compound in a database. For untargeted peak-level workflows, the relationship between peaks, candidate metabolites, and pathways is even more complex. Benchmarking studies show that annotation strategy and enrichment method can substantially alter the pathways reported, so method choice and assumptions must be documented (Lu et al., 2023).

4.3 How to Interpret MetaboAnalyst Pathway Results

In over-representation analysis, a significant result indicates that matched input metabolites occur in a pathway more frequently than expected under the selected background and mapping rules. Rank- or score-based enrichment methods use different statistical assumptions. In all cases, statistical significance does not by itself prove pathway activation, inhibition, altered enzyme activity, or changed metabolic flux. Pathway impact reflects the selected topology metric; it is not a causal-effect scale. Figure 1 shows a mummichog-based functional-analysis output from MetaboAnalyst 6.0 rather than an over-representation analysis result. Such plots should be interpreted alongside the matched metabolites, effect sizes, change directions, mapping quality, and orthogonal evidence rather than in isolation.

Mummichog-based pathway-level functional analysis in MetaboAnalyst 6.0 using LC-MS metabolomics case study

Figure 1. Example of mummichog-based pathway-level functional analysis in MetaboAnalyst 6.0 using an LC-MS metabolomics case study. Cropped from Figure 2C in Pang et al. (2024), Nucleic Acids Research, under CC BY 4.0; no other changes were made.

5. Complementary Metabolomics Databases and Repositories

KEGG, HMDB, and MetaboAnalyst cover many common tasks, but specialized resources add value when the biological scope, data type, or evidence need changes.

Reactome. This curated, reaction-centered knowledgebase is particularly useful for human biology, mechanistic event sequences, signaling, regulation, and cross-omics interpretation. It can complement KEGG when a human pathway requires more detailed reaction or compartment context (Ragueneau et al., 2026).

MetaCyc. MetaCyc emphasizes experimentally supported metabolic pathways and enzymes across all domains of life. Its coverage and evidence-oriented curation are useful for microbial, plant, comparative-metabolism, and metabolic-engineering questions (Caspi et al., 2020).

LIPID MAPS. LIPID MAPS provides standardized lipid nomenclature, classification, structures, databases, and lipid-focused tools. It is a useful companion for lipid annotation and lipid pathway context, but it does not replace broad polar-metabolite resources (Conroy et al., 2024).

MetaboLights. MetaboLights is a cross-species, cross-technique repository for metabolomics studies, raw experimental data, processed results, and metadata. It supports data reuse, method comparison, and reproducibility rather than pathway computation (Yurekten et al., 2024).

6. How to Choose a Metabolomics Database or Analysis Platform

Choose the resource according to the research question, organism, and evidence stage. Use HMDB to standardize candidate human metabolites and add biospecimen context; KEGG to map reactions and pathways; MetaboAnalyst to perform statistical pathway analysis; and MetaboLights to locate public studies and datasets. Reactome, MetaCyc, and LIPID MAPS add human reaction detail, cross-organism metabolic coverage, and lipid-specific annotation, respectively.

No platform covers the full workflow. Annotation, biological contextualization, pathway mapping, statistical analysis, and validation should remain separate steps, with uncertainty carried forward throughout the process.

Table 1. Metabolomics Databases and Analysis Platforms: Best Uses, Outputs, and Limitations

Resource Resource Type Best Used For Main Output Key Limitation
KEGG Pathway knowledgebase Reaction and pathway mapping; metabolite-gene integration Pathway maps, modules, and reaction links Coverage and interpretation depend on identifiers, species, and licensing conditions
HMDB Human metabolite knowledgebase Candidate metabolite standardization and human biological context MetaboCards, spectra, and cross-references Human-centered; database records are not identification proof
MetaboAnalyst Analysis platform Statistics, enrichment, topology, networks, and joint analysis Statistical and pathway analysis results Strongly dependent on input, mapping, and parameter choices
Reactome Curated pathway knowledgebase Human reaction-level and multi-omics interpretation Curated reactions, pathway hierarchy, and diagrams Primarily human-centered
MetaCyc Metabolic pathway database Experimentally supported metabolism across organisms Curated pathways, reactions, enzymes, and supporting evidence Less focused on clinical and biofluid context
LIPID MAPS Lipid knowledgebase Lipid naming, classification, structural annotation, and pathway context Lipid records, classes, tools, and cross-references Lipid-specific rather than general metabolomics
MetaboLights Study repository Public metabolomics datasets and metadata Raw and processed data, study metadata, and accessions Not a pathway analysis tool

7. Metabolomics Database Workflow: From Annotation to Pathway Interpretation

Effective use of metabolomics resources is sequential rather than platform-centered. Each step should produce a checked output before the information enters the next resource. Figure 2 presents one commonly used analytical framework; the database-specific decisions discussed below are concentrated in metabolite identification, pathway mapping, and biological interpretation.

Typical metabolomics workflow from compound detection and preprocessing to database-assisted annotation statistical and pathway analysis and multi-omics integration

Figure 2. Typical metabolomics workflow from compound detection and preprocessing to database-assisted annotation, statistical and pathway analysis, and multi-omics integration. Reproduced from Figure 1 in Chen et al. (2022), Metabolites, under CC BY 4.0.

7.1 Standardize Metabolite Identities and Record Confidence

Begin with analytical evidence rather than a pathway label. Record the original feature or analyte label, standardized metabolite name, molecular formula, a structure identifier such as InChIKey or SMILES, relevant database IDs, and identification confidence. Preserve the original identifiers during conversion so mapping errors can be traced. Confirmed metabolites should remain distinct from putatively annotated compounds, unresolved isomer groups, or accurate-mass candidates (Alseekh et al., 2021).

7.2 Add Biological Context and Select Reference Resources

Next, use HMDB for human biofluid, tissue, concentration, or disease context; LIPID MAPS for lipid-specific classification; and MetaboLights for comparable studies and metadata. Then select KEGG, Reactome, or MetaCyc according to species, pathway granularity, and whether the interpretation requires metabolite-gene integration, human reaction detail, or cross-organism metabolic evidence.

7.3 Perform Pathway Mapping, Enrichment, and Topology Analysis

Map compounds to reactions and pathways before interpreting enrichment. In MetaboAnalyst or another analysis environment, record identifier type, organism, pathway library, background set, enrichment method, topology measure, multiple-testing correction, and software version. Record unmatched and ambiguous mappings rather than discarding them. A high unmatched rate may indicate identifier problems, an unsuitable pathway library, limited database coverage, or biology outside the selected resource.

7.4 Cross-Validate Pathway Interpretation with Multi-Omics and Experiments

For each prioritized pathway, inspect which metabolites were actually matched, whether changes are directionally coherent, and whether the result is driven by one highly connected metabolite. Gene expression, protein abundance, enzyme activity, or post-translational modification data can provide independent support for the same reaction neighborhood. Targeted quantification and external datasets can assess analytical robustness and reproducibility. Stable-isotope tracing, enzyme-activity assays, genetic or pharmacological perturbation, and other functional experiments provide more direct evidence for metabolic mechanism or flux. Differences among tissues, biofluids, species, disease stages, and sampling times should remain explicit rather than being merged into one generic pathway narrative.

8. Common Pitfalls and Reporting Essentials

Most database-related errors arise from uncertain identities, incorrect identifier mapping, inappropriate organism or background choices, and overinterpretation of enrichment outputs. Database matches should retain confidence labels, original identifiers, and unresolved cases; pathway significance should not be described as activation, causal strength, or flux. For reproducibility, report the database or platform version or access date, organism, identifier type, pathway library, background set, enrichment and topology methods, multiple-testing correction, unmatched-metabolite rate, and key parameters (Alseekh et al., 2021; Pang et al., 2024). Table 2 summarizes the main checks.

Table 2. Common Metabolomics Database and Pathway Analysis Errors and Recommended Checks

Common Error Why It Matters Recommended Check
Database match treated as confirmed identification Inflates confidence and propagates false structures into pathways Report identification confidence; verify with standards, MS/MS, retention time, or other orthogonal evidence, as appropriate
Synonyms, isomers, or identifiers mapped incorrectly Creates false assignments, duplicates, or lost compounds Standardize names, preserve original identifiers, and cross-check mappings
Wrong organism or pathway library Uses biologically inappropriate pathway coverage Document the organism and pathway library before analysis
Undefined or biased background set Can distort enrichment significance Use the measured or otherwise testable metabolite universe when appropriate
Enrichment interpreted as pathway activation Overstates statistical association as mechanism or flux Inspect metabolite direction, effect size, reaction context, and orthogonal evidence
Database versions and analysis parameters not reported Prevents reproducibility and later comparison Record access date or version, methods, multiple-testing correction, settings, and unmatched rate

9. FAQ: Metabolomics Databases and Pathway Analysis

Is HMDB a metabolic pathway database?

HMDB is primarily a human metabolite knowledgebase. It contains pathway links and related biological information, but its central role is metabolite annotation and human context rather than KEGG-style reaction and pathway integration.

Is MetaboAnalyst a database or an analysis platform?

MetaboAnalyst is an analysis platform. It applies statistical, enrichment, topology, network, and multi-omics methods using metabolite data and pathway libraries drawn from relevant knowledge resources.

Can KEGG or MetaboAnalyst prove that a pathway is activated?

No. They provide mapping and statistical evidence. A pathway-activity claim requires consistent metabolite-level evidence and, where appropriate, direct support from enzyme-activity measurements, isotope tracing, perturbation experiments, or other functional assays. Gene-expression and protein-abundance data can provide complementary evidence but do not establish pathway activity on their own.

MetwareBio Support for Metabolite Annotation and Pathway Analysis

MetwareBio supports untargeted metabolomics, targeted metabolomics, and widely targeted metabolomics, together with statistical analysis, metabolite annotation, pathway interpretation, and integration with transcriptomics or proteomics. A staged project can distinguish exploratory annotation and pathway prioritization from follow-up targeted quantification, while reserving stronger mechanistic claims for appropriate functional validation.

Planning a metabolomics study that requires metabolite annotation, pathway analysis, or multi-omics integration? Contact MetwareBio to discuss an analytical strategy aligned with the study design, sample type, and biological question.

Contact Us

Read More: Pathway Analysis and Metabolomics Data Interpretation

These articles extend the discussion from database selection to practical pathway analysis workflows, enrichment methods, and data processing strategies for metabolomics research.

Beginner for KEGG Pathway Analysis: The Complete Guide

A practical companion to this article, focusing specifically on KEGG pathway analysis workflows. Learn how to navigate KEGG Mapper, interpret compound-to-pathway mapping, and use KEGG modules for metabolomics interpretation.

GO vs KEGG vs GSEA: How to Choose the Right Enrichment Analysis

When pathway databases are selected, the next question is which enrichment method to use. This guide compares Gene Ontology, KEGG, and GSEA approaches, helping researchers match the method to their study design and data type.

Common Lipidomics Databases and Software

LIPID MAPS is introduced here as a complementary resource. This article expands on lipid-specific databases and software tools, covering structural annotation, classification systems, and analytical platforms for lipidomics research.

A Summary of Commonly Used Public Metabolite Databases

A broader overview of publicly available metabolite databases beyond KEGG and HMDB. This resource helps researchers identify the right database for specific organisms, compound classes, and analytical platforms.

GSEA Enrichment Analysis: A Quick Guide

For researchers using MetaboAnalyst or similar platforms, this guide explains how Gene Set Enrichment Analysis works, when to choose rank-based methods over over-representation analysis, and how to interpret enrichment results.

Advanced Techniques in Metabolomics Data Processing

Before database annotation and pathway analysis, raw data must be properly processed. This article covers peak detection, alignment, normalization, and quality filtering steps that directly affect downstream pathway mapping results.

References

  1. Alseekh, S., Aharoni, A., Brotman, Y., et al. (2021). Mass spectrometry-based metabolomics: A guide for annotation, quantification and best reporting practices. Nature Methods, 18, 747-756. https://doi.org/10.1038/s41592-021-01197-1
  2. Caspi, R., Billington, R., Keseler, I. M., et al. (2020). The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Research, 48(D1), D445-D453. https://doi.org/10.1093/nar/gkz862
  3. Chen, Y., Li, E.-M., & Xu, L.-Y. (2022). Guide to metabolomics analysis: A bioinformatics workflow. Metabolites, 12(4), 357. https://doi.org/10.3390/metabo12040357
  4. Conroy, M. J., Andrews, R. M., Andrews, S., et al. (2024). LIPID MAPS: Update to databases and tools for the lipidomics community. Nucleic Acids Research, 52(D1), D1677-D1682. https://doi.org/10.1093/nar/gkad896
  5. Kanehisa, M., Furumichi, M., Sato, Y., et al. (2025). KEGG: Biological systems database as a model of the real world. Nucleic Acids Research, 53(D1), D672-D677. https://doi.org/10.1093/nar/gkae909
  6. KEGG. (2024, October 1). Copyright and disclaimer. https://www.kegg.jp/kegg/legal.html
  7. Lu, Y., Pang, Z., & Xia, J. (2023). Comprehensive investigation of pathway enrichment methods for functional interpretation of LC-MS global metabolomics data. Briefings in Bioinformatics, 24(1), bbac553. https://doi.org/10.1093/bib/bbac553
  8. Pang, Z., Lu, Y., Zhou, G., et al. (2024). MetaboAnalyst 6.0: Towards a unified platform for metabolomics data processing, analysis and interpretation. Nucleic Acids Research, 52(W1), W398-W406. https://doi.org/10.1093/nar/gkae253
  9. Ragueneau, E., Gong, C., Sinquin, P., et al. (2026). Reactome Knowledgebase 2026. Nucleic Acids Research, 54(D1), D673-D681. https://doi.org/10.1093/nar/gkaf1223
  10. Wishart, D. S., Guo, A., Oler, E., et al. (2022). HMDB 5.0: The Human Metabolome Database for 2022. Nucleic Acids Research, 50(D1), D622-D631. https://doi.org/10.1093/nar/gkab1062
  11. Yurekten, O., Payne, T., Tejera, N., et al. (2024). MetaboLights: Open data repository for metabolomics. Nucleic Acids Research, 52(D1), D640-D646. https://doi.org/10.1093/nar/gkad1045

 

Contact Us
Name can't be empty
Email error!
Message can't be empty
CONTACT FOR DEMO

Next-Generation Omics Solutions:
Proteomics & Metabolomics

Submit your inquiry to explore customized proteomics and metabolomics services for your research, or contact us at support-global@metwarebio.com..
Name can't be empty
Email error!
Message can't be empty
CONTACT FOR DEMO
+1(781)975-1541
LET'S STAY IN TOUCH
submit
Copyright © 2025 Metware Biotechnology Inc. All Rights Reserved.
support-global@metwarebio.com +1(781)975-1541
8A Henshaw Street, Woburn, MA 01801
Contact Us Now
Name can't be empty
Email error!
Message can't be empty
support-global@metwarebio.com +1(781)975-1541
8A Henshaw Street, Woburn, MA 01801
Register Now
Name can't be empty
Email error!
Message can't be empty