Microbiome Metagenomics
The culture-independent analysis of collective microbial genomes recovered from an environmental or host-associated sample.
- Revision
- 1
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 18.08.2026 10:51
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
Metagenomics studies DNA extracted directly from communities containing bacteria, archaea, fungi, viruses and other organisms. Amplicon surveys profile selected marker genes, whereas shotgun metagenomics sequences the complete mixture. The resulting data can estimate taxonomic composition, metabolic potential and variation among microbial strains without requiring every organism to be cultured.
Technical foundations
Shotgun metagenomic reads originate from genomes with highly unequal abundance and may include repeated elements shared across taxa. Assembly algorithms connect reads through overlap or de Bruijn graphs, but strain variation can fragment assemblies or collapse related sequences. Binning uses composition, abundance across samples and linkage information to group contigs into draft genomes. Completeness and contamination are estimated from expected marker genes, yet those metrics depend on lineage assumptions. Functional annotation compares predicted proteins with curated families and pathways; many genes remain hypothetical, and horizontal transfer complicates assignment of a function to one organism.
How it works
Sequencing reads are quality controlled, screened for host contamination and either mapped to reference genomes or assembled into longer contigs. Genes are predicted and annotated, contigs may be grouped into metagenome-assembled genomes, and abundance is estimated across samples. Statistical models then relate community features to environmental conditions or clinical phenotypes.
Measurement and research methods
Study design starts before sequencing. Replicated sampling, blanks, positive controls and randomised extraction reveal contamination and batch structure. Quantitative spike-ins, flow cytometry or PCR can complement relative abundance with approximate absolute load. Taxonomic profilers based on marker genes or k-mers differ in reference dependence and resolution. Differential-abundance methods must account for compositional data, excess zeros and multiple testing. Metatranscriptomics, metaproteomics and metabolomics supply evidence of activity, while stable-isotope probing and cultivation can connect a pathway to a cell and substrate experimentally.
Key ideas
- Relative abundance can change even when the absolute number of a taxon remains constant.
- A detected gene indicates potential function, not necessarily expression or biochemical activity.
- Extraction chemistry, storage, sequencing depth and reference databases shape the observed community.
Current research frontier
The frontier reconstructs strain-resolved genomes, plasmids, phages and spatial interactions across longitudinal cohorts. Long reads improve assembly and mobile-element linkage, but higher error or DNA-input requirements must be considered. Machine-learning models predict phenotype from community data, although geographic, dietary and technical confounding can limit transfer. Therapeutic microbiome manipulation demands defined mechanisms, manufacturing controls and surveillance for transferred resistance or virulence genes. Ecological models now examine colonisation resistance, resource competition and resilience after disturbance. A central challenge is moving from descriptive association to reproducible causal effects that remain valid across hosts and environments.
Why it matters
Metagenomics reveals uncultivated biodiversity and supports studies of human health, agriculture, oceans, soils and biogeochemical cycles. It can identify resistance genes, reconstruct pathways and guide targeted cultivation or functional experiments.
Limits and open questions
Contamination, incomplete databases and compositional statistics can produce misleading associations. Host-linked microbiome data raise privacy concerns, and causal claims usually require longitudinal, mechanistic or intervention evidence beyond a cross-sectional sequence survey.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.