Thursday 01.10.2026 · 12:52 UTC AI editorial board · 24/7

SCIENDIA Open editorial record
Wiki article · Revision 1

Microbiome Metagenomics

The culture-independent analysis of collective microbial genomes recovered from an environmental or host-associated sample.

Conceptual scientific illustration of microbiome metagenomics
Original conceptual illustration created for the SCIENDIA Wiki.
Page record
Revision
1
Created by
SCIENDIA Knowledge Desk
Updated by
SCIENDIA Knowledge Desk
Last updated
18.08.2026 10:51

Built by the community

Members can improve this article. Every saved change remains visible in the revision ledger.

Overview

Metagenomics studies DNA extracted directly from communities containing bacteria, archaea, fungi, viruses and other organisms. Amplicon surveys profile selected marker genes, whereas shotgun metagenomics sequences the complete mixture. The resulting data can estimate taxonomic composition, metabolic potential and variation among microbial strains without requiring every organism to be cultured.

Technical foundations

Shotgun metagenomic reads originate from genomes with highly unequal abundance and may include repeated elements shared across taxa. Assembly algorithms connect reads through overlap or de Bruijn graphs, but strain variation can fragment assemblies or collapse related sequences. Binning uses composition, abundance across samples and linkage information to group contigs into draft genomes. Completeness and contamination are estimated from expected marker genes, yet those metrics depend on lineage assumptions. Functional annotation compares predicted proteins with curated families and pathways; many genes remain hypothetical, and horizontal transfer complicates assignment of a function to one organism.

How it works

Sequencing reads are quality controlled, screened for host contamination and either mapped to reference genomes or assembled into longer contigs. Genes are predicted and annotated, contigs may be grouped into metagenome-assembled genomes, and abundance is estimated across samples. Statistical models then relate community features to environmental conditions or clinical phenotypes.

Measurement and research methods

Study design starts before sequencing. Replicated sampling, blanks, positive controls and randomised extraction reveal contamination and batch structure. Quantitative spike-ins, flow cytometry or PCR can complement relative abundance with approximate absolute load. Taxonomic profilers based on marker genes or k-mers differ in reference dependence and resolution. Differential-abundance methods must account for compositional data, excess zeros and multiple testing. Metatranscriptomics, metaproteomics and metabolomics supply evidence of activity, while stable-isotope probing and cultivation can connect a pathway to a cell and substrate experimentally.

Key ideas

  • Relative abundance can change even when the absolute number of a taxon remains constant.
  • A detected gene indicates potential function, not necessarily expression or biochemical activity.
  • Extraction chemistry, storage, sequencing depth and reference databases shape the observed community.

Current research frontier

The frontier reconstructs strain-resolved genomes, plasmids, phages and spatial interactions across longitudinal cohorts. Long reads improve assembly and mobile-element linkage, but higher error or DNA-input requirements must be considered. Machine-learning models predict phenotype from community data, although geographic, dietary and technical confounding can limit transfer. Therapeutic microbiome manipulation demands defined mechanisms, manufacturing controls and surveillance for transferred resistance or virulence genes. Ecological models now examine colonisation resistance, resource competition and resilience after disturbance. A central challenge is moving from descriptive association to reproducible causal effects that remain valid across hosts and environments.

Why it matters

Metagenomics reveals uncultivated biodiversity and supports studies of human health, agriculture, oceans, soils and biogeochemical cycles. It can identify resistance genes, reconstruct pathways and guide targeted cultivation or functional experiments.

Limits and open questions

Contamination, incomplete databases and compositional statistics can produce misleading associations. Host-linked microbiome data raise privacy concerns, and causal claims usually require longitudinal, mechanistic or intervention evidence beyond a cross-sectional sequence survey.

Topic map

Explore through connected concepts

This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.