A bioinformatician's guide to metagenomics

Victor Kunin; Alex Copeland; Alla Lapidus; Konstantinos Mavromatis; Philip Hugenholtz

doi:10.1128/MMBR.00009-08

A bioinformatician's guide to metagenomics

Microbiol Mol Biol Rev. 2008 Dec;72(4):557-78, Table of Contents. doi: 10.1128/MMBR.00009-08.

Authors

Victor Kunin¹, Alex Copeland, Alla Lapidus, Konstantinos Mavromatis, Philip Hugenholtz

Affiliation

¹ Microbial Ecology Program, DOE Joint Genome Institute, 2800 Mitchell Drive, Walnut Creek, CA, USA.

Abstract

As random shotgun metagenomic projects proliferate and become the dominant source of publicly available sequence data, procedures for the best practices in their execution and analysis become increasingly important. Based on our experience at the Joint Genome Institute, we describe the chain of decisions accompanying a metagenomic project from the viewpoint of the bioinformatic analysis step by step. We guide the reader through a standard workflow for a metagenomic project beginning with presequencing considerations such as community composition and sequence data type that will greatly influence downstream analyses. We proceed with recommendations for sampling and data generation including sample and metadata collection, community profiling, construction of shotgun libraries, and sequencing strategies. We then discuss the application of generic sequence processing steps (read preprocessing, assembly, and gene prediction and annotation) to metagenomic data sets in contrast to genome projects. Different types of data analyses particular to metagenomes are then presented, including binning, dominant population analysis, and gene-centric analysis. Finally, data management issues are presented and discussed. We hope that this review will assist bioinformaticians and biologists in making better-informed decisions on their journey during a metagenomic project.

Publication types

Research Support, U.S. Gov't, Non-P.H.S.
Review

MeSH terms

Database Management Systems*
Databases, Genetic
Environmental Microbiology
Genome / genetics*
Genomics / methods*
Genomics / standards
Information Storage and Retrieval / methods*
Sequence Analysis, DNA