A genome-wide survey of human pseudogenes

David Torrents; Mikita Suyama; Evgeny Zdobnov; Peer Bork

doi:10.1101/gr.1455503

A genome-wide survey of human pseudogenes

Genome Res. 2003 Dec;13(12):2559-67. doi: 10.1101/gr.1455503.

Authors

David Torrents¹, Mikita Suyama, Evgeny Zdobnov, Peer Bork

Affiliation

¹ EMBL, Heidelberg 69117, Germany.

Abstract

We screened all intergenic regions in the human genome to identify pseudogenes with a combination of homology searches and a functionality test using the ratio of silent to replacement nucleotide substitutions (KA/KS). We identified 19,724 regions of which 95% +/- 3% are estimated to evolve neutrally and thus are likely to encode pseudogenes. Half of these have no detectable truncation in their pseudocoding regions and therefore are not identifiable by methods that require the presence of truncations to prove nonfunctionality. A comparative analysis with the mouse genome showed that 70% of these pseudogenes have a retrotranspositional origin (processed), and the rest arose by segmental duplication (nonprocessed). Although the spread of both types of pseudogenes correlates with chromosome size, nonprocessed pseudogenes appear to be enriched in regions with high gene density. It is likely that the human pseudogenes identified here represent only a small fraction of the total, which probably exceeds the number of genes.

Publication types

Validation Study

MeSH terms

Benchmarking / methods
Benchmarking / statistics & numerical data
Computational Biology / methods*
Computational Biology / statistics & numerical data
DNA, Intergenic / genetics
Genome, Human*
Humans
Online Systems
Pseudogenes*
Sequence Homology, Nucleic Acid
Statistical Distributions

Substances

DNA, Intergenic