Comparing the Statistical Fate of Paralogous and Orthologous Sequences

F. Massip; M. Sheinman; S. Schbath; P.F. Arndt

doi:10.1534/genetics.116.193912

Article Dans Une Revue Genetics Année : 2016

Comparing the Statistical Fate of Paralogous and Orthologous Sequences

(1) , , ,

F. Massip

Fonction : Auteur

Statistique en grande dimension pour la génomique

M. Sheinman

Fonction : Auteur

S. Schbath

Fonction : Auteur

P.F. Arndt

Fonction : Auteur

Résumé

For several decades, sequence alignment has been a widely used tool in bioinformatics. For instance, finding homologous sequences with a known function in large databases is used to get insight into the function of nonannotated genomic regions. Very efficient tools like BLAST have been developed to identify and rank possible homologous sequences. To estimate the significance of the homology, the ranking of alignment scores takes a background model for random sequences into account. Using this model we can estimate the probability to find two exactly matching subsequences by chance in two unrelated sequences. For two homologous sequences, the corresponding probability is much higher, which allows us to identify them. Here we focus on the distribution of lengths of exact sequence matches between protein-coding regions of pairs of evolutionarily distant genomes. We show that this distribution exhibits a power-law tail with an exponent alpha = -5. Developing a simple model of sequence evolution by substitutions and segmental duplications, we show analytically and computationally that paralogous and orthologous gene pairs contribute differently to this distribution. Our model explains the differences observed in the comparison of coding and noncoding parts of genomes, thus providing a better understanding of statistical properties of genomic sequences and their evolution.

Mots clés

comparative genomics statistical genomics DNA duplications genome evolution

Domaines

Sciences du Vivant [q-bio]

Lauriane Pillet : Connectez-vous pour contacter le contributeur

https://univ-lyon1.hal.science/hal-02053633

Soumis le : vendredi 1 mars 2019-14:33:44

Dernière modification le : mercredi 24 janvier 2024-08:50:43

Dates et versions

hal-02053633 , version 1 (01-03-2019)

Identifiants

HAL Id : hal-02053633 , version 1
DOI : 10.1534/genetics.116.193912
PRODINRA : 382714
PUBMED : 27474728
PUBMEDCENTRAL : PMC5068840
WOS : 000385871400009

Citer

F. Massip, M. Sheinman, S. Schbath, P.F. Arndt. Comparing the Statistical Fate of Paralogous and Orthologous Sequences. Genetics, 2016, 204 (2), pp.475-482. ⟨10.1534/genetics.116.193912⟩. ⟨hal-02053633⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS INRIA UNIV-LYON1 BIOENVIS LBBE UDL

23 Consultations

0 Téléchargements

Comparing the Statistical Fate of Paralogous and Orthologous Sequences

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager