PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

CS-PSeq-Gen: simulating the evolution of protein sequence under constraints.

UNLABELLED: CS-PSeq-Gen is a program derived from PSeq-Gen, designed to perform simulations of the evolution of protein sequences under the constraints of a reconstructed phylogeny. It also provides a basis for the investigation of the correlated evolution of sites. AVAILABILITY: http://condor.urbb.jussieu.fr/CS-PSeq-Gen.html

Amino Acid Sequence↗

Intratumor chromosomal heterogeneity in advanced carcinomas of the uterine cervix.

Intratumor heterogeneity in chromosomal aberrations is believed to represent a major challenge in the treatment of cancer. The aim of our work was to assess the chromosomal heterogeneity of advanced cervical carcinomas and to distinguish aberrations that had occurred at a late stage of the disease from early events. A total of 55 biopsies, sampled from 2-4 different sites within 20 tumors, were analyzed by use of comparative genomic hybridization. Heterogeneous aberrations were identified as those present in at least 1 of the biopsies and which were not seen, nor seen as a tendency, in the others of the same tumor. The homogeneous aberrations were those seen in all biopsies of the tumor. The most frequent homogeneous aberrations were gain of 3q (65%), 20q (65%) and 5p (50%), indicating that these are early events in the development of the disease. Chromosomal heterogeneity was observed in 11 tumors. The most frequent heterogeneous aberrations were loss of 4p14-q25 (60% of 10 cases with this aberration), and gain of 2p22-pter (50% of 6 cases), 11qcen-q13 (33% of 9 cases) and 8q (27% of 11 cases), suggesting that these events promote progression at a later stage. Many of the heterogeneous regions contained genes known to influence the prognosis of cervical cancer, such as 7p (EGFR), 8q (c-MYC), 11qcen-q13 (CCND1) and 17q (ERBB2). Three evolution sequences for the subpopulations in the heterogeneous tumors were identified: a serial, a parallel and a mixed sequence. In 2 tumors with a serial sequence, it was indicated that the aberrations +8 and -X had occurred after the other heterogeneous aberrations and hence were the aberrations most recently formed. Our results suggest pronounced chromosomal instability in advanced cervical carcinomas. Moreover, aggressive and treatment-resistant subpopulations may emerge at a late stage and possibly contribute to a poor prognosis of the advanced stages.

Biopsy↗

Genetic algorithm-based model of evolutionary dynamics of class II transposable elements.

We propose a new conceptual framework to study the dynamics of transposable elements. Based on a genetic algorithm, our model is designed as a self-organizing system. Our results show that transposable elements could emerge from a single endonuclease gene. The DNA repair mechanisms appear to condition the emergence success of class II TEs. Antagonist selective forces acting on transposable elements and their hosts induce by their opposition differences in the sequence evolution of the functional domains and of the copies.

Algorithms↗

Pattern and timing of evolutionary divergences among hominoids based on analyses of complete mtDNAs.

We have examined and dated primate divergences by applying a newly established molecular/ paleontological reference, the evolutionary separation between artiodactyls and cetaceans anchored at 60 million years before present (MYBP). Owing to the morphological transformations coinciding with the transition from terrestrial to aquatic (marine) life and the large body size of the animals (which makes their fossils easier to find), this reference can be defined, paleontologically, within much narrower time limits compared to any local primate calibration marker hitherto applied for dating hominoid divergences. Application of the artiodactyl/ cetacean reference (A/C-60) suggests that hominoid divergences took place much earlier than has been concluded previously. According to a homogeneous-rate model of sequence evolution, the primary hominoid divergence, i.e., that between the families Hylobatidae (gibbons) and Hominidae, was dated at approximately 36 MYBP. The corresponding dating for the divergence between Pongo (orangutan) and Gorilla-Pan (chimpanzee) -Homo is approximately 24.5 MYBP, that for Gorilla vs Homo-Pan is approximately 18 MYBP, and that for Homo vs Pan approximately 13.5 MYBP. The split between Sumatran and Bornean orangutans was dated at approximately 10.5 MYBP and that between the common and pygmy chimpanzees at approximately 7 MYBP. Analyses of a single gene (cytochrome b) suggest that the divergence within the Catarrhini, i.e., between Hominoidea and Old World monkeys (Cercopithecoidea), took place > 40 MYBP; that within the Anthropoidea, i.e., between Catarrhini and Platyrrhini (New World monkeys), > 60 MYBP; and that between Anthropoidea and Prosimii (lemur), approximately 80 MYBP. These separation times are about two times more ancient than those applied previously as references for the dating of hominoid divergences. The present findings automatically imply a much slower evolution in hominoid DNA (both mitochondrial and nuclear) than commonly recognized.

Animals↗

Assessment of the C(4) phosphoenolpyruvate carboxylase gene diversity in grasses (Poaceae).

C(4) phosphoenolpyruvate carboxylase (PEPC) is a key enzyme in the C(4) photosynthetic pathway. To analyze the diversity of the corresponding gene in grasses, we designed PCR primers to specifically amplify C(4) PEPC cDNA fragments. Using RT-PCR, we generated partial PEPC cDNA sequences in several grasses displaying a C(4) photosynthetic pathway. All these sequences displayed a high homology (78-99%) with known grass C(4) PEPCs. PCR amplification did not occur in two grasses that display the C(3) photosynthetic pathway, and therefore we assumed that all generated sequences corresponded to C(4) PEPC transcripts. Based on one large cDNA segment, phylogenetic reconstruction enabled us to assess the relationships between 22 grass species belonging to the subfamilies Panicoideae, Arundinoideae and Chloridoideae. The phylogenetic relationships between species deduced from C(4) PEPC sequences were similar to those deduced from other molecular data. The sequence evolution of the C(4) PEPC isoform was faster than in the other PEPC isoforms. Finally, the utility of the C(4) PEPC gene phylogeny to study the evolution of C(4) photosynthesis in grasses is discussed.

Journal Article↗

Molecular phylogeny of conjugating green algae (Zygnemophyceae, Streptophyta) inferred from SSU rDNA sequence comparisons.

Nuclear-encoded SSU rDNA sequences have been obtained from 64 strains of conjugating green algae (Zygnemophyceae, Streptophyta, Viridiplantae). Molecular phylogenetic analyses of 90 SSU rDNA sequences of Viridiplantae (inciuding 78 from the Zygnemophyceae) were performed using complex evolutionary models and maximum likelihood, distance, and maximum parsimony methods. The significance of the results was tested by bootstrap analyses, deletion of long-branch taxa, relative rate tests, and Kishino-Hasegawa tests with user-defined trees. All results support the monophyly of the class Zygnemophyceae and of the order Desmidiales. The second order, Zygnematales, forms a series of early-branching clades in paraphyletic succession, with the two traditional families Mesotaeniaceae and Zygnemataceae not recovered as lineages. Instead, a long-branch Spirogyra/Sirogonium clade and the later-diverging Netrium and Roya clades represent independent clades. Within the order Desmidiales, the families Gonatozygaceae and Closteriaceae are monophyletic, whereas the Peniaceae (represented only by Penium margaritaceum) and the Desmidiaceae represent a single weakly supported lineage. Within the Desmidiaceae short internal branches and varying rates of sequence evolution among taxa reduce the phylogenetic resolution significantly. The SSU rDNA-based phylogeny is largely congruent with a published analysis of the rbcL phylogeny of the Zygnemophyceae (McCourt et al. 2000) and is also in general agreement with classification schemes based on cell wall ultrastructure. The extended taxon sampling at the subgenus level provides solid evidence that many genera in the Zygnemophyceae are not monophyletic and that the genus concept in the group needs to be revised.

Chlorophyta↗

Entanglement invariants and phylogenetic branching.

It is possible to consider stochastic models of sequence evolution in phylogenetics in the context of a dynamical tensor description inspired from physics. Approaching the problem in this framework allows for the well developed methods of mathematical physics to be exploited in the biological arena. We present the tensor description of the homogeneous continuous time Markov chain model of phylogenetics with branching events generated by dynamical operations. Standard results from phylogenetics are shown to be derivable from the tensor framework. We summarize a powerful approach to entanglement measures in quantum physics and present its relevance to phylogenetic analysis. Entanglement measures are found to give distance measures that are equivalent to, and expand upon, those already known in phylogenetics. In particular we make the connection between the group invariant functions of phylogenetic data and phylogenetic distance functions. We introduce a new distance measure valid for three taxa based on the group invariant function known in physics as the "tangle". All work is presented for the homogeneous continuous time Markov chain model with arbitrary rate matrices.

Evolution, Molecular↗

Comprehensive analysis of synonymous codon usage bias and evolutionary dynamics in the chloroplast genomes of eight Coptis species.

Coptis is a medically important genus renowned for producing valuable isoquinoline alkaloids. Although its chloroplast genomes encode key components for photosynthesis and plastid gene expression, the evolutionary constraints acting on their coding sequences and synonymous codon usage remain poorly resolved. Here, we combined a transparent taxon-level sampling strategy with comparative analyses of chloroplast CDSs from eight Coptis taxa. We quantified nucleotide composition, relative synonymous codon usage, effective number of codons, neutrality and PR2 patterns, and correspondence analysis, and then integrated these results with a core-CDS distance analysis and gene-wise pairwise dN/dS estimates. The chloroplast genomes showed a conserved AT-rich composition, especially at the third codon position (GC3 approximately 30.3-30.8%), with a consistent GC1 > GC2 > GC3 trend. Thirty preferred codons were detected, 28 ending in A/T, and eleven optimal codons were shared across the genus. The core-CDS distance analysis recovered a close relationship between C. chinensis and C. chinensis var. brevisepala, whereas most coding genes showed dN/dS values below one, consistent with pervasive purifying constraint. Across 48 consistently filtered CDSs, GC3s was negatively associated with mean dN (Spearman rho = -0.404, P = 0.00439) and CAI was positively associated with mean dN (rho = 0.303, P = 0.0361), whereas the remaining associations were not significant (all P > = 0.0972). These results extend codon-usage analysis by linking synonymous-site composition to coding-sequence evolution within Coptis, while providing a hypothesis-generating resource for future plastid engineering studies.

Genome, Chloroplast↗

Evolutionary rate variation in eukaryotic lineage specific human intronless proteins.

The present study examines 783 human-mouse orthologous gene pairs for their pattern of sequence evolution, contrasting mammalia, eukaryota, coelomata, and bilateria specific human intronless genes. Such comparisons may be of use in understanding the general evolution of human genome. Evolutionary rate analyses indicate that mammalia specific human intronless genes are evolving faster as compared to other intronless genes specific to eukaryotic lineage, indicating towards their rapid evolution. The observations indicates that the genes conserved in eukaryota, coelomata, and bilateria, that is, proteins that arose earlier in evolution as compared to mammalia specific genes evolve slowly and are subjected to negative selection. The cause underlying rate variations was also explored. Although mutational bias might slightly fasten the nonsynonymous rates in mammalia specific genes, it is unlikely to be major cause of rate difference between the various categories. Furthermore, rate of divergence of mammalia specific intronless genes has been related to functional classification using the protein family annotation. Protein function was found in some cases to have larger impact on the rate of evolution of genes. Also, the codon usage pattern of mammalia specific intronless genes do not seem to differ much from those of other intronless genes conserved solely in eukaryotic lineage.

Animals↗

A highly unexpected strong correlation between fixation probability of nonsynonymous mutations and mutation rate.

Under prevailing theories, the nonsynonymous-to-synonymous substitution ratio (i.e. K(a)/K(s)), which measures the fixation probability of nonsynonymous mutations, is correlated with the strength of selection. In this article, we report that K(a)/K(s) is also strongly correlated with the mutation rate as measured by K(s), and that this correlation appears to have a similar magnitude as the correlation between K(a)/K(s) and selective strength. This finding cannot be reconciled with current theories. It suggests that we should re-evaluate the current paradigms of coding-sequence evolution, and that the wide use of K(a)/K(s) as a measure of selective strength needs reassessment.

Animals↗

Did brain-specific genes evolve faster in humans than in chimpanzees?

One of the most distinctive characteristics of humans among primates is the size, organization and function of the brain. A recent study has proposed that there was widespread accelerated sequence evolution of genes functioning in the nervous system during human origins. Here we test this hypothesis by a genome-wide analysis of genes that are expressed predominantly or specifically in brain tissues and genes that have important roles in the brain, identified on the basis of five different definitions of brain specificity. Although there is little overlap among the five sets of brain-specific genes, none of them supports human acceleration. On the contrary, some datasets show significantly fewer nonsynonymous substitutions in humans than in chimpanzees for brain-specific genes relative to other genes in the genome. Our results suggest that the unique features of the human brain did not arise by a large number of adaptive amino acid changes in many proteins.

Animals↗

Rabbits, if anything, are likely Glires.

Rodentia (e.g., mice, rats, dormice, squirrels, and guinea pigs) and Lagomorpha (e.g., rabbits, hares, and pikas) are usually grouped into the Glires. Status of this controversial superorder has been evaluated using morphology, paleontology, and mitochondrial plus nuclear DNA sequences. This growing corpus of data has been favoring the monophyly of Glires. Recently, Misawa and Janke [Mol. Phylogenet. Evol. 28 (2003) 320] analyzed the 6441 amino acids of 20 nuclear proteins for six placental mammals (rat, mouse, rabbit, human, cattle, and dog) and two outgroups (chicken and xenopus), and observed a basal position of the two murine rodents among the former. They concluded that "the Glires hypothesis was rejected." We here reanalyzed [loc. cit.] data set under maximum likelihood and Bayesian tree-building approaches, using phylogenetic models that take into account among-site variation in evolutionary rates and branch-length variation among proteins. Our observations support both the association of rodents and lagomorphs and the monophyly of Euarchontoglires (=Supraprimates) as the most likely explanation of the protein alignments. We conducted simulation studies to evaluate the appropriateness of lissamphibian and avian outgroups to root the placental tree. When the outgroup-to-ingroup evolutionary distance increases, maximum parsimony roots the topology along the long Mus-Rattus branch. Maximum likelihood, in contrast, roots the topology along different branches as a function of their length. Maximum likelihood appears less sensitive to the "long-branch attraction artifact" than is parsimony. Our phylogenetic conclusions were confirmed by the analysis of a different protein data set using a similar sample of species but different outgroups. We also tested the effect of the addition of afrotherian and xenarthran taxa. Using the linearized tree method, [loc. cit.] estimated that mice and rats diverged about 35 million years ago. Molecular dating based on the Bayesian relaxed molecular clock method suggests that the 95% credibility interval for the split between mice and rats is 7-17 Mya. We here emphasize the need for appropriate models of sequence evolution (matrices of amino acid replacement, taking into account among-site rate variation, and independent parameters across independent protein partitions) and for a taxonomically broad sample, and conclude on the likelihood that rodents and lagomorphs together constitute a monophyletic group (Glires).

Animals↗

Phylogenetic estimation under codon models can be biased by codon usage heterogeneity.

In theory, codon models that account for the dependence of nucleotide substitutions between codon positions as well as differences between synonymous and non-synonymous changes best describe the sequence evolution in protein coding genes. However, in practice we know little about the degree to which violations of the assumptions of codon model-based estimates occur, and how significant these artifacts may be. In nucleotide-based phylogenies from first and second codon positions in a concatenated plastid gene data set, two distantly related taxa--dinoflagellate and haptophyte plastids--were robustly grouped together. This artifactual grouping is attributed to the parallel heterogeneity in leucine (Leu) and serine (Ser) codon usages in the data set. Here, by using this data set, we demonstrated that codon-based phylogenetic estimations are seriously biased, robustly uniting the dinoflagellate and haptophyte plastids into a monophyletic clade, when the model assumption of homogeneity of codon composition was violated. Our results suggest that similar phylogenetic artifacts may occur via codon usage heterogeneity in any amino acids in codon model-based estimations. We advise that homogeneity in codon usage across taxa in a data set be confirmed before codon model-based phylogenetic estimation is attempted.

Codon↗

Retention of enzyme gene duplicates by subfunctionalization.

Duplication-degeneration-complementation (DDC) describes a process by which evolving duplicates of a pleiotropic ancestral gene divide up the multiple functions of the ancestor between them (i.e. subfunctionalize), and this ultimately frustrates the rate of pseudogene formation. Focusing explicitly on enzyme-like pleiotropic function, we model DDC driven by sequence divergence between duplicates. The model incorporates an idealized sequence-function mapping in which enzyme-substrate binding affinity is related to hydrophobic versus polar (HP) amino-acid composition of tertiary structure about the binding pocket. In this sense, a transparent coupling between physical-chemical function of an enzyme and sequence evolution is presented.

Enzymes↗

Microsatellite mutations in the germline: implications for evolutionary inference.

Microsatellite DNA sequences mutate at rates several orders of magnitude higher than that of the bulk of DNA. Such high rates mean that spontaneous mutations that form new-length variants can realistically be seen in pedigree analysis. Data on observed mutation events from various organisms are now accumulating, allowing inferences on DNA sequence evolution to be made through an unusually direct approach. Here I discuss and integrate microsatellite mutation data in an evolutionary context. A striking feature of the mutation process is that it seems highly heterogeneous, with distinct differences between species, repeat types, loci and alleles. Age and sex also affect the mutation rate. Within genomes at equilibrium, the microsatellite-length distribution is a delicate balance between biased mutation processes and point mutations acting towards the decay of repetitive DNA. Indeed, simple repeats do not evolve simply.

Evolution, Molecular↗

Human SNP variability and mutation rate are higher in regions of high recombination.

Understanding the co-variation of nucleotide diversity and local recombination rates is important both for the mapping of disease-associated loci and in understanding the causes of sequence evolution. It is known that single nucleotide polymorphisms (SNPs) around protein coding genes show higher diversity in regions of high recombination. Here, we find that this correlation holds for SNPs across the entire human genome, the great majority of which are not near exons or control elements. Contrasting with results from coding regions, we provide evidence that the higher nucleotide diversity in regions of high recombination is most likely due, at least in part, to a higher mutation rate. One possible explanation for this is that recombination is mutagenic.

Animals↗

Mitochondrial genomics flies high.

A new pair of papers reports the complete coding sequences for 28 mitochondrial DNAs in Drosophila. By examining the patterns of polymorphism and divergence among functionally distinct classes of synonymous and nonsynonymous nucleotide sites, Bill Ballard provides a comprehensive whole-genome picture of how mtDNA sequence evolution can depart from the strictly neutral and nearly neutral models of molecular evolution.

Journal Article↗