PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

PASSML: combining evolutionary inference and protein secondary structure prediction.

MOTIVATION: Evolutionary models of amino acid sequences can be adapted to incorporate structure information; protein structure biologists can use phylogenetic relationships among species to improve prediction accuracy. Results : A computer program called PASSML ('Phylogeny and Secondary Structure using Maximum Likelihood') has been developed to implement an evolutionary model that combines protein secondary structure and amino acid replacement. The model is related to that of Dayhoff and co-workers, but we distinguish eight categories of structural environment: alpha helix, beta sheet, turn and coil, each further classified according to solvent accessibility, i.e. buried or exposed. The model of sequence evolution for each of the eight categories is a Markov process with discrete states in continuous time, and the organization of structure along protein sequences is described by a hidden Markov model. This paper describes the PASSML software and illustrates how it allows both the reconstruction of phylogenies and prediction of secondary structure from aligned amino acid sequences. AVAILABILITY: PASSML 'ANSI C' source code and the example data sets described here are available at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html and 'downstream' Web pages. CONTACT: P.Lio@gen.cam.ac.uk

Adenylate Kinase↗

Subgenome-specific markers in allopolyploid cotton Gossypium hirsutum: implications for evolutionary analysis of polyploids.

We developed a set of genetic markers specific to the A and D genome types of cotton using representational difference analysis (RDA). These markers produce amplification products with genomic DNA from allotetraploid cotton Gossypium hirsutum. One of the markers is a polymorphic amplified restriction fragment (PARF) - a sequence found in both A and D genomes but differently flanked by restriction sites. Results of phylogenetic analysis of the PARF sequences from diploid cottons and from allotetraploid G. hirsutum agree with a previous observation of the interlocus concerted evolution (sequences corresponding to A and D genomes are homogenized to a D genome-type sequence). Our study shows how RDA can be used to develop genome-specific markers that can be used to study molecular evolution of allopolyploids.

DNA, Plant↗

Recent horizontal transfer of mellifera subfamily mariner transposons into insect lineages representing four different orders shows that selection acts only during horizontal transfer.

We report the isolation and sequencing of genomic copies of mariner transposons involved in recent horizontal transfers into the genomes of the European earwig, Forficula auricularia; the European honey bee, Apis mellifera; the Mediterranean fruit fly, Ceratitis capitata; and a blister beetle, Epicauta funebris, insects from four different orders. These elements are in the mellifera subfamily and are the second documented example of full-length mariner elements involved in this kind of phenomenon. We applied maximum likelihood methods to the coding sequences and determined that the copies in each genome were evolving neutrally, whereas reconstructed ancestral coding sequences appeared to be under selection, which strengthens our previous hypothesis that the primary selective constraint on mariner sequence evolution is the act of horizontal transfer between genomes.

Amino Acid Sequence↗

Evidence for the adaptive evolution of the carbon fixation gene rbcL during diversification in temperature tolerance of a clade of hot spring cyanobacteria.

Determining the molecular basis of enzyme adaptation is central to understanding the evolution of environmental tolerance but is complicated by the fact that not all amino acid differences between ecologically divergent taxa are adaptive. Analysing patterns of nucleotide sequence evolution can potentially guide the investigation of protein adaptation by identifying candidate codon sites on which diversifying selection has been operating. Here, I test whether there is evidence for molecular adaptation of the carbon fixation gene rbcL for a clade of hot spring cyanobacteria in the genus Synechococcus that has diverged in thermotolerance. Amino acid replacements during Synechococcus radiation have resulted in an increase in the number of hydrophobic residues in the RbcLs of more thermotolerant strains. A similar increase in hydrophobicity has been observed for many thermostable proteins. Maximum likelihood models which allow for heterogeneity among codon sites in the ratio of nonsynonymous to synonymous nucleotide substitutions estimated a class of amino acid sites as a target of positive selection. Depending on the model, a single amino acid site that interacts with a flexible element involved in the opening and closing of the active site was estimated with either low or moderate support to be a member of this class. Site-directed mutagenesis approaches are being explored in order to directly test its adaptive significance.

Adaptation, Biological↗

Evidence for multiple functional copies of the male sex-determining locus, Sry, in African murine rodents.

Southern hybridization data suggest that the male sex-determining locus, Sry, is often duplicated in rodents. Here we explore DNA sequence evolution of orthologous and paralogous copies of Sry isolated from six species of African murines. PCR amplification followed by direct sequencing revealed from two to four copies of Sry per species. All copies include a long open reading frame, with a stop codon that coincides closely with the stop codon of the house mouse, Mus musculus, a species known to have a single copy of Sry. A phylogenetic analysis suggests that there are at least seven paralogous copies of Sry in this group of rodents. Putative orthologues are identical; sequence divergence among putative paralogues ranges from 1 to 8% (excluding the CAG repeat), with much lower levels of divergence in the high-mobility group (HMG-box) region than in the C-terminal region. A high proportion of nucleotide substitutions in both regions result in amino-acid replacement. The long open reading frame, conserved HMG-box, and pattern of evolution of the putative paralogues suggest that they are functional.

Amino Acid Sequence↗

Biogeography of Sulawesian shrews: testing for their origin with a parametric bootstrap on molecular data.

In order to identify the zoogeographic origin of shrews (genus Crocidura) living on the oceanic island of Sulawesi, 15 taxa from Southeast Asia and 1 from Europe were examined for sequence variation in a segment (617 bp) of the mitochondrial cytochrome b gene. The null hypothesis of a monophyletic origin of all Sulawesian shrews was investigated by a phylogenetic reconstruction using maximum parsimony. According to a parametric bootstrap which simulated sequence evolution for these taxa, the null hypothesis could be rejected as highly unlikely (P < 0.01). Therefore, the molecular phylogeny strongly suggests that overwater colonization of Sulawesi by shrews succeeded on at least two occasions. The first, relatively ancient wave of colonizers radiated and gave rise to a surprizingly diverse assemblage of at least five species which now coexist in perfect sympatry on Sulawesi. The second wave, of more recent origin, gave rise to Crocidura nigripes, a species which retained close genetic affinities with other Malay shrews.

Animals↗

Structure and sequence variation of the genes encoding the polymorphic, immunodominant molecule (PIM), an antigen of Theileria parva recognized by inhibitory monoclonal antibodies.

The polymorphic, immunodominant molecule (PIM) of Theileria parva is the predominant antigen recognized by sera from infected cattle and by monoclonal antibodies (mAb) used to differentiate parasite strains. As such, the antigen is under consideration as a diagnostic antigen, and since the mAbs can neutralize sporozoite infectivity in vitro, in immunization experiments. Initial comparison of two PIM cDNA sequences suggested that the PIM genes consist of conserved 5' and 3' termini flanking a central variable region. We present further evidence, based on sequence analysis, supporting this general structure for the PIM genes. Evidence is also presented for a single copy of the PIM gene per haploid genome, implying that the different versions of PIM are encoded by distinct alleles. The central variable region of the PIM allele from the T. parva (Marikebuni) stock was found to contain 13 copies of the tetrapeptide repeat Gln-Pro-Glu-Pro. We also detected point mutations in the 5' and 3' termini of the PIM alleles, including regions recognized by the neutralizing and typing mAb. This contrasted with the high sequence conservation of the two introns of the genes, suggesting that the protein is undergoing rapid evolution. Sequence comparison of PIM genes from buffalo- and cattle-derived parasites supported earlier results that the parasites infecting buffaloes constitute a more heterogeneous population than those from cattle.

Amino Acid Sequence↗

Heterozygosity, heteromorphy, and phylogenetic trees in asexual eukaryotes.

Little attention has been paid to the consequences of long-term asexual reproduction for sequence evolution in diploid or polyploid eukaryotic organisms. Some elementary theory shows that the amount of neutral sequence divergence between two alleles of a protein-coding gene in an asexual individual will be greater than that in a sexual species by a factor of 2tu, where t is the number of generations since sexual reproduction was lost and u is the mutation rate per generation in the asexual lineage. Phylogenetic trees based on only one allele from each of two or more species will show incorrect divergence times and, more often than not, incorrect topologies. This allele sequence divergence can be stopped temporarily by mitotic gene conversion, mitotic crossing-over, or ploidy reduction. If these convergence events are rare, ancient asexual lineages can be recognized by their high allele sequence divergence. At intermediate frequencies of convergence events, it will be impossible to reconstruct the correct phylogeny of an asexual clade from the sequences of protein coding genes. Convergence may be limited by allele sequence divergence and heterozygous chromosomal rearrangements which reduce the homology needed for recombination and result in aneuploidy after crossing-over or ploidy cycles.

Eukaryotic Cells↗

Evolution of repeated sequences in non-coding regions of the genome.

Repeated sequences are found ubiquitously in the eukaryotic genome. Population genetic studies on the evolution of such repeated sequences are reviewed while paying special attention to those sequences found in the non-coding regions of the genome. Specifically, the evolution of dispersed repeated sequences by the transposition as well as the evolution of short tandemly repeated sequences due to either replication slippage or unequal sister chromatid exchange are considered. The approach of combining both model and data analyses which has been successfully employed in the development of the neutral theory is also considered to be useful in better understanding the evolution and biological meaning of these sequences.

Animals↗

The amphioxus rab GDP-dissociation inhibitor (GDI) gene is neural-specific: implications for the evolution of chordate rab GDI genes.

The rab GDP-dissociation inhibitor (rab GDI) proteins are involved in the regulation of vesicle-mediated cellular transport. We isolated the amphioxus rab GDI gene, analyzed its expression during amphioxus development, and performed a phylogenetic analysis of the rab GDI family. In contrast to the two major rab GDI forms in mammals, the alpha and beta forms, there is only one rab GDI isoform in amphioxus. Our analysis indicates that the occurrence of the alpha and beta forms of rab GDI preceded the divergence of lineages leading to birds and mammals, and that the amphioxus rab GDI may have evolved directly from the common ancestor of both forms. While the mammalian rab GDI beta-genes are ubiquitously expressed, the rab GDI alpha genes are predominantly expressed in neural tissues. The expression analysis of the amphioxus rab GDI gene shows predominantly neural expression similar to that of the mammalian rab GDI alpha form, suggesting that the ancestral expression pattern of chordate rab GDI was neural. In addition, the chicken rab GDI beta-like gene also shows neural-specific expression, which indicates that the neural expression was retained in both early postduplication alpha and beta isoforms and that a novel function associated with ubiquitous expression may have evolved uniquely in mammals. These results reveal a likely scenario of functional divergence of the rab GDI genes after duplication of the ancestral gene. A similar pattern of evolution, in which one of the duplicated genes retained a role similar to that of the ancestral one while other genes were recruited into novel roles, was also observed in the analysis of chordate Otx and hedgehog genes. In the rab GDI, hedgehog, and Otx gene families, the gene retaining the ancestral role shows a lower rate of sequence evolution than its counterpart, which was recruited for a novel function.

Amino Acid Sequence↗

Purification and characterization of a peptide from amyloid-rich pancreases of type 2 diabetic patients.

Deposition of amyloid in pancreatic islets is a common feature in human type 2 diabetic subjects but because of its insolubility and low tissue concentrations, the structure of its monomer has not been determined. We describe a peptide, of calculated molecular mass 3905 Da, that was a major protein component of amyloid-rich pancreatic extracts of three type 2 diabetic patients. After collagenase treatment, an extract containing 20-50% amyloid was solubilized by sonication into 70% formic acid and the peptide was purified by gel filtration followed by reverse-phase high-performance liquid chromatography. We term this peptide diabetes-associated peptide, as it was not detected in extracts of pancreas from any of six normal subjects. Diabetes-associated peptide contains 37 amino acids and is 46% identical to the sequences of rat and human calcitonin gene-related peptide, indicating that these peptides are related in evolution. Sequence identities with conserved residues of the insulin A chain were also seen in a 16-residue segment. On extraction, the islet amyloid is particulate and insoluble like the core particles of Alzheimer disease. Their monomers have similar molecular masses, each having a hydropathic region that can probably form beta-pleated sheets. The accumulation of amyloid, including diabetes-associated peptide, in islets may impair islet function in type 2 diabetes mellitus.

Aged↗

A nonhyperthermophilic common ancestor to extant life forms.

The G+C nucleotide content of ribosomal RNA (rRNA) sequences is strongly correlated with the optimal growth temperature of prokaryotes. This property allows inference of the environmental temperature of the common ancestor to all life forms from knowledge of the G+C content of its rRNA sequences. A model of sequence evolution, assuming varying G+C content among lineages and unequal substitution rates among sites, was devised to estimate ancestral base compositions. This method was applied to rRNA sequences of various species representing the major lineages of life. The inferred G+C content of the common ancestor to extant life forms appears incompatible with survival at high temperature. This finding challenges a widely accepted hypothesis about the origin of life.

Animals↗

Evolutionary rate heterogeneity in proteins with long disordered regions.

The dominant view in protein science is that a three-dimensional (3-D) structure is a prerequisite for protein function. In contrast to this dominant view, there are many counterexample proteins that fail to fold into a 3-D structure, or that have local regions that fail to fold, and yet carry out function. Protein without fixed 3-D structure is called intrinsically disordered. Motivated by anecdotal accounts of higher rates of sequence evolution in disordered protein than in ordered protein we are exploring the molecular evolution of disordered proteins. To test whether disordered protein evolves more rapidly than ordered protein, pairwise genetic distances were compared between the ordered and the disordered regions of 26 protein families having at least one member with a structurally characterized region of disorder of 30 or more consecutive residues. For five families, there were no significant differences in pairwise genetic distances between ordered and disordered sequences. The disordered region evolved significantly more rapidly than the ordered region for 19 of the 26 families. The functions of these disordered regions are diverse, including binding sites for protein, DNA, or RNA and also including flexible linkers. The functions of some of these regions are unknown. The disordered regions evolved significantly more slowly than the ordered regions for the two remaining families. The functions of these more slowly evolving disordered regions include sites for DNA binding. More work is needed to understand the underlying causes of the variability in the evolutionary rates of intrinsically ordered and disordered protein.

Amino Acid Sequence↗

Constant relative rate of protein evolution and detection of functional diversification among bacterial, archaeal and eukaryotic proteins.

BACKGROUND: Detection of changes in a protein's evolutionary rate may reveal cases of change in that protein's function. We developed and implemented a simple relative rates test in an attempt to assess the rate constancy of protein evolution and to detect cases of functional diversification between orthologous proteins. The test was performed on clusters of orthologous protein sequences from complete bacterial genomes (Chlamydia trachomatis, C. muridarum and Chlamydophila pneumoniae), complete archaeal genomes (Pyrococcus horikoshii, P. abyssi and P. furiosus) and partially sequenced mammalian genomes (human, mouse and rat). RESULTS: Amino-acid sequence evolution rates are significantly correlated on different branches of phylogenetic trees representing the great majority of analyzed orthologous protein sets from all three domains of life. However, approximately 1% of the proteins from each group of species deviates from this pattern and instead shows variation that is consistent with an acceleration of the rate of amino-acid substitution, which may be due to functional diversification. Most of the putative functionally diversified proteins from all three species groups are predicted to function at the periphery of the cells and mediate their interaction with the environment. CONCLUSIONS: Relative rates of protein evolution are remarkably constant for the three species groups analyzed here. Deviations from this rate constancy are probably due to changes in selective constraints associated with diversification between orthologs. Functional diversification between orthologs is thought to be a relatively rare event. However, the resolution afforded by the test designed specifically for genomic-scale datasets allowed us to identify numerous cases of possible functional diversification between orthologous proteins.

Animals↗

Long-term evolution of the 5'UTR and a region of NS4 containing a CTL epitope of hepatitis C virus in two haemophilic patients.

Haemophilic patients exposed to unsterilized clotting factor concentrates prior to 1985 have become infected with hepatitis C virus (HCV). We have studied the sequence evolution of the 5'UTR and a region of NS4 over 12 years in one human immunodeficiency virus (HIV) positive haemophilic patient and 14 years for one HIV negative haemophilic patient. One sample each year from the date of HCV infection to 1994 was analysed for genotype, virus load and nucleotide sequence of the two genetic loci. Both patients were infected with HCV genotype 1 throughout the study period. The virus load profiles were similar except that the profile for the HIV infected patient was displaced 4 years earlier relative to the other patient. Mean divergence of the quasispecies at both the 5'UTR and NS4 loci was higher in the HIV coinfected patient. Phylogenetic analysis indicated that evolution of the 5'UTR was host independent, whereas the NS4 region containing a CD8 restricted CTL epitope evolved in a host specific fashion.

Amino Acid Sequence↗

Phylogenetic utility of the nuclear gene arginine decarboxylase: an example from Brassicaceae.

Arginine decarboxylase (ADC) is an important enzyme in the production of putrescine and polyamines in plants. It is encoded by a single or low-copy nuclear gene that lacks introns in sequences studied to date. The rate of Adc amino acid sequence evolution is similar to that of ndhF for the angiosperm family studied. Highly conserved regions provide several target sites for PCR priming and sequencing and aid in nucleotide and amino acid sequence alignment across a range of taxonomic levels, while a variable region provides an increased number of potentially informative characters relative to ndhF for the taxa surveyed. The utility of the Adc gene in plant molecular systematic studies is demonstrated by analysis of its partial nucleotide sequences obtained from 13 representatives of Brassicaceae and 3 outgroup taxa, 2 from the mustard oil clade (order Capparales) and 1 from the related order Malvales. Two copies of the Adc gene, Adc1 and Adc2, are found in all members of the Brassicaceae studied to data except the basal genus Aethionema. The resulting Adc gene tree provides robust phylogenetic data regarding relationships within the complex mustard family, as well as independent support for proposed tribal realignments based on other molecular data sets such as those from chloroplast DNA.

Arabidopsis↗