PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

A likelihood ratio test for evolutionary rate shifts and functional divergence among proteins.

Changes in protein function can lead to changes in the selection acting on specific residues. This can often be detected as evolutionary rate changes at the sites in question. A maximum-likelihood method for detecting evolutionary rate shifts at specific protein positions is presented. The method determines significance values of the rate differences to give a sound statistical foundation for the conclusions drawn from the analyses. A statistical test for detecting slowly evolving sites is also described. The methods are applied to a set of Myc proteins for the identification of both conserved sites and those with changing evolutionary rates. Those positions with conserved and changing rates are related to the structures and functions of their proteins. The results are compared with an earlier Bayesian method, thereby highlighting the advantages of the new likelihood ratio tests.

Amino Acid Sequence↗

Plant-insect interactions: double-dating associated insect and plant lineages reveals asynchronous radiations.

An increasing number of plant-insect studies using phylogenetic analysis suggest that cospeciation events are rare in plant-insect systems. Instead, nonrandom patterns of phylogenetic congruence are produced by phylogenetically conserved host switching (to related plants) or tracking of particular resources or traits (e.g., chemical). The dominance of host switching in many phytophagous insect groups may make the detection of genuine cospeciation events difficult. One important test of putative cospeciation events is to verify whether reciprocal speciation is temporally plausible. We explored techniques for double-dating of both plant and insect phylogenies. We use dated molecular phylogenies of a psyllid (Hemiptera)-Genisteae (Fabaceae) system, a predominantly monophagous insect-plant association widespread on the Atlantic Macaronesian islands. Phylogenetic reconciliation analysis suggests high levels of parallel cladogenesis between legumes and psyllids. However, dating using molecular clocks calibrated on known geological ages of the Macaronesian islands revealed that the legume and psyllid radiations were not contemporaneous but sequential. Whereas the main plant radiation occurred some 8 million years ago, the insect radiation occurred about 3 million years ago. We estimated that >60% of the psyllid speciation has resulted from host switching between related hosts. The only evidence for true cospeciation is in the much more recent and localized radiation of genistoid legumes in the Canary Islands, where the psyllid and legume radiations have been partially contemporaneous. The identification of specific cospeciation events over this time period, however, is hindered by the phylogenetic uncertainty in both legume and psyllid phylogenies due to the apparent rapidity of the species radiations.

Animals↗

The LIS1-related NUDF protein of Aspergillus nidulans interacts with the coiled-coil domain of the NUDE/RO11 protein.

The nudF gene of the filamentous fungus Aspergillus nidulans acts in the cytoplasmic dynein/dynactin pathway and is required for distribution of nuclei. NUDF protein, the product of the nudF gene, displays 42% sequence identity with the human protein LIS1 required for neuronal migration. Haploinsufficiency of the LIS1 gene causes a malformation of the human brain known as lissencephaly. We screened for multicopy suppressors of a mutation in the nudF gene. The product of the nudE gene isolated in the screen, NUDE, is a homologue of the nuclear distribution protein RO11 of Neurospora crassa. The highly conserved NH(2)-terminal coiled-coil domain of the NUDE protein suffices for protein function when overexpressed. A similar coiled-coil domain is present in several putative human proteins and in the mitotic phosphoprotein 43 (MP43) of X. laevis. NUDF protein interacts with the Aspergillus NUDE coiled-coil in a yeast two-hybrid system, while human LIS1 interacts with the human homologue of the NUDE/RO11 coiled-coil and also the Xenopus MP43 coiled-coil. In addition, NUDF coprecipitates with an epitope-tagged NUDE. The fact that NUDF and LIS1 interact with the same protein domain strengthens the notion that these two proteins are functionally related.

1-Alkyl-2-acetylglycerophosphocholine Esterase↗

Analysis of conserved microsatellite sequences suggests closer relationship between water buffalo Bubalus bubalis and sheep Ovis aries.

The distribution and evolutionary pattern of the conserved microsatellite repeat sequences (CA)n, (TGG)6, and (GGAT)4 were studied to determine the divergence time and phylogenetic position of the water buffalo, Bubalus bubalis. The mean allelic frequencies of these repeat loci showed a high level of heterozygosity among the euartiodactyls (buffalo, cattle, sheep, and goat). Genetic distances calculated from the allelic frequencies of these microsatellites were used to position Bubalus bubalis in the phylogenetic tree. The tree topology revealed a closer proximity of the Bubalus bubalis to the Ovis aries (sheep) genome than to other domestic species. The estimated time of divergence of the water buffalo genome relative to cattle, goat, sheep, pig, rabbit, and horse was found to be 21, 0.5, 0.7, 94, 20.3, and 408 million years (Myr), respectively. Although water buffaloes share morphological and biochemical similarities with cattle, our study using the microsatellite sequences places the bubaline species in an entirely new phylogenetic position. Our results also suggest that with respect to these repeat loci, the water buffalo genome shares a common ancestry with sheep and goat after the divergence of subfamily Bovinae (Bos taurus) from the family Bovidae.

Animals↗

Sequence alignment in molecular biology.

Molecular biology is becoming a computationally intense realm of contemporary science and faces some of the current grand scientific challenges. In its context, tools that identify, store, compare and analyze effectively large and growing numbers of bio-sequences are found of increasingly crucial importance. Biosequences are routinely compared or aligned, in a variety of ways, to infer common ancestry, to detect functional equivalence, or simply while searching for similar entries in a database. A considerable body of knowledge has accumulated on sequence alignment during the past few decades. Without pretending to be exhaustive, this paper attempts a survey of some criteria of wide use in sequence alignment and comparison problems, and of the corresponding solutions. The paper is based on presentations and literature given at the Workshop on Sequence Alignment held at Princeton, N.J., in November 1994, as part of the DIMACS Special Year on Mathematical Support for Molecular Biology.

Algorithms↗

Bellerophon: a program to detect chimeric sequences in multiple sequence alignments.

SUMMARY: Bellerophon is a program for detecting chimeric sequences in multiple sequence datasets by an adaption of partial treeing analysis. Bellerophon was specifically developed to detect 16S rRNA gene chimeras in PCR-clone libraries of environmental samples but can be applied to other nucleotide sequence alignments. AVAILABILITY: Bellerophon is available as an interactive web server at http://foo.maths.uq.edu.au/~huber/bellerophon.pl

Algorithms↗

The ASCH superfamily: novel domains with a fold related to the PUA domain and a potential role in RNA metabolism.

Several studies show that transcription coactivators are often bi-functional ribonucleoprotein complexes that also regulate pre-mRNA processing and splicing decisions. Using sensitive sequence profile searches and structural comparisons we show that the C-terminal domain of the human coactivator protein ASC-1 defines a novel superfamily, the ASC-1 homology (ASCH) domain. The approximately 110 amino acid long ASCH domains are widely represented in all the three superkingdoms of life and several prokaryotic viruses. We show that the ASCH superfamily adopts a beta-barrel fold similar to the PUA domain superfamily. Using multiple lines of evidence, we suggest that members of the ASCH superfamily are likely to function as RNA-binding domains in contexts related to coactivation, RNA-processing and possibly prokaryotic translation regulation. Structural analysis of ASCH domains reveals the presence of a potential RNA-binding cleft associated with a conserved sequence motif, which is characteristic of this superfamily. Despite their similar structure, the ASCH and PUA domains appear to occupy distinct functional niches, with the former domains typically occurring in a standalone form in polypeptides, and the latter domains showing fusions to a variety of RNA-modifying enzymes.

Animals↗

Morphological, molecular, and chromosomal discrimination of cryptic Anopheles (Nyssorhynchus) (Diptera: Culicidae) from South America.

Based on similarity of male genitalia, the malaria vector Anopheles trinkae Faran from the eastern Andean piedmont of Colombia, Ecuador, Peru, and Bolivia was determined by Peyton (1993) to be a junior synonym of An. dunhami Causey, then known from a single locality in Amazonian Brazil. Following an appraisal of molecular, chromosomal, and morphological characters, we conclude herein that the 2 taxa are specifically distinct and remove An. trinkae from synonymy with An. dunhami. Eggs of the 2 species are distinguished easily by the anterior crown, long floats, and closed deck that occur only in An. trinkae. The X chromosome of larval polytenes is divisible into R and L arms in An. dunhami, but not in An. trinkae. A phenogram based on banding pattern scores from 18 random amplified polymorphic DNA primers separated with 100% resolution An. dunhami, An. trinkae, Anopheles nuneztovari Gabaldón and Anopheles darlingi Root. In the ITS2 region of rDNA, 25% of base sites distinguished An. trinkae from An. dunhami and 21% from the related An. nuneztovari; males of these 3 species had accessory glands of significantly different sizes. Preliminary isoenzyme screening indicated that 3 of 11 loci were diagnostic for separating An. trinkae from An. dunhami. The results indicate that An. dunhami is related more closely to An. nuneztovari than to An. trinkae and illustrate the merits of a multidisciplinary approach to mosquito systematics.

Animals↗

An evolutionary model for protein-coding regions with conserved RNA structure.

Here we present a model of nucleotide substitution in protein-coding regions that also encode the formation of conserved RNA structures. In such regions, apparent evolutionary context dependencies exist, both between nucleotides occupying the same codon and between nucleotides forming a base pair in the RNA structure. The overlap of these fundamental dependencies is sufficient to cause "contagious" context dependencies which cascade across many nucleotide sites. Such large-scale dependencies challenge the use of traditional phylogenetic models in evolutionary inference because they explicitly assume evolutionary independence between short nucleotide tuples. In our model we address this by replacing context dependencies within codons by annotation-specific heterogeneity in the substitution process. Through a general procedure, we fragment the alignment into sets of short nucleotide tuples based on both the protein coding and the structural annotation. These individual tuples are assumed to evolve independently, and the different tuple sets are assigned different annotation-specific substitution models shared between their members. This allows us to build a composite model of the substitution process from components of traditional phylogenetic models. We applied this to a data set of full-genome sequences from the hepatitis C virus where five RNA structures are mapped within the coding region. This allowed us to partition the effects of selection on different structural elements and to test various hypotheses concerning the relation of these effects. Of particular interest, we found evidence of a functional role of loop and bulge regions, as these were shown to evolve according to a different and more constrained selective regime than the nonpairing regions outside the RNA structures. Other potential applications of the model include comparative RNA structure prediction in coding regions and RNA virus phylogenetics.

Base Sequence↗

The murine IgM secretory poly(A) site contains dual upstream and downstream elements which affect polyadenylation.

Regulation of polyadenylation efficiency at the secretory poly(A) site plays an essential role in gene expression at the immunoglobulin (IgM) locus. At this poly(A) site the consensus AAUAAA hexanucleotide sequence is embedded in an extended AU-rich region and there are two downstream GU-rich regions which are suboptimally placed. As these sequences are involved in formation of the polyadenylation pre-initiation complex, we examined their function in vivo and in vitro . We show that the upstream AU-rich region can function in the absence of the consensus hexanucleotide sequence both in vivo and in vitro and that both GU-rich regions are necessary for full polyadenylation activity in vivo and for formation of polyadenylation-specific complexes in vitro . Sequence comparisons reveal that: (i) the dual structure is distinct for the IgM secretory poly(A) site compared with other immunoglobulin isotype secretory poly(A) sites; (ii) the presence of an AU-rich region close to the consensus hexanucleotide is evolutionarily conserved for IgM secretory poly(A) sites. We propose that the dual structure of the IgM secretory poly(A) site provides a flexibility to accommodate changes in polyadenylation complex components during regulation of polyadenylation efficiency.

Alternative Splicing↗

Analysis of TALE superclass homeobox genes (MEIS, PBC, KNOX, Iroquois, TGIF) reveals a novel domain conserved between plants and animals.

A new Caenorhabditis elegans homeobox gene, ceh-25, is described that belongs to the TALE superclass of atypical homeodomains, which are characterized by three extra residues between helix 1 and helix 2. ORF and PCR analysis revealed a novel type of alternative splicing within the homeobox. The alternative splicing occurs such that two different homeodomains can be generated, which differ in their first 25 amino acids. ceh-25 is an orthologue of the vertebrate Meis genes and it shares a new conserved domain of 130 amino acids with them. A thorough analysis of all TALE homeobox genes was performed and a new classification is presented. Four TALE classes are identified in animals: PBC, MEIS, TGIF and IRO (Iroquois); two types in fungi: the mating type genes (M-ATYP) and the CUP genes; and two types in plants: KNOX and BEL. The IRO class has a new conserved motif downstream of the homeodomain. For the KNOX class, a conserved domain, the KNOX domain, was defined upstream of the homeodomain. Comparison of the KNOX domain and the MEIS domain shows significant sequence similarity revealing the existence of an archetypal group of homeobox genes that encode two associated conserved domains. Thus TALE homeobox genes were already present in the common ancestor of plants, fungi and animals and represent a branch distinct from the typical homeobox genes.

Alternative Splicing↗

Alternative splicing regulation at tandem 3' splice sites.

Alternative splicing (AS) constitutes a major mechanism creating protein diversity in humans. Previous bioinformatics studies based on expressed sequence tag and mRNA data have identified many AS events that are conserved between humans and mice. Of these events, approximately 25% are related to alternative choices of 3' and 5' splice sites. Surprisingly, half of all these events involve 3' splice sites that are exactly 3 nt apart. These tandem 3' splice sites result from the presence of the NAGNAG motif at the acceptor splice site, recently reported to be widely spread in the human genome. Although the NAGNAG motif is common in human genes, only a small subset of sites with this motif is confirmed to be involved in AS. We examined the NAGNAG motifs and observed specific features such as high sequence conservation of the motif, high conservation of approximately 30 bp at the intronic regions flanking the 3' splice site and overabundance of cis-regulatory elements, which are characteristic of alternatively spliced tandem acceptor sites and can distinguish them from the constitutive sites in which the proximal NAG splice site is selected. Our findings imply that AS at tandem splice sites and constitutive splicing of the distal NAG are highly regulated.

Alternative Splicing↗

A theoretical model of restriction endonuclease NlaIV in complex with DNA, predicted by fold recognition and validated by site-directed mutagenesis and circular dichroism spectroscopy.

Restriction enzymes (REases) are commercial reagents commonly used in DNA manipulations and mapping. They are regarded as very attractive models for studying protein-DNA interactions and valuable targets for protein engineering. Their amino acid sequences usually show no similarities to other proteins, with rare exceptions of other REases that recognize identical or very similar sequences. Hence, they are extremely hard targets for structure prediction and modeling. NlaIV is a Type II REase, which recognizes the interrupted palindromic sequence GGNNCC (where N indicates any base) and cleaves it in the middle, leaving blunt ends. NlaIV shows no sequence similarity to other proteins and virtually nothing is known about its sequence-structure-function relationships. Using protein fold recognition, we identified a remote relationship between NlaIV and EcoRV, an extensively studied REase, which recognizes the GATATC sequence and whose crystal structure has been determined. Using the 'FRankenstein's monster' approach we constructed a comparative model of NlaIV based on the EcoRV template and used it to predict the catalytic and DNA-binding residues. The model was validated by site-directed mutagenesis and analysis of the activity of the mutants in vivo and in vitro as well as structural characterization of the wild-type enzyme and two mutants by circular dichroism spectroscopy. The structural model of the NlaIV-DNA complex suggests regions of the protein sequence that may interact with the 'non-specific' bases of the target and thus it provides insight into the evolution of sequence specificity in restriction enzymes and may help engineer REases with novel specificities. Before this analysis was carried out, neither the three-dimensional fold of NlaIV, its evolutionary relationships or its catalytic or DNA-binding residues were known. Hence our analysis may be regarded as a paradigm for studies aiming at reducing 'white spaces' on the evolutionary landscape of sequence-function relationships by combining bioinformatics with simple experimental assays.

Amino Acid Sequence↗

Analysis of the Petunia TM6 MADS box gene reveals functional divergence within the DEF/AP3 lineage.

Antirrhinum majus DEFICIENS (DEF) and Arabidopsis thaliana APETALA3 (AP3) MADS box proteins are required to specify petal and stamen identity. Sampling of DEF/AP3 homologs revealed two types of DEF/AP3 proteins, euAP3 and TOMATO MADS BOX GENE6 (TM6), within core eudicots, and we show functional divergence in Petunia hybrida euAP3 and TM6 proteins. Petunia DEF (also known as GREEN PETALS [GP]) is expressed mainly in whorls 2 and 3, and its expression pattern remains unchanged in a blind (bl) mutant background, in which the cadastral C-repression function in the perianth is impaired. Petunia TM6 functions as a B-class organ identity protein only in the determination of stamen identity. Atypically, Petunia TM6 is regulated like a C-class rather than a B-class gene, is expressed mainly in whorls 3 and 4, and is repressed by BL in the perianth, thereby preventing involvement in petal development. A promoter comparison between DEF and TM6 indicates an important change in regulatory elements during or after the duplication that resulted in euAP3- and TM6-type genes. Surprisingly, although TM6 normally is not involved in petal development, 35S-driven TM6 expression can restore petal development in a def (gp) mutant background. Finally, we isolated both euAP3 and TM6 genes from seven solanaceous species, suggesting that a dual euAP3/TM6 B-function system might be the rule in the Solanaceae.

Base Sequence↗

The S15 self-incompatibility haplotype in Brassica oleracea includes three S gene family members expressed in stigmas.

Self-incompatibility in Brassica is controlled by a single, highly polymorphic locus that extends over several hundred kilobases and includes several expressed genes. Two stigma proteins, the S locus receptor kinase (SRK) and the S locus glycoprotein (SLG), are encoded by genes located at the S locus and are thought to be involved in the recognition of self-pollen by the stigma. We report here that two different SLG genes, SLGA and SLGB, are located at the S locus in the class II, pollen-recessive S15 haplotype. Both genes are interrupted by a single intron; however, SLGA encodes both soluble and membrane-anchored forms of SLG, whereas SLGB encodes only soluble SLG proteins. Thus, including SRK, the S locus in the S15 haplotype contains at least three members of the S gene family. The protein products of these three genes have been characterized, and each SLG glycoform was assigned to an SLG gene. Evidence is presented that the S2 and S5 haplotypes carry only one or the other of the SLG genes, indicating either that they are redundant or that they are not required for the self-incompatibility response.

Alleles↗

Complete genome sequence of the broad host range single-stranded RNA phage PRR1 places it in the Levivirus genus with characteristics shared with Alloleviviruses.

Single-stranded RNA (ssRNA) bacteriophages of the family Leviviridae infect gram-negative bacteria. They are restricted to a single host genus. Phage PRR1 is an exception, having a broad host range due to the promiscuity of the receptor encoded by the IncP plasmid. Here we report the complete genome sequence of PRR1. Three proteins homologous with those of other ssRNA phages, i.e., maturation, coat, and replicase proteins, were identified. A fourth protein has a lysis function. Comparison of PRR1 with other members of the Leviviridae family places PRR1 in the genus Levivirus with some characteristics more similar to those of members of the genus Allolevivirus.

Allolevivirus↗

Genetic diversity among Lassa virus strains.

The arenavirus Lassa virus causes Lassa fever, a viral hemorrhagic fever that is endemic in the countries of Nigeria, Sierra Leone, Liberia, and Guinea and perhaps elsewhere in West Africa. To determine the degree of genetic diversity among Lassa virus strains, partial nucleoprotein (NP) gene sequences were obtained from 54 strains and analyzed. Phylogenetic analyses showed that Lassa viruses comprise four lineages, three of which are found in Nigeria and the fourth in Guinea, Liberia, and Sierra Leone. Overall strain variation in the partial NP gene sequence was found to be as high as 27% at the nucleotide level and 15% at the amino acid level. Genetic distance among Lassa strains was found to correlate with geographic distance rather than time, and no evidence of a "molecular clock" was found. A method for amplifying and cloning full-length arenavirus S RNAs was developed and used to obtain the complete NP and glycoprotein gene (GP1 and GP2) sequences for two representative Nigerian strains of Lassa virus. Comparison of full-length gene sequences for four Lassa virus strains representing the four lineages showed that the NP gene (up to 23.8% nucleotide difference and 12.0% amino acid difference) is more variable than the glycoprotein genes. Although the evolutionary order of descent within Lassa virus strains was not completely resolved, the phylogenetic analyses of full-length NP, GP1, and GP2 gene sequences suggested that Nigerian strains of Lassa virus were ancestral to strains from Guinea, Liberia, and Sierra Leone. Compared to the New World arenaviruses, Lassa and the other Old World arenaviruses have either undergone a shorter period of diverisification or are evolving at a slower rate. This study represents the first large-scale examination of Lassa virus genetic variation.

Africa, Western↗

Evolutionary conserved chromosomal segments in the human karyotype are bounded by unstable chromosome bands.

In this paper an ancestral karyotype for primates, defining for the first time the ancestral chromosome morphology and the banding patterns, is proposed, and the ancestral syntenic chromosomal segments are identified in the human karyotype. The chromosomal bands that are boundaries of ancestral segments are identified. We have analyzed from data published in the literature 35 different primate species from 19 genera, using the order Scandentia, as well as other published mammalian species as out-groups, and propose an ancestral chromosome number of 2n = 54 for primates, which includes the following chromosomal forms: 1(a+c(1)), 1(b+c(2)), 2a, 2b, 3/21, 4, 5, 6, 7a, 7b, 8, 9, 10a, 10b, 11, 12a/22a, 12b/22b, 13, 14/15, 16a, 16b, 17, 18, 19a, 19b, 20 and X and Y. From this analysis, we have been able to point out the human chromosome bands more "prone" to breakage during the evolutionary pathways and/or pathology processes. We have observed that 89.09% of the human chromosome bands, which are boundaries for ancestral chromosome segments, contain common fragile sites and/or intrachromosomal telomeric-like sequences. A more in depth analysis of twelve different human chromosomes has allowed us to determine that 62.16% of the chromosomal bands implicated in inversions and 100% involved in fusions/fissions correspond to fragile sites, intrachromosomal telomeric-like sequences and/or bands significantly affected by X irradiation. In addition, 73% of the bands affected in pathological processes are co-localized in bands where fragile sites, intrachromosomal telomeric-like sequences, bands significantly affected by X irradiation and/or evolutionary chromosomal bands have been described. Our data also support the hypothesis that chromosomal breakages detected in pathological processes are not randomly distributed along the chromosomes, but rather concentrate in those important evolutionary chromosome bands which correspond to fragile sites and/or intrachromosomal telomeric-like sequences.

Alouatta↗