PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Hox cluster genomics in the horn shark, Heterodontus francisci.

Reconstructing the evolutionary history of Hox cluster origins will lead to insights into the developmental and evolutionary significance of Hox gene clusters in vertebrate phylogeny and to their role in the origins of various vertebrate body plans. We have isolated two Hox clusters from the horn shark, Heterodontus francisci. These have been sequenced and compared with one another and with other chordate Hox clusters. The results show that one of the horn shark clusters (HoxM) is orthologous to the mammalian HoxA cluster and shows a structural similarity to the amphioxus cluster, whereas the other shark cluster (HoxN) is orthologous to the mammalian HoxD cluster based on cluster organization and a comparison with noncoding and Hox gene-coding sequences. The persistence of an identifiable HoxA cluster over an 800-million-year divergence time demonstrates that the Hox gene clusters are highly integrated and structured genetic entities. The data presented herein identify many noncoding sequence motifs conserved over 800 million years that may function as genetic control motifs essential to the developmental process.

Animals↗

The root of the universal tree and the origin of eukaryotes based on elongation factor phylogeny.

The genes for the protein synthesis elongation factors Tu (EF-Tu) and G (EF-G) are the products of an ancient gene duplication, which appears to predate the divergence of all extant organismal lineages. Thus, it should be possible to root a universal phylogeny based on either protein using the second protein as an outgroup. This approach was originally taken independently with two separate gene duplication pairs, (i) the regulatory and catalytic subunits of the proton ATPases and (ii) the protein synthesis elongation factors EF-Tu and EF-G. Questions about the orthology of the ATPase genes have obscured the former results, and the elongation factor data have been criticized for inadequate taxonomic representation and alignment errors. We have expanded the latter analysis using a broad representation of taxa from all three domains of life. All phylogenetic methods used strongly place the root of the universal tree between two highly distinct groups, the archaeons/eukaryotes and the eubacteria. We also find that a combined data set of EF-Tu and EF-G sequences favors placement of the eukaryotes within the Archaea, as the sister group to the Crenarchaeota. This relationship is supported by bootstrap values of 60-89% with various distance and maximum likelihood methods, while unweighted parsimony gives 58% support for archaeal monophyly.

Adenosine Triphosphatases↗

Mouse cytoplasmic polyadenylylation element binding protein: an evolutionarily conserved protein that interacts with the cytoplasmic polyadenylylation elements of c-mos mRNA.

Cytoplasmic polyadenylylation is an essential process that controls the translation of maternal mRNAs during early development and depends on two cis elements in the 3' untranslated region: the polyadenylylation hexanucleotide AAUAAA and a U-rich cytoplasmic polyadenylylation element (CPE). In searching for factors that could mediate cytoplasmic polyadenylylation of mouse c-mos mRNA, which encodes a serine/threonine kinase necessary for oocyte maturation, we have isolated the mouse homolog of CPEB, a protein that binds to the CPEs of a number of mRNAs in Xenopus oocytes and is required for their polyadenylylation. Mouse CPEB (mCPEB) is a 62-kDa protein that binds to the CPEs of c-mos mRNA. mCPEB mRNA is present in the ovary, testis, and kidney; within the ovary, this RNA is restricted to oocytes. mCPEB shows 80% overall identity with its Xenopus counterpart, with a higher homology in the carboxyl-terminal portion, which contains two RNA recognition motifs and a cysteine/histidine repeat. Proteins from arthropods and nematodes are also similar to this region, suggesting an ancient and widely used mechanism to control polyadenylylation and translation.

Amino Acid Sequence↗

Evolutionary analyses of the 12-kDa acidic ribosomal P-proteins reveal a distinct protein of higher plant ribosomes.

The P-protein complex of eukaryotic ribosomes forms a lateral stalk structure in the active site of the large ribosomal subunit and is thought to assist in the elongation phase of translation by stimulating GTPase activity of elongation factor-2 and removal of deacylated tRNA. The complex in animals, fungi, and protozoans is composed of the acidic phosphoproteins P0 (35 kDa), P1 (11-12 kDa), and P2 (11-12 kDa). Previously we demonstrated by protein purification and microsequencing that ribosomes of maize (Zea mays L.) contain P0, one type of P1, two types of P2, and a distinct P1/P2 type protein designated P3. Here we implemented distance matrices, maximum parsimony, and neighbor-joining analyses to assess the evolutionary relationships between the 12 kDa P-proteins of maize and representative eukaryotic species. The analyses identify P3, found to date only in mono- and dicotyledonous plants, as an evolutionarily distinct P-protein. Plants possess three distinct groups of 12 kDa P-proteins (P1, P2, and P3), whereas animals, fungi, and protozoans possess only two distinct groups (P1 and P2). These findings demonstrate that the P-protein complex has evolved into a highly divergent complex with respect to protein composition despite its critical position within the active site of the ribosome.

Amino Acid Sequence↗

Box H and box ACA are nucleolar localization elements of U17 small nucleolar RNA.

The nucleolar localization elements (NoLEs) of U17 small nucleolar RNA (snoRNA), which is essential for rRNA processing and belongs to the box H/ACA snoRNA family, were analyzed by fluorescence microscopy. Injection of mutant U17 transcripts into Xenopus laevis oocyte nuclei revealed that deletion of stems 1, 2, and 4 of U17 snoRNA reduced but did not prevent nucleolar localization. The deletion of stem 3 had no adverse effect. Therefore, the hairpins of the hairpin-hinge-hairpin-tail structure formed by these stems are not absolutely critical for nucleolar localization of U17, nor are sequences within stems 1, 3, and 4, which may tether U17 to the rRNA precursor by base pairing. In contrast, box H and box ACA are major NoLEs; their combined substitution or deletion abolished nucleolar localization of U17 snoRNA. Mutation of just box H or just the box ACA region alone did not fully abolish the nucleolar localization of U17. This indicates that the NoLEs of the box H/ACA snoRNA family function differently from the bipartite NoLEs (conserved boxes C and D) of box C/D snoRNAs, where mutation of either box alone prevents nucleolar localization.

Animals↗

Genomic divergence between human and chimpanzee estimated from large-scale alignments of genomic sequences.

To study the genomic divergence between human and chimpanzee, large-scale genomic sequence alignments were performed. The genomic sequences of human and chimpanzee were first masked with the RepeatMasker and the repeats were excluded before alignments. The repeats were then reinserted into the alignments of nonrepetitive segments and entire sequences were aligned again. A total of 2.3 million base pairs (Mb) of genomic sequences, including repeats, were aligned and the average nucleotide divergence was estimated to be 1.22%. The Jukes-Cantor (JC) distances (nucleotide divergences) in nonrepetitive (1.44 Mb) and repetitive sequences (0.86 Mb) are 1.14% and 1.34%, respectively, suggesting a slightly higher average rate in repetitive sequences. Annotated coding and noncoding regions of homologous chimpanzee genes were also retrieved from GenBank and compared. The average synonymous and nonsynonymous divergences in 88 coding genes are 1.48% and 0.55%, respectively. The JC distances in intron, 5' flanking, 3' flanking, promoter, and pseudogene regions are 1.47%, 1.41%, 1.68%, 0.75%, and 1.39%, respectively. It is not clear why the genetic distances in most of these regions are somewhat higher than those in genomic sequences. One possible explanation is that some of the genes may be located in regions with higher mutation rates.

Animals↗

IQPNNI: moving fast through tree space and stopping in time.

An efficient tree reconstruction method (IQPNNI) is introduced to reconstruct a phylogenetic tree based on DNA or amino acid sequence data. Our approach combines various fast algorithms to generate a list of potential candidate trees. The key ingredient is the definition of so-called important quartets (IQs), which allow the computation of an intermediate tree in O(n(2)) time for n sequences. The resulting tree is then further optimized by applying the nearest neighbor interchange (NNI) operation. Subsequently a random fraction of the sequences is deleted from the best tree found so far. The deleted sequences are then re-inserted in the smaller tree using the important quartet puzzling (IQP) algorithm. These steps are repeated several times and the best tree, with respect to the likelihood criterion, is considered as the inferred phylogenetic tree. Moreover, we suggest a rule which indicates when to stop the search. Simulations show that IQPNNI gives a slightly better accuracy than other programs tested. Moreover, we applied the approach to 218 small subunit rRNA sequences and 500 rbcL sequences. We found trees with higher likelihood compared to the results by others. A program to reconstruct DNA or amino acid based phylogenetic trees is available online (http://www.bi.uni-duesseldorf.de/software/iqpnni).

Algorithms↗

QuasiMotiFinder: protein annotation by searching for evolutionarily conserved motif-like patterns.

Sequence signature databases such as PROSITE, which include amino acid segments that are indicative of a protein's function, are useful for protein annotation. Lamentably, the annotation is not always accurate. A signature may be falsely detected in a protein that does not carry out the associated function (false positive prediction, FP) or may be overlooked in a protein that does carry out the function (false negative prediction, FN). A new approach has emerged in which a signature is replaced with a sequence profile, calculated based on multiple sequence alignment (MSA) of homologous proteins that share the same function. This approach, which is superior to the simple pattern search, essentially searches with the sequence of the query protein against an MSA library. We suggest here an alternative approach, implemented in the QuasiMotiFinder web server (http://quasimotifinder.tau.ac.il/), which is based on a search with an MSA of homologous query proteins against the original PROSITE signatures. The explicit use of the average evolutionary conservation of the signature in the query proteins significantly reduces the rate of FP prediction compared with the simple pattern search. QuasiMotiFinder also has a reduced rate of FN prediction compared with simple pattern searches, since the traditional search for precise signatures has been replaced by a permissive search for signature-like patterns that are physicochemically similar to known signatures. Overall, QuasiMotiFinder and the profile search are comparable to each other in terms of performance. They are also complementary to each other in that signatures that are falsely detected in (or overlooked by) one may be correctly detected by the other.

Amino Acid Motifs↗

Secondary structure of mitochondrial 12S rRNA among fish and its phylogenetic applications.

The complete 12S ribosomal RNA(rRNA) sequences from 23 gobioid species and nine diverse assortments of other fish species were employed to establish a core secondary structure model for fish 12S rRNA. Of the 43 stems recognized, 41 were supported by at least some compensatory evidence among vertebrates. The rates of nucleotide substitution were lower in stems than in loops. This may produce less phylogenetic information in stems when recently diverged taxa are compared. An analysis of compensatory substitution shows that the percentage of covariation is 68%, and the weighting factor for phylogenetic analyses to account for the dependence of mutations should be 0.66. Different stem-loop weighting schemes applied to the analyses of phylogenetic relationships of the Gobioidei indicate that down-weighting paired regions because of nonindependence could not improve the present phylogenetic analysis. A biased nucleotide composition (adenine% [A%] > thymine% [T%], cytosine% [C%] > guanine% [G%]) in the loop regions was also observed in the mammalian counterpart. The excess of A and C in the loop regions may be because of the asymmetric mechanism of mtDNA replication, which leads to the spontaneous deamination of C and A. This process may also be responsible for a transition-transversion bias and the patterns of nucleotide substitutions in both stems and loops.

Animals↗

The plastid chromosome of Atropa belladonna and its comparison with that of Nicotiana tabacum: the role of RNA editing in generating divergence in the process of plant speciation.

The nuclear and plastid genomes of the plant cell form a coevolving unit which in interspecific combinations can lead to genetic incompatibility of compartments even between closely related taxa. This phenomenon has been observed for instance in Atropa-Nicotiana cybrids. We have sequenced the plastid chromosome of Atropa belladonna (deadly nightshade), a circular DNA molecule of 156,688 bp, and compared it with the corresponding published sequence of its relative Nicotiana tabacum (tobacco) to understand how divergence at the level of this genome can contribute to nuclear-plastid incompatibilities and to speciation. It appears that (1) regulatory elements, i.e., promoters as well as translational and replicational signal elements, are well conserved between the two species; (2) genes--including introns--are even more highly conserved, with differences residing predominantly in regions of low functional importance; and (3) RNA editotypes differ between the two species, which makes this process an intriguing candidate for causing rapid reproductive isolation of populations.

Amino Acid Sequence↗

A simple method for classifying genes and a bootstrap test for classifications.

A new simple method for classifying genes is proposed based on Klastorin's method. This method classifies genes into monophyletic groups which are made distinct from each other by evolutionary changes. The method is applicable as long as the phylogenetic tree of genes is obtained. There is a fast algorithm for obtaining the classification. A bootstrap test of a classification is also presented. As an example, we classified opsin genes. The classification obtained by this method is the same as the previous classification based on the function of opsins.

Algorithms↗

The evolutionary landscape of functional model proteins.

To study the distinct influences of structure and function on evolution, we propose a minimalist model for proteins with binding pockets, called functional model proteins, based on a shifted-HP model on a two-dimensional square lattice. These model proteins are not maximally compact and contain an empty lattice site surrounded by at least three nearest neighbors, thus providing a binding pocket. Functional model proteins possess a unique native state, cooperative folding and tolerance to mutation. Due to the explicit functionality in these models (by design), we have been able to explore their fitness or evolutionary landscapes, as characterized by the size and distribution of homologous families and by the complexity of the inter-relatedness of the functional model proteins. Mindful that these minimalist models are highly idealized and two-dimensional, functional model proteins should nevertheless provide a useful means for exploring the constraints of maintaining structure and function on the evolution of proteins.

Amino Acid Sequence↗

Evolutionary trace analysis of TGF-beta and related growth factors: implications for site-directed mutagenesis.

The TGF-beta family of growth factors contains a large number of homologous proteins, grouped in several subfamilies on the basis of sequence identity. These subgroups can be combined into three broader groups of related cytokines, with marked specificities for their cellular receptors: the TGF-betas, the activins and the BMPs/GDFs. Although structural information is available for some members of the TGF-beta family, very little is known about the way in which these growth factors interact with the extra-cellular domains of their multiple cell surface receptors or with the specific protein inhibitors thought to modulate their activity. In this paper, we use the evolutionary trace method [Lichtarge et al. (1996) J. Mol. Biol., 257, 342-358] to locate two functional patches on the surface of TGF-beta-like growth factors. The first of these is centred on a conserved proline (P(36) in TGF-betas 1-3) and contains two amino acids which could account for the receptor specificity of TGF-betas (H(34) and E(35)). The second patch is located on the other side of the growth factor protomer and surrounds a hydrophobic cavity, large enough to accommodate the side chain of an aromatic residue. In addition to two conserved tryptophans at positions 30 and 32, the main protagonists in this potential binding interface are found at positions 31, 92, 93 and 98. Several mutagenesis studies have highlighted the importance of the C-terminal region of the growth factor molecule in TGF-betas and of residues in activin A equivalent to positions 31 and 94 of the TGF-betas for the binding of type II receptors to these ligands. These data, together with our improved knowledge of possible functional residues, can be used in future structure-function analysis experiments.

Amino Acid Motifs↗

A phylogenetic hypothesis for passerine birds: taxonomic and biogeographic implications of an analysis of nuclear DNA sequence data.

Passerine birds comprise over half of avian diversity, but have proved difficult to classify. Despite a long history of work on this group, no comprehensive hypothesis of passerine family-level relationships was available until recent analyses of DNA-DNA hybridization data. Unfortunately, given the value of such a hypothesis in comparative studies of passerine ecology and behaviour, the DNA-hybridization results have not been well tested using independent data and analytical approaches. Therefore, we analysed nucleotide sequence variation at the nuclear RAG-1 and c-mos genes from 69 passerine taxa, including representatives of most currently recognized families. In contradiction to previous DNA-hybridization studies, our analyses suggest paraphyly of suboscine passerines because the suboscine New Zealand wren Acanthisitta was found to be sister to all other passerines. Additionally, we reconstructed the parvorder Corvida as a basal paraphyletic grade within the oscine passerines. Finally, we found strong evidence that several family-level taxa are misplaced in the hybridization results, including the Alaudidae, Irenidae, and Melanocharitidae. The hypothesis of relationships we present here suggests that the oscine passerines arose on the Australian continental plate while it was isolated by oceanic barriers and that a major northern radiation of oscines (i.e. the parvorder Passerida) originated subsequent to dispersal from the south.

Animals↗

Noncoding regulatory sequences of Ciona exhibit strong correspondence between evolutionary constraint and functional importance.

We show that sequence comparisons at different levels of resolution can efficiently guide functional analyses of regulatory regions in the ascidians Ciona savignyi and Ciona intestinalis. Sequence alignments of several tissue-specific genes guided discovery of minimal regulatory regions that are active in whole-embryo reporter assays. Using the Troponin I (TnI) locus as a case study, we show that more refined local sequence analyses can then be used to reveal functional substructure within a regulatory region. A high-resolution saturation mutagenesis in conjunction with comparative sequence analyses defined essential sequence elements within the TnI regulatory region. Finally, we found a significant, quantitative relationship between function and sequence divergence of noncoding functional elements. This work demonstrates the power of comparative sequence analysis between the two Ciona species for guiding gene regulatory experiments.

Animals↗

Phylogenetic analyses identify 10 classes of the protein disulfide isomerase family in plants, including single-domain protein disulfide isomerase-related proteins.

Protein disulfide isomerases (PDIs) are molecular chaperones that contain thioredoxin (TRX) domains and aid in the formation of proper disulfide bonds during protein folding. To identify plant PDI-like (PDIL) proteins, a genome-wide search of Arabidopsis (Arabidopsis thaliana) was carried out to produce a comprehensive list of 104 genes encoding proteins with TRX domains. Phylogenetic analysis was conducted for these sequences using Bayesian and maximum-likelihood methods. The resulting phylogenetic tree showed that evolutionary relationships of TRX domains alone were correlated with conserved enzymatic activities. From this tree, we identified a set of 22 PDIL proteins that constitute a well-supported clade containing orthologs of known PDIs. Using the Arabidopsis PDIL sequences in iterative BLAST searches of public and proprietary sequence databases, we further identified orthologous sets of 19 PDIL sequences in rice (Oryza sativa) and 22 PDIL sequences in maize (Zea mays), and resolved the PDIL phylogeny into 10 groups. Five groups (I-V) had two TRX domains and showed structural similarities to the PDIL proteins in other higher eukaryotes. The remaining five groups had a single TRX domain. Two of these (quiescin-sulfhydryl oxidase-like and adenosine 5'-phosphosulfate reductase-like) had putative nonisomerase enzymatic activities encoded by an additional domain. Two others (VI and VIII) resembled small single-domain PDIs from Giardia lamblia, a basal eukaryote, and from yeast. Mining of maize expressed sequence tag and RNA-profiling databases indicated that members of all of the single-domain PDIL groups were expressed throughout the plant. The group VI maize PDIL ZmPDIL5-1 accumulated during endoplasmic reticulum stress but was not found within the intracellular membrane fractions and may represent a new member of the molecular chaperone complement in the cell.

Base Sequence↗

Partial shotgun sequencing of the Boechera stricta genome reveals extensive microsynteny and promoter conservation with Arabidopsis.

Comparative genomics provides insight into the evolutionary dynamics that shape discrete sequences as well as whole genomes. To advance comparative genomics within the Brassicaceae, we have end sequenced 23,136 medium-sized insert clones from Boechera stricta, a wild relative of Arabidopsis (Arabidopsis thaliana). A significant proportion of these sequences, 18,797, are nonredundant and display highly significant similarity (BLASTn e-value < or = 10(-30)) to low copy number Arabidopsis genomic regions, including more than 9,000 annotated coding sequences. We have used this dataset to identify orthologous gene pairs in the two species and to perform a global comparison of DNA regions 5' to annotated coding regions. On average, the 500 nucleotides upstream to coding sequences display 71.4% identity between the two species. In a similar analysis, 61.4% identity was observed between 5' noncoding sequences of Brassica oleracea and Arabidopsis, indicating that regulatory regions are not as diverged among these lineages as previously anticipated. By mapping the B. stricta end sequences onto the Arabidopsis genome, we have identified nearly 2,000 conserved blocks of microsynteny (bracketing 26% of the Arabidopsis genome). A comparison of fully sequenced B. stricta inserts to their homologous Arabidopsis genomic regions indicates that indel polymorphisms >5 kb contribute substantially to the genome size difference observed between the two species. Further, we demonstrate that microsynteny inferred from end-sequence data can be applied to the rapid identification and cloning of genomic regions of interest from nonmodel species. These results suggest that among diploid relatives of Arabidopsis, small- to medium-scale shotgun sequencing approaches can provide rapid and cost-effective benefits to evolutionary and/or functional comparative genomic frameworks.

Arabidopsis↗