PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

ESTviewer: a web interface for visualizing mouse, rat, cattle, pig and chicken conserved ESTs in human genes and human alternatively spliced variants.

ESTviewer is a web application for interactively visualizing human gene structures, with emphasis on mammalian and avian expressed sequence tags (ESTs) that are conserved in the human genome and alternatively spliced (AS) variants. AS variants from the UCSC, Vega and PSEP annotations are presented in this application for comparison. EST data from six species, human, mouse, rat, cattle, pig and chicken, are mapped to the human genome to show cross-species EST conservation in annotated exonic and intronic regions. Cross-species EST conservation is evolutionarily and functionally important because it represents the effects of selection pressure on genic regions and transcriptome over evolutionary time. Emphatically, ESTviewer provides a convenient tool to compare highly conserved non-human ESTs and human AS variants. The application takes human gene accession Ids or coordinates of genomic sequences as inputs and presents annotated gene structures and their AS variants. In addition, the lengths and percentages of human genic regions covered by ESTs are displayed to show the level of EST coverage of different species. The percentages of the UCSC, Vega and PSEP annotated exons covered by ESTs of the six studied species are also displayed in the interface.

Animals↗

Improved pairwise alignments of proteins in the Twilight Zone using local structure predictions.

MOTIVATION: In recent years, advances have been made in the ability of computational methods to discriminate between homologous and non-homologous proteins in the 'twilight zone' of sequence similarity, where the percent sequence identity is a poor indicator of homology. To make these predictions more valuable to the protein modeler, they must be accompanied by accurate alignments. Pairwise sequence alignments are inferences of orthologous relationships between sequence positions. Evolutionary distance is traditionally modeled using global amino acid substitution matrices. But real differences in the likelihood of substitutions may exist for different structural contexts within proteins, since structural context contributes to the selective pressure. RESULTS: HMMSUM (HMMSTR-based substitution matrices) is a new model for structural context-based amino acid substitution probabilities consisting of a set of 281 matrices, each for a different sequence-structure context. HMMSUM does not require the structure of the protein to be known. Instead, predictions of local structure are made using HMMSTR, a hidden Markov model for local structure. Alignments using the HMMSUM matrices compare favorably to alignments carried out using the BLOSUM matrices or structure-based substitution matrices SDM and HSDM when validated against remote homolog alignments from BAliBASE. HMMSUM has been implemented using local Dynamic Programming and with the Bayesian Adaptive alignment method.

Algorithms↗

Bayesian search of functionally divergent protein subgroups and their function specific residues.

MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html

Algorithms↗

Unique mammalian tRNA-derived repetitive elements in dermopterans: the t-SINE family and its retrotransposition through multiple sources.

Short interspersed nuclear elements (SINEs) are dispersed repetitive DNA sequences that are major components of all mammalian genomes. They have been described in almost all lineages of Euarchontoglires (rodents, rabbits, primates, flying lemurs, and tree shrews), except in flying lemurs. Most SINE family members are composed of three distinct regions: a 5' tRNA-related region, a tRNA-unrelated region, and a short tandem repeat at the 3' end that is AT-rich. The newly discovered SINE family in Cynocephalus deviates from this common structure. All 30 SINE loci analyzed in this family lack a tRNA-unrelated region and are composed exclusively of tRNA-related elements. Therefore, this novel SINE structure, described for the first time in mammalian genomes, was designated as t-SINE. The t-SINE family exhibits a high copy number and is specific to flying lemurs. Three major t-SINE subfamilies could be distinguished on the basis of characteristic nucleotides, deletions, insertions, and duplications. These sequence-specific characteristics within subfamilies and sub-subfamilies reveal that they are derived copies of distinct progenitors. We present evolutionary relationships between subfamilies and compare relationships between the subfamilies and the isoleucine tRNA gene. t-SINE amplification occurred through multiple sources and is supposedly mobilized via the L1-encoded reverse transcriptase-dependent retrotranspositional mechanism in trans.

Animals↗

Exploring the relationship between sequence similarity and accurate phylogenetic trees.

We have characterized the relationship between accurate phylogenetic reconstruction and sequence similarity, testing whether high levels of sequence similarity can consistently produce accurate evolutionary trees. We generated protein families with known phylogenies using a modified version of the PAML/EVOLVER program that produces insertions and deletions as well as substitutions. Protein families were evolved over a range of 100-400 point accepted mutations; at these distances 63% of the families shared significant sequence similarity. Protein families were evolved using balanced and unbalanced trees, with ancient or recent radiations. In families sharing statistically significant similarity, about 60% of multiple sequence alignments were 95% identical to true alignments. To compare recovered topologies with true topologies, we used a score that reflects the fraction of clades that were correctly clustered. As expected, the accuracy of the phylogenies was greatest in the least divergent families. About 88% of phylogenies clustered over 80% of clades in families that shared significant sequence similarity, using Bayesian, parsimony, distance, and maximum likelihood methods. However, for protein families with short ancient branches (ancient radiation), only 30% of the most divergent (but statistically significant) families produced accurate phylogenies, and only about 70% of the second most highly conserved families, with median expectation values better than 10(-60), produced accurate trees. These values represent upper bounds on expected tree accuracy for sequences with a simple divergence history; proteins from 700 Giardia families, with a similar range of sequence similarities but considerably more gaps, produced much less accurate trees. For our simulated insertions and deletions, correct multiple sequence alignments did not perform much better than those produced by T-COFFEE, and including sequences with expressed sequence tag-like sequencing errors did not significantly decrease phylogenetic accuracy. In general, although less-divergent sequence families produce more accurate trees, the likelihood of estimating an accurate tree is most dependent on whether radiation in the family was ancient or recent. Accuracy can be improved by combining genes from the same organism when creating species trees or by selecting protein families with the best bootstrap values in comprehensive studies.

Animals↗

A yeast arginine specific tRNA is a remnant aspartate acceptor.

High specificity in aminoacylation of transfer RNAs (tRNAs) with the help of their cognate aminoacyl-tRNA synthetases (aaRSs) is a guarantee for accurate genetic translation. Structural and mechanistic peculiarities between the different tRNA/aaRS couples, suggest that aminoacylation systems are unrelated. However, occurrence of tRNA mischarging by non-cognate aaRSs reflects the relationship between such systems. In Saccharomyces cerevisiae, functional links between arginylation and aspartylation systems have been reported. In particular, it was found that an in vitro transcribed tRNAAsp is a very efficient substrate for ArgRS. In this study, the relationship of arginine and aspartate systems is further explored, based on the discovery of a fourth isoacceptor in the yeast genome, tRNA4Arg. This tRNA has a sequence strikingly similar to that of tRNAAsp but distinct from those of the other three arginine isoacceptors. After transplantation of the full set of aspartate identity elements into the four arginine isoacceptors, tRNA4Arg gains the highest aspartylation efficiency. Moreover, it is possible to convert tRNA4Arg into an aspartate acceptor, as efficient as tRNAAsp, by only two point mutations, C38 and G73, despite the absence of the major anticodon aspartate identity elements. Thus, cryptic aspartate identity elements are embedded within tRNA4Arg. The latent aspartate acceptor capacity in a contemporary tRNAArg leads to the proposal of an evolutionary link between tRNA4Arg and tRNAAsp genes.

Aspartic Acid↗

A systematic search for RNA editing sites in pea chloroplasts: an editing event causes diversification from the evolutionarily conserved amino acid sequence.

RNA editing in higher plant chloroplasts involves C-to-U conversion at specific sites in the transcripts. To examine whether pea shares editing sites with other angiosperms, a systematic search for editing sites in pea chloroplast transcripts was performed. Based on amino acid sequence alignment, 451 RNA editing sites were predicted from 60 transcripts. Sequence analysis of amplified cDNAs for these potential editing sites revealed 19 true editing sites from 13 transcripts. Together with those reported previously, the total number of editing sites is 27 from 16 transcripts in pea chloroplasts. Twenty-two sites are conserved among other plant species, whereas five sites are unique to pea. Among the 27 editing sites, seven are partially edited. The most interesting is the ndhG site 1, which has led to the diversification of the evolutionarily conserved amino acid sequence. This observation suggests that some of the editing events cause the diversity of amino acid sequences, and hence, that prediction of editing sites based on amino acid sequence alignment has its own limitations.

Amino Acid Sequence↗

A medaka gene map: the trace of ancestral vertebrate proto-chromosomes revealed by comparative gene mapping.

The mapping of Hox clusters and many duplicated genes in zebrafish indicated an extra whole-genome duplication in ray-fined fish. However, to reconstruct the preduplication chromosomes (proto-chromosomes), the comparative genomic studies of more distantly related teleosts are essential. Medaka and zebrafish are ideal for this purpose, because their lineages separated from their last common ancestor approximately 140 million years ago. To reconstruct ancient vertebrate chromosomes, including the chromosomes of the vertebrate ancestor of humans from 450 million years ago, we mapped 818 genes and expressed sequence tags (ESTs) on a single meiotic backcross panel obtained from inbred strains of the medaka, Oryzias latipes. Comparisons of linkage relationships of orthologous genes among three species of vertebrates (medaka, zebrafish, and human) indicate the number and content of the chromosomes of the last common ancestor of ray-fined fish and lobe-fined fish (including humans), and the extra whole genome duplication event in the ray-fin lineage occurred in the common ancestor of perhaps all teleosts.

Animals↗

Evaluation of monocot and eudicot divergence using the sugarcane transcriptome.

Over 40,000 sugarcane (Saccharum officinarum) consensus sequences assembled from 237,954 expressed sequence tags were compared with the protein and DNA sequences from other angiosperms, including the genomes of Arabidopsis and rice (Oryza sativa). Approximately two-thirds of the sugarcane transcriptome have similar sequences in Arabidopsis. These sequences may represent a core set of proteins or protein domains that are conserved among monocots and eudicots and probably encode for essential angiosperm functions. The remaining sequences represent putative monocot-specific genetic material, one-half of which were found only in sugarcane. These monocot-specific cDNAs represent either novelties or, in many cases, fast-evolving sequences that diverged substantially from their eudicot homologs. The wide comparative genome analysis presented here provides information on the evolutionary changes that underlie the divergence of monocots and eudicots. Our comparative analysis also led to the identification of several not yet annotated putative genes and possible gene loss events in Arabidopsis.

Arabidopsis↗

Proteomic characterization of evolutionarily conserved and variable proteins of Arabidopsis cytosolic ribosomes.

Analysis of 80S ribosomes of Arabidopsis (Arabidopsis thaliana) by use of high-speed centrifugation, sucrose gradient fractionation, one- and two-dimensional gel electrophoresis, liquid chromatography purification, and mass spectrometry (matrix-assisted laser desorption/ionization time-of-flight and electrospray ionization) identified 74 ribosomal proteins (r-proteins), of which 73 are orthologs of rat r-proteins and one is the plant-specific r-protein P3. Thirty small (40S) subunit and 44 large (60S) subunit r-proteins were confirmed. In addition, an ortholog of the mammalian receptor for activated protein kinase C, a tryptophan-aspartic acid-domain repeat protein, was found to be associated with the 40S subunit and polysomes. Based on the prediction that each r-protein is present in a single copy, the mass of the Arabidopsis 80S ribosome was estimated as 3.2 MD (1,159 kD 40S; 2,010 kD 60S), with the 4 single-copy rRNAs (18S, 26S, 5.8S, and 5S) contributing 53% of the mass. Despite strong evolutionary conservation in r-protein composition among eukaryotes, Arabidopsis 80S ribosomes are variable in composition due to distinctions in mass or charge of approximately 25% of the r-proteins. This is a consequence of amino acid sequence divergence within r-protein gene families and posttranslational modification of individual r-proteins (e.g. amino-terminal acetylation, phosphorylation). For example, distinct types of r-proteins S15a and P2 accumulate in ribosomes due to evolutionarily divergence of r-protein genes. Ribosome variation is also due to amino acid sequence divergence and differential phosphorylation of the carboxy terminus of r-protein S6. The role of ribosome heterogeneity in differential mRNA translation is discussed.

Amino Acid Sequence↗

Rec-I-DCM3: a fast algorithmic technique for reconstructing large phylogenetic trees.

Phylogenetic trees are commonly reconstructed based on hard optimization problems such as maximum parsimony (MP) and maximum likelihood (ML). Conventional MP heuristics for producing phylogenetic trees produce good solutions within reasonable time on small datasets (up to a few thousand sequences), while ML heuristics are limited to smaller datasets (up to a few hundred sequences). However, since MP (and presumably ML) is NP-hard, such approaches do not scale when applied to large datasets. In this paper, we present a new technique called Recursive-Iterative-DCM3 (Rec-I-DCM3), which belongs to our family of Disk-Covering Methods (DCMs). We tested this new technique on ten large biological datasets ranging from 1,322 to 13,921 sequences and obtained dramatic speedups as well as significant improvements in accuracy (better than 99.99%) in comparison to existing approaches. Thus, high-quality reconstructions can be obtained for datasets at least ten times larger than was previously possible.

Algorithms↗

Evolutionary relationships among extradiol dioxygenases.

A structure-validated alignment of 35 extradiol dioxygenase sequences including two-domain and one-domain enzymes was derived. Strictly conserved residues include the metal ion ligands and several catalytically essential active site residues, as well as a number of structurally important residues that are remote from the active site. Phylogenetic analyses based on this alignment indicate that the ancestral extradiol dioxygenase was a one-domain enzyme and that the two-domain enzymes arose from a single genetic duplication event. Subsequent divergence among the two-domain dioxygenases has resulted in several families, two of which are based on substrate preference. In several cases, the two domains of a given enzyme express different phylogenies, suggesting the possibility that such enzymes arose from the recombination of genes encoding different dioxygenases. A phylogeny-based classification system for extradiol dioxygenases is proposed.

Amino Acid Sequence↗

Structural and genetic characterization of the Shigella boydii type 10 and type 6 O antigens.

Comparison of the O antigens of Shigella boydii types 10 and 6 by chemical analysis and nuclear magnetic resonance spectroscopy showed that their structures are similar, with the only difference being the presence or absence of d-ribofuranose, which is the immunodominant sugar in S. boydii type 10. In S. boydii type 6, a residue previously reported as alpha-d-GlcpA, was shown to be beta-d-GlcpA as in S. boydii type 10. S. boydii types 10 and 6 are reported not to cross-react serologically, and the role of d-ribofuranose in the specificity of S. boydii was confirmed by making a mutant of type 10 that lacked d-ribofuranose. However, S. boydii type 11, which has a d-ribofuranose but with different linkage does show cross-reaction with type 10. The O-antigen gene loci of S. boydii types 10 and 6 were shown to be virtually identical except that orf8 (wbaM), which was confirmed as the ribofuranosyltransferase gene, is interrupted by IS629 in type 6. Therefore, it is proposed that the O-antigen gene cluster of S. boydii type 6 was derived from type 10 by an IS element insertion.

Amino Acid Sequence↗

Comparative sequencing of the serine-aspartate repeat-encoding region of the clumping factor B gene (clfB) for resolution within clonal groups of Staphylococcus aureus.

Molecular techniques such as spa typing and multilocus sequence typing use DNA sequence data for differentiating Staphylococcus aureus isolates. Although spa typing is capable of detecting both genetic micro- and macrovariation, it has less discriminatory power than the more labor-intensive pulsed-field gel electrophoresis (PFGE) and costly genomic DNA microarray analyses. This limitation hinders strain interrogation for newly emerging clones and outbreak investigations in hospital or community settings where robust clones are endemic. To overcome this constraint, we developed a typing system using DNA sequence analysis of the serine-aspartate (SD) repeat-encoding region within the gene encoding the keratin- and fibrinogen-binding clumping factor B (clfB typing) and tested whether it is capable of discriminating within clonal groups. We analyzed 116 S. aureus strains, and the repeat region was present in all isolates, varying in sequence and in length from 420 to 804 bp. In a sample of 36 well-characterized genetically diverse isolates, clfB typing subdivided identical spa and PFGE clusters which had been discriminated by whole-genome DNA microarray mapping. The combination of spa typing and clfB typing resulted in a discriminatory power (99.5%) substantially higher than that of spa typing alone and closely approached that of the whole-genome microarray (100.0%). clfB typing also successfully resolved genetic differences among isolates differentiated by PFGE that had been collected over short periods of time from single hospitals and that belonged to the most prevalent S. aureus clone in the United States. clfB typing demonstrated in vivo, in vitro, and interpatient transmission stability yet revealed that this locus may be recombinogenic in a primarily clonal population structure. Taken together, these data show that the SD repeat-encoding region of clfB is a highly stable marker of microvariation, that in conjunction with spa typing it may serve as a DNA sequence-based alternative to PFGE for investigating genetically similar strains, and that it is useful for analyzing collections of isolates in both long-term population-based and local epidemiologic studies.

Adhesins, Bacterial↗

LTR retrotransposons in the dioecious plant Silene latifolia.

Conserved domains of two types of LTR retrotransposons, Tyl-copia- and Ty3-gypsy-like retrotransposons, were isolated from the dioecious plant Silene latifolia, whose sex is determined by X and Y chromosomes. Southern hybridization analyses using these retrotransposons as probes resulted in identical patterns from male and female genomes. Fluorescence in situ hybridization indicated that these retrotransposons do not accumulate specifically in the sex chromosomes. These results suggest that recombination between the sex chromosomes of S. latifolia has not been severely reduced. Conserved reverse transcriptase regions of Ty1-copia-like retrotransposons were isolated from 13 different Silene species and classified into two major families. Their categorization suggests that parallel divergence of the Ty1-copia-like retrotransposons occurred during the differentiation of Silene species. Most functional retrotransposons from three dioecious species, S. latifolia, S. dioica, and S. diclinis, fell into two clusters. The evolutionary dynamics of retrotransposons implies that, in the genus Silene, dioecious species evolved recently from gynodioecious species.

Amino Acid Sequence↗

Identification, structure, and differential expression of members of a BURP domain containing protein family in soybean.

Expressed sequence tags (ESTs) exhibiting homology to a BURP domain containing gene family were identified from the Glycine max (L.) Merr. EST database. These ESTs were assembled into 16 contigs of variable sizes and lengths. Consistent with the structure of known BURP domain containing proteins, the translation products exhibit a modular structure consisting of a C-terminal BURP domain, an N-terminal signal sequence, and a variable internal region. The soybean family members exhibit 35-98% similarity in a -100-amino-acid C-terminal region, and a phylogenetic tree constructed using this region shows that some soybean family members group together in closely related pairs, triplets, and quartets, whereas others remain as singletons. The structure of these groups suggests that multiple gene duplication events occurred during the evolutionary history of this family. The depth and diversity of G. max EST libraries allowed tissue-specific expression patterns of the putative soybean BURPs to be examined. Consistent with known BURP proteins, the newly identified soybean BURPs have diverse expression patterns. Furthermore, putative paralogs can have both spatially and quantitatively distinct expression patterns. We discuss the functional and evolutionary implications of these findings, as well as the utility of EST-based analyses for identifying and characterizing gene families.

Amino Acid Sequence↗

Conserved synteny and gene order difference between human chromosome 12 and pig chromosome 5.

A comparative map of human chromosome 12 (HSA 12) and pig chromosome 5 (SSC 5) was constructed using ten pig expressed sequence tags (ESTs). These ESTs were isolated from primary granulosa cell cultures by differential display (EST b10b), or from a granulosa cDNA library (VIIIE1, DRIM, N*9, RIIID2 and RVIC1) or from a small intestine cDNA library (ATPSB, ITGB7, MYH9, and STAT2). Also used were two Traced Orthologous Amplified Sequence Tags (TOASTs) (LALBA, TRA1), one microsatellite-associated gene (IGF1) and finally five human YACs selected for their cytogenetic position, with a view to increasing the number of informative markers for the comparison. Large-insert clones were obtained by screening a pig bacterial artificial chromosome (BAC) library with specific primers for each EST and TOAST and for IGF1. These BACs were used as probes for fluorescent in situ hybridisation (FISH) both on porcine and human metaphases. In addition, the human YACs were FISH mapped on pig chromosomes. This allowed us to refine and, in some cases, to correct the previous mapping obtained with a somatic cell hybrid panel. While these data confirm chromosome painting results showing that the distal part of SSC 5p arm is conserved on HSA 22, while the rest of the chromosome corresponds to HSA 12, they also demonstrate gene-order differences between human and pig. In addition, it was also possible to determine the position of the synteny breakpoint.

Animals↗