PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Genomic features in the breakpoint regions between syntenic blocks.

MOTIVATION: We study the largely unaligned regions between the syntenic blocks conserved in humans and mice, based on data extracted from the UCSC genome browser. These regions contain evolutionary breakpoints caused by inversion, translocation and other processes. RESULTS: We suggest explanations for the limited amount of genomic alignment in the neighbourhoods of breakpoints. We discount inferences of extensive breakpoint reuse as artefacts introduced during the reconstruction of syntenic blocks. We find that the number, size and distribution of small aligned fragments in the breakpoint regions depend on the origin of the neighbouring blocks and the other blocks on the same chromosome. We account for this and for the generalized loss of alignment in the regions partially by artefacts due to alignment protocols and partially by mutational processes operative only after the rearrangement event. These results are consistent with breakpoints occurring randomly over virtually the entire genome.

Algorithms↗

Non-coding RNAs in Ciona intestinalis.

MOTIVATION: The analysis of animal genomes showed that only a minute part of their DNA codes for proteins. Recent experimental results agree, however, that a large fraction of these genomes are transcribed and hence are probably functional at the RNA level. A computational survey of vertebrate genomes has predicted thousands of previously unknown ncRNAs with evolutionarily conserved secondary structures. Extending these comparative studies beyond vertebrates is difficult, however, since most ncRNAs evolve quickly at the sequence level while conserving their characteristic secondarystructures. RESULTS: We report on a computational screen of structured ncRNAs in the urochordate lineage based on a comparison of the genomic data from Ciona intestinalis, Ciona savignyi and Oikopleura dioica. We predict >1000 ncRNAs with an evolutionarily conserved RNA secondary structure. Of these, about a quarter are located in introns of known protein coding sequences. A few RNA motifs can be identified as known RNAs, including approximately 300 tRNAs, some 100 snRNA genes and a few microRNAs and snoRNAs. AVAILABILITY: www.bioinf.uni-leipzig.de/Publications/SUPPLEMENTS/05-008/

Animals↗

Conserved structural features in eukaryotic and prokaryotic fucosyltransferases.

Fucosyltransferases are the enzymes transferring fucose from GDP-Fuc to Gal in an alpha1,2-linkage and to GlcNAc in alpha1,3-, alpha1,4-, or alpha1,6-linkages. Since all fucosyltransferases utilize the same nucleotide sugar, their specificity will probably reside in the recognition of the acceptor and in the type of linkage formed. A search of nucleotide and protein databases yielded more than 30 sequences of fucosyltransferases originating from mammals, chicken, nematode, and bacteria. On the basis of protein sequence similarities, these enzymes can be classified into four distinct families: (1) the alpha-2-fucosyltransferases, (2) the alpha-3-fucosyltransferases, (3) the mammalian alpha-6-fucosyltransferases, and (4) the bacterial alpha-6-fucosyltransferases. Nevertheless, using the sensitive hydrophobic cluster analysis (HCA) method, conserved structural features as well as a consensus peptide motif have been clearly identified in the catalytic domains of all alpha-2 and alpha-6-fucosyltranferases, from prokaryotic and eukaryotic origin, that allowed the grouping of these enzymes into one superfamily. In addition, a few amino acids were found strictly conserved in this family, and two of these residues have been reported to be essential for enzyme activity for a human alpha-2-fucosyltransferase. The alpha-3-fucosyltransferases constitute a distinct family as they lack the consensus peptide, but some regions display similarities with the alpha-2 and alpha-6-fucosyltranferases. All these observations strongly suggest that the fucosyltransferases share some common structural and catalytic features.

Amino Acid Sequence↗

Different evolutionary processes shaped the mouse and human olfactory receptor gene families.

We report a comprehensive comparative analysis of human and mouse olfactory receptor (OR) genes. The OR family is the largest mammalian gene family known. We identify approximately 93% of an estimated 1500 mouse ORs, exceeding previous estimates and the number of human ORs by 50%. Only 20% are pseudogenes, giving a functional OR repertoire in mice that is three times larger than that of human. The proteins encoded by intact human ORs are less highly conserved than those of mouse, in patterns that suggest that even some apparently intact human OR genes may encode non-functional proteins. Mouse ORs are clustered in 46 genomic locations, compared to a much more dispersed pattern in human. We find orthologous clusters at syntenic human locations for most mouse genes, indicating that most OR gene clusters predate primate-rodent divergence. However, many recent local OR duplications in both genomes obscure one-to-one orthologous relationships, thereby complicating cross-species inferences about OR-ligand interactions. Local duplications are the major force shaping the gene family. Recent interchromosomal duplications of ORs have also occurred, but much more frequently in human than in mouse. In addition to clarifying the evolutionary forces shaping this gene family, our study provides the basis for functional studies of the transcriptional regulation and ligand-binding capabilities of the OR gene family.

Amino Acid Sequence↗

The Plasmodium selenoproteome.

The use of selenocysteine (Sec) as the 21st amino acid in the genetic code has been described in all three major domains of life. However, within eukaryotes, selenoproteins are only known in animals and algae. In this study, we characterized selenoproteomes and Sec insertion systems in protozoan Apicomplexa parasites. We found that among these organisms, Plasmodium and Toxoplasma utilized Sec, whereas Cryptosporidium did not. However, Plasmodium had no homologs of known selenoproteins. By searching computationally for evolutionarily conserved selenocysteine insertion sequence (SECIS) elements, which are RNA structures involved in Sec insertion, we identified four unique Plasmodium falciparum selenoprotein genes. These selenoproteins were incorrectly annotated in PlasmoDB, were conserved in other Plasmodia and had no detectable homologs in other species. We provide evidence that two Plasmodium SECIS elements supported Sec insertion into parasite and endogenous selenoproteins when they were expressed in mammalian cells, demonstrating that the Plasmodium SECIS elements are functional and indicating conservation of Sec insertion between Apicomplexa and animals. Dependence of the plasmodial parasites on selenium suggests possible strategies for antimalarial drug development.

Amino Acid Sequence↗

Polymorphism and divergence in the beta-globin replication origin initiation region.

DNA sequence polymorphism and divergence was examined in the vicinity of the human beta-globin gene cluster origin of replication initiation region (IR), a 1.3-kb genomic region located immediately 5' of the adult-expressed beta-globin gene. DNA sequence variation in the replication origin IR and 5 kb of flanking DNA was surveyed in samples drawn from two populations, one African (from the Gambia, West Africa) and the other European (from Oxford, England). In these samples, levels of nucleotide and length polymorphism in the IR were found to be more than two times as high as adjacent non-IR-associated regions (estimates of per-nucleotide heterozygosity were 0.30% and 0.12%, respectively). Most polymorphic positions identified in the origin IR fall within or just adjacent to a 52-bp alternating purine-pyrimidine ((RY)n) sequence repeat. Within- and between-populations divergence is highest in this portion of the IR, and interspecific divergence in the same region, determined by comparison with an orthologous sequence from the chimpanzee, is also pronounced. Higher levels of diversity in this subregion are not, however, primarily attributable to slippage-mediated repeat unit changes, as nucleotide substitution contributes disproportionately to allelic heterogeneity. An estimate of helical stability in the sequenced region suggests that the hypervariable (RY)n constitutes the major DNA unwinding element (DUE) of the replication origin IR, the location at which the DNA duplex first unwinds and new strand synthesis begins. These findings suggest that the beta-globin IR experiences a higher underlying rate of neutral mutation than do adjacent genomic regions and that enzyme fidelity associated with the initiation of DNA replication at this origin may be compromised. The significance of these findings for our understanding of eukaryotic replication origin biology is discussed.

Animals↗

The tomato photomorphogenetic mutant, aurea, is deficient in phytochromobilin synthase for phytochrome chromophore biosynthesis.

The aurea mutants of tomato have been widely used as phytochrome-deficient mutants for photomorphogenetic and photobiological studies. By expressed sequence tag (EST)-based screening of sequence databases, we found a tomato gene that encodes a protein homologous to Arabidopsis HY2 for phytochromobilin synthase catalyzing the last step of phytochrome chromophore biosynthesis. The tomato protein expressed in Escherichia coli showed phytochromobilin synthase activity. The corresponding loci in all aurea mutants tested have nucleotide substitutions, deletions or DNA rearrangements. These results indicate that aurea is a mutant of phytochromobilin synthase in tomato. We also discuss a phylogenetic analysis of phytochromobilin synthases in the bilin reductase family.

Amino Acid Sequence↗

tRNA gene clusters at the 3' end of rRNA operons are specific to virulent subgroups of Streptococcus agalactiae strains, as demonstrated by molecular differential analysis at the population level.

The aim of this work was to characterize a 2.4 kb randomly amplified polymorphic DNA (RAPD) fragment described as a marker for a phylogenetic group of Streptococcus agalactiae strains significantly associated with neonatal meningitis. This fragment was analysed by cloning and sequencing, and showed that two types of tRNA gene cluster flank the 3' end of the rRNA operons in S. agalactiae strains. Both types of tRNA gene cluster act as markers for phylogenetic subgroups of strains within the species. One type could be used to distinguish two of the three virulent intraspecies subgroups to which most of the S. agalactiae strains able to invade the central nervous system of neonates belong. This raises the possibility that there is a link between these tRNA genes and the virulence of the bacterium. Based on this analysis, PCR primers were designed to determine whether S. agalactiae strains are likely to belong to lineages of organisms in which most of the highly virulent strains isolated from cerebrospinal fluid cluster. In addition, this work demonstrated that RAPD can be used to detect novel particularities within intraspecies variants of pathogens.

Bacterial Typing Techniques↗

Structure and function of a conserved DNA region coding for tartrate utilization in Agrobacterium vitis.

Three tartrate utilization regions from Agrobacterium vitis strains involved in host specificity have been compared, to clearly define the borders of these regions and eventually identify specific sequences that could provide a mechanism of duplication of this region. A 10.8-kb conserved DNA fragment called the TAR element, found in different genetic contexts, was defined. A comparison of the two tartrate dehydrogenase genes (ttuC and ttuC') in each of the three TAR elements suggests that these genes co-evolve.

Alcohol Oxidoreductases↗

Neuropeptide Y-family receptors Y6 and Y7 in chicken. Cloning, pharmacological characterization, tissue distribution and conserved synteny with human chromosome region.

The peptides of the neuropeptide Y (NPY) family exert their functions, including regulation of appetite and circadian rhythm, by binding to G-protein coupled receptors. Mammals have five subtypes, named Y1, Y2, Y4, Y5 and Y6, and recently Y7 has been discovered in fish and amphibians. In chicken we have previously characterized the first four subtypes and here we describe Y6 and Y7. The genes for Y6 and Y7 are located 1 megabase apart on chromosome 13, which displays conserved synteny with human chromosome 5 that harbours the Y6 gene. The porcine PYY radioligand bound the chicken Y6 receptor with a K(d) of 0.80 +/- 0.36 nm. No functional coupling was demonstrated. The Y6 mRNA is expressed in hypothalamus, gastrointestinal tract and adipose tissue. Porcine PYY bound chicken Y7 with a K(d) of 0.14 +/- 0.01 nm (mean +/- SEM), whereas chicken PYY surprisingly had a much lower affinity, with a Ki of 41 nm, perhaps as a result of its additional amino acid at the N terminus. Truncated peptide fragments had greatly reduced affinity for Y7, in agreement with its closest relative, Y2, in chicken and fish, but in contrast to Y2 in mammals. This suggests that in mammals Y2 has only recently acquired the ability to bind truncated PYY. Chicken Y7 has a much more restricted tissue distribution than other subtypes and was only detected in adrenal gland. Y7 seems to have been lost in mammals. The physiological roles of Y6 and Y7 remain to be identified, but our phylogenetic and chromosomal analyses support the ancient origin of these Y receptor genes by chromosome duplications in an early (pregnathostome) vertebrate ancestor.

Amino Acid Sequence↗

Conservation of the 15-kilodalton lipoprotein among Treponema pallidum subspecies and strains and other pathogenic treponemes: genetic and antigenic analyses.

The 15-kDa lipoprotein of Treponema pallidum is a major immunogen during natural syphilis infection in humans and experimental infection in other hosts. The humoral and cellular immune responses to this molecule appear late in infection as resistance to reinfection is developing. One therefore might hypothesize that this antigen is important for protective immunity. This possibility is explored by using both genetic and antigenic approaches. Limited or no cross-protection has been demonstrated between the T. pallidum subspecies and strains or between Treponema species. We therefore hypothesized that if the 15-kDa antigen was of major importance in protective immunity, it might be a likely site of antigenic diversity. To explore this possibility, the sequences of the open reading frames of the 15-kDa gene have been determined for Treponema pallidum subsp. pallidum (Nichols and Bal-3 strains), T. pallidum subsp. pertenue (Gauthier strain), T. pallidum subsp. endemicum (Bosnia strain), Treponema paraluiscuniculi (Cuniculi A, H, and K strains), and a little-characterized simian isolate of Treponema sp. (Fribourg-Blanc strain). No significant differences in DNA sequences of the genes for the coding region of the 15-kDa antigen were found among the different species and subspecies studied. In addition, all organisms showed expression of the 15-kDa antigen as determined by monoclonal antibody staining. The role of the 15-kDa antigen in protection against homologous infection with T. pallidum subsp. pallidum Nichols was examined in rabbits immunized with a purified recombinant 15-kDa fusion protein. No alteration in chancre development was observed in immunized, compared to unimmunized, rabbits, and the antisera induced by the immunization failed to enhance phagocytosis of T. pallidum subsp. pallidum by macrophages in vitro. These results do not support a major role for this antigen in protection against syphilis infection.

Animals↗

Genetic and functional interaction of evolutionarily conserved regions of the Prp18 protein and the U5 snRNA.

Both the Prp18 protein and the U5 snRNA function in the second step of pre-mRNA splicing. We identified suppressors of mutant prp18 alleles in the gene for the U5 snRNA (SNR7). The suppressors' U5 snRNAs have either a U4-to-A or an A8-to-C mutation in the evolutionarily invariant loop 1 of U5. Suppression is specific for prp18 alleles that encode proteins with mutations in a highly conserved region of Prp18 which forms an unstructured loop in crystals of Prp18. The snr7 suppressors partly restored the pre-mRNA splicing activity that was lost in the prp18 mutants. The close functional relationship of Prp18 and U5 is emphasized by the finding that two snr7 alleles, U5A and U6A, are dominant synthetic lethal with prp18 alleles. Our results support the idea that Prp18 and the U5 snRNA act in concert during the second step of pre-mRNA splicing and suggest a model in which the conserved loop of Prp18 acts to stabilize the interaction of loop 1 of the U5 snRNA with the splicing intermediates.

Alleles↗

Protein protein interactions, evolutionary rate, abundance and age.

BACKGROUND: Does a relationship exist between a protein's evolutionary rate and its number of interactions? This relationship has been put forward many times, based on a biological premise that a highly interacting protein will be more restricted in its sequence changes. However, to date several studies have voiced conflicting views on the presence or absence of such a relationship. RESULTS: Here we perform a large scale study over multiple data sets in order to demonstrate that the major reason for conflict between previous studies is the use of different but overlapping datasets. We show that lack of correlation, between evolutionary rate and number of interactions in a data set is related to the error rate. We also demonstrate that the correlation is not an artifact of the underlying distributions of evolutionary distance and interactions and is therefore likely to be biologically relevant. Further to this, we consider the claim that the dependence is due to gene expression levels and find some supporting evidence. A strong and positive correlation between the number of interactions and the age of a protein is also observed and we show this relationship is independent of expression levels. CONCLUSION: A correlation between number of interactions and evolutionary rate is observed but is dependent on the accuracy of the dataset being used. However it appears that the number of interactions a protein participates in depends more on the age of the protein than the rate at which it changes.

Aging↗

Improving the specificity of high-throughput ortholog prediction.

BACKGROUND: Orthologs (genes that have diverged after a speciation event) tend to have similar function, and so their prediction has become an important component of comparative genomics and genome annotation. The gold standard phylogenetic analysis approach of comparing available organismal phylogeny to gene phylogeny is not easily automated for genome-wide analysis; therefore, ortholog prediction for large genome-scale datasets is typically performed using a reciprocal-best-BLAST-hits (RBH) approach. One problem with RBH is that it will incorrectly predict a paralog as an ortholog when incomplete genome sequences or gene loss is involved. In addition, there is an increasing interest in identifying orthologs most likely to have retained similar function. RESULTS: To address these issues, we present here a high-throughput computational method named Ortholuge that further evaluates previously predicted orthologs (including those predicted using an RBH-based approach) - identifying which orthologs most closely reflect species divergence and may more likely have similar function. Ortholuge analyzes phylogenetic distance ratios involving two comparison species and an outgroup species, noting cases where relative gene divergence is atypical. It also identifies some cases of gene duplication after species divergence. Through simulations of incomplete genome data/gene loss, we show that the vast majority of genes falsely predicted as orthologs by an RBH-based method can be identified. Ortholuge was then used to estimate the number of false-positives (predominantly paralogs) in selected RBH-predicted ortholog datasets, identifying approximately 10% paralogs in a eukaryotic data set (mouse-rat comparison) and 5% in a bacterial data set (Pseudomonas putida - Pseudomonas syringae species comparison). Higher quality (more precise) datasets of orthologs, which we term "ssd-orthologs" (supporting-species-divergence-orthologs), were also constructed. These datasets, as well as Ortholuge software that may be used to characterize other species' datasets, are available at http://www.pathogenomics.ca/ortholuge/ (software under GNU General Public License). CONCLUSION: The Ortholuge method reported here appears to significantly improve the specificity (precision) of high-throughput ortholog prediction for both bacterial and eukaryotic species. This method, and its associated software, will aid those performing various comparative genomics-based analyses, such as the prediction of conserved regulatory elements upstream of orthologous genes.

Algorithms↗

Multi-species sequence comparison: the next frontier in genome annotation.

Multi-species comparisons of DNA sequences are more powerful for discovering functional sequences than pairwise DNA sequence comparisons. Most current computational tools have been designed for pairwise comparisons, and efficient extension of these tools to multiple species will require knowledge of the ideal evolutionary distance to choose and the development of new algorithms for alignment, analysis of conservation, and visualization of results.

Animals↗

Identification of BHB splicing motifs in intron-containing tRNAs from 18 archaea: evolutionary implications.

Most introns of archaeal tRNA genes (tDNAs) are located in the anticodon loop, between nucleotides 37 and 38, the unique location of their eukaryotic counterparts. However, in several Archaea, mostly in Crenarchaeota, introns have been found at many other positions of the tDNAs. In the present work, we revisit and extend all previous findings concerning the identification, exact location, size, and possible fit to the proposed bulge-helix-bulge structural motif (BHB, now renamed hBHBh') of the sequences spanning intron-exon junctions in intron-containing tRNAs of 18 archaea. A total of 103 introns were found located at the usual position 37/38 and 33 introns at 14 other different positions, that is, in the anticodon stem and loop, in the D-and T-loops, in the V-arm, or in the amino acid arm. For introns located at 37/38 and elsewhere in the pre-tRNA, canonical hBHBh' motifs were not always found. Instead, a relaxed hBH or HBh' motif including the constant central 4-bp helix H flanked by one helix (h or h') on either side generating only one bulge could be disclosed. Also, for introns located elsewhere than at position 37/38, the hBHBh' (or HBh') structure competes with the three-dimensional structure of the mature tRNA, attesting to important structural rearrangements during the complex multistep maturation-splicing processes. A homotetramer-type of splicing endonuclease (like in all Crenarchaeota) instead of a homodimeric-type of enzyme (as in most Euryarchaeota) appears to best fit the requirement for splicing introns at relaxed hBH or HBh' motifs, and may represent the most primitive form of this enzyme.

Base Sequence↗

An integrative genomic approach to uncover molecular mechanisms of prokaryotic traits.

With mounting availability of genomic and phenotypic databases, data integration and mining become increasingly challenging. While efforts have been put forward to analyze prokaryotic phenotypes, current computational technologies either lack high throughput capacity for genomic scale analysis, or are limited in their capability to integrate and mine data across different scales of biology. Consequently, simultaneous analysis of associations among genomes, phenotypes, and gene functions is prohibited. Here, we developed a high throughput computational approach, and demonstrated for the first time the feasibility of integrating large quantities of prokaryotic phenotypes along with genomic datasets for mining across multiple scales of biology (protein domains, pathways, molecular functions, and cellular processes). Applying this method over 59 fully sequenced prokaryotic species, we identified genetic basis and molecular mechanisms underlying the phenotypes in bacteria. We identified 3,711 significant correlations between 1,499 distinct Pfam and 63 phenotypes, with 2,650 correlations and 1,061 anti-correlations. Manual evaluation of a random sample of these significant correlations showed a minimal precision of 30% (95% confidence interval: 20%-42%; n = 50). We stratified the most significant 478 predictions and subjected 100 to manual evaluation, of which 60 were corroborated in the literature. We furthermore unveiled 10 significant correlations between phenotypes and KEGG pathways, eight of which were corroborated in the evaluation, and 309 significant correlations between phenotypes and 166 GO concepts evaluated using a random sample (minimal precision = 72%; 95% confidence interval: 60%-80%; n = 50). Additionally, we conducted a novel large-scale phenomic visualization analysis to provide insight into the modular nature of common molecular mechanisms spanning multiple biological scales and reused by related phenotypes (metaphenotypes). We propose that this method elucidates which classes of molecular mechanisms are associated with phenotypes or metaphenotypes and holds promise in facilitating a computable systems biology approach to genomic and biomedical research.

Algorithms↗