PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

PAS kinase: an evolutionarily conserved PAS domain-regulated serine/threonine kinase.

PAS domains regulate the function of many intracellular signaling pathways in response to both extrinsic and intrinsic stimuli. PAS domain-regulated histidine kinases are common in prokaryotes and control a wide range of fundamental physiological processes. Similarly regulated kinases are rare in eukaryotes and are to date completely absent in mammals. PAS kinase (PASK) is an evolutionarily conserved gene product present in yeast, flies, and mammals. The amino acid sequence of PASK specifies two PAS domains followed by a canonical serine/threonine kinase domain, indicating that it might represent the first mammalian PAS-regulated protein kinase. We present evidence that the activity of PASK is regulated by two mechanisms. Autophosphorylation at two threonine residues located within the activation loop significantly increases catalytic activity. We further demonstrate that the N-terminal PAS domain is a cis regulator of PASK catalytic activity. When the PAS domain-containing region is removed, enzyme activity is significantly increased, and supplementation of the purified PAS-A domain in trans selectively inhibits PASK catalytic activity. These studies define a eukaryotic signaling pathway suitable for studies of PAS domains in a purified in vitro setting.

Amino Acid Sequence↗

A hierarchical approach to aligning collinear regions of genomes.

MOTIVATION: As a first approximation, similarity between two long orthologous regions of genomes can be represented by a chain of local similarities. Within such a chain, pairs of successive similarities are collinear (non-conflicting), i.e. segments involved in the nth similarity precede in both sequences segments involved in the (n+1)th similarity. However, when all similarities between two long sequences are considered, usually there are many conflicts between them. Although some conflicts can be avoided by masking transposons or low-complexity sequences, selecting only those similarities that reflect orthology and, thus, belong to the evolutionarily true chain is not trivial. RESULTS: We propose a simple, hierarchical algorithm of finding the true chain of local similarities. Starting from similarities with low P-values, we resolve each pairwise conflict by deleting a similarity with a higher P-value. This greedy approach constructs a chain of similarities faster than when a chain optimal with respect to some global criterion is sought, and makes more sense biologically.

Algorithms↗

Into the heart of darkness: large-scale clustering of human non-coding DNA.

MOTIVATION: It is currently believed that the human genome contains about twice as much non-coding functional regions as it does protein-coding genes, yet our understanding of these regions is very limited. RESULTS: We examine the intersection between syntenically conserved sequences in the human, mouse and rat genomes, and sequence similarities within the human genome itself, in search of families of non-protein-coding elements. For this purpose we develop a graph theoretic clustering algorithm, akin to the highly successful methods used in elucidating protein sequence family relationships. The algorithm is applied to a highly filtered set of about 700 000 human-rodent evolutionarily conserved regions, not resembling any known coding sequence, which encompasses 3.7% of the human genome. From these, we obtain roughly 12 000 non-singleton clusters, dense in significant sequence similarities. Further analysis of genomic location, evidence of transcription and RNA secondary structure reveals many clusters to be significantly homogeneous in one or more characteristics. This subset of the highly conserved non-protein-coding elements in the human genome thus contains rich family-like structures, which merit in-depth analysis. AVAILABILITY: Supplementary material to this work is available at http://www.soe.ucsc.edu/~jill/dark.html

Animals↗

Comparative genomics reveals unusually long motifs in mammalian genomes.

MOTIVATION: The recent discovery of the first small modulatory RNA (smRNA) presents the challenge of finding other molecules of similar length and conservation level. Unlike short interfering RNA (siRNA) and micro-RNA (miRNA), effective computational and experimental screening methods are not currently known for this species of RNA molecule, and the discovery of the one known example was partly fortuitous because it happened to be complementary to a well-studied DNA binding motif (the Neuron Restrictive Silencer Element). RESULTS: The existing comparative genomics approaches (e.g., phylogenetic footprinting) rely on alignments of orthologous regions across multiple genomes. This approach, while extremely valuable, is not suitable for finding motifs with highly diverged "non-alignable" flanking regions. Here we show that several unusually long and well conserved motifs can be discovered de novo through a comparative genomics approach that does not require an alignment of orthologous upstream regions. These motifs, including Neuron Restrictive Silencer Element, were missed in recent comparative genomics studies that rely on phylogenetic footprinting. While the functions of these motifs remain unknown, we argue that some may represent biologically important sites. AVAILABILITY: Our comparative genomics software, a web-accessible database of our results and a compilation of experimentally validated binding sites for NRSE can be found at http://www.cse.ucsd.edu/groups/bioinformatics.

Algorithms↗

RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models.

UNLABELLED: RAxML-VI-HPC (randomized axelerated maximum likelihood for high performance computing) is a sequential and parallel program for inference of large phylogenies with maximum likelihood (ML). Low-level technical optimizations, a modification of the search algorithm, and the use of the GTR+CAT approximation as replacement for GTR+Gamma yield a program that is between 2.7 and 52 times faster than the previous version of RAxML. A large-scale performance comparison with GARLI, PHYML, IQPNNI and MrBayes on real data containing 1000 up to 6722 taxa shows that RAxML requires at least 5.6 times less main memory and yields better trees in similar times than the best competing program (GARLI) on datasets up to 2500 taxa. On datasets > or =4000 taxa it also runs 2-3 times faster than GARLI. RAxML has been parallelized with MPI to conduct parallel multiple bootstraps and inferences on distinct starting trees. The program has been used to compute ML trees on two of the largest alignments to date containing 25,057 (1463 bp) and 2182 (51,089 bp) taxa, respectively. AVAILABILITY: icwww.epfl.ch/~stamatak

Algorithms↗

Mutagenesis of phospholipase D defines a superfamily including a trans-Golgi viral protein required for poxvirus pathogenicity.

Phospholipase D (PLD) genes are members of a superfamily that is defined by several highly conserved motifs. PLD in mammals has been proposed to play a role in membrane vesicular trafficking and signal transduction. Using site-directed mutagenesis, 25 point mutants have been made in human PLD1 (hPLD1) and characterized. We find that a motif (HxKxxxxD) and a serine/threonine conserved in all members of the PLD superfamily are critical for PLD biochemical activity, suggesting a possible catalytic mechanism. Functional analysis of catalytically inactive point mutants for yeast PLD demonstrates that the meiotic phenotype ensuing from PLD deficiency in yeast derives from a loss of enzymatic activity. Finally, mutation of an HxKxxxxD motif found in a vaccinia viral protein expressed in the Golgi complex results in loss of efficient vaccinia virus cell-to-cell spreading, implicating the viral protein as a member of the superfamily and suggesting that it encodes a lipid modifying or binding activity. The results suggest that vaccinia virus and hPLD1 may act through analogous mechanisms to effect viral cellular egress and vesicular trafficking, respectively.

Amino Acid Sequence↗

The rate and character of spontaneous mutation in an RNA virus.

Estimates of spontaneous mutation rates for RNA viruses are few and uncertain, most notably due to their dependence on tiny mutation reporter sequences that may not well represent the whole genome. We report here an estimate of the spontaneous mutation rate of tobacco mosaic virus using an 804-base cognate mutational target, the viral MP gene that encodes the movement protein (MP). Selection against newly arising mutants was countered by providing MP function from a transgene. The estimated genomic mutation rate was on the lower side of the range previously estimated for lytic animal riboviruses. We also present the first unbiased riboviral mutational spectrum. The proportion of base substitutions is the same as that in a retrovirus but is lower than that in most DNA-based organisms. Although the MP mutant frequency was 0.02-0.05, 35% of the sequenced mutants contained two or more mutations. Therefore, the mutation process in populations of TMV and perhaps of riboviruses generally differs profoundly from that in populations of DNA-based microbes and may be strongly influenced by a subpopulation of mutator polymerases.

Base Sequence↗

The evolutionary origin of an altruistic gene.

Although the conditions favoring altruism are being increasingly understood, the evolutionary origins of the genetic basis for this behavior remain elusive. Here, we show that reproductive altruism (i.e., a sterile soma) in the multicellular green alga, Volvox carteri, evolved via the co-option of a life-history gene whose expression in the unicellular ancestor was conditioned on an environmental cue (as an adaptive strategy to enhance survival at an immediate cost to reproduction) through shifting its expression from a temporal (environmentally induced) into a spatial (developmental) context. The gene belongs to a diverged and structurally heterogeneous multigene family sharing a SAND-like domain (a DNA-binding module involved in gene transcription regulation). To our knowledge, this is the first example of a social gene specifically associated with reproductive altruism, whose origin can be traced back to a solitary ancestor. These findings complement recent proposals that the differentiation of sterile castes in social insects involved the co-option of regulatory networks that control sequential shifts between phases in the life cycle of solitary insects.

Algal Proteins↗

Human EFO1p exhibits acetyltransferase activity and is a unique combination of linker histone and Ctf7p/Eco1p chromatid cohesion establishment domains.

Proper segregation of chromosomes during mitosis requires that the products of chromosome replication are paired together-termed sister chromatid cohesion. In budding yeast, Ctf7p/Eco1p is an essential protein that establishes cohesion between sister chromatids during S phase. In fission yeast, Eso1p also functions in cohesion establishment, but is comprised of a Ctf7p/Eco1p domain fused to a Rad30p domain (a DNA polymerase) both of which are independently expressed in budding yeast. In this report, we identify and characterize the first candidate human ortholog of Ctf7p/Eco1p, which we term hEFO1p (human Establishment Factor Ortholog). As in fission yeast Eso1p, the hEFO1p open reading frame extends well upstream of the C-terminal Ctf7p/Eco1p domain. However, this N-terminal extension in hEFO1p is unlike Rad30p, but instead exhibits significant homology to linker histone proteins. Thus, hEFO1p is a unique fusion of linker histone and cohesion establishment domains. hEFO1p is widely expressed among the tissues tested. Consistent with a role in chromosome segregation, hEFO1p localizes exclusively to the nucleus when expressed in HeLa tissue culture cells. Moreover, biochemical analyses reveal that hEFO1p exhibits acetyltransferase activity. These findings document the first characterization of a novel human acetyltransferase, hEFO1p, that is comprised of both linker histone and Ctf7p/Eco1p domains.

Acetyltransferases↗

GenDecoder: genetic code prediction for metazoan mitochondria.

Although the majority of the organisms use the same genetic code to translate DNA, several variants have been described in a wide range of organisms, both in nuclear and organellar systems, many of them corresponding to metazoan mitochondria. These variants are usually found by comparative sequence analyses, either conducted manually or with the computer. Basically, when a particular codon in a query-species is linked to positions for which a specific amino acid is consistently found in other species, then that particular codon is expected to translate as that specific amino acid. Importantly, and despite the simplicity of this approach, there are no available tools to help predicting the genetic code of an organism. We present here GenDecoder, a web server for the characterization and prediction of mitochondrial genetic codes in animals. The analysis of automatic predictions for 681 metazoans aimed us to study some properties of the comparative method, in particular, the relationship among sequence conservation, taxonomic sampling and reliability of assignments. Overall, the method is highly precise (99%), although highly divergent organisms such as platyhelminths are more problematic. The GenDecoder web server is freely available from http://darwin.uvigo.es/software/gendecoder.html.

Amino Acid Sequence↗

Protein structure and evolutionary history determine sequence space topology.

Understanding the observed variability in the number of homologs of a gene is a very important unsolved problem that has broad implications for research into coevolution of structure and function, gene duplication, pseudogene formation, and possibly for emerging diseases. Here, we attempt to define and elucidate some possible causes behind the observed irregularity in sequence space. We present evidence that sequence variability and functional diversity of a gene or fold family is influenced by quantifiable characteristics of the protein structure. These characteristics reflect the structural potential for sequence plasticity, i.e., the ability to accept mutation without losing thermodynamic stability. We identify a structural feature of a protein domain-contact density-that serves as a determinant of entropy in sequence space, i.e., the ability of a protein to accept mutations without destroying the fold (also known as fold designability). We show that (log) of average gene family size exhibits statistical correlation (R(2) > 0.9.) with contact density of its three-dimensional structure. We present evidence that the size of individual gene families are influenced not only by the designability of the structure, but also by evolutionary history, e.g., the amount of time the gene family was in existence. We further show that our observed statistical correlation between gene family size and contact density of the structure is valid on many levels of evolutionary divergence, i.e., not only for closely related sequence, but also for less-related fold and superfamily levels of homology.

Amino Acid Sequence↗

The evolutionary history of the genus Acanthamoeba and the identification of eight new 18S rRNA gene sequence types.

The 18S rRNA gene (Rns) phylogeny of Acanthamoeba is being investigated as a basis for improvements in the nomenclature and taxonomy of the genus. We previously analyzed Rns sequences from 18 isolates from morphological groups 2 and 3 and found that they fell into four distinct evolutionary lineages we called sequence types T1-T4. Here, we analyzed sequences from 53 isolates representing 16 species and including 35 new strains. Eight additional lineages (sequence types T5-T12) were identified. Four of the 12 sequence types included strains from more than one nominal species. Thus, sequence types could be equated with species in some cases or with complexes of closely related species in others. The largest complex, sequence type T4, which contained six closely related nominal species, included 24 of 25 keratitis isolates. Rns sequence variation was insufficient for full phylogenetic resolution of branching orders within this complex, but the mixing of species observed at terminal nodes confirmed that traditional classification of isolates has been inconsistent. One solution to this problem would be to equate sequence types and single species. Alternatively, additional molecular information will be required to reliably differentiate species within the complexes. Three sequence types of morphological group 1 species represented the earliest divergence in the history of the genus and, based on their genetic distinctiveness, are candidates for reclassification as one or more novel genera.

Acanthamoeba↗

Relationships among the O-antigen gene clusters of Salmonella enterica groups B, D1, D2, and D3.

The O antigen is an important cell wall antigen of gram-negative bacteria, and the genes responsible for its biosynthesis are located in a gene cluster. We have cloned and sequenced the DNA segment unique to the O-antigen gene cluster of Salmonella enterica group D3. This segment includes a novel O-antigen polymerase gene (wzyD3). The polymerase gives alpha(1-->6) linkages but has no detectable sequence similarity to that of group D2, which confers the same linkage. We find the remnant of a D3-like wzy gene in the O-antigen gene clusters of groups D1 and B and suggest that this is the original wzy gene of these O-antigen gene clusters.

Antibody Specificity↗

Analysis of the type 1 pilin gene cluster fim in Salmonella: its distinct evolutionary histories in the 5' and 3' regions.

The type 1 pilin encoded by fim is present in both Escherichia coli and Salmonella natural isolates, but several lines of evidence indicate that similarities at the fim locus may be an example of independent acquisition rather than common ancestry. For example, the fim gene cluster is found at different chromosomal locations and with distinct gene orders in these closely related species. In this work we examined the fim gene cluster of Salmonella, the genes of which show high nucleotide sequence divergence from their E. coli counterparts, as well as a different G+C content and codon usage. DNA hybridization analysis revealed that, among the salmonellae, the fim gene cluster is present in all isolates of S. enterica but is absent from S. bongori. Molecular phylogenetic analyses of the fimA and fimI genes yield an estimate of phylogeny that is in satisfactory congruence with housekeeping and other virulence genes examined in this species. In contrast, phylogenetic analyses of the fimZ, fimY, and fimW genes indicate that horizontal transfer of this region has occurred more than once. There is also size variation in the fimZ, fimY, and fimW intergenic regions in the 3' region, and these genes are absent in isolate S2983 of subspecies IIIa. Interestingly, the G+C contents of the fimZ, fimY, and fimW genes are less than 46%, which is considerably lower than those of the other six genes of the fim cluster. This study demonstrates that horizontal transmission of all or part of the same gene cluster can occur repeatedly, with the result that different regions of a single gene cluster may have different evolutionary histories.

Adhesins, Escherichia coli↗

Genome-wide analysis of core promoter elements from conserved human and mouse orthologous pairs.

BACKGROUND: The canonical core promoter elements consist of the TATA box, initiator (Inr), downstream core promoter element (DPE), TFIIB recognition element (BRE) and the newly-discovered motif 10 element (MTE). The motifs for these core promoter elements are highly degenerate, which tends to lead to a high false discovery rate when attempting to detect them in promoter sequences. RESULTS: In this study, we have performed the first analysis of these core promoter elements in orthologous mouse and human promoters with experimentally-supported transcription start sites. We have identified these various elements using a combination of positional weight matrices (PWMs) and the degree of conservation of orthologous mouse and human sequences--a procedure that significantly reduces the false positive rate of motif discovery. Our analysis of 9,010 orthologous mouse-human promoter pairs revealed two combinations of three-way synergistic effects, TATA-Inr-MTE and BRE-Inr-MTE. The former has previously been putatively identified in human, but the latter represents a novel synergistic relationship. CONCLUSION: Our results demonstrate that DNA sequence conservation can greatly improve the identification of functional core promoter elements in the human genome. The data also underscores the importance of synergistic occurrence of two or more core promoter elements. Furthermore, the sequence data and results presented here can help build better computational models for predicting the transcription start sites in the promoter regions, which remains one of the most challenging problems.

Animals↗

Application of a sensitive collection heuristic for very large protein families: evolutionary relationship between adipose triglyceride lipase (ATGL) and classic mammalian lipases.

BACKGROUND: Manually finding subtle yet statistically significant links to distantly related homologues becomes practically impossible for very populated protein families due to the sheer number of similarity searches to be invoked and analyzed. The unclear evolutionary relationship between classical mammalian lipases and the recently discovered human adipose triglyceride lipase (ATGL; a patatin family member) is an exemplary case for such a problem. RESULTS: We describe an unsupervised, sensitive sequence segment collection heuristic suitable for assembling very large protein families. It is based on fan-like expanding, iterative database searches. To prevent inclusion of unrelated hits, additional criteria are introduced: minimal alignment length and overlap with starting sequence segments, finding starting sequences in reciprocal searches, automated filtering for compositional bias and repetitive patterns. This heuristic was implemented as FAMILYSEARCHER in the ANNIE sequence analysis environment and applied to search for protein links between the classical lipase family and the patatin-like group. CONCLUSION: The FAMILYSEARCHER is an efficient tool for tracing distant evolutionary relationships involving large protein families. Although classical lipases and ATGL have no obvious sequence similarity and differ with regard to fold and catalytic mechanism, homology links detected with FAMILYSEARCHER show that they are evolutionarily related. The conserved sequence parts can be narrowed down to an ancestral core module consisting of three beta-strands, one alpha-helix and a turn containing the typical nucleophilic serine. Moreover, this ancestral module also appears in numerous enzymes with various substrate specificities, but that critically rely on nucleophilic attack mechanisms.

Adipose Tissue↗

Grain of truth.

Explore the source record for details and available documents.

Base Pairing↗

Refinement and prediction of protein prenylation motifs.

We refined the motifs for carboxy-terminal protein prenylation by analysis of known substrates for farnesyltransferase (FT), geranylgeranyltransferase I (GGT1) and geranylgeranyltransferase II (GGT2). In addition to the CaaX box for the first two enzymes, we identify a preceding linker region that appears constrained in physicochemical properties, requiring small or flexible, preferably hydrophilic, amino acids. Predictors were constructed on the basis of sequence and physical property profiles, including interpositional correlations, and are available as the Prenylation Prediction Suite (PrePS, http://mendel.imp.univie.ac.at/sat/PrePS) which also allows evaluation of evolutionary motif conservation. PrePS can predict partially overlapping substrate specificities, which is of medical importance in the case of understanding cellular action of FT inhibitors as anticancer and anti-parasite agents.

Alkyl and Aryl Transferases↗