PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

A putative lichenysin A synthetase operon in Bacillus licheniformis: initial characterization.

Certain Bacillus licheniformis strains isolated from oil wells have been shown to produce a very effective biosurfactant, lichenysin A, which is structurally similar to another less active lipopeptide, surfactin. Surfactin, like many small peptides in prokaryotes and lower eukaryotes, is synthesized non-ribosomally by multi-enzyme peptide synthetase complex. Analysis of several peptide synthetases of bacterial and fungal origin has revealed a high degree of sequence conservation. Two 35-mer oligonucleotides derived from highly conserved motifs ('core I' and 'core II') of surfactin synthetase were used to identify the cloned putative operon of lichenysin A synthetase lchA from B. licheniformis BNP29, a strain not amenable to genetic manipulation in a BAC system (F-plasmid-based bacterial artificial chromosome) based on Escherichia coli and its single-copy plasmid F-factor. A 32.4 kb fragment containing lichenysin A biosynthesis locus was sequenced and analysed. The structural architecture of putative lichenysin A synthetase protein containing seven amino acid (aa) activation-thiolation, two epimerization and one thioesterase domains is discussed in terms of its similarity to surfactin and other peptide synthetases. The 100 aa peptide chain situated between the highly conserved signature sequences FDXX and NXYGPTE(IV)X within amino acid binding domains of peptide synthetases is proposed to be a minimal block dictating the substrate specificity of the enzymes. A new operon-type structure has been localized directly upstream from the lichenysin A synthetase genes which, on the basis of sequence determination, potentially encode a four-member ABC-type transport system involved in product secretion.

Amino Acid Sequence↗

Roles of conserved residues in the arginase family.

Arginases and related enzymes metabolize arginine or similar nitrogen-containing compounds to urea or formamide. In the present report a sequence alignment of 31 members of this family was generated. The alignment, together with the crystal structure of rat liver arginase, allowed the assignment of possible functional or structural roles to 32 conserved residues and conservative substitutions. Two of these residues were previously identified as functionally essential by analysis of inherited defects in the type I arginase gene. Nearly half of the conserved residues are either glycines or prolines located at critical bends in the protein structure. Most metal-coordinating residues, including one histidine and four aspartic acid residues, are strictly conserved. Two additional histidines involved in metal-binding and catalysis are conserved in all arginases and in almost all other family members. Two positions with invariant similarities may serve as indirect metal ligands. Evolutionary relationships within this family were also suggested. Vertebrate type I and II arginases appear to have developed independently from an early gene duplication event. A ureohydrolase sequence from Caenorhabditis elegans is more closely related to other arginases than previously appreciated, while unclassified enzymes from Methanococcus jannaschii and Methanothermus fervidus appear more similar to arginase-related enzymes. In addition, enzymes from Arabidopsis thaliana and Synechocystis, previously identified as arginases, more closely resemble arginase-related enzymes than currently known arginases.

Amino Acid Sequence↗

An evolutionary approach reveals a high protein-coding capacity of the human genome.

We developed a new evolutionary method for identifying exons from genomic sequences and found 19000 potential coding exons that are absent from all existing annotations of the human genome. Of these, 13700 satisfied very stringent criteria and can with confidence be considered as novel exons. Evidently, a large number of new human genes can be identified using evolutionary approaches.

Animals↗

CADp44: a novel regulatory subunit of the 26S proteasome and the mammalian homolog of yeast Sug2p.

We have identified a novel protein, CADp44, based on the analysis of cDNAs derived from the brainstem of the 13-lined ground squirrel, Spermophilus tridecemlineatus. CADp44 has an unmodified molecular mass of 44,178 Da and contains multiple functional domains, including a conserved ATPase domain (CAD) and a leucine zipper motif. We show that distinct regions of the CADp44 sequence are identical to a set of peptides prepared from a recently identified bovine protein, referred to as p42, which is found in the PA700 regulatory complex of the 26S proteasome (DeMartino et al., 1996). We also show that CADp44 is the functional homolog of the newly characterized Sug2 protein from the budding yeast, Saccharomyces cerevisiae (Russell et al., 1996). Consistent with its role as a component of the 26S proteasome, CADp44 mRNA is found in all ground squirrel tissues examined. Evolutionary relationships based on sequence analysis show that both CADp44 and yeast Sug2p are distinct from the other five CAD ATPases found in the PA700, and together comprise the sixth and newest CAD subunit of the regulatory complex of the 26S proteasome.

ATPases Associated with Diverse Cellular Activitie↗

Genomic characterization of the Neurofibromatosis Type 1 gene of Fugu rubripes.

The genomic structure of the Neurofibromatosis Type1 (NF1) gene of Fugu rubripes was investigated by sequence analysis of two overlapping cosmids. The Fugu NF1 gene spans 27 kb and is 13 times smaller than the human counterpart owing primarily to reduced intron size. The predicted amino acid sequence is highly related to that of human neurofibromin, exhibiting an overall similarity of 91.5%. Nearly all exons described for the human NF1 gene could be identified, except exon 12b and the alternatively spliced exons 9br and 48a. With the exception of the splice acceptor site in front of exon 16, all splice sites are in identical positions to those found in the human gene. Intron 1, which is 100-140 kb long in humans, spans 2575 bp in the Fugu NF1 gene. Another large intron of the human NF1 gene, intron 27b (45-50 kb), is 3942 bp of size in Fugu. Sequences related to the OMgp gene (Oligodendrocyte-Myelin-glycoprotein) or the EVI2A gene (ecotropic viral integration site), which are inserted into human NF1 intron 27b, were not detected in the corresponding Fugu intron. However, a single exon gene with similarity to the human EVI2B gene has been found on the reverse strand of Fugu intron 27b. This suggests that the human EVI2B gene and the Fugu gene in intron 27b have a common ancestor. We found the expression of this inserted gene in liver and kidney, but not in brain tissue of Fugu rubripes.

Amino Acid Sequence↗

Multispecies comparative analysis of a mammalian-specific genomic domain encoding secretory proteins.

The mammalian-specific casein gene cluster comprises 3 or 4 evolutionarily related genes and 1 physically linked gene with a functional association. To gain a better understanding of the mechanisms regulating the entire casein cluster at the genomic level we initiated a multispecies comparative sequence analysis. Despite the high level of divergence at the coding level, these studies have identified uncharacterized family members within two species and the presence at orthologous positions of previously uncharacterized genes. Also the previous suggestion that the histatin/statherin gene family, located in this region, was primate specific was ruled out. All 11 genes identified in this region appear to encode secretory proteins. Conservation of a number of noncoding regions was observed; one coincides with an element previously suggested to be important for beta-casein gene expression in human and cow. The conserved regions might have biological importance for the regulation of genes in this genomic "neighborhood."

Amino Acid Sequence↗

Casein kinase I: spatial organization and positioning of a multifunctional protein kinase family.

The casein kinase I family of serine/threonine protein kinases is highly conserved from yeast to humans. Until only recently, both the function and regulation of these enzymes remained poorly uncharacterised in that they appeared to be constitutively active and were capable of phosphorylating an untold number of other proteins. While relatively little was known regarding the exact function of the higher eukaryotic isoforms, the casein kinase I (CKI) isoforms from yeast have been genetically linked to vesicular trafficking, DNA repair, cell cycle progression and cytokinesis. All five S. cerevisiae isoforms are known to be associated with discrete cellular compartments and this localization has been shown to be absolutely essential for their respective functions. New evidence now suggests that the CKI isoforms in more complex systems also exhibit non-homogeneous subcellular distributions that may prove vital to defining the function and regulation of these enzymes. In particular, CKIalpha, the most-characterized vertebrate isoform, is associated with cytosolic vesicles, the mitotic spindle and structures within the nucleus. Functions associated with these localizations coincide with those previously reported in yeast, suggesting a conservation of function. Other reports have indicated that each of the remaining CKI isoforms have the capacity to make associations with components of several signal transduction pathways, thereby channeling CKI function toward specific regulatory events. This review will examine what is now known about the higher eukaryotic CKI family members from the perspective localization as a means of gaining a better understanding of the function and regulation of these kinases.

Amino Acid Sequence↗

Complete genome sequences of cellular life forms: glimpses of theoretical evolutionary genomics.

The availability of complete genome sequences of cellular life forms creates the opportunity to explore the functional content of the genomes and evolutionary relationships between them at a new qualitative level. With the advent of these sequences, the construction of a minimal gene set sufficient for sustaining cellular life and reconstruction of the genome of the last common ancestor of bacteria, eukaryotes, and archaea become realistic, albeit challenging, research projects. A version of the minimal gene set for modern-type cellular life derived by comparative analysis of two bacterial genomes, those of Haemophilus influenzae and Mycoplasma genitalium, consists of approximately 250 genes. A comparison of the protein sequences encoded in these genes with those of the proteins encoded in the complete yeast genome suggests that the last common ancestor of all extant life might have had an RNA genome.

Bacterial Proteins↗

Signal sequences control gating of the protein translocation channel in a substrate-specific manner.

N-terminal signal sequences mediate targeting of nascent chains to the endoplasmic reticulum and facilitate opening of the protein translocation channel to the passage of substrate. We have assessed each of these steps for a diverse set of mammalian signals. While minimal differences were seen in their targeting function, signal sequences displayed a remarkable degree of variation in initiating nascent chain access to the lumenal environment. Such substrate-specific properties of signals were evolutionarily conserved, functionally matched to their respective mature domains, and important for the proper biogenesis of some proteins. Thus, the sequence variations of signals do not simply represent functional degeneracy, but instead encode critical differences in translocon gating that are coordinated with their respective passengers to facilitate efficient translocation.

3T3 Cells↗

The R protein of SARS-CoV: analyses of structure and function based on four complete genome sequences of isolates BJ01-BJ04.

The R (replicase) protein is the uniquely defined non-structural protein (NSP) responsible for RNA replication, mutation rate or fidelity, regulation of transcription in coronaviruses and many other ssRNA viruses. Based on our complete genome sequences of four isolates (BJ01-BJ04) of SARS-CoV from Beijing, China, we analyzed the structure and predicted functions of the R protein in comparison with 13 other isolates of SARS-CoV and 6 other coronaviruses. The entire ORF (open-reading frame) encodes for two major enzyme activities, RNA-dependent RNA polymerase (RdRp) and proteinase activities. The R polyprotein undergoes a complex proteolytic process to produce 15 function-related peptides. A hydrophobic domain (HOD) and a hydrophilic domain (HID) are newly identified within NSP1. The substitution rate of the R protein is close to the average of the SARS-CoV genome. The functional domains in all NSPs of the R protein give different phylogenetic results that suggest their different mutation rate under selective pressure. Eleven highly conserved regions in RdRp and twelve cleavage sites by 3CLP (chymotrypsin-like protein) have been identified as potential drug targets. Findings suggest that it is possible to obtain information about the phylogeny of SARS-CoV, as well as potential tools for drug design, genotyping and diagnostics of SARS.

Amino Acid Sequence↗

Yeast Rrp9p is an evolutionarily conserved U3 snoRNP protein essential for early pre-rRNA processing cleavages and requires box C for its association.

Pre-rRNA processing in eukaryotic cells requires participation of several snoRNPs. These include the highly conserved and abundant U3 snoRNP, which is essential for synthesis of 18S rRNA. Here we report the characterization of Rrp9p, a novel yeast U3 protein, identified via its homology to the human U3-55k protein. Epitope-tagged Rrp9p specifically precipitates U3 snoRNA, but Rrp9p is not required for the stable accumulation of this snoRNA. Genetic depletion of Rrp9p inhibits the early cleavages of the primary pre-rRNA transcript at A0, A1, and A2 and, consequently, production of 18S, but not 25S and 5.8S, rRNA. The hU3-55k protein can partially complement a yeast rrp9 null mutant, indicating that the function of this protein has been conserved. Immunoprecipitation of extracts from cells that coexpress epitope-tagged Rrp9p and various mutant forms of U3 snoRNA limits the region required for association of Rrp9p to the U3-specific box B/C motif. Box C is essential, whereas box B plays a supportive role.

Amino Acid Sequence↗

The evolutionarily conserved region of the U snRNA export mediator PHAX is a novel RNA-binding domain that is essential for U snRNA export.

In metazoa, a subset of spliceosomal U snRNAs are exported from the nucleus after transcription. This export occurs in a large complex containing a U snRNA, the nuclear cap binding complex (CBC), the leucine-rich nuclear export signal receptor CRM1/Xpo1, RanGTP, and the recently identified phosphoprotein PHAX (phosphorylated adaptor for RNA export). Previous results indicated that PHAX made direct contact with RNA, CBC, and Xpo1 in the U snRNA export complex. We have now performed a systematic characterization of the functional domains of PHAX. The most evolutionarily conserved region of PHAX is shown to be a novel RNA-binding domain that is essential for U snRNA export. In addition, PHAX contains two major nuclear localization signals (NLSs) that are required for its recycling to the nucleus after export. The interaction domain of PHAX with CBC is at least partly distinct from the RNA-binding domain and the NLSs. Thus, the different interaction domains of PHAX allow it to act as a scaffold for the assembly of U snRNA export complexes.

Active Transport, Cell Nucleus↗

Complete sequence and comparative genome analysis of the dairy bacterium Streptococcus thermophilus.

The lactic acid bacterium Streptococcus thermophilus is widely used for the manufacture of yogurt and cheese. This dairy species of major economic importance is phylogenetically close to pathogenic streptococci, raising the possibility that it has a potential for virulence. Here we report the genome sequences of two yogurt strains of S. thermophilus. We found a striking level of gene decay (10% pseudogenes) in both microorganisms. Many genes involved in carbon utilization are nonfunctional, in line with the paucity of carbon sources in milk. Notably, most streptococcal virulence-related genes that are not involved in basic cellular processes are either inactivated or absent in the dairy streptococcus. Adaptation to the constant milk environment appears to have resulted in the stabilization of the genome structure. We conclude that S. thermophilus has evolved mainly through loss-of-function events that remarkably mirror the environment of the dairy niche resulting in a severely diminished pathogenic potential.

Bacterial Proteins↗

Compact and ordered collapse of randomly generated RNA sequences.

As the raw material for evolution, arbitrary RNA sequences represent the baseline for RNA structure formation and a standard to which evolved structures can be compared. Here, we set out to probe, using physical and chemical methods, the structural properties of RNAs having randomly generated oligonucleotide sequences that were of sufficient length and information content to encode complex, functional folds, yet were unbiased by either genealogical or functional constraints. Typically, these unevolved, nonfunctional RNAs had sequence-specific secondary structure configurations and compact magnesium-dependent conformational states comparable to those of evolved RNA isolates. But unlike evolved sequences, arbitrary sequences were prone to having multiple competing conformations. Thus, for RNAs the size of small ribozymes, natural selection seems necessary to achieve uniquely folding sequences, but not to account for the well-ordered secondary structures and overall compactness observed in nature.

Base Sequence↗

Parafibromin is a nuclear protein with a functional monopartite nuclear localization signal.

Parafibromin is a nuclear protein with a tumour suppressor role in the development of non-hereditary and hereditary parathyroid carcinomas, and the hyperparathyroidism-jaw tumour (HPT-JT) syndrome, which is associated with renal and uterine tumours. Nuclear localization signal(s), (NLS(s)), of the 61 kDa parafibromin remain to be defined. Utilization of computer-prediction programmes, identified five NLSs (three bipartite (BP) and two monopartite (MP)). To investigate their functionality, wild-type (WT) and mutant parafibromin constructs tagged with enhanced green fluorescent protein or cMyc were transiently expressed in COS-7 cells, or human embryonic kidney 293 (HEK293) cells, and their subcellular locations determined by confocal fluorescence microscopy. Western blot analyses of nuclear and cytoplasmic fractions from the transfected cells were also performed. WT parafibromin localized to the nucleus and deletions or mutations of the three predicted BP and one of the predicted MP NLSs did not affect this localization. In contrast, deletions or mutations of a MP NLS, at residues 136-139, resulted in loss of nuclear localization. Furthermore, the critical basic residues, KKXR, of this MP NLS were found to be evolutionarily conserved, and over 60% of all parafibromin mutations lead to a loss of this NLS. Thus, an important functional domain of parafibromin, consisting of an evolutionarily conserved MP NLS, has been identified.

Amino Acid Sequence↗

Phenotype-genotype correlation in Hirschsprung disease is illuminated by comparative analysis of the RET protein sequence.

The ability to discriminate between deleterious and neutral amino acid substitutions in the genes of patients remains a significant challenge in human genetics. The increasing availability of genomic sequence data from multiple vertebrate species allows inclusion of sequence conservation and physicochemical properties of residues to be used for functional prediction. In this study, the RET receptor tyrosine kinase serves as a model disease gene in which a broad spectrum (> or = 116) of disease-associated mutations has been identified among patients with Hirschsprung disease and multiple endocrine neoplasia type 2. We report the alignment of the human RET protein sequence with the orthologous sequences of 12 non-human vertebrates (eight mammalian, one avian, and three teleost species), their comparative analysis, the evolutionary topology of the RET protein, and predicted tolerance for all published missense mutations. We show that, although evolutionary conservation alone provides significant information to predict the effect of a RET mutation, a model that combines comparative sequence data with analysis of physiochemical properties in a quantitative framework provides far greater accuracy. Although the ability to discern the impact of a mutation is imperfect, our analyses permit substantial discrimination between predicted functional classes of RET mutations and disease severity even for a multigenic disease such as Hirschsprung disease.

Amino Acid Sequence↗