PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Proteome composition in Plasmodium falciparum: higher usage of GC-rich nonsynonymous codons in highly expressed genes.

The parasite Plasmodium falciparum, responsible for the most deadly form of human malaria, is one of the extremely AT-rich genomes sequenced so far and known to possess many atypical characteristics. Using multivariate statistical approaches, the present study analyzes the amino acid usage pattern in 5038 annotated protein-coding sequences in P. falciparum clone 3D7. The amino acid composition of individual proteins, though dominated by the directional mutational pressure, exhibits wide variation across the proteome. The Asn content, expression level, mean molecular weight, hydropathy, and aromaticity are found to be the major sources of variation in amino acid usage. At all stages of development, frequencies of residues encoded by GC-rich codons such as Gly, Ala, Arg, and Pro increase significantly in the products of the highly expressed genes. Investigation of nucleotide substitution patterns in P. falciparum and other Plasmodium species reveals that the nonsynonymous sites of highly expressed genes are more conserved than those of the lowly expressed ones, though for synonymous sites, the reverse is true. The highly expressed genes are, therefore, expected to be closer to their putative ancestral state in amino acid composition, and a plausible reason for their sequences being GC-rich at nonsynonymous codon positions could be that their ancestral state was less AT-biased. Negative correlation of the expression level of proteins with respective molecular weights supports the notion that P. falciparum, in spite of its intracellular parasitic lifestyle, follows the principle of cost minimization.

Amino Acids↗

Sequence and evolutionary analysis of the human trypsin subfamily of serine peptidases.

Serine peptidases (SP) are peptidases with a uniquely activated serine residue in the substrate-binding site. SP can be classified into clans with distinct evolutionary histories and each clan further subdivided into families. We analyzed 79 proteins representing the S1A subfamily of human SP, obtained from different databases. Multiple alignment identified 87 highly conserved amino acid residues. In most cases of substitution, a residue of similar character was inserted, implying that the overall character of the local region was conserved. We also identified several conserved protein motifs. 7-13 cysteine positions, potentially forming disulfide bridges, were also found to be conserved. Most members are secreted as inactive (pro) forms with a trypsin-like cleavage site for activation. Substrate specificity was predicted to be trypsin-like for most members, with few chymotrypsin-like proteins. Phylogenetic analysis enabled us to classify members of the S1A subfamily into structurally related groups; this might also help to functionally sort members of this subfamily and give an idea about their possible functions.

Amino Acid Motifs↗

Identification of an evolutionary conserved SURF-6 domain in a family of nucleolar proteins extending from human to yeast.

The mammalian SURF-6 protein is localized in the nucleolus, yet its function remains elusive in the recently characterized nucleolar proteome. We discovered by searching the Protein families database that a unique evolutionary conserved SURF-6 domain is present in the carboxy-terminal of a novel family of eukaryotic proteins extending from human to yeast. By using the enhanced green fluorescent protein as a fusion protein marker in mammalian cells, we show that proteins from distantly related taxonomic groups containing the SURF-6 domain are localized in the nucleolus. Deletion sequence analysis shows that multiple regions of the SURF-6 protein are capable of nucleolar targeting independently of the evolutionary conserved domain. We identified that the Saccharomyces cerevisiae member of the SURF-6 family, named rrp14 or ykl082c, has been categorized in yeast databases to interact with proteins involved in ribosomal biogenesis and cell polarity. These results classify SURF-6 as a new family of nucleolar proteins in the eukaryotic kingdom and point out that SURF-6 has a distinct domain within the known nucleolar proteome that may mediate complex protein-protein interactions for analogous processes between yeast and mammalian cells.

Amino Acid Sequence↗

Conservation and diversity of ancient hemoglobins in Bacteria.

A group of single-domain proteins in Bacteria similar to thermoglobin, an oxygen-avid hemoglobin representative of the ancestral form, reveals the primordial structure, function, and evolvability of the family. Conserved residues at specific positions function to bind ligand or participate in hydrophobic packing of the protein core during protein folding. A potential hydrogen bond network consisting of a tyrosine and glutamine residue in the distal ligand-binding site of most hemoglobins suggests that the ancestral protein bound oxygen avidly. Two divergent hemoglobins with mutations at generally conserved positions contain non-canonical ligand-binding sites, illustrating plasticity of the fold. One binds heme in a manner similar to cytochromes and may represent an evolutionary link to the precursor of the hemoglobin fold. Conservation suggests specific biochemical properties of the ancestral protein; diversity suggests an evolvability of this group of hemoglobins tolerant of mutations that perturb conserved biochemical properties for adaptation to novel functions.

Amino Acid Sequence↗

A splice variant of the human CCA-adding enzyme with modified activity.

The human CCA-adding enzyme (tRNA nucleotidyltransferase) is an essential enzyme that catalyzes the addition of the CCA terminus to the 3' end of tRNA precursors, a reaction which is a fundamental prerequisite for mature tRNAs to become aminoacylated and to participate in protein biosynthesis. To date only one form of this enzyme has been identified in humans. Here, we describe the sequence and activity of a splice variant of the human CCA-adding enzyme identified in public cDNA databases. The in silico analyses performed on this splice variant indicate that there is conservation of the alternative splice donor site among species and indicate that it seems to be used in vivo. Moreover, the recombinantly expressed protein is active in vitro and accepts tRNA transcripts as substrates incorporating the dinucleotide sequence CC to their 3' end, in contrast to the activity of the full length enzyme. These findings strongly suggest that the splice variant of the human CCA-adding enzyme is expressed in the cell although the in vivo function remains unclear.

Adenosine Triphosphate↗

Evidence from the evolutionary analysis of nucleotide sequences for a recombinant history of SARS-CoV.

The origins and evolutionary history of the Severe Acute Respiratory Syndrome (SARS) coronavirus (SARS-CoV) remain an issue of uncertainty and debate. Based on evolutionary analyses of coronavirus DNA sequences, encompassing an approximately 13kb stretch of the SARS-TOR2 genome, we provide evidence that SARS-CoV has a recombinant history with lineages of types I and III coronavirus. We identified a minimum of five recombinant regions ranging from 83 to 863bp in length and including the polymerase, nsp9, nsp10, and nsp14. Our results are consistent with a hypothesis of viral host jumping events, concomitant with the reassortment of bird and mammalian coronaviruses, a scenario analogous to earlier outbreaks of influenzae.

Animals↗

Characterization and tissue distribution of multiple agouti-family genes in pufferfish, Takifugu rubripes.

Four types of agouti-family genes (AGRP1, AGRP2, ASIP1 and ASIP2) were obtained from torafugu, Takifugu rubripes. Their characterization and structure were analyzed to elucidate the relationship among the torafugu agouti-family genes. Both AGRP1 and AGRP2 showed genomic synteny with the human AGRP gene. Phylogenetic tree analysis showed that AGRP1 formed a cluster with human AGRP. We inferred that torafugu AGRP1 and AGRP2 are orthologs of human AGRP and that they are paralogous genes derived from genome duplication occurred in the teleost phylogeny. Torafugu ASIP1 showed genomic synteny with the human ASIP, but ASIP2 did not. The ASIP1 expression level was about five times higher in the white ventral skin than in the black dorsal skin. Therefore, we concluded that torafugu ASIP1 is an ortholog of human ASIP, nevertheless, we are unable to determine if torafugu ASIP2 is a paralog of ASIP1 or not.

Agouti Signaling Protein↗

The myotubularin family: from genetic disease to phosphoinositide metabolism.

The myotubularin-related genes define a large family of eukaryotic proteins, most of them initially characterized by the presence of a ten-amino acid consensus sequence related to the active sites of tyrosine phosphatases, dual-specificity protein phosphatases and the lipid phosphatase PTEN. Myotubularin (hMTM1), the founder member, is mutated in myotubular myopathy, and a close homolog (hMTMR2) was recently found mutated in a recessive form of Charcot-Marie-Tooth neuropathy. Although myotubularin was thought to be a dual-specificity protein phosphatase, recent results indicate that it is primarily a lipid phosphatase, acting on phosphatidylinositol 3-monophosphate, and might be involved in the regulation of phosphatidylinositol 3-kinase (PI 3-kinase) pathway and membrane trafficking.

Amino Acid Sequence↗

Chlamydomonas U2, U4 and U6 snRNAs. An evolutionary conserved putative third interaction between U4 and U6 snRNAs which has a counterpart in the U4atac-U6atac snRNA duplex.

The spliceosomal UsnRNAs U2, U4 and U6 from the green alga Chlamydomonas reinhardtii (Cre) were sequenced using a combination of RNA and cDNA sequencing methods and were compared to other sequenced UsnRNAs. The lengths of Cre U6 and Cre U2 RNAs are similar to those of their higher plant equivalents. Cre U4 RNA is shorter (139 nt) than its counterpart from higher plants (150-154 nt), and contains stem IV and loop D which are absent, with the exception of the Tetrahymena U4 RNA, from the U4 RNAs of other unicellular organisms studied to date. Base-pairing interactions between U6 and U4 RNAs and between U6 and U2 RNAs, identical to those described for mammalian and yeast systems, are structurally feasible in the Cre system. In addition, based on comparative analyses of the predicted U4/U6 RNA duplex from various species, an evolutionary conserved third putative U6-U4 interaction was found. Interestingly, it can also be formed with the recently discovered U6atac and U4atac RNAs. This is a strong support in favor of the possible biological significance of this third putative interaction. Based on comparative analysis, an extension of the earlier described U6-U2 interaction patterns is also proposed.

Alternative Splicing↗

The tyrosine decarboxylase operon of Lactobacillus brevis IOEB 9809: characterization and conservation in tyramine-producing bacteria.

Bacterial genes of tyrosine decarboxylases were recently identified. Here we continued the sequencing of the tyrosine decarboxylase locus of Lactobacillus brevis IOEB 9809 and determined a total of 7979 bp. The sequence contained four complete genes encoding a tyrosyl-tRNA synthetase, the tyrosine decarboxylase, a probable tyrosine permease and a Na+/H+ antiporter. Rapid amplification of cDNA ends (RACE) was employed to determine the 5'-end of mRNAs containing the tyrosine decarboxylase gene. It was located only 34-35 nucleotides upstream of the start codon, suggesting that the preceding tyrosyl-tRNA synthetase gene was transcribed separately. In contrast, reverse transcription-polymerase chain reactions (RT-PCRs) carried out with primers designed to amplify regions spanning gene junctions showed that some mRNAs contained the four genes. Homology searches revealed similar clusters of four genes in the genome sequences of Enterococcus faecalis and Enterococcus faecium. Phylogenetic analyses supported the hypothesis that these genes evolved all together. These data suggest that bacterial tyrosine decarboxylases are encoded in an operon containing four genes.

Amino Acid Sequence↗

The human genome: an immuno-centric view of evolutionary strategies.

A hallmark of modern biology is the realization of the fundamental unity of biological processes in all life forms. Consequently, the complete genome sequencing of various bacteria, yeast (Saccharomyces cerevisiae), fly (Drosophila melanogaster) and worm (Caenorhabditis elegans) over the past five years has already had an impact on all of biology. "Model organisms" have contributed a great deal to immunology; for example, the Toll receptors of the fly provided the impetus for the investigation of Toll-like receptors, which proved to be fundamental elements in the mammalian innate immune system. The recent release of a draft sequence of the human genome provides the first panoramic view of the 30000-35000 human genes in the human genetic blueprint and provides a plethora of new details, the significance of which will take some time to appreciate. The over-riding concepts that emerge from these studies relate primarily to general evolutionary processes that are equally as relevant to immunology as they are to other disciplines of biology.

Amino Acid Sequence↗

Rio2p, an evolutionarily conserved, low abundant protein kinase essential for processing of 20 S Pre-rRNA in Saccharomyces cerevisiae.

Saccharomyces cerevisiae Rio2p (encoded by open reading frame Ynl207w) is an essential protein of unknown function that displays significant sequence similarity to Rio1p/Rrp10p. The latter was recently shown to be an evolutionarily conserved, predominantly cytoplasmic serine/threonine kinase whose presence is required for the final cleavage at site D that converts 20 S pre-rRNA into mature 18 S rRNA. A data base search identified homologs of Rio2p in a wide variety of eukaryotes and Archaea. Detailed sequence comparison and in vitro kinase assays using recombinant protein demonstrated that Rio2p defines a subfamily of protein kinases related to, but both structurally and functionally distinct from, the one defined by Rio1p. Failure to deplete Rio2p in cells containing a GAL-rio2 gene and direct analysis of Rio2p levels by Western blotting indicated the protein to be low abundant. Using a GAL-rio2 gene carrying a point mutation that reduces the kinase activity, we found that depletion of this mutant protein blocked production of 18 S rRNA due to inhibition of the cleavage of cytoplasmic 20 S pre-rRNA at site D. Production of the large subunit rRNAs was not affected. Thus, Rio2p is the second protein kinase that is essential for cleavage at site D and the first in which the processing defect can be linked to its enzymatic activity. Contrary to Rio1p/Rrp10p, however, Rio2p appears to be localized predominantly in the nucleus.

Base Sequence↗