PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Initiation of mucin-type O-glycosylation in dictyostelium is homologous to the corresponding step in animals and is important for spore coat function.

Like animal cells, many unicellular eukaryotes modify mucin-like domains of secretory proteins with multiple O-linked glycans. Unlike animal mucin-type glycans, those of some microbial eukaryotes are initiated by alpha-linked GlcNAc rather than alpha-GalNAc. Based on sequence similarity to a recently cloned soluble polypeptide hydroxyproline GlcNAc-transferase that modifies Skp1 in the cytoplasm of the social ameba Dictyostelium, we have identified an enzyme, polypeptide alpha-N-acetylglucosaminyltransferase (pp alpha-GlcNAc-T2), that attaches GlcNAc to numerous secretory proteins in this organism. Unlike the Skp1 GlcNAc-transferase, pp alpha-GlcNAc-T2 is predicted to be a type 2 transmembrane protein. A highly purified, soluble, recombinant fragment of pp alpha-GlcNAc-T2 efficiently transfers GlcNAc from UDP-GlcNAc to synthetic peptides corresponding to mucin-like domains in two proteins that traverse the secretory pathway. pp alpha-GlcNAc-T2 is required for addition of GlcNAc to peptides in cell extracts and to the proteins in vivo. Mass spectrometry and Edman degradation analyses show that pp alpha-GlcNAc-T2 attaches GlcNAc in alpha-linkage to the Thr residues of all the synthetic mucin repeats. pp alpha-GlcNAc-T2 is encoded by the previously described modB locus defined by chemical mutagenesis, based on sequence analysis and complementation studies. This finding establishes that the many phenotypes of modB mutants, including a permeability defect in the spore coat, can now be ascribed to defects in mucin-type O-glycosylation. A comparison of the sequences of pp alpha-GlcNAc-T2 and the animal pp alpha-GalNAc-transferases reveals an ancient common ancestry indicating that, despite the different N-acetylhexosamines involved, the enzymes share a common mechanism of action.

Acetylglucosamine↗

Structural clues to Rab GTPase functional diversity.

Rab GTPases are key regulators of membrane trafficking in eukaryotes. Recent structural analysis of a number of Rabs, either alone or in complex with partner proteins, has provided new insight into the importance of both conserved and non-conserved features of these proteins that specify their unique functions and localizations. This review will highlight what we have learned from crystallographic analysis of this important protein family.

Amino Acid Motifs↗

GeneSyn: a tool for detecting conserved gene order across genomes.

UNLABELLED: GeneSyn is a software tool that allows automatic detection of conserved gene order from annotated genomes. AVAILABILITY: Available free of charge for Unix/Linux/Cygwin platforms at ftp://159.149.110.11/pub/GeneSyn_1.0/ SUPPLEMENTARY INFORMATION: ftp://159.149.110.11/pub/GeneSyn_1.0/

Algorithms↗

Human mitochondrial DNA variation and evolution: analysis of nucleotide sequences from seven individuals.

We have analyzed nucleotide sequence variation in an approximately 900-base pair region of the human mitochondrial DNA molecule encompassing the heavy strand origin of replication and the D-loop. Our analysis has focused on nucleotide sequences available from seven humans. Average nucleotide diversity among the sequences is 1.7%, several-fold higher than estimates from restriction endonuclease site variation in mtDNA from these individuals and previously reported for other humans. This disparity is consistent with the rapidly evolving nature of this noncoding region. However, several instances of convergent or parallel gain and loss of restriction sites due to multiple substitutions were observed. In addition, other results suggest that restriction site (as well as pairwise sequence) comparisons may underestimate the total number of substitutions that have occurred since the divergence of two mtDNA sequences from a common ancestral sequence, even at low levels of divergence. This emphasizes the importance of recognizing the large standard errors associated with estimates of sequence variability, particularly when constructing phylogenies among closely related sequences. Analysis of the observed number and direction of substitutions revealed several significant biases, most notably a strand dependence of substitution type and a 32-fold bias favoring transitions over transversions. The results also revealed a significantly nonrandom distribution of nucleotide substitutions and sequence length variation. Significantly more multiple substitutions were observed than expected for these closely related sequences under the assumption of uniform rates of substitution. The bias for transitions has resulted in predominantly convergent or parallel changes among the observed multiple substitutions. There is no convincing evidence that recombination has contributed to the mtDNA sequence diversity we have observed.

Base Sequence↗

Ancient SINEs from African endemic mammals.

Afrotheria is a newly recognized taxon comprising elephants, hyraxes, sea cows, aardvarks, golden moles, tenrecs, and elephant shrews, each of which originated in Africa. Although some members of this taxon were once classified into distantly related groups, recent molecular studies have demonstrated their close relationships. It was suggested that this group emerged as a result of physical isolation of the African continent during the successive breakup events of Gondowanaland. In this study, a novel family of SINEs, designated AfroSINEs, was isolated and characterized from the genomes of afrotherians. This SINE family is distributed exclusively among the afrotherian species, confirming their monophyletic relationships. Furthermore, a distinct subfamily, which shares a deletion in the middle region of the SINE, was identified. The distribution of this subfamily is apparently restricted to the genomes of hyraxes, elephants, and sea cows, suggesting monophyly of these three groups, which was previously proposed as Paenungulata. We characterized the structures of the AfroSINEs from all afrotherian representatives by PCR, and we discuss how they were generated as well as the phylogenetic relationships of their host species.

Africa↗

Species-specific differences in the operational RNA code for aminoacylation of tRNA(Trp).

Identity elements play essential roles in the recognition of tRNAs by their cognate aminoacyl-tRNA synthetase. An operational RNA code relates amino acids to specific sequences and structural features of tRNA acceptor stems. In this study, a series of tRNA(Trp) variants was prepared by in vitro transcription and their efficiencies of aminoacylation by tryptophan (k(cat)/K(m)) were measured with the aid of Bacillus subtilis and human tryptophanyl-tRNA synthetases (TrpRS). The identity elements in the operational RNA code of human tRNA(Trp) were found to be: major element, discriminator base A73; minor elements, G1/C72 and U5/G68. From the cross-species aminoacylation assays, we conclude that the identity elements in tRNA(Trp) from B.subtilis and human all contribute to species-specific aminoacylation by TrpRS. Analyses of 22 TrpRS sequences covering three taxonomic domains (bacteria, eukarya and archaea) reveal that the sequences are divided into two evolutionarily distant groups. The same partition is also observed in the analyses of tRNA(Trp) acceptor stem sequences. Our data suggest that the two TrpRS groups may reflect co-adaptations needed to accommodate changes in the operational RNA code for tryptophan.

Acylation↗

Conserved sequence motifs, alignment, and secondary structure for the third domain of animal 12S rRNA.

Secondary structure models are an important step for aligning sequences, understanding probabilities of nucleotide substitutions, and evaluating the reliability of phylogenetic reconstructions. A set of conserved sequence motifs is derived from comparative sequence analysis of 184 invertebrate and vertebrate taxa (including many taxa from the same genera, families, and orders) with reference to a secondary structure model for domain III of animal mitochondrial small subunit (12S) ribosomal RNA. A template is presented to assist with secondary structure drawing. Our model is similar to previous models but is more specific to mitochondrial DNA, fitting both invertebrate and vertebrate groups, including taxa with markedly different nucleotide compositions. The second half of the domain III sequence can be difficult to align precisely, even when secondary structure information is considered. This is especially true for comparisons of anciently diverged taxa, but well-conserved motifs assist in determining biologically meaningful alignments. Patterns of conservation and variability in both paired and unpaired regions make differential phylogenetic weighting in terms of "stems" and "loops" unsatisfactory. We emphasize looking carefully at the sequence data before and during analyses, and advocate the use of conserved motifs and other secondary structure information for assessing sequencing fidelity.

Animals↗

A nuclear gene for higher level phylogenetics: phosphoenolpyruvate carboxykinase tracks mesozoic-age divergences within Lepidoptera (Insecta).

The sequence of phosphoenolpyruvate carboxykinase (PEPCK) has been previously identified as a promising candidate for reconstructing Mesozoic-age divergences (Friedlander, Regier, and Mitter 1992, 1994). To test this hypothesis more rigorously, 597 nucleotides of aligned PEPCK coding sequence (approximately 30% of the coding region) were generated from 18 species representing Mesozoic-age lineages of moths (Insecta: Lepidoptera) and outgroup taxa. Relationships among basal Lepidoptera are well established by morphological analysis, providing a strong test for the utility of a gene which has not previously been used in systematics. Parsimony and other phylogenetic analyses were conducted on nucleotides by codon positions (nt1, nt2, nt3) separately and in combination, and on amino acids, for comparison to the test phylogeny. The highest concordance was achieved with nt1 + nt2, for which one of two most-parsimonious trees was identical to the test phylogeny, and with all nucleotides when nt3 was down-weighted sevenfold or higher, for which a single most-parsimonious tree identical to the test phylogeny resulted. Substitutions in nt3 approached saturation in many, but not all, pairwise comparisons and their exclusion or severe downweighting greatly increased the degree of concordance with the test phylogeny. Neighbor-joining analysis confirms this finding. The utility of PEPCK for phylogenetics is demonstrated over a time span for which few other suitable genes are currently available.

Amino Acid Sequence↗

Uricase protein sequences: conserved during vertebrate evolution but absent in humans.

Uricase is a peroxisomal liver enzyme that catalyzes the oxidation of uric acid to allantoin during purine catabolism. It is present in vertebrates in most species of fish, amphibians, and mammals but its enzymatic activity is absent in hominoids. We have used Western blot analysis in a comparative study to establish a homology among uricases from different species of vertebrates. Using antibodies against denatured rat liver uricase, we have been able to detect for the first time cross-reactivity with the uricase of species ranging in the evolutionary scale from fish to primates (macaque). Our results suggest that these uricases have a common evolutionary origin. Our conclusion is also supported by the fact that uricase from different species exhibits identical tissue, subcellular localization, and similarity of molecular weights. This study was extended to include human liver samples. Using the same approach but with a more sensitive detection system (alkaline phosphatase instead of peroxidase), we did not detect polypeptide species related to rat uricase in human fetal or adult liver samples, which indicates that during hominoid evolution, the mutational event responsible for the loss of uricase activity in humans precluded formation of a translatable uricase mRNA.

Animals↗

The largest subunit of human RNA polymerase III is closely related to the largest subunit of yeast and trypanosome RNA polymerase III.

In both yeast and mammalian systems, considerable progress has been made toward the characterization of the transcription factors required for transcription by RNA polymerase III. However, whereas in yeast all of the RNA polymerase III subunits have been cloned, relatively little is known about the enzyme itself in higher eukaryotes. For example, no higher eukaryotic sequence corresponding to the largest RNA polymerase III subunit is available. Here we describe the isolation of cDNAs that encode the largest subunit of human RNA polymerase III, as suggested by the observations that (1) antibodies directed against the cloned protein immunoprecipitate an active enzyme whose sensitivity to different concentrations of alpha-amanitin is that expected for human RNA polymerase III; and (2) depletion of transcription extracts with the same antibodies results in inhibition of transcription from an RNA polymerase III, but not from an RNA polymerase II, promoter. Sequence comparisons reveal that regions conserved in the RNA polymerase I, II, and III largest subunits characterized so far are also conserved in the human RNA polymerase III sequence, and thus probably perform similar functions for the human RNA polymerase III enzyme.

Amino Acid Sequence↗

Genome-wide analysis of the ERF gene family in Arabidopsis and rice.

Genes in the ERF family encode transcriptional regulators with a variety of functions involved in the developmental and physiological processes in plants. In this study, a comprehensive computational analysis identified 122 and 139 ERF family genes in Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa L. subsp. japonica), respectively. A complete overview of this gene family in Arabidopsis is presented, including the gene structures, phylogeny, chromosome locations, and conserved motifs. In addition, a comparative analysis between these genes in Arabidopsis and rice was performed. As a result of these analyses, the ERF families in Arabidopsis and rice were divided into 12 and 15 groups, respectively, and several of these groups were further divided into subgroups. Based on the observation that 11 of these groups were present in both Arabidopsis and rice, it was concluded that the major functional diversification within the ERF family predated the monocot/dicot divergence. In contrast, some groups/subgroups are species specific. We discuss the relationship between the structure and function of the ERF family proteins based on these results and published information. It was further concluded that the expansion of the ERF family in plants might have been due to chromosomal/segmental duplication and tandem duplication, as well as more ancient transposition and homing. These results will be useful for future functional analyses of the ERF family genes.

Amino Acid Motifs↗

Salmonella typhi contains identical intervening sequences in all seven rrl genes.

Salmonella typhi Ty2 rrl genes contain intervening sequences (IVSs) in helix-25 but not in helix-45 on the basis of observed 23S rRNA fragmentation caused by IVS excision. We have confirmed this and shown all seven IVSs to be identical by isolating genomic DNA fragments containing each of the seven rrl genes from S. typhi Ty2 by use of pulsed-field gel electrophoresis; each rrl gene was amplified by PCR in the helix-25 and helix-45 regions and cycle sequenced. Thirty independent wild-type S. typhi strains, tested by genomic PCR and DraI restriction, also have seven rrl genes with helix-25 IVSs and no helix-45 IVSs. We propose that IVS homogeneity in S. typhi occurs because gene conversion drives IVS sequence maintenance and because adaptation to human hosts results in limited clonal diversity.

Base Sequence↗

Absence of mutations in the interspecies conserved regions of the CFTR promoter region in cystic fibrosis (CF) and CF related patients.

This study was aimed at testing if a 5.2 kb untranslated region on both sides of the first CFTR exon, shown to contain regulatory elements, could carry mutations responsible for cystic fibrosis (CF) or CF related phenotypes. Selection of the DNA segments studied within this region was based upon the identification of conserved sequences throughout evolution (phylogenetic footprints, PFs). Comparison of the CFTR sequences in eight species representing four orders of mammals (man, gibbon, rhesus monkey, squirrel, monkey, rabbit, cow, rat, and mouse) identified four clusters of PFs within the 3.9 kb of DNA sequence upstream from the initiation codon, as well as two nearby PFs at +1 kb within intron 1. Six DNA segments containing PFs were scanned for mutations by denaturing gradient gel electrophoresis (DGGE) in patients with CF (n = 29), congenital bilateral absence of the vas deferens (n = 143), or disseminated bronchiectasis (n = 33), for whom only one or no mutations had been identified despite extensive DGGE analysis of the 27 CFTR exons and exon/intron boundaries. Only one polymorphism (-966 T-->G) was identified with a frequency of 2.2% and no other sequence variations were found. This study reinforces the idea that the promoter region in the CFTR is not frequently mutated.

Animals↗

Trend of amino acid composition of proteins of different taxa.

Archaea, bacteria and eukaryotes represent the main kingdoms of life. Is there any trend for amino acid compositions of proteins found in full genomes of species of different kingdoms? What is the percentage of totally unstructured proteins in various proteomes? We obtained amino acid frequencies for different taxa using 195 known proteomes and all annotated sequences from the Swiss-Prot data base. Investigation of the two data bases (proteomes and Swiss-Prot) shows that the amino acid compositions of proteins differ substantially for different kingdoms of life, and this difference is larger between different proteomes than between different kingdoms of life. Our data demonstrate that there is a surprisingly small selection for the amino acid composition of proteins for higher organisms (eukaryotes) and their viruses in comparison with the "random" frequency following from a uniform usage of codons of the universal genetic code. On the contrary, lower organisms (bacteria and especially archaea) demonstrate an enhanced selection of amino acids. Moreover, according to our estimates, 12%, 3% and 2% of the proteins in eukaryotic, bacterial and archaean proteomes are totally disordered, and long (> 41 residues) disordered segments are found to occur in 16% of arhaean, 20% of eubacterial and 43% of eukaryotic proteins for 19 archaean, 159 bacterial and 17 eukaryotic proteomes, respectively. A correlation between amino acid compositions of proteins of various taxa, show that the highest correlation is observed between eukaryotes and their viruses (the correlation coefficient is 0.98), and bacteria and their viruses (the correlation coefficient is 0.96), while correlation between eukaryotes and archaea is 0.85 only.

Amino Acid Sequence↗

Structural maintenance of chromosomes (SMC) proteins, a family of conserved ATPases.

SUMMARY: The structural maintenance of chromosomes (SMC) proteins are essential for successful chromosome transmission during replication and segregation of the genome in all organisms. SMCs are generally present as single proteins in bacteria, and as at least six distinct proteins in eukaryotes. The proteins range in size from approximately 110 to 170 kDa, and each has five distinct domains: amino- and carboxy-terminal globular domains, which contain sequences characteristic of ATPases, two coiled-coil regions separating the terminal domains and a central flexible hinge. SMC proteins function together with other proteins in a range of chromosomal transactions, including chromosome condensation, sister-chromatid cohesion, recombination, DNA repair and epigenetic silencing of gene expression. Recent studies are beginning to decipher molecular details of how these processes are carried out.

Adenosine Triphosphatases↗

Identification of a parathyroid hormone in the fish Fugu rubripes.

UNLABELLED: A PTH gene has been isolated from the fish Fugu rubripes. The encoded protein of 80 amino acid has the lowest homology with any of the PTH family members. Fugu PTH(1-34) had 5-fold lower potency than human PTH(1-34) in a mammalian cell system. INTRODUCTION: Parathyroid hormone (PTH) is the major hypercalcemic hormone in higher vertebrates. Fish lack parathyroid glands, but there have numerous attempts to identify and isolate PTH from fish. MATERIALS AND METHODS: Polymerase chain reaction (PCR) was performed with primers based on preliminary data from the Joint Genome Institute database. PCR amplification was performed on genomic DNA isolated from Fugu rubripes. PCR products were purified and DNA was sequenced. All sequence was confirmed from more than one independently amplified PCR product. Multiple sequence alignments were carried out, and the percentage of identities and similarities were calculated. An unrooted phylogenetic tree, using all the known PTH and PTH-related protein (PTHrP) amino acid sequences, was determined. Synthetic peptides were tested in a biological assay that measured cyclic adenosine 3',5'-monophosphate formation in UMR106.1 cells. Rabbit polyclonal antisera specific for N-terminal human PTHrP and one rabbit polyclonal antiserum specific for N terminus hPTH were used to test the cross-reactivity with fPTH(1-34) in immunoblots.

Amino Acid Sequence↗

[Application of gene sequence cluster in research for H3 antigenic evolution of influenza A virus].

OBJECTIVE: Gene sequence data were clustered to explore evolution lineages of H3 antigen of influenza A virus. METHODS: All data of H3 RNA sequence in NCBI Genbank and Influenza sequence database were downloaded and aligned in ClustalX while two step cluster method were applied to explore the data. RESULTS: All sequences were aggregated into ten clusters, while seven of them mainly were human virus. Human virus and avian/other mammal virus were separated into different clusters distinctively, but coexisted into same clusters with swine virus. Time and host distribution were very distinctive in these clusters, but no geographic distribution features were found. CONCLUSION: With the interaction of human immunity system, H3 antigen mutated significantly every 5 - 7 years, and the speed of mutation had accelerated with the application of influenza vaccines in recent years. Mean while, human and swine influenza virus were not separated distinctly between clusters indicating that they had short inheritance distance. Result showed again that swine served as the mixer for antigenic recombination of different influenza virus.

Antigenic Variation↗