PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

The human melanocortin-1 receptor locus: analysis of transcription unit, locus polymorphism and haplotype evolution.

The complete sequence of the MC1R locus has been assembled, the coding region of the gene is intronless and placed within a 12 kb region flanked by the NULP1 and TUBB4 genes. The immediate promoter region has an E-box site with homology to the M-box consensus known to bind the microphthalmia transcription factor (MITF); however, promoter deletion analysis and transactivation studies have failed to show activation through this element by MITF. Polymorphism within the coding region, immediate 5' promoter region and a variable number tandem repeat (VNTR) minisatellite within the locus have been examined in a collection of Caucasian families and African individuals. Haplotype analysis shows linkage disequilibrium between the VNTR and MC1R coding region red hair variant alleles which can be used to estimate the age of these missense changes. Assuming a mean VNTR mutation rate of 1% and a star phylogeny, we estimate the Arg151Cys variant arose 7500 years before the present day, suggesting these variants may have arisen in the Caucasian population more recently than previously thought.

Amino Acid Sequence↗

Glutathione transferase-like proteins encoded in genomes of yeasts and fungi: insights into evolution of a multifunctional protein superfamily.

Most fungal glutathione transferases (GSTs) do not fit easily into any of the previously characterised classes by immunological, sequence or catalytic criteria. In contrast to the paucity of studies on GSTs cloned or isolated from fungal sources, a screen of databases revealed 67 GST-like sequences from 21 fungal species. Comparison by multiple sequence alignment generated a dendrogram revealing five clusters of GST-like proteins designated clusters 1, 2, EFIBgamma, Ure2p and MAK16, the last three of which have previously been related to the GST superfamily. Surprisingly, a relatively small number of fungal GSTs belong to mainstream classes and the previously-described fungal Gamma class is not widespread in the 21 species studied. Representative crystal structures are available for the EFIBgamma and Ure2p classes and the domain structures of representative sequences are compared with these. In addition, there are some "orphan" sequences that do not fit into any previously-described class, but show similarity to genes implicated in fungal biosynthetic gene clusters. We suggest that GST-like sequences are widespread in fungi, participating in a wide range of functions. They probably evolved by a process similar to domain "shuffling".

Amino Acid Sequence↗

Evolutionally conserved intermediates between ubiquitin and NEDD8.

The investigation of common structural motifs provides additional information on why proteins conserve similar topologies yet may have non-conserved amino acid sequences. Proteins containing the ubiquitin superfold have similar topologies, although the sequence conservation is rather poor. Here, we present novel similarities and differences between the proteins ubiquitin and NEDD8. They have 57% identical sequence, almost identical backbone topology and similar functional strategy, although their physiological functions are mutually different. Using variable pressure NMR spectroscopy, we found that the two proteins have similar conformational fluctuation in the evolutionary conserved enzyme-binding region and contain a structurally similar locally disordered conformer (I) in equilibrium with the basic folded conformer (N). A notable difference between the two proteins is that the equilibrium population of I is far greater for NEDD8 (DeltaG(0)(NI)<5 kJ/mol) than for ubiquitin (DeltaG(0)(NI)=15.2(+/-1.0) kJ/mol), and that the tendency for overall unfolding (U) is also far higher for NEDD8 (DeltaG(0)(NU)=11.0(+/-1.5) kJ/mol) than for ubiquitin (DeltaG(0)(NU)=31.3(+/-4.7) kJ/mol). These results suggest that the marked differences in thermodynamic stabilities of the locally disordered conformer (I) and the overall unfolding species (U) are a key to determine the functional differences of the two structurally similar proteins in physiology.

Amino Acid Sequence↗

Tracing the evolution of H-2 D region genes using sequences associated with a repetitive element.

The class I genes in the murine MHC are genetically divided into the K, D, Qa, and T1a region subfamilies. These genes presumably arose by duplication from a common class I ancestor. Oligonucleotide probes specific for sequences associated with a moderately repetitive B2 SINE element, which is inserted into the 3' untranslated region of the H-2D and H-2L genes, were used to examine the evolutionary relationship between these classically defined D region genes (H-2D and H-2L) and the other members of the class I gene family. Hybridization analyses of recombinant cosmid and genomic DNA indicated that the D region genes separated genetically from the other members of the class I gene family 12 to 14 million years ago. The evidence suggests that during this time frame the chromosomal segment harboring the characteristic insertion became fixed in the ancestral population which gave rise to Mus domesticus. Previous studies have shown that the number of genes present in the Qa and T1a regions varies among inbred strains and among laboratory stocks of wild mice derived from more distant species on the genus Mus. No evidence was found in this study to support the hypothesis that variation in class I gene number is the result of recent duplications of the functionally defined class I genes of the D region, H-2D and H-2L.

Animals↗

Evolution of tRNA-like sequences and genome variability.

Transfer RNA (tRNA)-like sequences were searched for in the nine basic taxonomic divisions of GenBank-121 (viruses, phages, bacteria, plants, invertebrates, vertebrates, rodents, mammals, and primates) by an original program package implementing a dynamic profile alignment approach for the genetic texts' analysis, in using 22 profiles of tRNAs of different isotypes. In total, 175,901 previously unknown tRNA-like sequences were revealed. The locations of the tRNA-likes were considered over the regions whose functional meaning is described by standard Feature Keys in GenBank. Many regions containing the tRNA-like sequences were recognized as known repeats. A mode of distribution of the tRNA-like sequences in a genome was proposed as expansion in a content of the various transposable elements. An analysis of the integrity of RNA polymerase III inner promoters in the tRNA-like sequences over the GenBank divisions has shown a high possibility of generating new copies of short interspersed nuclear element (SINE) repeats in all divisions, excepting primates. The numerous tRNA-likes found in the regions of RNA polymerase II promoters have suggested an adaptation of RNA polymerase III promoter to a binding of RNA polymerase II.

Algorithms↗

Evolution of transcribed and spacer sequences in the ribosomal RNA genes of Drosophila.

Examination of the ribosomal RNA (rRNA) gene of six sibling species that make up the D. melanogaster subgroup reveals that the nontranscribed spacer is highly conserved during evolution. Indeed, the spacer is at least as conserved as the transcribed rRNA sequence in four of the six species and only slightly less conserved in the others. These data support the hypothesis previously suggested (Tartof and Dawid, 1976) that selection has a significant role in maintaining the parallel evolution of genetically separate but homologous redundant gene clusters.

Animals↗

Evolution of plant microRNA gene families.

MicroRNAs (miRNAs) are important post-transcriptional regulators of their target genes in plants and animals. miRNAs are usually 20-24 nucleotides long. Despite their unusually small sizes, the evolutionary history of miRNA gene families seems to be similar to their protein-coding counterparts. In contrast to the small but abundant miRNA families in the animal genomes, plants have fewer but larger miRNA gene families. Members of plant miRNA gene families are often highly similar, suggesting recent expansion via tandem gene duplication and segmental duplication events. Although many miRNA genes are conserved across plant species, the same gene family varies significantly in size and genomic organization in different species, which may cause dosage effects and spatial and temporal differences in target gene regulations. In this review, we summarize the current progress in understanding the evolution of plant miRNA gene families.

Animals↗

Aperiodic quantum random walks.

We generalize the quantum random walk protocol for a particle in a one-dimensional chain, by using several types of biased quantum coins, arranged in aperiodic sequences, in a manner that leads to a rich variety of possible wave-function evolutions. Quasiperiodic sequences, following the Fibonacci prescription, are of particular interest, leading to a sub-ballistic wave-function spreading. In contrast, random sequences lead to diffusive spreading, similar to the classical random walk behavior. We also describe how to experimentally implement these aperiodic sequences.

Journal Article↗

Comparative study of the ribosomal RNA from Leishmania and Trypanosoma.

The ribosomal RNA from several stocks of the genera Leishmania and Trypanosoma were studied by gel electrophoresis, sedimentation on sucrose density gradients and RNA/DNA hybridization experiments. Three major components were observed after electrophoresis in polyacrylamide gels (PAGE-SDS), the relative molecular masses being respectively: X1 = 0.83 megadaltons, X2 = 0.63 megadaltons and X3 = 0.54 megadaltons for Leishmania RNA; and X1 = 0.86 megaldaltons, X2 = 0.78 megadaltons, and X3 = 0.58 megadaltons for Trypanosoma RNA. Depending upon the isolation procedure, a fourth component, X0 = 1.2 megadaltons (26S), became evident. The later component was purified from Leishmania brasiliensis (Y) by centrifugation on a linear 15-30% sucrose density gradient. This component, after heat denaturation and PAGE-SDS, gave rise to two bands coinciding in molecular mass with those of X2 and X3, indicating that these components are part of the large ribosomal subunit whereas X1 belongs to the small one. The above mentioned differences in mobilities of components X1 and X2 between the two genera were no longer observed after electrophoresis in denaturing agarose-formaldehyde gels, suggesting secondary structural differences among these RNA species. Hybridization experiments with L. brasiliensis (Y) DNA showed that both RNA types compete equally well for the ribosomal sites in this DNA, and that L. brasiliensis (Y) rRNA recognizes the ribosomal sites in DNA of Trypanosoma cruzi (EP), thus indicating that no gross changes occurred in their nucleotide sequences during evolution.

Animals↗

A 7-bp insertion in the 3' untranslated region suggests the duplication and concerted evolution of the rabbit SRY gene.

In this work we report the genetic polymorphism of a 7-bp insertion in the 3' untranslated region of the rabbit SRY gene. The polymorphic GAATTAA motif was found exclusively in one of the two divergent rabbit Y-chromosomal lineages, suggesting that its origin is more recent than the separation of the O. c. algirus and O. c. cuniculus Y-chromosomes. In addition, the remarkable observation of haplotypes exhibiting 0, 1 and 2 7-bp inserts in essentially all algirus populations suggests that the rabbit SRY gene is duplicated and evolving under concerted evolution.

3' Untranslated Regions↗

The vascular endothelial growth factor (VEGF) family: angiogenic factors in health and disease.

Vascular endothelial growth factors (VEGFs) are a family of secreted polypeptides with a highly conserved receptor-binding cystine-knot structure similar to that of the platelet-derived growth factors. VEGF-A, the founding member of the family, is highly conserved between animals as evolutionarily distant as fish and mammals. In vertebrates, VEGFs act through a family of cognate receptor kinases in endothelial cells to stimulate blood-vessel formation. VEGF-A has important roles in mammalian vascular development and in diseases involving abnormal growth of blood vessels; other VEGFs are also involved in the development of lymphatic vessels and disease-related angiogenesis. Invertebrate homologs of VEGFs and VEGF receptors have been identified in fly, nematode and jellyfish, where they function in developmental cell migration and neurogenesis. The existence of VEGF-like molecules and their receptors in simple invertebrates without a vascular system indicates that this family of growth factors emerged at a very early stage in the evolution of multicellular organisms to mediate primordial developmental functions.

Alternative Splicing↗

The evolution of proteins from random amino acid sequences: II. Evidence from the statistical distributions of the lengths of modern protein sequences.

This paper continues an examination of the hypothesis that modern proteins evolved from random heteropeptide sequences. In support of the hypothesis, White and Jacobs (1993, J Mol Evol 36:79-95) have shown that any sequence chosen randomly from a large collection of nonhomologous proteins has a 90% or better chance of having a lengthwise distribution of amino acids that is indistinguishable from the random expectation regardless of amino acid type. The goal of the present study was to investigate the possibility that the random-origin hypothesis could explain the lengths of modern protein sequences without invoking specific mechanisms such as gene duplication or exon splicing. The sets of sequences examined were taken from the 1989 PIR database and consisted of 1,792 "super-family" proteins selected to have little sequence identity, 623 E. coli sequences, and 398 human sequences. The length distributions of the proteins could be described with high significance by either of two closely related probability density functions: The gamma distribution with parameter 2 or the distribution for the sum of two exponential random independent variables. A simple theory for the distributions was developed which assumes that (1) protoprotein sequences had exponentially distributed random independent lengths, (2) the length dependence of protein stability determined which of these protoproteins could fold into compact primitive proteins and thereby attain the potential for biochemical activity, (3) the useful protein sequences were preserved by the primitive genome, and (4) the resulting distribution of sequence lengths is reflected by modern proteins. The theory successfully predicts the two observed distributions which can be distinguished by the functional form of the dependence of protein stability on length. The theory leads to three interesting conclusions. First, it predicts that a tetra-nucleotide was the signal for primitive translation termination. This prediction is entirely consistent with the observations of Brown et al. (1990a,b, Nucleic Acids Res 18:2079-2086 and 18: 6339-6345) which show that tetra-nucleotides (stop codon plus following nucleotide) are the actual signals for termination of translation in both prokaryotes and eukaryotes. Second, the strong dependence of statistical length distributions on sequence-termination signaling codes implies that the evolution of stop codons and translation-termination processes was as important as gene splicing in early evolution. Third, because the theory is based upon a simple no-exon stochastic model, it provides a plausible alternative to a limited universe of exons from which all proteins evolved by gene duplication and exon splicing (Dorit et al. 1990, Science 250:1377-1382).

Amino Acid Sequence↗

Genomic organization and evolution of the soybean SB92 satellite sequence.

Repetitive DNA sequences comprise a large percentage of plant genomes, and their characterization provides information about both species and genome evolution. We have isolated a recombinant clone containing a highly repeated DNA element (SB92) that is homologous to ca. 0.9% of the soybean genome or about 10(5) copies. This repeated sequence is tandemly arranged and is found in four or five major genomic locations. FISH analysis of metaphase chromosomes suggests that two of these locations are centromeric. We have determined the sequence of two cloned repeats and performed genomic sequencing to obtain a consensus sequence. The consensus repeat size was 92 bp and exhibited an average of 10% nucleotide substitution relative to the two cloned repeats. This high level of sequence diversity suggests an ancient origin but is inconsistent with the limited phylogenetic distribution of SB92, which is found at high copy number only in the annual soybeans. It therefore seems likely that this sequence is undergoing very rapid evolution.

Base Sequence↗

The molecular epidemiology and evolution of Epstein-Barr virus: sequence variation and genetic recombination in the latent membrane protein-1 gene.

The phylogeny and evolution of Epstein-Barr virus (EBV) genetic variation are poorly understood. EBV latent membrane protein-1 (LMP-1) gene sequences are especially heterogeneous and may be useful as a tool for EBV genotype identification. Therefore, LMP-1 sequences obtained directly from EBV-infected human tissues were examined by PCR amplification and cloning. EBV genotypes were defined as "strains" from among 22 identified LMP-1 sequence patterns. Three molecular mechanisms were identified by which genetic diversity arises in the LMP-1 gene: point mutation, sequence deletion or duplication, and homologous recombination. The rate of LMP-1 gene evolution was found to be accelerated by coinfection with multiple EBV strains. The results of this study refine our understanding of LMP-1 sequence variation and enable accurate discrimination between independent EBV infection events and the consequence of intrahost EBV evolution. Thus, this LMP-1 sequence-based approach to EBV molecular epidemiology will facilitate the study of intrahost EBV infection, coinfection, and persistence.

Amino Acid Sequence↗

Characterization of candidate class A, B and E floral homeotic genes from the perianthless basal angiosperm Chloranthus spicatus (Chloranthaceae).

The classic ABC model explains the activities of each class of floral homeotic genes in specifying the identity of floral organs. Thus, changes in these genes may underlay the origin of floral diversity during evolution. In this study, three MADS-box genes were isolated from the perianthless basal angiosperm Chloranthus spicatus. Sequence and phylogenetic analyses revealed that they are AP1-like, AP3-like and SEP3-like genes, and hence these genes were termed CsAP1, CsAP3 and CsSEP3, respectively. Due to these assignments, they represent candidate class A, class B and class E genes, respectively. Expression patterns suggest that the CsAP1, CsAP3 and CsSEP3 genes function during flower development of C. spicatus. CsAP1 is expressed broadly in the flower, which may reflect the ancestral function of SQUA-like genes in the specification of inflorescence and floral meristems rather than in patterning of the flower. CsAP3 is exclusively expressed in male floral organs, providing the evidence that AP3-like genes have ancestral function in differentiation between male and female reproductive organs. CsSEP3 expression is not detectable in spike meristems, but its mRNA accumulates throughout the flower, supporting the view that SEP-like genes have conserved expression pattern and function throughout angiosperm. Studies of synonymous vs nonsynonymous nucleotide substitutions indicate that these genes have not evolved under changes in evolutionary forces. All the data above suggest that the genes may have maintained at least some ancestral functions despite the lack of perianth in the flowers of C. spicatus.

Amino Acid Sequence↗

Inferring protein-protein interacting sites using residue conservation and evolutionary information.

This paper proposes a novel method using protein residue conservation and evolution information, i.e., spatial sequence profile, sequence information entropy and evolution rate, to infer protein binding sites. Some predictors based on support vector machines (SVMs) algorithm are constructed to predict the role of surface residues in protein-protein interface. By combining protein residue characters, the prediction performance can be improved obviously. We then made use of the predicted labels of neighbor residues to improve the performance of the predictors. The efficiency and the effectiveness of our proposed approach are verified by its better prediction performance based on a non-redundant data set of heterodimers.

Amino Acid Sequence↗

Fish possess multiple copies of fgfrl1, the gene for a novel FGF receptor.

FGFRL1 is a novel FGF receptor that lacks the intracellular tyrosine kinase domain. While mammals, including man and mouse, possess a single copy of the FGFRL1 gene, fish have at least two copies, fgfrl1a and fgfrl1b. In zebrafish, both genes are located on chromosome 14, separated by about 10 cM. The two genes show a similar expression pattern in several zebrafish tissues, although the expression of fgfrl1b appears to be weaker than that of fgfrl1a. A clear difference is observed in the ovary of Fugu rubripes, which expresses fgfrl1a but not fgfrl1b. It is therefore possible that subfunctionalization has played a role in maintaining the two fgfrl1 genes during the evolution of fish. In human beings, the FGFRL1 gene is located on chromosome 4, adjacent to the SPON2, CTBP1 and MEAEA genes. These genes are also found adjacent to the fgfrl1a gene of Fugu, suggesting that FGFRL1, SPON2, CTBP1 and MEAEA were preserved as a coherent block during the evolution of Fugu and man.

Amino Acid Sequence↗