PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Recco: recombination analysis using cost optimization.

MOTIVATION: Recombination plays an important role in the evolution of many pathogens, such as HIV or malaria. Despite substantial prior work, there is still a pressing need for efficient and effective methods of detecting recombination and analyzing recombinant sequences. RESULTS: We introduce Recco, a novel fast method that, given a multiple sequence alignment, scores the cost of obtaining one of the sequences from the others by mutation and recombination. The algorithm comes with an illustrative visualization tool for locating recombination breakpoints. We analyze the sequence alignment with respect to all choices of the parameter alpha weighting recombination cost against mutation cost. The analysis of the resulting cost curve yields additional information as to which sequence might be recombinant. On random genealogies Recco is comparable in its power of detecting recombination with the algorithm Geneconv (Sawyer, 1989). For specific relevant recombination scenarios Recco significantly outperforms Geneconv.

Algorithms↗

A globin gene of ancient evolutionary origin in lower vertebrates: evidence for two distinct globin families in animals.

Hemoglobin, myoglobin, neuroglobin, and cytoglobin are four types of vertebrate globins with distinct tissue distributions and functions. Here, we report the identification of a fifth and novel globin gene from fish and amphibians, which has apparently been lost in the evolution of higher vertebrates (Amniota). Because its function is presently unknown, we tentatively call it globin X (GbX). Globin X sequences were obtained from three fish species, the zebrafish Danio rerio, the goldfish Carassius auratus, and the pufferfish Tetraodon nigroviridis, and the clawed frog Silurana tropicalis. Globin X sequences are distinct from vertebrate hemoglobins, myoglobins, neuroglobins, and cytoglobins. Globin X displays the highest identity scores with neuroglobin (approximately 26% to 35%), although it is not a neuronal protein, as revealed by RT-PCR experiments on goldfish RNA from various tissues. The distal ligand-binding and the proximal heme-binding histidines (E7 and F8), as well as the conserved phenylalanine CD1 are present in the globin X sequences, but because of extensions at the N-terminal and C-terminal, the globin X proteins are longer than the typical eight alpha-helical globins and comprise about 200 amino acids. In addition to the conserved globin introns at helix positions B12.2 and G7.0, the globin X genes contain two introns in E10.2 and H10.0. The intron in E10.2 is shifted by 1 bp in respect to the vertebrate neuroglobin gene (E11.0), providing possible evidence for an intron sliding event. Phylogenetic analyses confirm an ancient evolutionary relationship of globin X with neuroglobin and suggest the existence of two distinct globin types in the last common ancestor of Protostomia and Deuterostomia.

Amino Acid Sequence↗

Evolution of virus-derived sequences for high-level replication of a subviral RNA.

Turnip crinkle virus (TCV) and its 356-nt satellite RNA satC share 151 nt of 3'-terminal sequence, which contain 8 positional differences and are predicted to fold into virtually identical structures, including a series of four phylogenetically inferred hairpins. SatC and TCV containing reciprocal exchanges of this region accumulate to only 15% or 1% of wild-type levels, respectively. Step-wise conversion of satC and TCV 3'-terminal sequences into the counterpart's sequence revealed the importance of having the cognate core promoter (Pr), which is composed of a single hairpin that differs in both sequence and stability, and an adjacent short 3'-terminal segment. The negative impact of the more stable TCV Pr on satC could not be attributed to lack of formation of a known tertiary interaction involving the 3'-terminal bases, nor an effect of coat protein, which binds specifically to TCV-like Pr and not the satC Pr. The satC Pr was a substantially better promoter than the TCV Pr when assayed in vitro using purified recombinant TCV RdRp, either in the context of satC or when assayed downstream of non-TCV-related sequence. Poor activity of the TCV Pr in vitro occurred despite solution structure probing indicating that its conformation in the context of satC is similar to the active form of the satC Pr, which is thought to form following a required conformational switch. These results suggest that evolution of satC following its initial formation generated a Pr that can function more efficiently in the absence of additional TCV sequence that may be required for full functionality of the TCV Pr.

Base Sequence↗

The complete structure of the rat VIP gene.

Vasoactive intestinal polypeptide (VIP) is a regulatory neuropeptide/neurotransmitter of 28 amino acids involved in a wide variety of physiological functions. Using synthetic oligodeoxynucleotide probes related to the rat VIP-cDNA, we have isolated and characterized the gene encoding the rat pre-pro VIP/PHI-27 and compared it to the human VIP gene. The rat VIP gene spanned 7400 base pairs, and contained 7 exons interrupted by 6 introns. 100% identity was found between the gene exons and the cDNA sequence. Differences in sizes of introns 2, 4 and 5 (shorter in the rat gene) are the reason for the shorter rat gene compared with the human gene of 8837 base pairs. Comparison of the genes in the two species showed a high homology in the exon sequences, 80-90% in exons 2, 4, 5, 6 and 30-50% in exons 1 and 7. In addition, the exon-intron junctions shared high identity between the genes. The rat untranslated exon 1 had little homology (30%) with human exon 1 and was 13 base pairs shorter. Interestingly, the 160 base pairs at the 5'-flanking region upstream of the cap-site share more than 75% identity between the two genes, including the exact position of TATA-boxes in positions -28, -145, -155, a cAMP-responsive element in position -80 and a CAAT sequence in position -127. The conservation of the 5'-flanking region of the VIP gene in parallel with the conservation of its coding exons emphasize the importance of these sequences during evolution.

Amino Acid Sequence↗

Further examples of evolution by gene duplication revealed through DNA sequence comparisons.

To test the theory that evolution by gene duplication occurs as a result of positive Darwinian selection that accompanies the acceleration of mutant substitutions, DNA sequences of recent duplication were analyzed by estimating the numbers of synonymous and nonsynonymous substitutions. For the troponin C family, at the period of differentiation of the fast and slow isoforms, amino acid substitutions were shown to have been accelerated relative to synonymous substitutions. Comparison of the first exon of alpha-actin genes revealed that amino acid substitutions were accelerated when the smooth muscle, skeletal and cardiac isoforms differentiated. Analysis of members of the heat shock protein 70 gene family of mammals indicates that heat shock responsive genes including duplicated copies are evolving rapidly, contrary to the cognitive genes which have been evolutionarily conservative. For the alpha 1-antitrypsin reactive center, the acceleration of amino acid substitution has been found for gene paris of recent duplication.

Actins↗

Evolution from primordial oligomeric repeats to modern coding sequences.

It seems as though nature was most innovative at the very beginning of life on this Earth a few billion years ago. For example, the functional competence of most, if not all, of the sugar-metabolizing enzymes was clearly established before the division of eukaryotes from prokaryotes eons ago, each critical active-site amino acid sequence being conserved ever since by bacteria as well as by mammals. I contend that this initial innovativeness was due to the first set of coding sequences being repeats of base oligomers, thus encoding polypeptide chains of various periodicities; such periodical polypeptide chains can easily acquire alpha-helical and beta-sheet-forming segments. In fact, the entire length of sugar-metabolizing enzymes is comprised of alternating alpha-helical and beta-sheet-forming segments. In the prebiotic (therefore nonenzymatic) replication of nucleic acids, what was in short supply was long templates, for there apparently was no inherent obstacle in copying of long templates, if such existed, in the presence of Zn2+. I submit that in this prebiotic condition, only those nucleotide oligomers that were internal doubles were automatically assured of progressive elongation to become long templates. For example, a decamer that was a pentameric repeat and its complementary sequence may pair unequally to initiate the next round of replication: first unit pairing with second, and a paired segment serving as a primer. As a consequence of this unequal pairing, decameric templates managed to become pentadecameric templates only after one round of replication, and this elongation process had no inherent limit.

Amino Acid Sequence↗

Evolution of c4 phosphoenolpyruvate carboxylase. Genes and proteins: a case study with the genus Flaveria.

C4 photosynthesis is characterized by a division of labour between two different photosynthetic cell types, mesophyll and bundle-sheath cells. Relying on phosphoenolpyruvate carboxylase (PEPC) as the primary carboxylase in the mesophyll cells a CO2 pump is established in C4 plants that concentrates CO2 at the site of ribulose 1,5-bisphosphate carboxylase/oxygenase in the bundle-sheath cells. The C4 photosynthetic pathway evolved polyphyletically implying that the genes encoding the C4 PEPC originated from non-photosynthetic PEPC progenitor genes that were already present in the C3 ancestral species. The dicot genus Flaveria (Asteraceae) is a unique system in which to investigate the molcular changes that had to occur in order to adapt a C3 ancestral PEPC gene to the special conditions of C4 photosynthesis. Flaveria contains not only C3 and C4 species but also a large number of C3-C4 intermediates which vary to the degree in which C4 photosynthetic traits are expressed. The C4 PEPC gene of Flaveria trinervia, which is encoded by the ppcA gene class, is highly expressed but only in mesophyll cells. The encoded PEPC protein possesses the typical kinetic and regulatory features of a C4-type PEPC. The orthologous ppcA gene of the C3 species Flaveria pringlei encodes a typical non-photosynthetic, C3-type PEPC and is weakly expressed with no apparent cell or organ specificity. PEPCs of the ppcA type have been detected also in C3-C4 intermediate Flaveria species. These orthologous PEPCs have been used to determine the molecular basis for C4 enzyme characteristics and to understand their evolution. Comparative and functional analyses of the ppcA promoters from F. trinervia and F. pringlei make it possible to identity the cis-regulatory sequences for mesophyll-specific gene expression and to search for the corresponding trans-regulatory factors.

Amino Acid Sequence↗

Structure and evolution of the Cinful retrotransposon family of maize.

A maize cDNA clone was isolated by virtue of its intense hybridization to total maize genomic DNA, indicating homology to highly repetitive sequences. Genomic homologues were identified and subcloned from an adh1-bearing maize yeast artificial chromosome (YAC). Sequencing revealed that the expressed sequence was part of a Ty3-gypsy-type retrotransposon. We discovered and sequenced two complete retrotransposons of this family, and named them Cinful elements because they are members of a family of maize retrotransposons including Zeon-1 and the first plant transposable element sequenced, the solo long terminal repeat (LTR) called Cin1. All are defective, as Cinful-1 and Cinful-2 elements lack gag and Zeon-1 lacks pol homology. Despite the apparent lack of an intact "autonomous" element, the Cinful family has expanded to a copy number of about 18 000, representing just under 9% of the maize genome. Both point mutations and major rearrangements, including possible gene acquisition, differentiate members of the Cinful family. Cinful family members were found to have an unusual feature that we also observed in two other Ty3-class retrotransposons of teosinte and tobacco: related tandem repeats that separate their internal domains with a gag- or pol-containing homology from a 3' segment of unknown function. The conserved and variable features identified provide insights into the origin, mutational history, and functional components of this major constituent of the maize genome.

Amino Acid Sequence↗

Evolution of human immunodeficiency virus type 1 nucleotide sequence diversity among close contacts.

The degree of change in the nucleotide sequence of human immunodeficiency virus type 1 (HIV-1) that occurs when it is transmitted sexually from one individual to another or vertically from mother to child is unknown. Previous studies have shown that most cultured HIV-1 isolates from the same individuals differed in the entire envelope gene nucleotide sequence by up to 2%, although most isolates from unrelated individuals differed by 6-22%. To examine diversity among HIV-1 isolates from close contacts, we determined the nucleotide sequences of viruses from a family with a known epidemiologic profile, in which a woman transmitted HIV-1 heterosexually to her partner and vertically to her daughter. Direct DNA sequence analysis of primary HIV-1 isolates amplified by PCR was used to distinguish the major and minor viral sequences, termed quasispecies, to rapidly determine the predominant sequences and their phylogenetic relationships. The nucleotide sequence diversity of a major portion of the HIV-1 envelope gene was 3.7% between isolates from the woman and her heterosexual partner and 8.5% between isolates from this woman and her daughter, who had been infected for a longer period than the partner. The configuration of the phylogenetic tree demonstrated that the daughter's predominant isolate evolved from a progenitor of her mother's current strain. This study provides evidence of a continuous spectrum of sequence diversity between any two isolates ranging from those derived from the same person to those from close contacts and, ultimately, those from unrelated individuals. These data and methods can be applied to epidemiologic investigations of possible HIV-1 transmission between health care workers and their patients.

Acquired Immunodeficiency Syndrome↗

Structure analysis of two Toxoplasma gondii and Neospora caninum satellite DNA families and evolution of their common monomeric sequence.

A family of repetitive DNA elements of approximately 350 bp-Sat350-that are members of Toxoplasma gondii satellite DNA was further analyzed. Sequence analysis identified at least three distinct repeat types within this family, called types A, B, and C. B repeats were divided into the subtypes B1 and B2. A search for internal repetitions within this family permitted the identification of conserved regions and the design of PCR primers that amplify almost all these repetitive elements. These primers amplified the expected 350-bp repeats and a novel 680-bp repetitive element (Sat680) related to this family. Two additional tandemly repeated high-order structures corresponding to this satellite DNA family were found by searching the Toxoplasma genome database with these sequences. These studies were confirmed by sequence analysis and identified: (1). an arrangement of AB1CB2 350-bp repeats and (2). an arrangement of two 350-bp-like repeats, resulting in a 680-bp monomer. Sequence comparison and phylogenetic analysis indicated that both high-order structures may have originated from the same ancestral 350-bp repeat. PCR amplification, sequence analysis and Southern blot showed that similar high-order structures were also found in the Toxoplasma-sister taxon Neospora caninum. The Toxoplasma genome database (http://ToxoDB.org ) permitted the assembly of a contig harboring Sat350 elements at one end and a long nonrepetitive DNA sequence flanking this satellite DNA. The region bordering the Sat350 repeats contained two differentially expressed sequence-related regions and interstitial telomeric sequences.

Animals↗

Comparative modeling of the three-dimensional structures of family 3 glycoside hydrolases.

There are approximately 100 known members of the family 3 group of glycoside hydrolases, most of which are classified as beta-glucosidases and originate from microorganisms. The only family 3 glycoside hydrolase for which a three-dimensional structure is available is a beta-glucan exohydrolase from barley. The structural coordinates of the barley enzyme is used here to model representatives from distinct phylogenetic clusters within the family. The majority of family 3 hydrolases have an NH(2)-terminal (alpha/beta)(8) barrel connected by a short linker to a second domain, which adopts an (alpha/beta)(6) sandwich fold. In two bacterial beta-glucosidases, the order of the domains is reversed. The catalytic nucleophile, equivalent to D285 of the barley beta-glucan exohydrolase, is absolutely conserved across the family. It is located on domain 1, in a shallow site pocket near the interface of the domains. The likely catalytic acid in the barley enzyme, E491, is on domain 2. Although similarly positioned acidic residues are present in closely related members of the family, the equivalent amino acid in more distantly related members is either too far from the active site or absent. In the latter cases, the role of catalytic acid is probably assumed by other acidic amino acids from domain 1.

Amino Acid Sequence↗

M13 endopeptidases: New conserved motifs correlated with structure, and simultaneous phylogenetic occurrence of PHEX and the bony fish.

M13 endopeptidase alignments have focused mainly on mammalian sequences and on the active site region defining the catalytic sequence signatures. Aligning all available M13 from bacteria to human on a full-length basis, we have performed a sequence analysis. This enabled us to highlight the origin and function of the M13 PHEX subtype family endopeptidase (phosphate regulating gene with homologies to endopeptidases on the X chromosome). New evolutionary conserved regions in both prokaryotes and eukaryotes have been detected and eukaryotic-specific regions clearly delineated. Using the recently solved neprilysin structure, we have observed that all new motifs, except one, localize in the spatial vicinity of the previously reported catalytic signatures. Interestingly, a highly hydrophobic pocket containing three newly reported motifs is centered by the C-terminal tryptophan residue. Extensive M13 searches in complete and in progress higher eukaryotic genomes have lead to the identification of Danio rerio as the simplest organism having PHEX. Finally, the human PHEX substrate, the parathyroid hormone-related peptide, PTHrP(107-139), is absent in bony fish: this suggests the existence of further PHEX substrates common to both bony fishes and higher vertebrates.

Amino Acid Motifs↗

Optimization of multiple-sequence alignment based on multiple-structure alignment.

Routinely used multiple-sequence alignment methods use only sequence information. Consequently, they may produce inaccurate alignments. Multiple-structure alignment methods, on the other hand, optimize structural alignment by ignoring sequence information. Here, we present an optimization method that unifies sequence and structure information. The alignment score is based on standard amino acid substitution probabilities combined with newly computed three-dimensional structure alignment probabilities. The advantage of our alignment scheme is in its ability to produce more accurate multiple alignments. We demonstrate the usefulness of the method in three applications: 1) computing more accurate multiple-sequence alignments, 2) analyzing protein conformational changes, and 3) computation of amino acid structure-sequence conservation with application to protein-protein docking prediction. The method is available at http://bioinfo3d.cs.tau.ac.il/staccato/.

Amino Acid Sequence↗

Modeling based on the structure of vicilins predicts a histidine cluster in the active site of oxalate oxidase.

It is known that germin, which is a marker of the onset of growth in germinating wheat, is an oxalate oxidase, and also that germins possess sequence similarity with legumin and vicilin seed storage proteins. These two pieces of information have been combined in order to generate a 3D model of germin based on the structure of vicilin and to examine the model with regard to a potential oxalate oxidase active site. A cluster of three histidine residues has been located within the conserved beta-barrel structure. While there is a relatively low level of overall sequence similarity between the model and the vicilin structures, the conservation of amino acids important in maintaining the scaffold of the beta-barrel lends confidence to the juxtaposition of the histidine residues. The cluster is similar structurally to those found in copper amine oxidase and other proteins, leading to the suggestion that it defines a metal-binding location within the oxalate oxidase active site. It is also proposed that the structural elements involved in intermolecular interactions in vicilins may play a role in oligomer formation in germin/oxalate oxidase.

Amino Acid Sequence↗

Cross-amplification and sequence variation of microsatellite loci in Eurasian hard pines.

Microsatellite transfer across coniferous species is a valued methodology because de novo development for each species is costly and there are many species with only a limited commodity value. Cross-species amplification of orthologous microsatellite regions provides valuable information on mutational and evolutionary processes affecting these loci. We tested 19 nuclear microsatellite markers from Pinus taeda L. (subsection Australes) and three from P. sylvestris L. (subsection Pinus) on seven Eurasian hard pine species ( P. uncinata Ram., P. sylvestris L., P. nigra Arn., P. pinaster Ait., P. halepensis Mill., P. pinea L. and P. canariensis Sm.). Transfer rates to species in subsection Pinus (36-59%) were slightly higher than those to subsections Pineae and Pinaster (32-45%). Half of the trans-specific microsatellites were found to be polymorphic over evolutionary times of approximately 100 million years (ten million generations). Sequencing of three trans-specific microsatellites showed conserved repeat and flanking regions. Both a decrease in the number of perfect repeats in the non-focal species and a polarity for mutation, the latter defined as a higher substitution rate in the flanking sequence regions close to the repeat motifs, were observed in the trans-specific microsatellites. The transfer of microsatellites among hard pine species proved to be useful for obtaining highly polymorphic markers in a wide range of species, thereby providing new tools for population and quantitative genetic studies.

Alleles↗

Purification and characterization of the methylene tetrahydromethanopterin dehydrogenase MtdB and the methylene tetrahydrofolate dehydrogenase FolD from Hyphomicrobium zavarzinii ZV580.

Recently, it has been shown that heterotrophic methylotrophic Proteobacteria contain tetrahydrofolate (H(4)F)- and tetrahydromethanopterin (H(4)MPT)-dependent enzymes. Here we report on the purification of two methylene tetrahydropterin dehydrogenases from the methylotroph Hyphomicrobium zavarzinii ZV580. Both dehydrogenases are composed of one type of subunit of 31 kDa. One of the dehydrogenases is NAD(P)-dependent and specific for methylene H(4)MPT (specific activity: 680 U/mg). Its N-terminal amino acid sequence showed sequence identity to NAD(P)-dependent methylene H(4)MPT dehydrogenase MtdB from Methylobacterium extorquens AM1. The second dehydrogenase is specific for NADP and methylene H(4)F (specific activity: 180 U/mg) and also exhibits methenyl H(4)F cyclohydrolase activity. Via N-terminal amino acid sequencing this dehydrogenase was identified as belonging to the classical bifunctional methylene H(4)F dehydrogenases/cyclohydrolases (FolD) found in many bacteria and eukarya. Apparently, the occurrence of methylene tetrahydrofolate and methylene tetrahydromethanopterin dehydrogenases is not uniform among different methylotrophic alpha-Proteobacteria. For example, FolD was not found in M. extorquens AM1, and the NADP-dependent methylene H(4)MPT dehydrogenase MtdA was present in the bacterium that also shows H(4)F activity.

Amino Acid Sequence↗

Characterization of the T-cell receptor gamma locus and analysis of the variable gene segment expression in rabbit.

The genomic organization and expression of genes of the T-cell receptor gamma (TRG) locus are described for mice and humans, but not for species such as rabbits (Oryctolagus cuniculus), in which gammadelta T cells compose a sizeable proportion of T cells in the periphery. We cloned 200 kb of the rabbit TRG locus and determined the TRGV gene usage in adult and newborn rabbits by RT-PCR. We identified two TRGJ genes, one TRGC gene, and 22 TRGV genes, all of which encoded functional variable regions. One TRGV gene is the unique member of the TRGV2 subgroup, whereas the other genes belong to the TRGV1 subgroup. Evolutionary analyses of TRGV1 genes identified three distinct groups that can be explained by separate duplication events in the rabbit genome. Evidence of gene conversion between TRGV1.1 and TRGV1.6 was observed. Both TRGV1 and TRGV2 subgroup genes were expressed in the spleen, intestine, and appendix of adult rabbits, and the repertoire of TRGV genes expressed in these tissues was similar. In these tissues from newborns, and in skin from adults, only the genes from the TRGV1 subgroup were expressed. Greater TRGV-J junctional diversity was found in tissues from adult compared to newborn rabbits. Our analyses indicate rabbits have a larger germ line encoded TRG repertoire compared with that of mice and humans. In addition, we found TRGV gene usage is alike in most tissues of rabbits similar to that found in humans but in contrast to that found in mice.

Age Factors↗