PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

The whole structure of the human nonfunctional L-gulono-gamma-lactone oxidase gene--the gene responsible for scurvy--and the evolution of repetitive sequences thereon.

L-Gulono-gamma-lactone oxidase (GULO), which catalyzes the last step of ascorbic acid biosynthesis, is missing in humans. The whole structure of the human gene homologue for this enzyme was disclosed by a computer-assisted search. Only five exons, as compared to 12 exons constituting the functional rat GULO gene, remain in the human genome. A comparison of these exons with those of their functional counterparts in rat showed that there are two single nucleotide deletions, one triple nucleotide deletion, and one single nucleotide insertion in the human sequence. When compared in terms of codons, the human sequence has a deletion of a single amino acid, two stop codons, and two aberrant codons missing one nucleotide besides many amino acid substitutions. A comparison of the remaining human exon sequences with the corresponding sequences of the guinea pig nonfunctional GULO gene revealed that the same substitutions from rats to both species occurred at a large number of nucleotide positions. From analyses of the molecular evolution of Alu sequences in the human GULO gene homologue, it is thought that two Alu sequences were inserted in the vicinity of a presumed position of lost exon 11 during the same period as GULO lost its function. It is predicted that six LINE-1 sequences located in and near the gene homologue were inserted not during that period.

Alu Elements↗

Male-driven evolution of DNA sequences.

It is commonly believed that the mutation rate is much higher in the human male germ line than in the female germ line because the number of germ-cell divisions per generation is much larger in males than in females. But direct estimation of mutation rates is difficult, relying mainly on sex-linked genetic diseases, so the ratio (alpha m) of male to female mutation rates is not clear. It has been noted that if alpha m is very large, then the rate of synonymous substitution in X-linked genes should be only 2/3 of that in autosomal genes, and comparison of human and rodent genes supported this prediction. As the number of X-linked genes used in the study was small and the X-linked and autosomal sequences were non-homologous, and given that the synonymous rate varies among genes, we sequenced the last intron (approximately 1 kb) of the Y-linked and X-linked zinc-finger-protein genes (ZFY and ZFX) in humans, orang-utans, baboons and squirrel monkeys. The ratio Y/X of the substitution rate in the Y-linked intron to that in the X-linked intron is approximately 2.3, which is close to that estimated from synonymous rates in the ZFY and ZFX genes and implies alpha m approximately 6. This estimate of alpha m supports the view that the evolution of DNA sequences in higher primates is male-driven. It is, however, much lower than the previous estimate and therefore raises a number of issues.

Animals↗

Lipases and esterases: a review of their sequences, structure and evolution.

This chapter aims to provide a brief review on the enzyme family of lipases and esterases. The sequences, 3D structures and pH dependent electrostatic signatures are presented and analyzed. Since the family comprises more than 100 sequences, we have tried to focus on the most interesting features from our perspective, which translates into finding similarities and differences between members of this family, in particular in and around the active sites, and to identify residues that are partially or totally conserved. Such residues we believe are either important for maintaining the structural scaf-fold of the protein or to maintain activity or specificity. The structure function relationship for these proteins is therefore of central interest. Can we uniquely identify a protein from this large family of sequences--and if so, what is the identifier? The protein family displays some highly complex features: many of the proteins are interfacially activated, i.e. they need to be in physical contact with the aggregated substrate. Access to the active site is blocked with either a loop fragment or an alpha-helical fragment in the absence of interfacial contact. Although the number of known, relevant protein 3D structures is growing steadily, we are nevertheless faced with a virtual explosion in the number of known or deduced amino acid sequences. It is therefore unrealistic to expect that all protein sequences within the foreseeable future will have their 3D structure determined by X-ray diffractional analysis or through other methods. When feasible the gene and/or the amino acid sequences will be analyzed from an evolutionary perspective. As the 3D folds are often remarkably similar, both among the triglyceride lipases as well as among the esterases, the functional diversities (e.g. specificity) must originate in differences in surface residue utilization, in particular of charged residues. The pH variations in the isopotential surfaces of some of the most interesting lipases are presented and a qualitative interpretation proposed. Finally we illustrate that NMR has potential for becoming an important tool in the study of lipases, esterases and their kinetics.

Amino Acid Sequence↗

Conserved sequences and the evolution of gene regulatory signals.

Studies of evolutionary conservation of gene regulatory signals have led to a paradox: extensive sequence similarity implies functional conservation in non-coding regions across mammalian species; however, this stands in contrast to our understanding of transcriptional regulatory sites composed of degenerate recognition sequences for transcription factors that can maintain functional equivalence despite considerable sequence divergence. The latter observation provides an explanation for the rapid evolution of new traits through the gain and loss of transcription factor binding sites that bring new genes under the control of an existing genetic regulatory network. The former observation might point to novel mechanisms of gene regulation and/or chromosome function that are currently unappreciated. Recent comparative genome analysis has highlighted extensive conserved sequences in mammalian genomes that are beginning to be functionally characterized.

Animals↗

Evolution of DNA sequence organization in mitochondrial genomes of Zea.

Molecular mechanisms underlying the evolution of mtDNA in maize and three species of teosinte have been investigated. By using DNA transfer techniques and cloned fragments of maize mtDNA, changes in the position of homologous sequences in restriction digests were analyzed in related taxa. The resulting patterns indicate general conservation of sequence homology in all four species. One-third of the cloned fragments showed sequences conserved in homology and position of BamHI restriction fragments. Other fragments produced patterns indicating that extensive rearrangement of DNA sequences has taken place. In two cases, hybridization patterns revealed major changes in relative abundance of specific sequences that we believe are related to the molecular heterogeneity phenomenon of maize mtDNA. The evolution of mtDNA in these plants appears dramatically different from that in animal mtDNA, where base substitutions may account for most of the observed changes in restriction patterns.

Journal Article↗

The boundary of macaque rDNA is constituted by low-copy sequences conserved during evolution.

In Macaca mulatta, the single rDNA array is flanked by a patchwork of sequences including subregions of human Yp11.2, 4q35.2, and 10p15.3. This composite DNA region is characterized by unique or low-copy sequences, resembling a potentially transcribed region. The analysis of Cercopithecus aethiops, Presbytis cristata, and Hylobates lar suggests that this complex sequence organization could be shared by Old World monkey and lesser ape species. After the lesser apes/great apes divergence, the unique or nonduplicated DNA region underwent amplification and spreading, preferentially marking the p arm of acrocentric chromosomes bearing the rDNA. The molecular analysis of human acrocentric chromosomes revealed some extent of remodeling of the rDNA boundary: near the human NOR, a large 4q35.2 duplication partially resembles that found in MMU; conversely, infrequently represented Yp11.2 sequences totally differed from those of the macaque, and 10p15.3 sequences were lacking. Thus, although evolutionary events modified the sequence organization of the MMU rDNA boundary, its overall sequence feature and the preferential location in vicinity to the NOR have been conserved.

Animals↗

Innovation from reduction: gene loss, domain loss and sequence divergence in genome evolution.

Analyses of genome sequences have revealed a surprisingly variable distribution of genes, reflecting the generation of novel genes, lateral gene transfer and gene loss. The impact of gene loss on organisms has been difficult to examine, but the loss of protein coding genes, the loss of domains within proteins and the divergence of genes have made surprising contributions to the differences among organisms. This paper reviews surveys of gene loss and divergence in fungal and archaeal genomes that indicate suites of functionally related genes tend to undergo loss and divergence. Instances of fungal gene loss highlighted here suggest that specific cellular systems have changed, such as Ca 2+ biology in Saccharomyces cerevisiae and peroxisome function in Schizosaccharomyces pombe. Analyses of loss and divergence can provide specific predictions regarding protein-protein interactions, and the relationship between networks of protein interactions and loss may form a part of a parametric model of genome evolution.

Chromosome Mapping↗

Molecular evolution and multilocus sequence typing of 145 strains of SARS-CoV.

In this study, we have identified 876 polymorphism sites in 145 complete or partial genomes of SARS-CoV available in the NCBI GenBank. One hundred and seventy-four of these sites existed in two or more SARS-CoV genome sequences. According to the sequence polymorphism, all SARS-CoVs can be divided into three groups: (I) group 1, animal-origin viruses (such as SARS-CoV SZ1, SZ3, SZ13 and SZ16); (II) group 2, all viruses with clinical origin during first epidemic; and (III) group 3, SARS-CoV GD03T0013. According to 10 special loci, group 2 again can be divided into genotypes C and T, which can be further divided into sub-genotypes C1-C4 and T1-T4. Positive Darwinian selections were identified between any pair of these three groups. Genotype C gives neutral selection. Genotype T, however, shows negative selection. By comparing the death rates of SARS patients in the different regions, it was found that the death rate caused by the viruses of the genotype C was lower than that of the genotype T. SARS-CoVs might originate from an unknown ancestor.

Base Sequence↗

Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.

Recognizing structural similarity without significant sequence identity has proved to be a challenging task. Sequence-based and structure-based methods as well as their combinations have been developed. Here, we propose a fold-recognition method that incorporates structural information without the need of sequence-to-structure threading. This is accomplished by generating sequence profiles from protein structural fragments. The structure-derived sequence profiles allow a simple integration with evolution-derived sequence profiles and secondary-structural information for an optimized alignment by efficient dynamic programming. The resulting method (called SP(3)) is found to make a statistically significant improvement in both sensitivity of fold recognition and accuracy of alignment over the method based on evolution-derived sequence profiles alone (SP) and the method based on evolution-derived sequence profile and secondary structure profile (SP(2)). SP(3) was tested in SALIGN benchmark for alignment accuracy and Lindahl, PROSPECTOR 3.0, and LiveBench 8.0 benchmarks for remote-homology detection and model accuracy. SP(3) is found to be the most sensitive and accurate single-method server in all benchmarks tested where other methods are available for comparison (although its results are statistically indistinguishable from the next best in some cases and the comparison is subjected to the limitation of time-dependent sequence and/or structural library used by different methods.). In LiveBench 8.0, its accuracy rivals some of the consensus methods such as ShotGun-INBGU, Pmodeller3, Pcons4, and ROBETTA. SP(3) fold-recognition server is available on http://theory.med.buffalo.edu.

Algorithms↗

Sequence, organization, and evolution of the A+T region of Drosophila melanogaster mitochondrial DNA.

The long (4.6-kb) A+T region of Drosophila melanogaster mitochondrial DNA has been cloned and sequenced. The A+T region is organized in two large arrays of tandemly repeated DNA sequence elements, with nonrepetitive intervening and flanking sequences comprising only 22% of its length. The first repeat array consists of five repeats of 338-373 bp. The second consists of four intact 464-bp repeats and a fifth partial repeat of 137 bp. Three DNA sequence elements are found to be highly conserved in D. melanogaster and in several Drosophila species with short A+T regions. These include a 300-bp DNA sequence element that overlaps the DNA replication origin and two thymidylate stretches identified on opposite DNA strands. We conclude that the length heterogeneity observed in the A+T regulatory region in mitochondrial DNAs from the genus Drosophila results from the expansion (and contraction) of the number of repeated DNA sequence elements. We also propose that the 300-bp conserved DNA sequence element, in conjunction with another primary sequence determinant, perhaps the adjacent thymidylate stretch, functions in the regulation of mitochondrial DNA replication.

Animals↗

Evolution of envelope sequences from the genital tract and peripheral blood of women infected with clade A human immunodeficiency virus type 1.

The development of viral diversity during the course of human immunodeficiency virus type 1 (HIV-1) infection may significantly influence viral pathogenesis. The paradigm for HIV-1 evolution is based primarily on studies of male cohorts in which individuals were presumably infected with a single virus variant of subtype B HIV-1. In this study, we evaluated virus evolution based on sequence information of the V1, V2, and V3 portions of HIV-1 clade A envelope genes obtained from peripheral blood and cervical secretions of three women with genetically heterogeneous viral populations near seroconversion. At the first sample following seroconversion, the number of nonsynonymous substitutions per potential nonsynonymous site (dn) significantly exceeded substitutions at potential synonymous sites (ds) in plasma viral sequences from all individuals. Generally, values of dn remained higher than values of ds as sequences from blood or mucosa evolved. Mutations affected each of the three variable regions of the envelope gene differently; insertions and deletions dominated changes in V1, substitutions involving charged amino acids occurred in V2, and sequential replacement of amino acids over time at a small subset of positions distinguished V3. The relationship among envelope nucleotide sequences obtained from peripheral blood mononuclear cells, plasma, and cervical secretions was evaluated for each individual by both phylogenetic and phenetic analyses. In all subjects, sequences from within each tissue compartment were more closely related to each other than to sequences from other tissues (phylogenetic tissue compartmentalization). At time points after seroconversion in two individuals, there was also greater genetic identity among sequences from the same tissue compartment than among sequences from different tissue compartments (phenetic tissue compartmentalization). Over time, temporal phylogenetic and phenetic structure was detectable in mucosal and plasma viral samples from all three women, suggesting a continual process of migration of one or a few infected cells into each compartment followed by localized expansion and evolution of that population.

Amino Acid Sequence↗

The maximum information principle and the evolution of nucleotide sequences.

The probability distributions of bases in nucleotide sequences are deduced from the maximum information principle by maximizing the entropy (due to random mutation of bases) under certain constraints (Markovian entropy, G + C content, etc, due to selection). Two formulations are given with respect to different selective constraints. The deviations of theoretical distributions from experimental data are lower than 10% for most sequences. It is shown that the Lagrange multipliers change from species to species systematically--i.e. selective constraints correlate with evolution.

Animals↗

Sequence organization and evolution, in all extant whalebone whales, of a DNA satellite with terminal chromosome localization.

A heavy (GC rich) DNA satellite with terminal chromosomal localization is characteristic for all mysticete (whalebone whale) genomes. Sequences of 58 repeats of the satellite were compared in all ten extant mysticete species. In three families comprising eight species, the typical repeat length was 422 (421) bp. In two species, the northern right whale and the bowhead, of family Balaenidae (right whales) the repeats were much longer, typically ca. 900 and ca. 1200 bp. In all species the repeats were composed of a unique portion of constant length (212/211 bp), and a subrepeat portion, the length of which was variable. The evolutionary rigidity of the unique portion of the repeat is contrasted by the pronounced length variability of the subrepeat portion. The subrepeat portion consists essentially of 6 bp motifs, such that length differences are usually in multiples of 6 bp. The motif TTAGGG constituted 35%-50% of the subrepeats. Comparison between the unique portion of the 58 sequenced repeats revealed that the repeats divided into two primary groups, one comprising the two balaenids, the other including the eight remaining species. The mean difference between the two groups averaged 8.4%. In this sequence comparison the repeats of the pygmy right whale constituted a group that was separated from repeats of the other species. In all other cases repeats were intermingled to some extent between species. Comparison of individual repeats suggests that the unique portion evolves in concert, at a slow rate. A neighbor-joining comparison between the consensuses of all species suggests that the unique portion of the repeats evolves at a somewhat different rate in different lineages.

Animals↗

Nucleotide sequence and molecular evolution of two tomato genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase.

We have isolated and sequenced two cDNA clones (LESS5 and LESS17) encoding the small subunit of ribulose-1,5-bisphosphate carboxylase of tomato (Lycopersicon esculentum). At the nucleotide level, the protein-coding regions of these genes are 85% conserved, while the untranslated 3' regions are only 55% conserved. Comparison with rbcS genes from other species of Solanaceae suggests that the tomato LESS5 gene, the Nicotiana tabacum NTSS23 gene and the Petunia hybrida SSU8 gene are orthologous members of the rbcS gene family. In addition, the tomato gene LESS17, and the Petunia hybrida gene SSU611, may also be orthologous, since their untranslated 3' regions are related. There is a large difference between the two tomato rbcS genes in the frequency of the CG dinucleotide. This difference may reflect different levels of methylation, and therefore expression, of the tomato genes. Many of the differences involving the CG dinucleotide can be represented as transitions between C and T on the noncoding strand. Such changes are consistent with observations that methylated cytosines are hot-spots for transitions.

Base Sequence↗

Sequence, structure and evolution of the gene coding for sn-glycerol-3-phosphate dehydrogenase in Drosophila melanogaster.

We present the complete nucleotide and deduced amino acid sequence for the gene encoding Drosophila sn-glycerol-3-phosphate dehydrogenase. A transcription unit of 5kb was identified which is composed of eight protein encoding exons. Three classes of transcripts were shown to differ only in the 3'-end and to code for three protein isoforms each with a different C-terminal amino acid sequence. Each transcript is shown to arise through the differential expression of three isotype-specific exons at the 3'-end of the gene by a developmentally regulated process of 3'-end formation and alternate splicing pathways of the pre-mRNA. In contrast, the 5'-end of the gene is simple in structure and each mRNA is transcribed from the same promoter sequence. A comparison of the organization of the Drosophila and murine genes and the primary amino acid sequence between a total of four species indicates that the GPDH gene-enzyme system is highly conserved and is evolving slowly.

Amino Acid Sequence↗

The human growth hormone locus: nucleotide sequence, biology, and evolution.

The human chromosomal growth hormone locus contained on cloned DNA and spanning approximately 66,500 bp was sequenced in its entirety to provide a framework for the analysis of its biology and evolution. This locus evolved by a series of duplications and contains in its present form five genes which display a remarkably high degree of sequence identity (approximately 95%) in all their domains. The DNA sequence of the locus reveals the presence of 48 middle repetitive sequence elements of the Alu type and one member of the KpnI family, all located in the intergenic regions. The expression of each gene was examined by screening pituitary and placental cDNA libraries by using gene-specific oligonucleotides. According to this analysis, the hGH-N gene is transcribed exclusively in the pituitary, whereas the other four genes (hCS-L, hCS-A, hGH-V, hCS-B) are expressed only in placental tissue, at levels characteristic for each gene. Particular DNA sequences found upstream of the individual promoter regions might account for the observed tissue specificity and different transcriptional activity of the genes. The hCS-L gene carries a G to A transition in a sequence used by the other four genes as an intronic 5' splice donor site. This mutation results in a different splicing pattern and, hence, in a novel sequence of the hCS-L gene mRNA and the deduced polypeptide.

Amino Acid Sequence↗

Heat shock protein 70 family: multiple sequence comparisons, function, and evolution.

The heat shock protein 70 kDa sequences (HSP70) are of great importance as molecular chaperones in protein folding and transport. They are abundant under conditions of cellular stress. They are highly conserved in all domains of life: Archaea, eubacteria, eukaryotes, and organelles (mitochondria, chloroplasts). A multiple alignment of a large collection of these sequences was obtained employing our symmetric-iterative ITERALIGN program (Brocchieri and Karlin 1998). Assessments of conservation are interpreted in evolutionary terms and with respect to functional implications. Many archaeal sequences (methanogens and halophiles) tend to align best with the Gram-positive sequences. These two groups also miss a signature segment [about 25 amino acids (aa) long] present in all other HSP70 species (Gupta and Golding 1993). We observed a second signature sequence of about 4 aa absent from all eukaryotic homologues, significantly aligned in all prokaryotic sequences. Consensus sequences were developed for eight groups [Archaea, Gram-positive, proteobacterial Gram-negative, singular bacteria, mitochondria, plastids, eukaryotic endoplasmic reticulum (ER) isoforms, eukaryotic cytoplasmic isoforms]. All group consensus comparisons tend to summarize better the alignments than do the individual sequence comparisons. The global individual consensus "matches" 87% with the consensus of consensuses sequence. A functional analysis of the global consensus identifies a (new) highly significant mixed charge cluster proximal to the carboxyl terminus of the sequence highlighting the hypercharge run EEDKKRRER (one-letter aa code used). The individual Archaea and Gram-positive sequences contain a corresponding significant mixed charge cluster in the location of the charge cluster of the consensus sequence. In contrast, the four Gram-negative proteobacterial sequences of the alignment do not have a charge cluster (even at the 5% significance level). All eukaryotic HSP70 sequences have the analogous charge cluster. Strikingly, several of the eukaryotic isoforms show multiple mixed charged clusters. These clusters were interpreted with supporting data related to HSP70 activity in facilitating chaperone, transport, and secretion function. We observed that the consensus contains only a single tryptophan residue and a single conserved cysteine. This is interpreted with respect to the target rule for disaggregating misfolded proteins. The mitochondrial HSP70 connections to bacterial HSP70 are analyzed, suggesting a polyphyletic split of Trypanosoma and Leishmania protist mitochondrial (Mt) homologues separated from Mt-animal/fungal/plant homologues. Moreover, the HSP70 sequences from the amitochondrial Entamoeba histolytica and Trichomonas vaginalis species were analyzed. The E. histolytica HSP70 is most similar to the higher eukaryotic cytoplasmic sequences, with significantly weaker alignments to ER sequences and much diminished matching to all eubacterial, mitochondrial, and chloroplast sequences. This appears to be at variance with the hypothesis that E. histolytica rather recently lost its mitochondrial organelle. T. vaginalis contains two HSP70 sequences, one Mt-like and the second similar to eukaryotic cytoplasmic sequences suggesting two diverse origins.

Amino Acid Sequence↗