PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “evolutionary analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Identification of a nuclear variant of MGEA5, a cytoplasmic hyaluronidase and a beta-N-acetylglucosaminidase.

MGEA5 was originally identified to be a novel human hyaluronidase, which is immunogenic in meningioma patients. Recently an N-acetylglucosaminidase was reported with identical sequence. Here, we define the origin of a splice variant by determining the genomic organization of the mgea5 gene. We find the splice variant missing a putative acetyltransferase domain of MGEA5. As for evolutionary analysis, we show that the MGEA5 is highly conserved in higher eukaryotes. As for expression analysis, we find both mRNA variants ubiquitously expressed in various human tissues and throughout mouse development. We generated polyclonal antibodies against MGEA5s/5 and identified proteins of 75 and 130 kDa, indicating posttranslational modifications of the larger protein. Cell fractionation revealed the cytoplasmic/cytoskeletal localization of the 130-kDa protein and the nuclear localization of the 75-kDa protein. We propose a model in which MGEA5 functions both as a hyaluronidase and an N-acetylglucosaminidase.

Acetylglucosaminidase↗

Evolutionarily conserved, "acatalytic" carbonic anhydrase-related protein XI contains a sequence motif present in the neuropeptide sauvagine: the human CA-RP XI gene (CA11) is embedded between the secretor gene cluster and the DBP gene at 19q13.3.

Conserved amino acid motifs are found in numerous expressed genes. Proteins and peptides with functional relationships may be identified using probes designed to hybridize with these motifs. An oligonucleotide probe was prepared to match the sequence of the expected active region of a frog corticotropin-releasing factor-like peptide sauvagine and used to screen a sheep brain cDNA library. A novel 1331-bp cDNA encoding a putative 328-residue protein with a theoretical mass of 36 kDa was identified. The presence of a strong signal sequence indicates that it is a secreted protein. The amino- and carboxy-terminal regions are characterized by several potential phosphorylation sites and binding motifs, suggesting a role in intracellular signal transduction. Although the protein possesses a 7-residue sequence identical to that found in sauvagine, its overall primary structure most closely resembles those of the alpha-carbonic anhydrases (alpha-CAs). Moreover, the detection of the human and mouse orthologues in the EST databases, together with an evolutionary analysis, indicates that the protein represents a new member of the alpha-CA gene family, which we designate carbonic anhydrase-related protein XI (CA-RP XI), encoded by CA11 (human) and Car11 (mouse, rat). The human CA11 gene appears to be located between the secretor type alpha(1,2)-fucosyltransferase gene cluster (FUT1-FUT2-FUT2P) and the D-site binding protein gene (DBP) on chromosome 19q13.3. Despite potentially inactivating changes in the active-site residues, CA-RP XI is evolving very slowly in mammals, a property indicative of an important function, which has also been observed in the two other "acatalytic" CA isoforms, CA-RP VIII and CA-RP X, whose functions are unknown.

Amino Acid Sequence↗

Genomic position analyses and the transcription machinery.

Position analyses have been devised to extract additional transcriptional information from rapidly expanding genomic data bases. The locations of promoter regulatory sites and also the locations of transcription factor DNA-binding domains are analyzed. Strongly preferred positions of activator binding sites occur in both Escherichia coli and eukaryotes, suggesting specific common features of transcription in the two systems. In both systems, regulatory proteins are found to have their DNA-binding domains near termini and the data suggest an evolutionary analysis that complements a phylogenetic analysis based on sequence alignments. The results indicate that positional information can be an important adjunct to sequence comparisons in analyzing genomic information.

Animals↗

Role of enzyme-substrate flexibility in catalytic activity: an evolutionary perspective.

Site-directed mutagenesis has proved an effective experimental technique to investigate catalytic mechanisms and to determine relations between enzyme structure and function. This article invokes an analytical model based on evolution by mutation and natural selection-Nature's analogue of site-directed mutagenesis-to derive a set of general rules relating enzyme structure and activity. The catalysts are described in terms of the structural parameters, rigidity and flexibility, and the functional variables, reaction rate and substrate specificity. The evolutionary model predicts the following structure-activity relations: (a) rigid enzyme-flexible substrate: large variation in reaction rates, broad substrate specificity; (b) rigid enzyme-rigid substrate: diffusion controlled rates, absolute specificity; (c) flexible enzyme-rigid substrate: intermediate reaction rates, group specificity; (d) flexible enzyme-flexible substrate: slow rates, absolute specificity. Spectroscopic methods and X-ray crystallography now yield important characteristics of enzyme-substrate complexes such as molecular flexibility. The evolutionary analysis we have exploited provides general principles for inferring catalytic activity from structural studies of enzyme-substrate complexes.

Animals↗

Molecular evolution of aerobic energy metabolism in primates.

As part of our goal to reconstruct human evolution at the DNA level, we have been examining changes in the biochemical machinery for aerobic energy metabolism. We find that protein subunits of two of the electron transfer complexes, complex III and complex IV, and cytochrome c, the protein carrier that connects them, have all undergone a period of rapid protein evolution in the anthropoid lineage that ultimately led to humans. Indeed, subunit IV of cytochrome c oxidase (COX; complex IV) provides one of the best examples of positively selected changes of any protein studied. The rate of subunit IV evolution accelerated in our catarrhine ancestors in the period between 40 to 18 million years ago and then decelerated in the descendant hominid lineages, a pattern of rate changes indicative of positive selection of adaptive changes followed by purifying selection acting against further changes. Besides clear evidence that adaptive evolution occurred for cytochrome c and subunits of complexes III (e.g., cytochrome c(1)) and IV (e.g., COX2 and COX4), modest rate accelerations in the lineage that led to humans are seen for other subunits of both complexes. In addition the contractile muscle-specific isoform of COX subunit VIII became a pseudogene in an anthropoid ancestor of humans but appears to be a functional gene in the nonanthropoid primates. These changes in the aerobic energy complexes coincide with the expansion of the energy-dependent neocortex during the emergence of the higher primates. Discovering the biochemical adaptations suggested by molecular evolutionary analysis will be an exciting challenge.

Animals↗

Adaptive evolution of larvae and life cycles.

The larval patterns of marine invertebrates pose intriguing questions for both evolutionary and developmental biologists. However, combined investigations have been rare. Quantitative models analyze the selective factors that drive evolutionary change in larval nutrition and timing of metamorphosis. Developmental studies describe the morphogenesis characterizing ancestral and derived larval patterns. Rigorous evolutionary analysis of the transition to derived modes of development is lacking and detailed developmental and ecological data are needed to test and refine theoretical models. A major challenge facing studies of life cycle evolution is the elucidation of the genetic structure and covariance of important developmental and larval traits.

Adaptation, Physiological↗

Evolutionary origin of a Kunitz-type trypsin inhibitor domain inserted in the amyloid beta precursor protein of Alzheimer's disease.

The Kunitz-type protease inhibitor is one of the serine protease inhibitors. It is found in blood, saliva, and all tissues in mammals. Recently, a Kunitz-type sequence was found in the protein sequence of the amyloid beta precursor protein (beta APP). It is known that beta APP accumulates in the neuritic plaques and cerebrovascular deposits of patients with Alzheimer's disease. Collagen type VI in chicken also has an insertion of a Kunitz-type sequence. To elucidate the evolutionary origin of these insertion sequences, we constructed a phylogenetic tree by use of all the available sequences of Kunitz-type inhibitors. The tree shows that the ancestral gene of the Kunitz-type inhibitor appeared about 500 million years ago. Thereafter, this gene duplicated itself many times, and some of the duplicates were inserted into other protein-coding genes. During this process, the Kunitz-type sequence in the present beta APP gene diverged from its ancestral gene about 270 million years ago and was inserted into the gene soon after duplication. Although the function of the insertion sequences is unknown, our molecular evolutionary analysis shows that these insertion sequences in beta APP have an evolutionarily close relationship with the inter-alpha-trypsin inhibitor or trypstatin, which inhibits the activity of tryptase, a novel membrane-bound serine protease in human T4+ lymphocytes.

Alzheimer Disease↗

Self peptides bound by HLA class I molecules are derived from highly conserved regions of a set of evolutionarily conserved proteins.

An evolutionary analysis of self peptides reported to be bound by HLA class I molecules showed that these peptides are largely derived from proteins that have been highly conserved in the history of mammals. These proteins also often have universal tissue expression and have a higher than average frequency of highly hydrophilic residues. The peptides themselves are generally still more highly conserved than the source proteins and have a higher frequency of highly hydrophobic residues, evidently often being derived from conserved hydrophobic cores of the source proteins. These results suggest that the mechanism by which peptides are derived for MHC presentation may preferentially select peptides from conserved protein regions. In the case of parasite-derived peptides, such a mechanism would be adaptive in that it would reduce the likelihood of escape mutants.

Animals↗

Molecular evolution and secondary structural conservation in the B-cell lymphoma leukemia 2 (bcl-2) family of proto-oncogene products.

The nature of the bcl-2 family of proto-oncogenes was analyzed by sequence alignment, secondary structure prediction, and phylogenetic techniques. Phylogenies were inferred from both the nucleic acid and amino acid sequences of the human, murine, rat, and chicken sequences for BCL-2 and BCL-X, human MCL1, murine A1, the nematode Caenorhabditis elegans and Caenorhabditis briggsiae ced-9 proteins, and the sequences BHRF1 from Epstein-Barr and LMW5-HL from African swine fever viruses. Both sequence alignment and secondary structure prediction techniques supported the conservation of both the overall secondary structure and the carboxy-terminal transmembrane domain in all members of the family. All the treeing methods employed (distance matrix, maximum likelihood, and parsimony) supported a tree in which the proapoptotic proteins BCL-2 and BCL-X represent the most recent additions to the group. All the trees also indicated that the viral proteins BHRF1 and LMW-HL arose from a common ancestor, an ancestor they shared in common with the pro-apoptotic control protein BAX, indicating that this function of BAX evolved only recently. The most ancient branches are represented by the nematode ced-9 protein and by the control genes MCL1 and A1, which in the treeing methods employed represent separate lineages within the most ancient grouping. These results demonstrate the evolution of a highly conserved family of developmental control genes from nematode to man--genes that encode proteins essential for normal development but which are highly conserved in terms of predicted structure and possible cellular localization. The evolutionary analysis also indicates that the family may be even larger than originally predicted and that other members are waiting to be discovered.

Amino Acid Sequence↗

Immunoglobulin lambda light chain evolution: Igl and Igl-like sequences form three major groups.

The nucleotide sequences, and the derived protein sequences, of immunoglobulin (Ig) Igl, Igl-like VpreB genes and the protein sequences of Igl-C regions were aligned and compared. A classification of the Igl and Igl-like VpreB sequences into three categories, designated groups I, II, and III, is proposed. Group I contains the human and mouse Igl-like VpreB genes. Group II contains Igl-V genes of the rabbit and the recently described mouse Igl-Vx gene. Group III includes the Igl-V genes, encoding all other known Igl-V region protein sequences, of mouse, rat, human, pig, sheep, and chicken. An evolutionary analysis of the three groups is presented, and suggests that the group III genes are evolving at a faster rate than those of the other groups and that within this group a further subdivision is possible: the V lambda-encoding genes of mouse, rat, and one human subgroup evolve faster than other group III genes. It is suggested that all mammalian species contain Igl-V genes of each group. A similar comparison between the protein sequences encoded by the known Igl-C genes indicates that the duplication of the Igl-J-C gene pairs occurred independently in each species, after mammalian speciation, and that the Igl-V-(J-C)(J-C) gene clusters of the mouse may not have their homologues in other species.

Amino Acid Sequence↗

A gene-culture model of human handedness.

A model of handedness incorporating both genetic and cultural processes is proposed, based on an evolutionary analysis, and maximum-likelihood estimates of its parameters are generated. This model has the characteristics that (i) no genetic variation underlies variation in handedness, and (ii) variation in handedness among humans is the result of a combination of cultural and developmental factors, but (iii) a genetic influence remains since handedness is a facultative trait. The model fits the data from 17 studies of handedness in families and 14 studies of handedness in monozygotic and dizygotic twins. This model has the additional advantages that it can explain why monozygotic and dizygotic twins and siblings have similar concordance rates, and no hypothetical selection regimes are required to explain the persistence of left handedness.

Biological Evolution↗

Genotype distribution in Nagoya and new genotype (genotype 3a) in Japanese patients with hepatitis C virus.

We evaluated hepatitis C virus (HCV) genotype distribution among Japanese patients in the city of Nagoya and the possible existence of any other genotype not determined by Okamoto's method. Eighty-five of 93 (91.4%) anti-HCV-positive patients had detectable HCV RNA. The genotype of the HCV isolate was determined in 84 of 85 (98.8%) of these HCV RNA-positive patients by Okamoto's method but determination was not possible in one (1.2%). Genotype 1b was detected in 58 of the 85 patients (68.2%), genotype 2a in 20 (23.5%), genotype 2b in 3 (3.5%), and genotype 1b + 2a in 3 (3.5%). In the remaining 1 patient in whom the genotype could not be determined, we determined the nucleotide sequence of the core region in HCV RNA extracted from this patient and evaluated it by molecular evolutionary analysis. This HCV isolate was then classified as genotype 3a. These results suggest that genotype 3a is rare among Japanese patients with HCV; thus, when classifying Japanese isolates, we should take more care because genotype 3a is not determined by current typing systems.

Amino Acid Sequence↗

cDNA and amino acid sequences of rainbow trout (Oncorhynchus mykiss) lysozymes and their implications for the evolution of lysozyme and lactalbumin.

The complete 129-amino-acid sequences of two rainbow trout lysozymes (I and II) isolated from kidney were established using protein chemistry microtechniques. The two sequences differ only at position 86, I having aspartic acid and II having alanine. A cDNA clone coding for rainbow trout lysozyme was isolated from a cDNA library made from liver mRNA. Sequencing of the cloned cDNA insert, which was 1 kb in length, revealed a 432-bp open reading frame encoding an amino-terminal peptide of 15 amino acids and a mature enzyme of 129 amino acids identical in sequence to II. Forms I and II from kidney and liver were also analyzed using enzymatic amplification via PCR and direct sequencing; both organs contain mRNA encoding the two lysozymes. Evolutionary trees relating DNA sequences coding for lysozymes c and alpha-lactalbumins provide evidence that the gene duplication giving rise to conventional vertebrate lysozymes c and to lactalbumin preceded the divergence of fishes and tetrapods about 400 Myr ago. Evolutionary analysis also suggests that amino acid replacements may have accumulated more slowly on the lineage leading to fish lysozyme than on those leading to mammal and bird lysozymes.

Amino Acid Sequence↗

Novel chaperonins in a prokaryote.

Group II chaperonins belong to the Hsp60 family occurring in archaea and eukaryotes. The archaeal chaperonins build the thermosome, which is similar to the eukaryotic CCT (chaperonin-containing TCP-1). Eukaryotes have eight subunits, and up until now, it was thought that archaea had between one and three subunits, depending on the species. We now report two novel subunits, termed Hsp60-4 and Hsp60-5, in the archaeon Methanosarcina acetivorans, which also has Hsp60-1, Hsp60-2, and Hsp60-3 with orthologs in Methanosarcinae. Hsp60-4 and Hsp60-5 occur only in M. acetivorans, which makes this organism unique in that it has the highest number of chaperonin subunits ever described for an archaeon. Evolutionary analysis suggests that either Hsp60-4 or Hsp60-5 paralogs have arisen by gene duplication with vastly increased accepted substitution rates or that they represent ancestral types found only in this species.

Amino Acid Sequence↗

Molecular evolution of the GapC gene family in Amsinckia spectabilis populations that differ in outcrossing rate.

Molecular evolutionary analysis of the glyceraldehyde 3-phosphate dehydrogenase (GapC) gene family was conducted in the plant genus Amsinckia (Boraginaceae), a group that exhibits marked variation in the mating system. GapC genes in this group differ from those of Arabidopsis thaliana in terms of both intron size and number. Phylogenetic and Southern hybridization analyses suggest the presence of multiple GapC loci, each defined by a set of base substitutions that are in strong linkage disequilibrium. One species of Amsinckia, A. spectabilis, was studied in some detail. This species consists of selfing (A. s. spectabilis) and outcrossing (A. s. microcarpa) varieties. Two selfing populations and one outcrossing population sample were analyzed in detail for variation at one of the members of this gene family. GapC3. A reduction in number of GapC3 haplotypes and level of genetic diversity was observed in the selfing populations of A. spectabilis. GapC3 in the outcrossing population (but not the two selfing populations) exhibited a significant departure from neutrality in the direction of an excess of singletons. These results are discussed in the context of forces acting on sequence evolution in populations with different mating systems.

Amsinckia↗

Comparative genomics reveals long, evolutionarily conserved, low-complexity islands in yeast proteins.

Eukaryotic proteomes abound in low-complexity sequences, including tandem repeats and regions with significantly biased amino acid compositions. We assessed the functional importance of compositionally biased sequences in the yeast proteome using an evolutionary analysis of 2838 orthologous open reading frame (ORF) families from three Saccharomyces species (S. cerevisiae, S. bayanus, and S. paradoxus). Sequence conservation was measured by the amino acid sequence variability and by the ratio of nonsynonymous-to-synonymous nucleotide substitutions (K(a)/K(s)) between pairs of orthologous ORFs. A total of 1033 ORF families contained one or more long (at least 45 residues), low-complexity islands as defined by a measure based on the Shannon information index. Low-complexity islands were generally less conserved than ORFs as a whole; on average they were 50% more variable in amino acid sequences and 50% higher in K(a)/K(s) ratios. Fast-evolving low-complexity sequences outnumbered conserved low-complexity sequences by a ratio of 10 to 1. Sequence differences between orthologous ORFs fit well to a selectively neutral Poisson model of sequence divergence. We therefore used the Poisson model to identify conserved low-complexity sequences. ORFs containing the 33 most conserved low-complexity sequences were overrepresented by those encoding nucleic acid binding proteins, cytoskeleton components, and intracellular transporters. While a few conserved low-complexity islands were known functional domains (e.g., DNA/RNA-binding domains), most were uncharacterized. We discuss how comparative genomics of closely related species can be employed further to distinguish functionally important, shorter, low-complexity sequences from the vast majority of such sequences likely maintained by neutral processes.

Amino Acid Sequence↗

Evolutionary correlation between linker histones and microtubular structures.

Histones of the H1 group (linker histones) are abundant components of chromatin in eukaryotes, occurring on average at one molecule per nucleosome. The recent reports on the lack of a clear phenotypic effect of knock-out mutations as well as overexpression of histone H1 genes in different organisms have seriously undermined the long-held view that linker histones are essential for the basic functions of eukaryotic cells. In an attempt to resolve the paradox of an abundant conserved protein without a clear function, we re-examined the molecular and phylogenetic data on linker histones to see if they could reveal any correlation between the features of H1 and the functional or morphological characteristics of cells or organisms. Because of an earlier demonstration that in sea urchin the chromatin-type histone H1 is also found in the flagellar microtubules (Multigner et al. 1992), we focused on the correlation between the features of H1 and those of microtubular structures. A phylogenetic tree based on multiple alignment of over 100 available HI sequences suggests that the first divergence of the globular domain of H1 (GH1) resulted in branching into separate types characteristic for plants/Dictyostelium and for animals/ascomycetes, respectively. The GH1s of these two types differ by a short region (usually 5 amino acids) placed at a specific location within the C-terminal wing subdomain of GH1. Evolutionary analysis of the diversification of H1 mRNA into cell-cycle-dependent (polyA-) and independent (polyA+) forms showed a mosaic occurrence of these two forms in plants and animals, despite the fact that the H1 proteins of plants and animals belong to two well-distinguished groups. However, among organisms from both animal and plant kingdom, only those with H1 mRNA of a polyA- type have flagellated gametes. This correlation as well as the demonstration that in Volvox carteri the accumulation of polyA- mRNA of H1 occurs concurrently with the production of new flagella (Lindauer et al. 1993), suggests a direct link between polyA- phenotype of histone H1 mRNA and flagellogenesis.

Animals↗

A phylogenetic framework for the aquaporin family in eukaryotes.

A comprehensive evolutionary analysis of aquaporins, a family of intrinsic membrane proteins that function as water channels, was conducted to establish groups of homology (i.e., to identify orthologues and paralogues) within the family and to gain insights into the functional constraints acting on the structure of the aquaporin molecule structure. Aquaporins are present in all living organisms, and therefore, they provide an excellent opportunity to further our understanding of the broader biological significance of molecular evolution by gene duplication followed by functional and structural specialization. Based on the resulting phylogeny, the 153 channel proteins analyzed were classified into six major paralogous groups: (1) GLPs, or glycerol-transporting channel proteins, which include mammalian AQP3, AQP7, and AQP9, several nematode paralogues, a yeast paralogue, and Escherichia coli GLP; (2) AQPs, or aquaporins, which include metazoan AQP0, AQP1, AQP2, AQP4, AQP5, and AQP6; (3) PIPs, or plasma membrane intrinsic proteins of plants, which include PIP1 and PIP2; (4) TIPs, or tonoplast intrinsic proteins of plants, which include alphaTIP, gammaTIP, and deltaTIP; (5) NODs, or nodulins of plants; and (6) AQP8s, or metazoan aquaporin 8 proteins. Of these groups, AQPs, PIPs, and TIPs cluster together. According to the results, the capacity to transport glycerol shown by several members of the family was acquired only early in the history of the family. The new phylogeny reveals that several water channel proteins are misclassified and require reassignment, whereas several previously undetermined ones can now be classified with confidence. The deduced phylogenetic framework was used to characterize the molecular features of water channel proteins. Three motifs are common to all family members: AEF (Ala-Glu-Phe), which is located in the N-terminal domain; and two NPA (Asp-Pro-Ala) boxes, which are located in the center and C-terminal domains, respectively. Other residues are found to be conserved within the major groups but not among them. Overall, the PIP subfamily showed the least variation. In general, no radical amino acid replacements affecting tertiary structure were identified, with the exception of Ala-->Ser in the TIP subfamily. Constancy of rates of evolution was demonstrated within the different paralogues but rejected among several of them (GLP and NOD).

Amino Acid Motifs↗