PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Comparative genomics of the late gene cluster from Lactobacillus phages.

Three prophage sequences were identified in the Lactobacillus johnsoni strain NCC533. Prophage Lj965 predicted a gene map very similar to those of pac-site Streptococcus thermophilus phages over its DNA packaging and head and tail morphogenesis modules. Sequence similarity linked the putative DNA packaging and head morphogenesis genes at the protein level. Prophage Lj965/S. thermophilus phage Sfi11/Lactococcus lactis phage TP901-1 on one hand and Lactobacillus delbrueckii phage LL-H/Lactobacillus plantarum phage phig1e/Listeria monocytogenes phage A118 on the other hand defined two sublines of structural gene clusters in pac-site Siphoviridae from low-GC Gram-positive bacteria. Bacillus subtilis phage SPP1 linked both sublines. The putative major head and tail proteins from Lj965 shared weak sequence similarity with phages from Gram-negative bacteria. A clearly independent line of structural genes in Siphoviridae from low-GC Gram-positive bacteria is defined by temperate cos-site phages including Lactobacillus gasseri phage adh, which also shared sequence similarity with phage D3 infecting a Gram-negative bacterium. A phylogenetic tree analysis demonstrated that the ClpP-like protein identified in four cos-site Siphoviridae from Lactobacillus, Lactococcus, Streptococcus, and Pseudomonas showed graded sequence relationships. The tree suggested that the ClpP-like proteins from the phages were not acquired by horizontal gene transfer from their corresponding bacterial hosts.

Amino Acid Sequence↗

Variation in synonymous substitution rates among mammalian genes and the correlation between synonymous and nonsynonymous divergences.

Using mammalian gene sequences, the variances in the numbers of synonymous and nonsynonymous substitutions among genes were estimated together with the correlation coefficient between the two. The expected correlation coefficient can be obtained under the neutral theory using these estimated values of the variances. The expected coefficient is found to often be one-half to two-thirds of the observed value. Possible causes for the disagreement were discussed, such as correlated selective constraints on the two types of substitutions and excess doublet mutations. The variance of mutation rate and that of selective constraint were also estimated. The results show that the coefficient of variation of the former is 0.2-0.3, whereas that of the latter is 0.7-0.9.

Animals↗

Seed storage proteins of spermatophytes share a common ancestor with desiccation proteins of fungi.

The legumin- and vicilin-like seed storage globulins of spermatophytes are specifically accumulated during embryogenesis and seed development. Previous studies have shown that a precursor common to both legumin and vicilin genes might have evolved by duplication from a single-domain ancestral gene. We here report that amino acid sequences of legumin and vicilin domains share statistically significant similarity to the germination-specific germins of wheat as well as to the spherulation-specific spherulins of myxomycetes. This conclusion is further supported by the derived intron-exon structure of a spherulin gene. Spherulins are thought to be involved in tissue desiccation or hydration. It is suggested that the present-day seed globulins of spermatophytes have evolved from a group of ancient proteins functional in cellular desiccation/hydration processes.

Amino Acid Sequence↗

Highly repetitive component alpha and related alphoid DNAs in man and monkeys.

The genomes of Old-World, New-World, and prosimian primates contain members of a large class of highly repetitive DNAs that are related to one another and to component alpha DNA of the African green monkey by their sequence homologies and restriction site periodicities. The members of this class of highly repetitive DNAs are termed the alphoid DNAs, after the prototypical member, component alpha of the African green monkey which was the first such DNA to be identified (Maio, 1971) and sequenced (Rosenberg et al., 1978). The alphoid DNAs appear to be uniquely primate sequences.--From the restriction enzyme cleavage patterns and Southern blot hybridizations under different stringency conditions, the alphoid DNAs comprise multiple sequence families exhibiting varying degrees of homology to component alpha DNA. They also share common elements in their restriction site periodicities (172 . n base-pairs), in the long-range organization of their repeating units, and in their banding behavior in CsCl and Cs2SO4 bouyant density gradients, in which they band within the bulk DNA as cryptic repetitive components.--In the three species from the Family Cercopithecidae examined, the alphoid DNAs represent the most abundant, tandemly repetitive sequence components, comprising about 24% of the African green monkey genome and 8 to 10% of the Rhesus monkey and baboon genomes. In restriction digests, the bulk of the alphoid DNAs among the Cercopithecidae appeared quantitatively reduced to a simple series of arithmetic segments based on a 172 base-pair (bp) repeat. In contrast with these simple restriction patterns, complex patterns were observed when human alphoid DNAs were cleaved with restriction enzymes. Detailed analysis revealed that the human genome contains multiple alphoid sequence families which differ from one another both in their repeat sequence organization and in their degree of homology to the African green monkey component alpha DNA.--The finding of alphoid sequences in other Old-World primate families, in a New-World monkey, and in a prosimian primate attests to the antiquity of these sequences in primate evolution and to the sequence conservatism of a large class of mammalian highly repetitive DNA. In addition, the relative conservatism exhibited by these sequences may distinguish the alphoid DNAs from more recently evolved highly repetitive components and satellite DNAs which have a more restricted taxonomical distribution.

Animals↗

Characterization of reiterated human DNA with respect to mammalian X chromosome homology.

Recombinants containing human repetitive DNA sequences were identified by dot hybridization and classified with respect to presence on the X chromosome and homology to mouse DNA. Using genomic probes that differ in number of X chromosomes, we observed extensive homology between human autosomal and X sequences. Hybridization to genomic probes that differ in species of origin indicate that these reiterated sequences have diverged between mouse and man. Eleven recombinants, each containing a different reiterated sequence(s), were hybridized in situ to metaphase chromosomes of mouse and man. These studies indicate that reiterated DNA which is homologous to the human X chromosome is more similar to DNA of human autosomes than to any murine chromosome. Therefore, it seems that reiterated DNA sequences on the human X chromosome have diverged as much during mammalian evolution as sequences on human autosomes. Moreover, the extensive modification of the original mammalian X has not interferred with the X inactivation process.

Animals↗

Mammalian genes as molecular clocks?

In analyzing the silent nucleotide substitutions in some mammalian mitochondrial mRNA coding genes, we had found that the frequency of each of the four nucleotides in rat, mouse, and cow, but not in humans, is the same in the silent third codon position (Lanave C, Preparata G, Saccone C, Serio G (1984) J Mol Evol 20:86-93). Because our findings for these three species were compatible with a stationary Markov process for the evolution of nucleotide sequences, we applied such a model to calculate the effective evolutionary silent substitution rate (vs) and the divergence times among the species. In this paper we have analyzed the first and second codon positions in the same mammalian mitochondrial genes. We found that in the first and second codon positions the human mitochondrial genes satisfy the stationarity conditions. This has allowed us to use the stochastic model mentioned above to calculate the divergence times among mouse, rat, cow, and human. Furthermore, we have analyzed the silent substitution rate in one nuclear gene for these four mammals. We found that in this gene the effective silent substitution rate is about 3 times lower than in mitochondrial genes, and that humans are in this case stationary with respect to the other three mammals in the third codon position as well. Application of our Markov model to this latter gene yields divergence times consistent with our previous determinations.

Animals↗

Identification of protein-tyrosine phosphatases in Archaea.

Protein-tyrosine dephosphorylation is a major mechanism in cellular regulation. A large number of protein-tyrosine phosphatases is known from Eukarya, and more recently bacterial homologues have also been identified. By employing conserved sequence patterns from both eukaryotic and bacterial protein-tyrosine phosphatases, we have identified three homologous sequences in two of the four complete archaeal genomes. Two hypothetical open reading frames in the genome of Methanococcus jannaschii (MJ0215 and MJECL20) and one in the genome of Pyrococcus horikoshii (PH1732) clearly bear all the conserved residues of this family. No homologues were found in the genomes of Archaeoglobus fulgidus and Methanobacterium thermoautotrophicum. This is the first report of protein-tyrosine phosphatase sequences in Archaea.

Amino Acid Sequence↗

The antigen receptor (NCCRP-1) on catfish and zebrafish nonspecific cytotoxic cells belongs to a new gene family characterized by an F-box-associated domain.

The catfish nonspecific cytotoxic cell receptor protein (NCCRP-1) provides an important function in target cell recognition and activation of cytotoxicity. This report identifies and characterizes a zebrafish orthologue of the catfish NCCRP-1. The zebrafish NCCRP-1 cDNA contains an open reading frame that encodes a predicted protein of 237 amino acids with a MW of 27 kDa and a pI of 5.5. Sequence similarities comparisons show that the NCCRP-1 receptors from these two phylogenetically distant species share a high degree of identity. These results suggested that NCCRP-1 performs a crucial function in innate immunity in teleosts. Further, a zebrafish 17-mer peptide corresponding to the catfish NCCRP-1 antigen-binding domain inhibited (catfish) cytotoxicity toward conventional tumor target cells (HL-60). These data appeared to indicate that the zebrafish NCCRP-1 protein may function as an antigen recognition molecule and, as such, may participate in innate immunity in teleosts. A homology search of the zebrafish NCCRP-1 protein revealed that it shares a significant level of identity with another group of proteins belonging to an F-box subfamily. These proteins share an F-box domain in the N terminus (not present in NCCRP-1) and an extremely conserved C-terminal region that has been termed the F-box-associated domain (FBA). The FBA is currently of unknown function. A new gene family is proposed in this work, based on similarities in the FBA sequences with the catfish and zebrafish NCCRP-1 peptides. This new gene family includes several F-box domain-containing proteins and a predicted C. elegans protein.

Amino Acid Motifs↗

Evidence of diversifying selection in human papillomavirus type 16 E6 but not E7 oncogenes.

Human papillomavirus type 16 is a common sexually transmitted pathogen capable of giving rise to cervical intraepithelial neoplasia and invasive carcinoma through the expression and activity of two adjacent oncogenes: E6 and E7. Naturally occurring amino acid variation is commonly observed in the E6 protein but to a much lesser extent in E7. In order to investigate the evolutionary mechanisms involved in the generation and maintenance of this variation, we examine 42 distinct E6-E7 haplotypes using codon-based genealogical techniques. These techniques involve estimation of the ratio of nonsynonymous to synonymous substitutions (dn/ds) and allow testing for directional (positive) natural selection. Positive selection was detected for four codon sites within the E6 oncogene but not in any E7 codons. The amino acid compositions and locations of selected sites are described. Possible sources of natural selection including antiviral immune pressure and polymorphism of host cellular proteins are discussed.

Amino Acid Sequence↗

Potential for retroposition by old Alu subfamilies.

Alu elements sharing sequence characteristics of the "old" subfamilies are thought to currently be retrotranspositionally inactive. We analyzed one of these old subfamilies of Alu elements, Sx, for sequence conservation relative to the consensus and the length of the "A-tail" as parameters to define the presence of potential Alu Sx source genes in the human genome. Sequence identity to the left half or the right half of the Alu Sx consensus sequence was evaluated for 4424 complete elements obtained from the human genome draft sequence. A small subset of Alu Sx left halves were found to be more conserved than any of the Alu Sx right halves. Selection for promoter function in active elements may explain the slightly higher conservation of the left half. In order to determine whether this sequence identity was the result of recent activity, or simply sequence conservation for older elements, PCR amplification of some of the loci containing Sx elements with conserved left/right halves from different primate genomes was carried out. Several of these Sx Alus were found to have amplified at a later evolutionary period (<35 mya) than expected based on previous studies of Sx elements. Analysis of "A-tail" length, a feature correlated with current retroposition activity, varied between Alu Sx element loci in different primates, where the length increased in specific Alu elements in the human genome. The presence of few conserved Alu Sx elements and the dynamic expansion/contraction of the A-tail suggests that some of these older subfamilies may still be active at very low levels or in a few individuals.

Alu Elements↗

The shark HoxN cluster is homologous to the human HoxD cluster.

The statistical analysis of phylogenetic footprints in the two known horn shark Hox clusters and the four mammalian clusters shows that the shark HoxN cluster is HoxD-like. This finding implies that the most recent common ancestor of jawed vertebrates had at least four Hox clusters, including those which are orthologous to the four mammalian Hox clusters.

Amino Acid Sequence↗

Getting the proto-Pax by the tail.

Pax genes encode transcription factors governing the determination of different cell types and even organs in the development of multicellular animals. Pax proteins are characterized by the presence of three evolutionarily conserved elements: two DNA-binding domains, the paired domain (PD) and paired-type homeodomain (PtHD), and the short octopeptide sequence (OP) located between PD and PtHD. PD is the defining feature of this class of genes, while OP and/or PtHD may be divergent or absent in some members of the family. Phylogenetic analyses of the PD and PtHD sequences do not distinguish which particular type of the extant Pax genes more resembles the ancestral type. Here we present evidence for the existence of a fourth evolutionarily conserved domain in the Pax proteins, the paired-type homeodomain tail (PHT). Our data also imply that the hypothetical proto-Pax protein most probably exhibited a complex structure, PD-OP-PtHD-PHT, which has been retained in the extant proteins Pax3/7 of the ascidia and lancelet, and Pax7 of the vertebrates. Finally, based on structural considerations, a scenario for the evolutionary emergence of the proto-Pax gene is proposed.

Amino Acid Sequence↗

Do pheromone binding proteins converge in amino acid sequence when pheromones converge?

Convergence in amino acid sequences between proteins can be strong evidence for selection. Here, I look for evidence of convergence in the amino acid sequences of pheromone binding protein (PBP) in response to convergence in pheromones. PBPs are involved in sex pheromone reception by the antennae of male moths. In this role PBPs may selectively bind pheromone components and experience convergent selection in response to convergence in pheromone components. However, examination of the PBPs of the taxa that have converged upon the use of (E)- or (Z)-11-tetradecenyl acetate as their major pheromone component reveals little evidence for convergence in the PBPs identified from these taxa. A few sites show a pattern consistent with convergence or parallelism; however, it cannot be ruled out that these sites share the ancestral state. Two of these sites fall within the proposed binding region of PBPs. These results suggest that PBPs either have not converged in sequence or have converged at very few sites in response to convergence on the same pheromone component.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of interleukin 16 from nonhuman primates and from the mouse.

Interleukin 16 (IL-16) is synthesized as a 67 000 Mr precursor (pro-IL-16), but only a carboxy terminal part of 12 000-14 000 Mr is secreted by CD8(+) lymphocytes. This lymphokine binds to CD4 and has been shown to induce migration, affect the activation state of T cells, and inhibit immunodeficiency virus replication. It has been suggested that CD8(+) cell-derived soluble factors play a pivotal role in protecting natural-host nonhuman primates from developing immunodeficiency following SIV infection. In a first attempt to address this question, we cloned and sequenced the IL-16 cDNA from different primates. Here we report the pro-IL-16 sequence from chimpanzees, African green monkeys (AGM), rhesus macaques, and cynomolgus macaques. In order to compare and analyze structural motifs possibly involved in processing, intracellular targeting, or secretion, we extended our study to the New World monkeys saimiri and aotus and to the mouse. Alignments of deduced amino acids reveal that the human protein shares 99% similarity to that of chimpanzees, approximately 95% to rhesus, cynomolgus and AGM, about 90% to aotus and saimiri, and 77.5% to the mouse. Phylogenetic analyses revealed the expected evolutionary groupings.

Amino Acid Sequence↗

Conserved organization of the ILT/LIR gene family within the polymorphic human leukocyte receptor complex.

The human leukocyte receptor complex (LRC) at Chromosome 19q13.4 encodes Ig superfamily proteins which regulate the function of various hematopoietic cell types. We investigated characteristics of the Ig-like transcript (ILT)/leukocyte Ig-like receptor (LIR) group of LRC genes in comparison with the other major LRC loci encoding the killer cell Ig-like receptors (KIRs). In direct contrast to KIR genes, the ILT/LIR loci of ethnically diverse individuals did not display haplotypic variations in gene number. Investigation of gene expression identified novel cDNA sequences related to the ILT2/LIR1, ILT4/LIR2, ILT3/LIR5, and ILT7 loci, while phylogenetic analysis revealed two distinct lineages of ILT/LIR genes. These two lineages differ in both the nature and extent of their sequence polymorphism. The presence of certain transcription factor-related motifs in the 5' untranslated region of ILT/LIR cDNAs correlates with the specific cell types in which particular ILT/LIR genes are expressed. Although extensive gene duplications and conversion events have apparently forged the LRC, our results indicate striking conservation in the organization of the ILT/LIR genes when compared with the related and closely linked KIR genes. This suggests the evolutionary maintenance of a significant function consistent with the cellular distribution of the ILT/LIR proteins.

Base Sequence↗

Gene structure of Taenia solium paramyosin.

Paramyosin is a muscle protein that probably plays a role in the survival of the larval stage of Taenia solium during its prolonged host-parasite relationship. Here we describe the structure of the gene coding for the paramyosin of T. solium. The characterization of two clones obtained from a genomic library showed that the complete gene of paramyosin contains 13 introns delimited by conventional eukaryotic splice signals. Comparison with the paramyosin genes of Drosophila melanogaster and Caenorhabditis elegans showed a lack of conservation of the exon/intron organization in contrast to other muscle genes. No evidence of alternative splicing sites were found, excluding the possibility that T. solium expresses a mini-paramyosin like D. melanogaster.

Amino Acid Sequence↗

Direct repeats in the flavivirus 3' untranslated region; a strategy for survival in the environment?

Previously, direct repeats (DRs) of 20-70 nucleotides were identified in the 3' untranslated regions (3'UTR) of flavivirus sequences. To address their functional significance, we have manually generated a pan-flavivirus 3'UTR alignment and correlated it with the corresponding predicted RNA secondary structures. This approach revealed that intra-group-conserved DRs evolved from six long repeated sequences (LRSs) which, as approximately 200-nucleotide domains were preserved only in the genomes of the slowly evolving tick-borne flaviviruses. We propose that short DRs represent the evolutionary remnants of LRSs rather than distinct molecular duplications. The relevance of DRs to virus replication enhancer function, and thus survival, is discussed.

3' Untranslated Regions↗