PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

String analysis and energy minimization in the partition of DNA sequences.

Two approaches to the understanding of biological sequences are confronted. While the recognition of particular signals in sequences relies on complex physical interactions, the problem is often analysed in terms of the presence or absence of literal motifs (strings) in the sequence. We present here a test-case for evaluating the potential of this approach. We classify DNA sequences as positive or negative depending on whether they contain a single melted domain in the middle of the sequence, which is a global physical property. Two sets of positive "biological" sequences were generated by a computer simulation of evolutionary divergence along the branches of a phylogenetic tree, under the constraint that each intermediate sequence be positive. These two sets and a set of random positive sequences were subjected to pattern analysis. The observed local patterns were used to construct expert systems to discriminate positive from negative sequences. The experts achieved 79% to 90% success on random positive sequences and up to 99% on the biological sets, while making less than 2% errors on negative sequences. Thus, the global constraints imposed on sequences by a physical process may generate local patterns that are sufficient to predict, with a reasonable probability, the behaviour of the sequences. However, rather large sets of biological sequences are required to generate patterns free of illegitimate constraints. Furthermore, depending upon the initial sequence, the sets of sequences generated on a phylogenetic tree may be amenable or refractory to string analysis, while obeying identical physical constraints. Our study clarifies the relationship between experts' errors on positive and negative sequences, and the contributions of legitimate and illegitimate patterns to these errors. The test-case appears suitable both for further investigations of problems in the theory of sequence evolution and for further testing of pattern analysis techniques.

Base Sequence

Role of immunity in maternal-infant HIV-1 transmission.

Factors influencing human immunodeficiency virus type 1 (HIV-1) mother-to-child transmission include both immunological and virological parameters: higher viral loads have been associated with clinical stage of HIV-1-infected individuals as well as higher risk of mother-to-child transmission. Furthermore, we have shown that transmitting mothers more frequently harbour HIV-1 isolates with rapid/high syncytium-inducing (SI) biological phenotype than non-transmitting mothers do. Genetically homogeneous virus populations have been found in HIV-1-infected children at birth, in contrast to the heterogeneous virus populations often found in their infected mothers. This observation suggests that a few virus variants are transmitted or initially are replicating in the child. By comparing the HIV-1 gp120 V3 region of sequentially obtained samples from infected children with samples obtained from their mothers at delivery we found, however, that multiple variants of HIV-1 with different outgrowth kinetics can be transmitted. In addition, we have obtained results indicating an impaired ability of the immune response to adapt to the sequence evolution of HIV-1 in transmitting mothers, as assessed by measuring serum reactivities to peptides representing selected yet closely related V3 sequences. By analysing the presence of antibodies in maternal serum at delivery, which neutralize autologous isolates as well as other primary virus isolates, we have indications that a protective immunity in HIV-1 mother-to-child transmission might exist. Immunotherapy has been assessed in infected adult individuals by passive immunization with a variety of HIV-1-specific antibody products. Data from these studies indicated a differential response to therapy according to the stage of the disease. Active vaccine strategies, including envelope glycoproteins, pursued so far in seronegative adult subjects have shown limitations because broadly neutralizing antibodies, such as can be found in infected individuals, have not been evoked. Further investigations are therefore needed to give support for the potential use of either passive and/or active immunization for the prevention of HIV-1 mother-to-child transmission.

Female

A maximum-likelihood approach to analyzing nonoverlapping and overlapping reading frames.

A model is presented for sequence evolution on the basis of which one can analyze combinations of noncoding, singly coding, and multiply coding regions of aligned homologous DNA sequences. It is a generalization of Kimura's (J. Mol. Evol. 16:111-120, 1980) and Li et al.'s (J. Mol. Evol. 36:96-99, 1985) transition-transversion models with selection on replacement substitutions. Based on a hierarchy of hypotheses, one will be able to estimate selection factors and transition and transversion distances for different combinations of regions ranging from many regions, each with their private set of parameters, to one set of parameters for all regions. The method is demonstrated on two aligned HIV1 retroviruses.

Amino Acid Sequence

Phylogenetic analysis using parsimony and likelihood methods.

The assumptions underlying the maximum-parsimony (MP) method of phylogenetic tree reconstruction were intuitively examined by studying the way the method works. Computer simulations were performed to corroborate the intuitive examination. Parsimony appears to involve very stringent assumptions concerning the process of sequence evolution, such as constancy of substitution rates between nucleotides, constancy of rates across nucleotide sites, and equal branch lengths in the tree. For practical data analysis, the requirement of equal branch lengths means similar substitution rates among lineages (the existence of an approximate molecular clock), relatively long interior branches, and also few species in the data. However, a small amount of evolution is neither a necessary nor a sufficient requirement of the method. The difficulties involved in the application of current statistical estimation theory to tree reconstruction were discussed, and it was suggested that the approach proposed by Felsenstein (1981, J. Mol. Evol. 17: 368-376) for topology estimation, as well as its many variations and extensions, differs fundamentally from the maximum likelihood estimation of a conventional statistical parameter. Evidence was presented showing that the Felsenstein approach does not share the asymptotic efficiency of the maximum likelihood estimator of a statistical parameter. Computer simulations were performed to study the probability that MP recovers the true tree under a hierarchy of models of nucleotide substitution; its performance relative to the likelihood method was especially noted. The results appeared to support the intuitive examination of the assumptions underlying MP. When a simple model of nucleotide substitution was assumed to generate data, the probability that MP recovers the true topology could be as high as, or even higher than, that for the likelihood method. When the assumed model became more complex and realistic, e.g., when substitution rates were allowed to differ between nucleotides or across sites, the probability that MP recovers the true topology, and especially its performance relative to that of the likelihood method, generally deteriorates. As the complexity of the process of nucleotide substitution in real sequences is well recognized, the likelihood method appears preferable to parsimony. However, the development of a statistical methodology for the efficient estimation of the tree topology remains a difficult open problem.

Animals

The phylogenetic position of the pterobranch hemichordates based on 18S rDNA sequence data.

Pterobranchs are a class of deuterostome metazoans that are sessile marine suspension feeders. Although this group has been poorly studied, understanding their phylogenetic affinities is central to understanding early metazoan evolution. Sequence data from the 5' end of the 18S rDNA gene was collected from a pterobranch, Rhabdopleura normani, and combined with other available 18S sequences. Using standard phylogenetic methods, the evolutionary relationships of deuterostome metazoans were reconstructed. The pterobranchs are most closely related to the enteropneust hemichordates. This was confirmed by bootstrap analyses and a topology-dependent cladistic permutation tail probability (T-PTP) test. My analysis agrees with Turbeville et al.'s (1994) and Wada and Satoh's (1994) finding that hemichordates are more closely related to echinoderms than to chordates, and it is proposed that Metschnikoff's (1881) name Ambulacraria be adopted for the clade defined by the last common ancestor of the hemichordates and echinoderms. These findings suggest that ciliated gill slits and the dorsal hollow nerve chord are pleisomorphic features of the Deuterostomia.

Animals

Stock structure and homing fidelity in Gulf of Mexico sturgeon (Acipenser oxyrinchus desotoi) based on restriction fragment length polymorphism and sequence analyses of mitochondrial DNA.

Efforts have been proposed worldwide to restore sturgeon populations through the use of hatcheries to supplement natural reproduction and to reintroduce sturgeon where they have become extinct. We examined the population structure and inferred the extent of homing in the anadromous Gulf of Mexico (Gulf) sturgeon (Acipenser oxyrinchus desotoi). Restriction fragment length polymorphism and control region sequence analyses of mitochondrial DNA (mtDNA) were used to identify haplotypes of Gulf sturgeon specimens obtained from eight drainages spanning the subspecies' entire distribution from Louisiana to Florida. Significant differences in haplotype frequencies indicated substantial geographic structuring of populations. A minimum of four regional or river-specific populations were identified (from west to east): (1) Pearl River, LA and Pascagoula River, MS, (2) Escambia and Yellow rivers, FI, (3) Choctawbatchee River, FL and (4) Apalachicola Ochlockonee, and Suwannee rivers, FL. Estimates of maternally mediated gene flow between any pair of the four regional or river-specific stocks ranged between 0.15 to 1.2. Tandem repeats in the mtDNA control region of Gulf sturgeon were not perfectly conserved. This result, together with an absence of heteroplasmy and length variation in Gulf sturgeon mtDNA, indicates that the molecular mechanisms of mtDNA control region sequence evolution differ among acipenserids.

Animals

Sex determination.

The cloning of the 'testis determining gene', SRY, promised a revolution in the understanding of sex determination in humans. The failure to isolate further genes involved in sex determination has been a disappointment. The biology of SRY, however, has kept the field exciting. The discoveries of sex reversing SRY mutations with variable penetrance, but with full expressivity, rapid SRY sequence evolution and circular Sry transcripts could not have been predicted.

Amino Acid Sequence

Exploring the Mitochondrial Genomes of Phoebe Species (Lauraceae): Structural Dynamics and Functional Conservation.

Plant mitochondrial genomes (mitogenomes) vary markedly in size and architecture despite generally slow rates of sequence evolution. Phoebe is an ecologically and economically valuable genus of Lauraceae, yet its mitogenome diversity remains poorly characterized. In this study, we newly sequenced, assembled, and annotated the mitogenomes of three nationally protected Class II wild plants (P. bournei, P. chekiangensis, P. zhennan) from China and compared their mitogenomic characteristics. The three assemblies were resolved into representative circular configurations ranging from 808 to 864 kb, with similar GC contents and conserved protein-coding capacity. Each mitogenome contained distinct 41 protein-coding genes, 27-28 transfer RNAs, and three ribosomal RNAs. Synteny analysis revealed extensive changes in homologous-block order and orientation despite substantial sequence homology among the three species. Abundant repeats occurred predominantly in noncoding regions, while plastid-derived fragments documented historical intracellular DNA transfer. The three species exhibited similar codon usage and predicted RNA-editing patterns, whereas low synonymous divergence limited inference from pairwise ratios. Phylogenetic analysis based on mitochondrial protein-coding genes recovered Phoebe as a well-supported monophyletic lineage. These results reveal substantial structural divergence accompanied by conserved nucleotide composition and coding capacity, providing valuable data for further understanding the evolutionary variation of plant mitogenomes of Phoebe and the Lauraceae.

Phoebe

Comparative mitogenomics of Ocnus glacialis reveals lineage-specific evolutionary rates and complex gene rearrangements in Dendrochirotida.

The order Dendrochirotida (Class Holothuroidea) is a species-rich echinoderm group, yet its internal evolutionary history remains poorly resolved due to limited mitogenomic resources. In this study, we characterized the first complete mitochondrial genome of Ocnus glacialis and conducted comparative analyses to elucidate its phylogenetic position and molecular evolutionary patterns. The circular mitogenome of O. glacialis is 16,776 bp in length, containing the canonical set of 37 genes. Among the analyzed dendrochirotids, O. glacialis exhibited the highest A + T content (70.88%) and a near-zero AT-skew, a compositional profile often linked to lineage-specific evolution in specialized environments. Selection pressure analyses, including branch-model tests, revealed that these compositional features are associated with relaxed purifying selection and an accelerated rate of sequence evolution. Branch-site analyses further identified specific codon sites in cytb, nad2, nad4l, nad5, and nad6 under positive or relaxed constraints. Structurally, O. glacialis displayed the most complex gene rearrangement pattern among the studied species, characterized by multiple tandem duplication-random loss (TDRL) events and extensive intergenic sequences. Furthermore, divergence time estimation suggests that these structural and compositional shifts occurred in tandem with the lineage's diversification. We propose that these mitogenomic signatures reflect a synergistic outcome of habitat transition toward Arctic cold-water and deep-sea environments, coupled with demographic factors such as reduced effective population sizes inherent to its benthic life history. By resolving taxonomic uncertainties, this study provides a robust temporal and molecular framework for understanding the evolutionary history and ecological diversification of the Ocnus lineage.

Animals

Modeling residue usage in aligned protein sequences via maximum likelihood.

A computational method is presented for characterizing residue usage, i.e., site-specific residue frequencies, in aligned protein sequences. The method obtains frequency estimates that maximize the likelihood of the sequences in a simple model for sequence evolution, given a tree or a set of candidate trees computed by other methods. These maximum-likelihood frequencies constitute a profile of the sequences, and thus the method offers a rigorous alternative to sequence weighting for constructing such a profile. The ability of this method to discard misleading phylogenetic effects allows the biochemical propensities of different positions in a sequence to be more clearly observed and interpreted.

Amino Acids

Length variation and secondary structure of introns in the Mlc1 gene in six species of Drosophila.

A nearly universal feature of intron sequences is that even closely related species exhibit a large number of insertion/deletion differences. The goal of the analysis described here is to test whether the observed pattern of insertion/deletion events in the genealogy of the myosin alkali light chain (Mlc1) gene is consistent with neutrality, and if not, to determine the underlying forces of evolutionary change. Mlc1 pre-mRNA is alternatively spliced, and one constraint is that signals necessary for tissue-specificity of directed splicing must be conserved. If the total length of an intron is functionally constrained, then the distribution of indels on branches of the gene genealogy should reflect a departure from randomness. Here we perform a phylogenetic analysis, inferring ancestral states wherever possible on a phylogeny of 29 alleles of Mlc1 from six species of Drosophila. Observed patterns of indels on the genealogy were compared to those from simulated data, with the result that we cannot reject the null hypothesis of neutrality. A clear departure from a neutral prediction was seen in the excess folding free energy predicted for the introns flanking the alternatively spliced exon. Relative rate tests also suggest a retardation in the rate of Mlc1 sequence evolution in the simulans clade.

Animals

DIPLOMO: the tool for a new type of evolutionary analysis.

A package of computer programs called DIPLOMO (DIstance PLOt MOnitor) has been developed for making pairwise comparisons of different estimates of the distances between a set of taxa by plotting them against each other in a simple scatter plot. Taxa with similar relative distance characteristics are thereby grouped graphically. Groupings of different taxa may be directly identified, and the distance characteristics of chosen groups visualised and compared using devices to give them different colours or symbols. The program is particularly useful for detecting and analysing subtle trends in gene sequence evolution. This is done by comparing different components of change, for example synonymous versus non-synonymous nucleotide changes, transversions versus transitions and changes in different genes of the same set of taxa, etc. The program has a wide range of other uses, for example comparing different methods of sequence analysis, assessing which components of genetic change correlate best with phenotypic change or with geographical separation. This paper describes the DIPLOMO package, and illustrates typical DIPLOMO analyses using lentivirus gene sequence data.

Biological Evolution

Strand asymmetries in DNA evolution.

The complementary strands of DNA differ with respect to replication and transcription. Both of these processes are asymmetric and can bias the occurrence of mutations between the strands: during replication, the discontinuous lagging strand undergoes certain errors at higher rates, and transcription overexposes the nontranscribed strand to DNA damage while targeting repair enzymes to the transcribed strand. While biases introduced during replication apparently have little impact on sequence evolution, the effects of transcription are observed in the asymmetric patterns of substitution in bacterial genes and might be influencing genome-wide patterns of base composition.

DNA

Molecular evolution of ruminant lysozymes.

The evolution of a new digestive enzyme, stomach lysozyme, from an antibacterial host defense enzyme provides a link between molecular evolution and organismal evolution. Lysozymes have been recruited at least three times (twice from a conventional lysozyme c and once from a calcium-binding lysozyme c) in vertebrates for functioning in the stomach. The recruitment of lysozyme for its new biological function involved many molecular changes, beyond those required to adapt the protein to function in the stomach. The evolution of the stomach lysozyme gene has been extensively studied in ruminant artiodactyls. In ruminants, the lysozyme c gene has duplicated to yield a family of about ten genes. These duplications allowed: (1) specialization of gene function and (2) increased levels of expression. The ruminant stomach lysozyme genes have evolved in an episodic fashion - there was a period of rapid adaptive sequence evolution, driven by positive selection in the early ruminant, that was followed by an increase in purifying selection upon the well-adapted stomach lysozyme sequence among modern species. Recombination of small portions (exons) of the genes between members of the lysozyme gene family may have aided in adaptive evolution. Evolution to a stomach lysozyme is not irreversible; at least one member of the ruminant stomach lysozyme gene family appears to have reverted to a more ancestral function, yet retains hallmarks of its history as a stomach lysozyme.

Animals

Origin of the NEFA and Nuc signal sequences.

The human protein NEFA binds calcium, contains a leucine zipper repeat that does not form a homodimer, and is proposed (along with the homologous Nuc protein) to have a common evolutionary history with an EF-hand ancestor. We have isolated and characterized the N-terminal domain of NEFA that contains a signal sequence inferred from both endoproteinase Asp-N (Asp-N) and tryptic digests. Analysis of this N-terminal sequence shows significant similarity to the conserved multiple domains of the mitochondrial carrier family (MCF) proteins. The leader sequence of Nuc is, however, most similar to the signal sequences of membrane and/or secreted proteins (e.g., mouse insulin-like growth factor receptor). We suggest that the divergent NEFA and Nuc N-terminal sequences may have independent origins and that the common high hydrophobicity governs their targeting to the ER. These results provide insights into signal sequence evolution and the multiple origins of protein targeting.

Amino Acid Sequence

Ross River virus genetic variants in Australia and the Pacific Islands.

HaeIII and TaqI restriction digest profiles of cDNA to infected cell RNA or virion RNA were used as a guide to genetic relationships between fourteen isolates of Ross River virus (RRV) obtained from mosquitoes collected in various localities in eastern Australia where the virus is endemic. RRV isolates from Fiji, American Samoa, the Cook Islands and the Wallis Islands where major outbreaks of epidemic polyarthritis took place in 1979-1980 were also examined. Among these RRV isolates we have identified three genetic types (I-III) on the basis of differences between their restriction digest profiles. We estimate that 1.5-5% nucleotide sequence diversity exists between genetic types. Within each genetic type strain differentiation gave rise to small but significant differences in restriction digest profiles. No clear pattern of geographic distribution of RRV genetic types could be established from the limited number of RRV isolates examined. Genetic types I, II and III, respectively, were isolated from three, three and one different mosquito species, indicating there is no strong association between genetic type and the species of mosquito vector. HaeIII restriction digest analysis did not detect any genetic difference between the four Pacific Island isolates, suggesting that a single RRV variant was involved in the epidemics. Genetically, this variant was closely related to isolates of genetic type II. Virtually identical HaeIII restriction digest profiles were observed for isolates obtained at various stages of the Pacific Island epidemics, suggesting that extensive sequence evolution did not accompany Ross River virus spread.

Alphavirus

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density

Restriction fragment polymorphism in the sex-determining region of the Y chromosomal DNA of European wild mice.

Using 32P-labeled probe consisting mainly of (GATA)n we have shown that a male specific Alu1 DNA blot pattern which defines the Y chromosome sex-determining locus in inbred mice is highly polymorphic in wild mice, indicating substantial sequence evolution in this region under field conditions. In all cases examined by in situ hybridization, the region concerned is paracentromeric. In contrast, the blot pattern of another probe (M 34) which detects repeated sequences specific to the mouse Y chromosome but outside the sex-determining locus, remains constant between different isolates.

Animals