PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Precise sequence assignment of replication origin in the control region of chick mitochondrial DNA relative to 5' and 3' D-loop ends, secondary structure, DNA synthesis, and protein binding.

The data reported identify for the first time the sequence of an avian mitochondrial heavy-strand replication origin, OH, located only about 12 nucleotides (nt) downstream from the conserved sequence block CSB-1, as well as the sequence of premature synthesis arrest of the 781 (+/-1) nt D-loop strand, only 6-7 nt downstream from a TAS-like (termination-associated) element. Both sites are associated with putative cruciform secondary structures. A major sequence-specific DNA-binding/cleavage site of a potential regulatory protein, the approximately 36-kDa aMDP1 (shown previously to stimulate mtDNA synthesis), is located about 90 nt upstream of OH. Correlated in vivo analysis of avian genome-length mtDNA replication provides missing evidence on the functional equivalence of D-loop origin with nascent initiation, and on the direction, asymmetry and temporal aspects of a full round of replication. The importance of the results to understanding the regulation of linked replication/transcription and the unusual sequence evolution of avian mtDNA is

Animals

Kinetoplast DNA minicircles: regions of extensive sequence divergence.

Previous work has shown that the kinetoplast minicircle DNA of Leishmania species exhibits species-specific sequence divergence and this observation has led to the development of a DNA probe-based diagnostic test for leishmaniasis. In the work reported here, we demonstrate that the minicircle is composed of three types of DNA sequences with differing specificities reflecting different rates of DNA sequence change. A library of cloned fragments of kinetoplast DNA (kDNA) from Leishmania mexicana amazonensis was prepared and the cloned subfragments were found to contain DNA sequences with different taxonomic specificities based on hybridization analysis with various species of Leishmania. Four groups of subfragments were found, those that hybridized with a large number of Leishmania sp. as well as sequences unique to the species, subspecies, or isolate. Analysis of nested deletions of a single, full-length minicircle demonstrates that these different taxonomic specificities are contained within a single minicircle. This implies that different regions of a single minicircle have DNA sequences that diverge at different rates. These sequences represent potentially valuable tools in diagnostic, epidemiologic, and ecological studies of leishmaniasis and provide the basis for a model of kDNA sequence evolution.

Animals

Analytical expression of the purine/pyrimidine autocorrelation function after and before random mutations.

The mutation process is a classical evolutionary genetic process. The type of mutations studied here is the random substitutions of a purine base R (adenine or guanine) by a pyrimidine base Y (cytosine or thymine) and reciprocally (transversions). The analytical expressions derived allow us to analyze in genes the occurrence probabilities of motifs and d-motifs (two motifs separated by any d bases) on the R/Y alphabet under transversions. These motif probabilities can be obtained after transversions (in the evolutionary sense; from the past to the present) and, unexpectedly, also before transversions (after back transversions, in the inverse evolutionary sense, from the present to the past). This theoretical part in Section 2 is a first generalization of a particular formula recently derived. The application in Section 3 is based on the analytical expression giving the autocorrelation function (the d-motif probabilities) before transversions. It allows us to study primitive genes from actual genes. This approach solves a biological problem. The protein coding genes of chloroplasts and mitochondria have a preferential occurrence of the 6-motif YRY(N)6YRY (maximum of the autocorrelation function for d = 6, N = R or Y) with a periodicity modulo 3. The YRY(N)6YRY preferential occurrence without the periodicity modulo 3 is also observed in the RNA coding genes (ribosomal, transfer, and small nuclear RNA genes) and in the noncoding genes (introns and 5' regions of eukaryotic nuclei). However, there are two exceptions to this YRY(N)6YRY rule: the protein coding genes of eukaryotic nuclei, and prokaryotes, where YRY(N)6YRY has the second highest value after YRY(N)0YRY (YRYYRY) with a periodicity modulo 3. When we go backward in time with the analytical expression, the protein coding genes of both eukaryotic nuclei and prokaryotes retrieve the YRY(N)6YRY preferential occurrence with a periodicity modulo 3 after 0.2 back transversions per base. In other words, the actual protein coding genes of chloroplasts and mitochondria are similar to the primitive protein coding genes of eukaryotic nuclei and prokaryotes. On the other hand, this application represents the first result concerning the mutation process in the model of DNA sequence evolution we recently proposed. According to this model, the actual genes on the R/Y alphabet derive from two successive evolutionary genetic processes: an independent mixing of a few nonrandom types of oligonucleotides leading to genes called primitive followed by a mutation process in these primitive genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence

Statistical tests of models of DNA substitution.

Penny et al. have written that "The most fundamental criterion for a scientific method is that the data must, in principle, be able to reject the model. Hardly any [phylogenetic] tree-reconstruction methods meet this simple requirement." The ability to reject models is of such great importance because the results of all phylogenetic analyses depend on their underlying models--to have confidence in the inferences, it is necessary to have confidence in the models. In this paper, a test statistic suggested by Cox is employed to test the adequacy of some statistical models of DNA sequence evolution used in the phylogenetic inference method introduced by Felsenstein. Monte Carlo simulations are used to assess significance levels. The resulting statistical tests provide an objective and very general assessment of all the components of a DNA substitution model; more specific versions of the test are devised to test individual components of a model. In all cases, the new analyses have the additional advantage that values of phylogenetic parameters do not have to be assumed in order to perform the tests.

Animals

Structure and sequence variation of the genes encoding the polymorphic, immunodominant molecule (PIM), an antigen of Theileria parva recognized by inhibitory monoclonal antibodies.

The polymorphic, immunodominant molecule (PIM) of Theileria parva is the predominant antigen recognized by sera from infected cattle and by monoclonal antibodies (mAb) used to differentiate parasite strains. As such, the antigen is under consideration as a diagnostic antigen, and since the mAbs can neutralize sporozoite infectivity in vitro, in immunization experiments. Initial comparison of two PIM cDNA sequences suggested that the PIM genes consist of conserved 5' and 3' termini flanking a central variable region. We present further evidence, based on sequence analysis, supporting this general structure for the PIM genes. Evidence is also presented for a single copy of the PIM gene per haploid genome, implying that the different versions of PIM are encoded by distinct alleles. The central variable region of the PIM allele from the T. parva (Marikebuni) stock was found to contain 13 copies of the tetrapeptide repeat Gln-Pro-Glu-Pro. We also detected point mutations in the 5' and 3' termini of the PIM alleles, including regions recognized by the neutralizing and typing mAb. This contrasted with the high sequence conservation of the two introns of the genes, suggesting that the protein is undergoing rapid evolution. Sequence comparison of PIM genes from buffalo- and cattle-derived parasites supported earlier results that the parasites infecting buffaloes constitute a more heterogeneous population than those from cattle.

Amino Acid Sequence

Heterozygosity, heteromorphy, and phylogenetic trees in asexual eukaryotes.

Little attention has been paid to the consequences of long-term asexual reproduction for sequence evolution in diploid or polyploid eukaryotic organisms. Some elementary theory shows that the amount of neutral sequence divergence between two alleles of a protein-coding gene in an asexual individual will be greater than that in a sexual species by a factor of 2tu, where t is the number of generations since sexual reproduction was lost and u is the mutation rate per generation in the asexual lineage. Phylogenetic trees based on only one allele from each of two or more species will show incorrect divergence times and, more often than not, incorrect topologies. This allele sequence divergence can be stopped temporarily by mitotic gene conversion, mitotic crossing-over, or ploidy reduction. If these convergence events are rare, ancient asexual lineages can be recognized by their high allele sequence divergence. At intermediate frequencies of convergence events, it will be impossible to reconstruct the correct phylogeny of an asexual clade from the sequences of protein coding genes. Convergence may be limited by allele sequence divergence and heterozygous chromosomal rearrangements which reduce the homology needed for recombination and result in aneuploidy after crossing-over or ploidy cycles.

Eukaryotic Cells

Evolution of repeated sequences in non-coding regions of the genome.

Repeated sequences are found ubiquitously in the eukaryotic genome. Population genetic studies on the evolution of such repeated sequences are reviewed while paying special attention to those sequences found in the non-coding regions of the genome. Specifically, the evolution of dispersed repeated sequences by the transposition as well as the evolution of short tandemly repeated sequences due to either replication slippage or unequal sister chromatid exchange are considered. The approach of combining both model and data analyses which has been successfully employed in the development of the neutral theory is also considered to be useful in better understanding the evolution and biological meaning of these sequences.

Animals

Purification and characterization of a peptide from amyloid-rich pancreases of type 2 diabetic patients.

Deposition of amyloid in pancreatic islets is a common feature in human type 2 diabetic subjects but because of its insolubility and low tissue concentrations, the structure of its monomer has not been determined. We describe a peptide, of calculated molecular mass 3905 Da, that was a major protein component of amyloid-rich pancreatic extracts of three type 2 diabetic patients. After collagenase treatment, an extract containing 20-50% amyloid was solubilized by sonication into 70% formic acid and the peptide was purified by gel filtration followed by reverse-phase high-performance liquid chromatography. We term this peptide diabetes-associated peptide, as it was not detected in extracts of pancreas from any of six normal subjects. Diabetes-associated peptide contains 37 amino acids and is 46% identical to the sequences of rat and human calcitonin gene-related peptide, indicating that these peptides are related in evolution. Sequence identities with conserved residues of the insulin A chain were also seen in a 16-residue segment. On extraction, the islet amyloid is particulate and insoluble like the core particles of Alzheimer disease. Their monomers have similar molecular masses, each having a hydropathic region that can probably form beta-pleated sheets. The accumulation of amyloid, including diabetes-associated peptide, in islets may impair islet function in type 2 diabetes mellitus.

Aged

Nucleotide sequence, polymorphism, and evolution of ovine MHC class II DQA genes.

The nucleotide sequence of all exons and introns, excluding exon 1, of the ovine major histocompatibility complex (MhcOvar) genes analogous to the HLA-DQA1 and -DQA2 genes has been determined and the gene structure found to be similar to that reported for other species. The predicted amino acid sequences of the Ovar-DQA genes have been compared with the equivalent DQA genes in man, mouse, rat, rabbit, and cattle and used to determine the evolutionary relationships of the sheep class II genes to these other species. Northern blot analysis of sheep mRNA using exon specific probes for each of the two Ovar-DQA genes show that both genes are transcribed, whereas in humans there is no evidence that HLA-DQA2 is transcriptionally active. Restriction fragment length polymorphisms (RFLPs) have been used to define a polymorphic series of alleles in both Ovar-DQA genes and have indicated that the number of DQA genes is not constant in sheep as it is in humans, but varies with the haplotype.

Amino Acid Sequence

String analysis and energy minimization in the partition of DNA sequences.

Two approaches to the understanding of biological sequences are confronted. While the recognition of particular signals in sequences relies on complex physical interactions, the problem is often analysed in terms of the presence or absence of literal motifs (strings) in the sequence. We present here a test-case for evaluating the potential of this approach. We classify DNA sequences as positive or negative depending on whether they contain a single melted domain in the middle of the sequence, which is a global physical property. Two sets of positive "biological" sequences were generated by a computer simulation of evolutionary divergence along the branches of a phylogenetic tree, under the constraint that each intermediate sequence be positive. These two sets and a set of random positive sequences were subjected to pattern analysis. The observed local patterns were used to construct expert systems to discriminate positive from negative sequences. The experts achieved 79% to 90% success on random positive sequences and up to 99% on the biological sets, while making less than 2% errors on negative sequences. Thus, the global constraints imposed on sequences by a physical process may generate local patterns that are sufficient to predict, with a reasonable probability, the behaviour of the sequences. However, rather large sets of biological sequences are required to generate patterns free of illegitimate constraints. Furthermore, depending upon the initial sequence, the sets of sequences generated on a phylogenetic tree may be amenable or refractory to string analysis, while obeying identical physical constraints. Our study clarifies the relationship between experts' errors on positive and negative sequences, and the contributions of legitimate and illegitimate patterns to these errors. The test-case appears suitable both for further investigations of problems in the theory of sequence evolution and for further testing of pattern analysis techniques.

Base Sequence

A maximum-likelihood approach to analyzing nonoverlapping and overlapping reading frames.

A model is presented for sequence evolution on the basis of which one can analyze combinations of noncoding, singly coding, and multiply coding regions of aligned homologous DNA sequences. It is a generalization of Kimura's (J. Mol. Evol. 16:111-120, 1980) and Li et al.'s (J. Mol. Evol. 36:96-99, 1985) transition-transversion models with selection on replacement substitutions. Based on a hierarchy of hypotheses, one will be able to estimate selection factors and transition and transversion distances for different combinations of regions ranging from many regions, each with their private set of parameters, to one set of parameters for all regions. The method is demonstrated on two aligned HIV1 retroviruses.

Amino Acid Sequence

Phylogenetic analysis using parsimony and likelihood methods.

The assumptions underlying the maximum-parsimony (MP) method of phylogenetic tree reconstruction were intuitively examined by studying the way the method works. Computer simulations were performed to corroborate the intuitive examination. Parsimony appears to involve very stringent assumptions concerning the process of sequence evolution, such as constancy of substitution rates between nucleotides, constancy of rates across nucleotide sites, and equal branch lengths in the tree. For practical data analysis, the requirement of equal branch lengths means similar substitution rates among lineages (the existence of an approximate molecular clock), relatively long interior branches, and also few species in the data. However, a small amount of evolution is neither a necessary nor a sufficient requirement of the method. The difficulties involved in the application of current statistical estimation theory to tree reconstruction were discussed, and it was suggested that the approach proposed by Felsenstein (1981, J. Mol. Evol. 17: 368-376) for topology estimation, as well as its many variations and extensions, differs fundamentally from the maximum likelihood estimation of a conventional statistical parameter. Evidence was presented showing that the Felsenstein approach does not share the asymptotic efficiency of the maximum likelihood estimator of a statistical parameter. Computer simulations were performed to study the probability that MP recovers the true tree under a hierarchy of models of nucleotide substitution; its performance relative to the likelihood method was especially noted. The results appeared to support the intuitive examination of the assumptions underlying MP. When a simple model of nucleotide substitution was assumed to generate data, the probability that MP recovers the true topology could be as high as, or even higher than, that for the likelihood method. When the assumed model became more complex and realistic, e.g., when substitution rates were allowed to differ between nucleotides or across sites, the probability that MP recovers the true topology, and especially its performance relative to that of the likelihood method, generally deteriorates. As the complexity of the process of nucleotide substitution in real sequences is well recognized, the likelihood method appears preferable to parsimony. However, the development of a statistical methodology for the efficient estimation of the tree topology remains a difficult open problem.

Animals

The phylogenetic position of the pterobranch hemichordates based on 18S rDNA sequence data.

Pterobranchs are a class of deuterostome metazoans that are sessile marine suspension feeders. Although this group has been poorly studied, understanding their phylogenetic affinities is central to understanding early metazoan evolution. Sequence data from the 5' end of the 18S rDNA gene was collected from a pterobranch, Rhabdopleura normani, and combined with other available 18S sequences. Using standard phylogenetic methods, the evolutionary relationships of deuterostome metazoans were reconstructed. The pterobranchs are most closely related to the enteropneust hemichordates. This was confirmed by bootstrap analyses and a topology-dependent cladistic permutation tail probability (T-PTP) test. My analysis agrees with Turbeville et al.'s (1994) and Wada and Satoh's (1994) finding that hemichordates are more closely related to echinoderms than to chordates, and it is proposed that Metschnikoff's (1881) name Ambulacraria be adopted for the clade defined by the last common ancestor of the hemichordates and echinoderms. These findings suggest that ciliated gill slits and the dorsal hollow nerve chord are pleisomorphic features of the Deuterostomia.

Animals

Stock structure and homing fidelity in Gulf of Mexico sturgeon (Acipenser oxyrinchus desotoi) based on restriction fragment length polymorphism and sequence analyses of mitochondrial DNA.

Efforts have been proposed worldwide to restore sturgeon populations through the use of hatcheries to supplement natural reproduction and to reintroduce sturgeon where they have become extinct. We examined the population structure and inferred the extent of homing in the anadromous Gulf of Mexico (Gulf) sturgeon (Acipenser oxyrinchus desotoi). Restriction fragment length polymorphism and control region sequence analyses of mitochondrial DNA (mtDNA) were used to identify haplotypes of Gulf sturgeon specimens obtained from eight drainages spanning the subspecies' entire distribution from Louisiana to Florida. Significant differences in haplotype frequencies indicated substantial geographic structuring of populations. A minimum of four regional or river-specific populations were identified (from west to east): (1) Pearl River, LA and Pascagoula River, MS, (2) Escambia and Yellow rivers, FI, (3) Choctawbatchee River, FL and (4) Apalachicola Ochlockonee, and Suwannee rivers, FL. Estimates of maternally mediated gene flow between any pair of the four regional or river-specific stocks ranged between 0.15 to 1.2. Tandem repeats in the mtDNA control region of Gulf sturgeon were not perfectly conserved. This result, together with an absence of heteroplasmy and length variation in Gulf sturgeon mtDNA, indicates that the molecular mechanisms of mtDNA control region sequence evolution differ among acipenserids.

Animals

Sex determination.

The cloning of the 'testis determining gene', SRY, promised a revolution in the understanding of sex determination in humans. The failure to isolate further genes involved in sex determination has been a disappointment. The biology of SRY, however, has kept the field exciting. The discoveries of sex reversing SRY mutations with variable penetrance, but with full expressivity, rapid SRY sequence evolution and circular Sry transcripts could not have been predicted.

Amino Acid Sequence

Exploring the Mitochondrial Genomes of Phoebe Species (Lauraceae): Structural Dynamics and Functional Conservation.

Plant mitochondrial genomes (mitogenomes) vary markedly in size and architecture despite generally slow rates of sequence evolution. Phoebe is an ecologically and economically valuable genus of Lauraceae, yet its mitogenome diversity remains poorly characterized. In this study, we newly sequenced, assembled, and annotated the mitogenomes of three nationally protected Class II wild plants (P. bournei, P. chekiangensis, P. zhennan) from China and compared their mitogenomic characteristics. The three assemblies were resolved into representative circular configurations ranging from 808 to 864 kb, with similar GC contents and conserved protein-coding capacity. Each mitogenome contained distinct 41 protein-coding genes, 27-28 transfer RNAs, and three ribosomal RNAs. Synteny analysis revealed extensive changes in homologous-block order and orientation despite substantial sequence homology among the three species. Abundant repeats occurred predominantly in noncoding regions, while plastid-derived fragments documented historical intracellular DNA transfer. The three species exhibited similar codon usage and predicted RNA-editing patterns, whereas low synonymous divergence limited inference from pairwise ratios. Phylogenetic analysis based on mitochondrial protein-coding genes recovered Phoebe as a well-supported monophyletic lineage. These results reveal substantial structural divergence accompanied by conserved nucleotide composition and coding capacity, providing valuable data for further understanding the evolutionary variation of plant mitogenomes of Phoebe and the Lauraceae.

Phoebe

Comparative mitogenomics of Ocnus glacialis reveals lineage-specific evolutionary rates and complex gene rearrangements in Dendrochirotida.

The order Dendrochirotida (Class Holothuroidea) is a species-rich echinoderm group, yet its internal evolutionary history remains poorly resolved due to limited mitogenomic resources. In this study, we characterized the first complete mitochondrial genome of Ocnus glacialis and conducted comparative analyses to elucidate its phylogenetic position and molecular evolutionary patterns. The circular mitogenome of O. glacialis is 16,776 bp in length, containing the canonical set of 37 genes. Among the analyzed dendrochirotids, O. glacialis exhibited the highest A + T content (70.88%) and a near-zero AT-skew, a compositional profile often linked to lineage-specific evolution in specialized environments. Selection pressure analyses, including branch-model tests, revealed that these compositional features are associated with relaxed purifying selection and an accelerated rate of sequence evolution. Branch-site analyses further identified specific codon sites in cytb, nad2, nad4l, nad5, and nad6 under positive or relaxed constraints. Structurally, O. glacialis displayed the most complex gene rearrangement pattern among the studied species, characterized by multiple tandem duplication-random loss (TDRL) events and extensive intergenic sequences. Furthermore, divergence time estimation suggests that these structural and compositional shifts occurred in tandem with the lineage's diversification. We propose that these mitogenomic signatures reflect a synergistic outcome of habitat transition toward Arctic cold-water and deep-sea environments, coupled with demographic factors such as reduced effective population sizes inherent to its benthic life history. By resolving taxonomic uncertainties, this study provides a robust temporal and molecular framework for understanding the evolutionary history and ecological diversification of the Ocnus lineage.

Animals

Length variation and secondary structure of introns in the Mlc1 gene in six species of Drosophila.

A nearly universal feature of intron sequences is that even closely related species exhibit a large number of insertion/deletion differences. The goal of the analysis described here is to test whether the observed pattern of insertion/deletion events in the genealogy of the myosin alkali light chain (Mlc1) gene is consistent with neutrality, and if not, to determine the underlying forces of evolutionary change. Mlc1 pre-mRNA is alternatively spliced, and one constraint is that signals necessary for tissue-specificity of directed splicing must be conserved. If the total length of an intron is functionally constrained, then the distribution of indels on branches of the gene genealogy should reflect a departure from randomness. Here we perform a phylogenetic analysis, inferring ancestral states wherever possible on a phylogeny of 29 alleles of Mlc1 from six species of Drosophila. Observed patterns of indels on the genealogy were compared to those from simulated data, with the result that we cannot reject the null hypothesis of neutrality. A clear departure from a neutral prediction was seen in the excess folding free energy predicted for the introns flanking the alternatively spliced exon. Relative rate tests also suggest a retardation in the rate of Mlc1 sequence evolution in the simulans clade.

Animals