PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Origins of chromosomal rearrangement hotspots in the human genome: evidence from the AZFa deletion hotspots.

BACKGROUND: The origins of the recombination hotspots that are a common feature of both allelic and non-allelic homologous recombination in the human genome are poorly understood. We have investigated, by comparative sequencing, the evolution of two hotspots of non-allelic homologous recombination on the Y chromosome that lie within paralogous sequences known to sponsor deletions resulting in male infertility. RESULTS: These recombination hotspots are characterized by signatures of concerted evolution, which indicate that gene conversion between paralogs has been predominant in shaping their recent evolution. By contrast, the paralogous sequences that surround the hotspots exhibit little evidence of gene conversion. A second feature of these rearrangement hotspots is the extreme interspecific sequence divergence (around 2.5%) that places them among the most divergent orthologous sequences between humans and chimpanzees. CONCLUSIONS: Several hominid-specific gene conversion events have rendered these hotspots better substrates for chromosomal rearrangements in humans than in chimpanzees or gorillas. Monte Carlo simulations of sequence evolution suggest that extreme sequence divergence is a direct consequence of gene conversion between paralogs. We propose that the coincidence of signatures of concerted evolution and recurrent breakpoints of chromosomal rearrangement (mapped at the sequence level) may enable the identification of putative rearrangement hotspots from analysis of comparative sequences from great apes.

Alleles↗

The amelogenin story: origin and evolution.

Genome sequencing and gene mapping have permitted the identification of HEVIN (SPARC-Like1) as the probable ancestor of the enamel matrix proteins (EMPs), amelogenin (AMEL), ameloblastin (AMBN) and enamelin (ENAM). We have undertaken a phylogenetic analysis to elucidate their relationships. AMEL genes available in databases, and new sequences obtained in blast searching genomes or expressed sequence tags, were compiled (22 full-length sequences), aligned, and the ancestral sequence calculated and used to search for similarities using psi-blast. Hits were obtained with the N-terminal region of AMBN, ENAM, and HEVIN. We retrieved all available AMBN (n=8), ENAM (n=3), and HEVIN (n=4) sequences. The sequences of the four proteins were aligned and analyzed phylogenetically. AMEL and AMBN are sister genes, which diverged after duplication of a common ancestor issued from ENAM. The latter derived from a copy of HEVIN. Comparisons of gene organization, amino acid sequences and location of ENAM and AMBN, adjacent on the same chromosome, suggest that AMBN is closer to ENAM than AMEL. This supports AMEL as being derived from AMBN duplication. This duplication occurred long before tetrapod differentiation, probably in an ancestral osteichthyan. The story of AMEL origin is completed as follows: SPARC-->HEVIN-->ENAM-->AMBN-->AMEL.

Amelogenin↗

Evolution of nuclear rDNA ITS sequences in the Cladophora albida/sericea clade (Chlorophyta).

Ribosomal DNA ITS sequences were compared among 13 different species and biogeographic isolates from the monophyletic "albida/sericea clade" in the green algal genus Cladophora. Six distinct ITS sequence types were found, characterized by multiple insertions and deletions and high levels of nucleotide substitution. Conserved domains within the ITS regions indicate the presence of ITS secondary structure. Low transition/transversion ratios among the six types and nearly symmetrical tree-length frequency distributions indicate some saturation, and low phylogenetic signal. Although branching order among five of the six ITS sequence types could not be resolved, estimates of ITS sequence divergence as compared with 18S divergence in a subset of the taxa suggests that the origin of the different ITS types is probably in the mid-Miocene (12 Ma ago) but that biogeographic isolates within a single ITS type (including both Pacific and Atlantic representatives) have probably dispersed on a time scale of thousands rather than millions of years.

Base Sequence↗

Molecular evolution of human paramyxoviruses. Nucleotide sequence analyses of the human parainfluenza type 1 virus NP and M protein genes and construction of phylogenetic trees for all the human paramyxoviruses.

The nucleotide sequences of the NP and M genes of human parainfluenza type 1 virus (HPIV-1) were determined. The NP gene was 1677 nucleotides long excluding polyadenylic acid. The NP gene contained a single large open reading frame (ORF), which encoded a polypeptide of 524 amino acids with a calculated molecular weight of 57,736. The M gene 1173 nucleotides long excluding the poly(A) tract and the sequence also contained a single large ORF which encoded a polypeptide of 348 amino acid with a molecular weight of 38,445, which was inconsistent with 28 kDa previously determined by SDS-PAGE. We aligned the deduced HPIV-1 NP and M protein sequences with 12 and 13 other paramyxoviruses, respectively, suggesting that a common tertiary structure was found in the NPs or Ms of HPIV-1, Sendai virus (SV), HPIV-3 and BPIV-3 and that other common structure was also maintained in these proteins of HPIV-2, SV 41 and 5, MuV, HPIV-4. Phylogenetic trees were constructed for the NP and M proteins of all the paramyxoviruses of which nucleotide sequences had been previously reported. Paramyxoviruses could be subdivided into two groups, i.e., PIV-1 group and PIV-2 group; the former group is composed of HPIV-1, SV, HPIV-3 and BPIV-3, and the latter group consists of HPIV-2, SV 41, SV 5, MuV, HPIV-4 A and HPIV-4 B.

Amino Acid Sequence↗

Evolution of influenza polymerase: nucleotide sequence of the PB2 gene of A/Chile/1/83 (H1 N1).

The complete nucleotide sequence of the PB2 gene of influenza virus A/Chile/1/83 (H1 N1) is presented. Sequence comparison between A/Chile PB2 protein and the known PB2 sequences of the influenza strains A/WSN/33 (H1 N1), A/PR/8/34 (H1 N1), A/NT/60/68 (H3 N2), A/Kiev/59/79 (H1 N1), A/FPV/Rostock/34 (H7 N1), and B/Ann Arbor/1/66 indicates extensive amino acid homology for the influenza A virus PB2 proteins. Small clusters of basic amino acids are conserved in all PB2 proteins including the influenza B PB2 protein which has only 39% sequence homology overall to the PB2 polypeptides of type A influenza viruses. The evolutionary rate of 5.7 x 10(-3) nucleotide substitutions per site per year and 0.25% amino acid changes per year between the A/Chile/1/83 and A/NT/60/68 PB2 appears to be higher than that calculated earlier for A/NT, A/PR/8 and A/WSN. An unusually high degree of sequence change between A/Chile/1/83 and A/Kiev/59/79 PB2 polymerase was revealed and this is discussed in terms of its probable origin.

Amino Acid Sequence↗

Recurrent duplication-driven transposition of DNA during hominoid evolution.

The underlying mechanism by which the interspersed pattern of human segmental duplications has evolved is unknown. Based on a comparative analysis of primate genomes, we show that a particular segmental duplication (LCR16a) has been the source locus for the formation of the majority of intrachromosomal duplications blocks on human chromosome 16. We provide evidence that this particular segment has been active independently in each great ape and human lineage at different points during evolution. Euchromatic sequence that flanks sites of LCR16a integration are frequently lineage-specific duplications. This process has mobilized duplication blocks (15-200 kb in size) to new genomic locations in each species. Breakpoint analysis of lineage-specific insertions suggests coordinated deletion of repeat-rich DNA at the target site, in some cases deleting genes in that species. Our data support a model of duplication where the probability that a segment of DNA becomes duplicated is determined by its proximity to core duplicons, such as LCR16a.

Animals↗

Nucleotide sequence, function, activation, and evolution of the cryptic asc operon of Escherichia coli K12.

The cryptic asc (previous called "SAC") operon of Escherichia coli K12 has been completely sequenced. It encodes a repressor (ascG); a PTS enzyme IIasc for the transport of arbutin, salicin, and cellobiose (ascF); and a phospho-beta-glucosidase that hydrolyzes the sugars which are phosphorylated during transport (ascB). ascG and ascFB are transcribed from divergent promoters. The cryptic operon is activated by the insertion of IS186 into the ascG (repressor) gene. The ascFB genes are paralogous to the cryptic bglFB genes, and ascG is paralogous to galR. The duplications that gave rise to these paralogous genes are estimated to have occurred approximately 320 Mya, a time that predates the divergence of E. coli and Salmonella typhimurium.

Amino Acid Sequence↗

[Structural characteristics of osteopontin mRNA].

Conservatism of primary and secondary structure of osteopontin mRNA in vertebrates of various taxes was studied using computer analysis and blot-hybridization procedure. Relatively high rate of evolution of a nucleotide sequence of osteopontin gene was revealed. Formation of hairpin-loop structures by adjacent sequences and a high content of AU-pairs indicate more flexible structure of the mammalian osteopontin mRNA as compared with the avian one. It has been detected some AU-enriched regions in 3'-terminal untranslated sequences, which presence is characteristic for short-lived mRNA. Evolutionary conserved subsequences in 3'- and 5'-terminal regions were suggested to be determinants of RNA-protein interaction in cytoplasm.

3' Untranslated Regions↗

Complete mitochondrial DNA sequences of six snakes: phylogenetic relationships and molecular evolution of genomic features.

Complete mitochondrial DNA (mtDNA) sequences were determined for representative species from six snake families: the acrochordid little file snake, the bold boa constrictor, the cylindrophiid red pipe snake, the viperid himehabu, the pythonid ball python, and the xenopeltid sunbeam snake. Thirteen protein-coding genes, 22 tRNA genes, 2 rRNA genes, and 2 control regions were identified in these mtDNAs. Duplication of the control region and translocation of the tRNALeu gene were two notable features of the snake mtDNAs. The duplicate control regions had nearly identical nucleotide sequences within species but they were divergent among species, suggesting concerted sequence evolution of the two control regions. In addition, the duplicate control regions appear to have facilitated an interchange of some flanking tRNA genes in the viperid lineage. Phylogenetic analyses were conducted using a large number of sites (9570 sites in total) derived from the complete mtDNA sequences. Our data strongly suggested a new phylogenetic relationship among the major families of snakes: ((((Viperidae, Colubridae), Acrochordidae), (((Pythonidae, Xenopeltidae), Cylindrophiidae), Boidae)), Leptotyphlopidae). This conclusion was distinct from a widely accepted view based on morphological characters in denying the sister-group relationship of boids and pythonids, as well as the basal divergence of nonmacrostomatan cylindrophiids. These results imply the significance to reconstruct the snake phylogeny with ample molecular data, such as those from complete mtDNA sequences.

Animals↗

Multilocus sequence analysis and comparative evolution of virulence-associated genes and housekeeping genes of Clostridium difficile.

A multilocus sequence analysis of ten virulence-associated genes was performed to study the genetic relationships between 29 Clostridium difficile isolates of various origins, hosts and clinical presentations, and selected from the main lineages previously defined by multilocus sequence typing (MLST) of housekeeping genes. Colonization-factor-encoding genes (cwp66, cwp84, fbp68, fliC, fliD, groEL and slpA), toxin A and B genes (tcdA and tcdB), and the toxin A and B positive regulator gene (tcdD) were investigated. Binary toxin genes (cdtA and cdtB) were also detected, and internal fragments were sequenced for positive isolates. Virulence-associated genes exhibited a moderate polymorphism, comparable to the polymorphism of housekeeping genes, whereas cwp66 and slpA genes appeared highly polymorphic. Isolates recovered from human pseudomembranous colitis cases did not define a specific lineage. The presence of binary toxin genes, detected in five of the 29 isolates (17 %), was also not linked to clinical presentation. Conversely, toxigenic A-B+ isolates defined a very homogeneous lineage, which is distantly related to other isolates. By clustering analysis, animal isolates were intermixed with human isolates. Multilocus sequence analysis of virulence-associated genes is consistent with a clonal population structure for C. difficile and with the lack of host specificity. The data suggest a co-evolution of several of the virulence-associated genes studied (including toxins A and B and the binary toxin genes) with housekeeping genes, reflecting the genetic background of C. difficile, whereas flagellin, cwp66 and slpA genes may undergo recombination events and/or environmental selective pressure.

Alleles↗

Evolution of apolipoprotein E: mouse sequence and evidence for an 11-nucleotide ancestral unit.

Apolipoprotein E (apo E) is responsible for the binding of very low density lipoprotein and chylomicron remnants to cellular receptors thereby removing them from circulation. We have isolated and determined the sequence of a cDNA encoding 285 amino acids and the entire 3' untranslated region of 112 nucleotides of mouse apo E. The remaining coding sequence was determined by sequencing mouse liver mRNA. Comparisons with rat and human apo E sequences showed a high degree of conservation although there were regions in each species that were characterized by unique insertions and deletions. Analysis of the sequence homologies within apo E revealed that the entire sequence is made up of repetitive units. The most primitive unit appeared to be an 11-nucleotide repeat within higher order repeats of 22 or 33 nucleotides. The 11-nucleotide unit -TCGGACGAGGC- is read in all three reading frames, and when tandemly repeated, it encodes the highly conserved amino acid sequence Xaa-(Glu/Asp)-(Glu/Asp)-Xaa-Arg-Xaa-Arg-Leu-Gly-Xaa-Xaa. We postulate that apo E and those other apolipoproteins related to it have arisen by duplications and subsequent modifications of this or a closely related 11-nucleotide ancestral sequence.

Amino Acid Sequence↗

Variability and evolution of highly repeated DNA sequences in the genus Beta.

Satellite DNA from wild beet species was separated from restriction endonuclease digested genomic DNA by polyacrylamide gel electrophoresis. Two nonhomologous HaeIII satellite DNA repeats were cloned from the wild beet Beta trigyna. The type I repeat is 140-149 bp long and AT rich, while the type II is 162 bp in size and GC rich. A third repetitive HaeIII element cloned from the related wild beet B. corolliflora was shown to be organized as a HinfI satellite DNA family in the cultivated beet B. vulgaris ssp. vulgaris and the wild beet B. vulgaris ssp. maritima. This type III satellite monomer is 149 bp long and contains a high number of short direct subrepeats. The monomer was found in different genomic organizations and copy numbers in all sections of the genus Beta indicating an amplification early in the phylogeny. The HaeIII repeats from B. trigyna are characterized by a lower variability and form long tandem arrays in the genomes of Corollinae species. The investigation of the distribution of all three sequence families provided data that may contribute to the solution of taxonomic problems of the genus Beta and be useful in the characterization of hybrids and derived lines with alien wild beet chromosomes.

Base Sequence↗

Sequence diversity and molecular evolution of the leukotoxin (lktA) gene in bovine and ovine strains of Mannheimia (Pasteurella) haemolytica.

The molecular evolution of the leukotoxin structural gene (lktA) of Mannheimia (Pasteurella) haemolytica was investigated by nucleotide sequence comparison of lktA in 31 bovine and ovine strains representing the various evolutionary lineages and serotypes of the species. Eight major allelic variants (1.4 to 15.7% nucleotide divergence) were identified; these have mosaic structures of varying degrees of complexity reflecting a history of horizontal gene transfer and extensive intragenic recombination. The presence of identical alleles in strains of different genetic backgrounds suggests that assortative (entire gene) recombination has also contributed to strain diversification in M. haemolytica. Five allelic variants occur only in ovine strains and consist of recombinant segments derived from as many as four different sources. Four of these alleles consist of DNA (52.8 to 96.7%) derived from the lktA gene of the two related species Mannheimia glucosida and Pasteurella trehalosi, and four contain recombinant segments derived from an allele that is associated exclusively with bovine or bovine-like serotype A2 strains. The two major lineages of ovine serotype A2 strains possess lktA alleles that have very different evolutionary histories and encode divergent leukotoxins (5.3% amino acid divergence), but both contain segments derived from the bovine allele. Homologous segments of donor and recipient alleles are identical or nearly identical, indicating that the recombination events are relatively recent and probably postdate the domestication of cattle and sheep. Our findings suggest that host switching of bovine strains from cattle to sheep, together with inter- and intraspecies recombinational exchanges, has played an important role in generating leukotoxin diversity in ovine strains. In contrast, there is limited allelic diversity of lktA in bovine strains, suggesting that transmission of strains from sheep to cattle has been less important in leukotoxin evolution.

Alleles↗

Universal constraint on evolution of all coding sequences.

What often escapes notice is the inherent paradox in the role of natural selection in evolution. Inasmuch as a coding sequence encoding a functionless protein shall be ignored by natural selection, the initial acquisition of a function by a protein has to be intrinsic in the construction principle of coding sequences. The very fact that many proteins of very divergent functions tend to have rather similar amino acid compositions suggests that there must be a universal rule in the construction of all coding sequences, and this rule was previously defined as the TA/CG-deficiency-TG/CT-excess rule. When manifesting itself in coding sequences of rather balanced base composition, this rule establishes C T G as one of the most numerous base trimers. While Leu is the major residue of most proteins; C T G is the most frequently utilized of 6 Leu codons in E. coli to man. Yet, the two most common base tetramers, G C T G and C C T G, need not be translated in their second reading frame to yield Leu. Their preferential translation in the first reading frame increases Ala and Pro at the expense of Leu, while that in the third reading frame yielded extremely Cys-rich metallothionein.

Journal Article↗

Rice PHYC gene: structure, expression, map position and evolution.

Although sequences representing members of the phytochrome (phy) family of photoreceptors have been reported in numerous species across the phylogenetic spectrum, relatively few phytochrome genes (PHY) have been fully characterized. Using rice, we have cloned and characterized the first PHYC gene from a monocot. Comparison of genomic and cDNA PHYC sequences shows that the rice PHYC gene contains three introns in the protein-coding region typical of most angiosperm PHY genes, in contrast to Arabidopsis PHYC, which lacks the third intron. Mapping of the transcription start site and 5'-untranslated region of the rice PHYC transcript indicates that it contains an unusually long, intronless, 5'-untranslated leader sequence of 715 bp. PHYC mRNA levels are relatively low compared to PHYA and PHYB mRNAs in rice seedlings, and are similar in dark- and light-treated seedlings, suggesting relatively low constitutive expression. Genomic mapping shows that the PHYA, PHYB, and PHYC genes are all located on chromosome 3 of rice, in synteny with these genes in linkage group C (sometimes referred to as linkage group A) of sorghum. Phylogenetic analysis indicates that rice phyC is closely related to sorghum phyC, but relatively strongly divergent from Arabidopsis phyC, the only full-length dicot phyC sequence available.

Amino Acid Sequence↗

A structure and evolution-guided Monte Carlo sequence selection strategy for multiple alignment-based analysis of proteins.

MOTIVATION: Various multiple sequence alignment-based methods have been proposed to detect functional surfaces in proteins, such as active sites or protein interfaces. The effect that the choice of sequences has on the conclusions of such analysis has seldom been discussed. In particular, no method has been discussed in terms of its ability to optimize the sequence selection for the reliable detection of functional surfaces. RESULTS: Here we propose, for the case of proteins with known structure, a heuristic Metropolis Monte Carlo strategy to select sequences from a large set of homologues, in order to improve detection of functional surfaces. The quantity guiding the optimization is the clustering of residues which are under increased evolutionary pressure, according to the sample of sequences under consideration. We show that we can either improve the overlap of our prediction with known functional surfaces in comparison with the sequence similarity criteria of selection or match the quality of prediction obtained through more elaborate non-structure based-methods of sequence selection. For the purpose of demonstration we use a set of 50 homodimerizing enzymes which were co-crystallized with their substrates and cofactors.

Algorithms↗

Dynamics of mitochondrial DNA evolution in animals: amplification and sequencing with conserved primers.

With a standard set of primers directed toward conserved regions, we have used the polymerase chain reaction to amplify homologous segments of mtDNA from more than 100 animal species, including mammals, birds, amphibians, fishes, and some invertebrates. Amplification and direct sequencing were possible using unpurified mtDNA from nanogram samples of fresh specimens and microgram amounts of tissues preserved for months in alcohol or decades in the dry state. The bird and fish sequences evolve with the same strong bias toward transitions that holds for mammals. However, because the light strand of birds is deficient in thymine, thymine to cytosine transitions are less common than in other taxa. Amino acid replacement in a segment of the cytochrome b gene is faster in mammals and birds than in fishes and the pattern of replacements fits the structural hypothesis for cytochrome b. The unexpectedly wide taxonomic utility of these primers offers opportunities for phylogenetic and population research.

Amino Acid Sequence↗