PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

The UniMarker (UM) method for synteny mapping of large genomes.

MOTIVATION: Synteny mapping, or detecting regions that are orthologous between two genomes, is a key step in studies of comparative genomics. For completely sequenced genomes, this is increasingly accomplished by whole-genome sequence alignment. However, such methods are computationally expensive, especially for large genomes, and require rather complicated post-processing procedures to filter out non-orthologous sequence matches. RESULTS: We have developed a novel method that does not require sequence alignment for synteny mapping of two large genomes, such as the human and mouse. In this method, the occurrence spectra of genome-wide unique 16mer sequences present in both the human and mouse genome are used to directly detect orthologous genomic segments. Being sequence alignment-free, the method is very fast and able to map the two mammalian genomes in one day of computing time on a single Pentium IV personal computer. The resulting human-mouse synteny map was shown to be in excellent agreement with those produced by the Mouse Genome Sequencing Consortium (MGSC) and by the Ensembl team; furthermore, the syntenic relationship of segments found only by our method was supported by BLASTZ sequence alignment.

Algorithms↗

Accurate identification of alternatively spliced exons using support vector machine.

MOTIVATION: Alternative splicing is a major component of the regulatory action on mammalian transcriptomes. It is estimated that over half of all human genes have more than one splice variant. Previous studies have shown that alternatively spliced exons possess several features that distinguish them from constitutively spliced ones. Recently, we have demonstrated that such features can be used to distinguish alternative from constitutive exons. In the current study, we used advanced machine learning methods to generate robust classifier of alternative exons. RESULTS: We extracted several hundred local sequence features of constitutive as well as alternative exons. Using feature selection methods we find seven attributes that are dominant for the task of classification. Several less informative features help to slightly increase the performance of the classifier. The classifier achieves a true positive rate of 50% for a false positive rate of 0.5%. This result enables one to reliably identify alternatively spliced exons in exon databases that are believed to be dominated by constitutive exons.

Algorithms↗

Automatic clustering of orthologs and inparalogs shared by multiple proteomes.

MOTIVATION: The complete sequencing of many genomes has made it possible to identify orthologous genes descending from a common ancestor. However, reconstruction of evolutionary history over long time periods faces many challenges due to gene duplications and losses. Identification of orthologous groups shared by multiple proteomes therefore becomes a clustering problem in which an optimal compromise between conflicting evidences needs to be found. RESULTS: Here we present a new proteome-scale analysis program called MultiParanoid that can automatically find orthology relationships between proteins in multiple proteomes. The software is an extension of the InParanoid program that identifies orthologs and inparalogs in pairwise proteome comparisons. MultiParanoid applies a clustering algorithm to merge multiple pairwise ortholog groups from InParanoid into multi-species ortholog groups. To avoid outparalogs in the same cluster, MultiParanoid only combines species that share the same last ancestor. To validate the clustering technique, we compared the results to a reference set obtained by manual phylogenetic analysis. We further compared the results to ortholog groups in KOGs and OrthoMCL, which revealed that MultiParanoid produces substantially fewer outparalogs than these resources. AVAILABILITY: MultiParanoid is a freely available standalone program that enables efficient orthology analysis much needed in the post-genomic era. A web-based service providing access to the original datasets, the resulting groups of orthologs, and the source code of the program can be found at http://multiparanoid.cgb.ki.se.

Algorithms↗

Conservation of polyamine regulation by translational frameshifting from yeast to mammals.

Regulation of ornithine decarboxylase in vertebrates involves a negative feedback mechanism requiring the protein antizyme. Here we show that a similar mechanism exists in the fission yeast Schizosaccharomyces pombe. The expression of mammalian antizyme genes requires a specific +1 translational frameshift. The efficiency of the frameshift event reflects cellular polyamine levels creating the autoregulatory feedback loop. As shown here, the yeast antizyme gene and several newly identified antizyme genes from different nematodes also require a ribosomal frameshift event for their expression. Twelve nucleotides around the frameshift site are identical between S.pombe and the mammalian counterparts. The core element for this frameshifting is likely to have been present in the last common ancestor of yeast, nematodes and mammals.

Amino Acid Sequence↗

An analysis of signatures of selective sweeps in natural populations of the house mouse.

Population and locus-specific reduction of variability of polymorphic loci could be an indication of positive selection at a linked site (selective sweep) and therefore point toward genes that have been involved in recent adaptations. Analysis of microsatellite variability offers a way to identify such regions and to ask whether they occur more often than expected by chance. We studied four populations of the house mouse (Mus musculus) to assess the frequency of such signatures of selective sweeps under natural conditions. Three samples represent the subspecies Mus m. dometicus [corrected] and came from Germany, France, and Cameroon. One sample came from Kazakhstan and constitutes a population of the subspecies Mus m. [corrected] musculus. Mitochondrial D-loop sequences from all animals confirm their respective assignments. Approximately 200 microsatellite loci were typed for up to 60 unrelated individuals from each population and evaluated for signs of selective sweeps on the basis of Schlötterer's ln RV and ln RH statistics. Our data suggest that there are slightly more signs of selective sweeps than would have been expected by chance alone in each of the populations and also highlights some of the statistical challenges faced in genome scans for detecting selection. Single-nucleotide polymorphism typing of one sweep signature in the M. m. domesticus populations around the beta-defensin 6 locus confirms a lowered nucleotide diversity in this region and limits the potential sweep region to about 20 kb. However, no amino acid exchange has occurred in the coding region when compared to M. m. musculus. If this sweep signature is due to a recent adaptation, it is expected that a regulatory change would have caused it. Our data provide a framework for conducting a systematic whole genome scan for signatures of selective sweeps in the mouse genome.

Amino Acid Sequence↗

Evidence for evolutionarily conserved secondary structure in the H19 tumor suppressor RNA.

The molecular basis for function of the mammalian H19 as a tumor suppressor is poorly understood. Large, conserved open reading frames (ORFs) are absent from both the human and mouse cDNAs, suggesting that it may act as an RNA. Contradicting earlier reports, however, recent studies have shown that the H19 transcript exists in polysomal form and is likely translated. To distinguish between possible functional roles for the gene product, we have characterized the sequence requirements for H19-mediated in vitro suppression of tumor cell clonogenicity and analyzed the sequence of the gene cloned from a range of mammals. A cDNA version of the human gene, lacking the unusually short introns characteristic of imprinted genes, is as effective as a genomic copy in blocking anchorage-independent growth by G401 cells. The first 710 nucleotides of the gene can be deleted with no effect on in vitro activity. Further truncations from either the 5'- or 3'-end, however, cause a loss of suppression of clonogenicity. Using conserved sequences within the H19 gene as PCR primers, genomic DNA fragments were amplified from a range of mammalian species that span the functional domain defined by deletion analysis. Sequences from cat, lynx, elephant, gopher and orangutan complement the previous database of sequences from human, mouse, rat and rabbit. Hypothetical translation of the resulting sequences shows an absence of conserved ORFs of any size. Free energy and covariational analysis of the RNA sequences was used to identify potential helical pairings within the H19 transcript. A set of 16 helices are supported by covariation (i.e. conservation of base pairing potential in the absence of primary sequence conservation). The predicted RNA pairings consist largely of local hairpins but also include several long range interactions that bridge the 5'- and 3'-ends of the functional domain. Given the evolutionary conservation of structure at the RNA level and the absence of conservation at the protein level, we presume that the functional product of the H19 gene is a structured RNA.

Animals↗

Plasmodium interspersed repeats: the major multigene superfamily of malaria parasites.

Functionally related homologues of known genes can be difficult to identify in divergent species. In this paper, we show how multi-character analysis can be used to elucidate the relationships among divergent members of gene superfamilies. We used probabilistic modelling in conjunction with protein structural predictions and gene-structure analyses on a whole-genome scale to find gene homologies that are missed by conventional similarity-search strategies and identified a variant gene superfamily in six species of malaria (Plasmodium interspersed repeats, pir). The superfamily includes rif in P.falciparum, vir in P.vivax, a novel family kir in P.knowlesi and the cir/bir/yir family in three rodent malarias. Our data indicate that this is the major multi-gene family in malaria parasites. Protein localization of products from pir members to the infected erythrocyte membrane in the rodent malaria parasite P.chabaudi, demonstrates phenotypic similarity to the products of pir in other malaria species. The results give critical insight into the evolutionary adaptation of malaria parasites to their host and provide important data for comparative immunology between malaria parasites obtained from laboratory models and their human counterparts.

Amino Acid Motifs↗

Widespread occurrence of spliceosomal introns in the rDNA genes of ascomycetes.

Spliceosomal (pre-mRNA) introns have previously been found in eukaryotic protein-coding genes, in the small nuclear RNAs of some fungi, and in the small- and large-subunit ribosomal DNA genes of a limited number of ascomycetes. How the majority of these introns originate remains an open question because few proven cases of recent and pervasive intron origin have been documented. We report here the widespread occurrence of spliceosomal introns (69 introns at 27 different sites) in the small- and large-subunit nuclear-encoded rDNA of lichen-forming and free-living members of the Ascomycota. Our analyses suggest that these spliceosomal introns are of relatively recent origin, i.e., within the Euascomycetes, and have arisen through aberrant reverse-splicing (in trans) of free pre-mRNA introns into rRNAs. The spliceosome itself, and not an external agent (e.g., transposable elements, group II introns), may have given rise to these introns. A nonrandom sequence pattern was found at sites flanking the rRNA spliceosomal introns. This pattern (AG-intron-G) closely resembles the proto-splice site (MAG-intron-R) postulated for intron insertions in pre-mRNA genes. The clustered positions of spliceosomal introns on secondary structures suggest that particular rRNA regions are preferred sites for insertion through reverse-splicing.

Ascomycota↗

The neomuran origin of archaebacteria, the negibacterial root of the universal tree and bacterial megaclassification.

Prokaryotes constitute a single kingdom, Bacteria, here divided into two new subkingdoms: Negibacteria, with a cell envelope of two distinct genetic membranes, and Unibacteria, comprising the new phyla Archaebacteria and Posibacteria, with only one. Other new bacterial taxa are established in a revised higher-level classification that recognizes only eight phyla and 29 classes. Morphological, palaeontological and molecular data are integrated into a unified picture of large-scale bacterial cell evolution despite occasional lateral gene transfers. Archaebacteria and eukaryotes comprise the clade neomura, with many common characters, notably obligately co-translational secretion of N-linked glycoproteins, signal recognition particle with 7S RNA and translation-arrest domain, protein-spliced tRNA introns, eight-subunit chaperonin, prefoldin, core histones, small nucleolar ribonucleoproteins (snoRNPs), exosomes and similar replication, repair, transcription and translation machinery. Eubacteria (posibacteria and negibacteria) are paraphyletic, neomura having arisen from Posibacteria within the new subphylum Actinobacteria (possibly from the new class Arabobacteria, from which eukaryotic cholesterol biosynthesis probably came). Replacement of eubacterial peptidoglycan by glycoproteins and adaptation to thermophily are the keys to neomuran origins. All 19 common neomuran character suites probably arose essentially simultaneously during the radical modification of an actinobacterium. At least 11 were arguably adaptations to thermophily. Most unique archaebacterial characters (prenyl ether lipids; flagellar shaft of glycoprotein, not flagellin; DNA-binding protein lob; specially modified tRNA; absence of Hsp90) were subsequent secondary adaptations to hyperthermophily and/or hyperacidity. The insertional origin of protein-spliced tRNA introns and an insertion in proton-pumping ATPase also support the origin of neomura from eubacteria. Molecular co-evolution between histones and DNA-handling proteins, and in novel protein initiation and secretion machineries, caused quantum evolutionary shifts in their properties in stem neomura. Proteasomes probably arose in the immediate common ancestor of neomura and Actinobacteria. Major gene losses (e.g. peptidoglycan synthesis, hsp90, secA) and genomic reduction were central to the origin of archaebacteria. Ancestral archaebacteria were probably heterotrophic, anaerobic, sulphur-dependent hyperthermoacidophiles; methanogenesis and halophily are secondarily derived. Multiple lateral gene transfers from eubacteria helped secondary archaebacterial adaptations to mesophily and genome re-expansion. The origin from a drastically altered actinobacterium of neomura, and the immediately subsequent simultaneous origins of archaebacteria and eukaryotes, are the most extreme and important cases of quantum evolution since cells began. All three strikingly exemplify De Beer's principle of mosaic evolution: the fact that, during major evolutionary transformations, some organismal characters are highly innovative and change remarkably swiftly, whereas others are largely static, remaining conservatively ancestral in nature. This phenotypic mosaicism creates character distributions among taxa that are puzzling to those mistakenly expecting uniform evolutionary rates among characters and lineages. The mixture of novel (neomuran or archaebacterial) and ancestral eubacteria-like characters in archaebacteria primarily reflects such vertical mosaic evolution, not chimaeric evolution by lateral gene transfer. No symbiogenesis occurred. Quantum evolution of the basic neomuran characters, and between sister paralogues in gene duplication trees, makes many sequence trees exaggerate greatly the apparent age of archaebacteria. Fossil evidence is compelling for the extreme antiquity of eubacteria [over 3500 million years (My)] but, like their eukaryote sisters, archaebacteria probably arose only 850 My ago. Negibacteria are the most ancient, radiating rapidly into six phyla. Evidence from molecular sequences, ultrastructure, evolution of photosynthesis, envelope structure and chemistry and motility mechanisms fits the view that the cenancestral cell was a photosynthetic negibacterium, specifically an anaerobic green non-sulphur bacterium, and that the universal tree is rooted at the divergence between sulphur and non-sulphur green bacteria. The negibacterial outer membrane was lost once only in the history of life, when Posibacteria arose about 2800 My ago after their ancestors diverged from Cyanobacteria.

Archaea↗

Luc7p, a novel yeast U1 snRNP protein with a role in 5' splice site recognition.

The characterization of a novel yeast-splicing factor, Luc7p, is presented. The LUC7 gene was identified by a mutation that causes lethality in a yeast strain lacking the nuclear cap-binding complex (CBC). Luc7p is similar in sequence to metazoan proteins that have arginine-serine and arginine-glutamic acid repeat sequences characteristic of a family of splicing factors. We show that Luc7p is a component of yeast U1 snRNP and is essential for vegetative growth. The composition of yeast U1 snRNP is altered in luc7 mutant strains. Extracts of these strains are unable to support any of the defined steps of splicing unless recombinant Luc7p is added. Although the in vivo defect in splicing wild-type reporter introns in a luc7 mutant strain is comparatively mild, splicing of introns with nonconsensus 5' splice site or branchpoint sequences is more defective in the mutant strain than in wild-type strains. By use of reporters that have two competing 5' splice sites, a loss of efficient splicing to the cap proximal splice site is observed in luc7 cells, analogous to the defect seen in strains lacking CBC. CBC can be coprecipitated with U1 snRNP from wild-type, but not from luc7, yeast strains. These data suggest that the loss of Luc7p disrupts U1 snRNP-CBC interaction, and that this interaction contributes to normal 5' splice site recognition.

Alternative Splicing↗

The automatic detection of homologous regions (ADHoRe) and its application to microcolinearity between Arabidopsis and rice.

It is expected that one of the merits of comparative genomics lies in the transfer of structural and functional information from one genome to another. This is based on the observation that, although the number of chromosomal rearrangements that occur in genomes is extensive, different species still exhibit a certain degree of conservation regarding gene content and gene order. It is in this respect that we have developed a new software tool for the Automatic Detection of Homologous Regions (ADHoRe). ADHoRe was primarily developed to find large regions of microcolinearity, taking into account different types of microrearrangements such as tandem duplications, gene loss and translocations, and inversions. Such rearrangements often complicate the detection of colinearity, in particular when comparing more anciently diverged species. Application of ADHoRe to the complete genome of Arabidopsis and a large collection of concatenated rice BACs yields more than 20 regions showing statistically significant microcolinearity between both plant species. These regions comprise from 4 up to 11 conserved homologous gene pairs. We predict the number of homologous regions and the extent of microcolinearity to increase significantly once better annotations of the rice genome become available.

Arabidopsis↗

Development of Y-chromosomal microsatellite markers for nonhuman primates.

We have analysed 136 newly identified human Y-chromosomal microsatellites in five (sub)species of nonhuman primates. We identified 83 male-specific loci for central chimpanzees, 82 for western chimpanzees, 67 for gorillas, 45 for orangutans and 19 loci for mandrills. Polymorphism was detected at 56 loci in central chimpanzees, 29 in western chimpanzees, 24 in western gorillas, 17 in orangutans and at three in mandrills. Success in male-specific amplification of human Y-chromosomal microsatellites in nonhuman primates was significantly negatively correlated with divergence time from the human lineage. We observed significantly more Y-chromosomal microsatellite diversity in central chimpanzees than in western chimpanzees. There were significantly more male-specific loci with longer alleles in humans than with longer alleles in the nonhuman primates; however, this significant difference disappeared when only the loci which are polymorphic in nonhuman primates were analysed, suggesting that ascertainment bias is responsible. This study provides primatologists with a large number of polymorphic, male-specific microsatellite markers that will be valuable for investigating relevant questions in behavioural ecology such as male reproductive strategies, kin-based cooperation among males and male-specific dispersal patterns in wild groups of nonhuman primates.

Animals↗

The evolution of the Vahlkampfiidae as deduced from 16S-like ribosomal RNA analysis.

The amoebae, a phenotypically diverse, paraphyletic group of protists, have been largely neglected by molecular phylogeneticists. To better understand the evolution of amoebae, we sequenced and analyzed the 16S-like ribosomal RNA genes of three vahlkampfiid amoebae: Paratetramitus jugosus, Tetramitus rostratus and Vahlkampfia lobospinosa. The Vahlkampfiidae lineage is monophyletic, branches early along the eukaryotic line of descent, and is not a close relative of the multicellular amoebae that also reversibly transform from amoebae to flagellates.

Animals↗

Phylogenetic position of the menaquinone-containing acidophilic chemo-organotroph Acidobacterium capsulatum.

The phylogenetic position of an acidophilic chemo-organotrophic menaquinone-containing bacterium, Acidobacterium capsulatum, was studied on the basis of 16S rRNA gene sequence information. A. capsulatum showed the highest level of sequence similarity to Heliobacterium chlorum, a member of the Gram-positive group, yet this level was only 81%. Distance matrix tree analysis suggested that A. capsulatum belongs to a unique lineage deeply branching from the Chlamydia-Planctomyces group or from the Gram-positive line.

Base Sequence↗

Comparative analysis of the LPS biosynthetic loci of the genetic subtypes of serovar Hardjo: Leptospira interrogans subtype Hardjoprajitno and Leptospira borgpetersenii subtype Hardjobovis.

Although Leptospira borgpetersenii subtype Hardjobovis and L. interrogans subtype Hardjoprajitno belong to different species, they are serologically indistinguishable and are therefore classified as serovar Hardjo. Since LPS is the major antigen involved in serological classification, this implies that the LPS of these subtypes is identical. Comparison of the LPS biosynthetic loci (rfb) of the subtypes revealed remarkable similarity, with 32 and 31 origins of replication (orfs) in the Hardjoprajitno and Hardjobovis rfb loci, respectively. The order and orientation of these orfs were identical with the exception of an additional orf in Hardjoprajitno between orfs 4 and 5 and intergenic sequences differing between the subtypes. The Hardjoprajitno rfb locus has been divided into four intercalated regions based on sequence similarity to other leptospiral rfb loci. orfJ1-orfJ14 as well as orfJ21-orfJ22 are more similar to regions of the rfb locus of L. borgpetersenii subtype Hardjobovis. orfJ15-orfJ20 as well as orfJ23-orfJ31 are almost identical to the corresponding orfs in L. interrogans serovar Copenhageni. We propose that the progenitor Hardjoprajitno strain, containing an rfb locus which closely resembled the Copenhageni locus, acquired orfs 1-14 and orfs 21-22 from subtype Hardjobovis resulting in two serologically indistinguishable subtypes of serovar Hardjo which in turn constituted the main bovine-adapted leptospiral serovar.

Bacterial Proteins↗

Genetic divergence and evolutionary instability in ospE-related members of the upstream homology box gene family in Borrelia burgdorferi sensu lato complex isolates.

A series of related genes that are flanked at their 5' ends by a conserved upstream sequence element called the upstream homology box (UHB) have been identified in Borrelia burgdorferi. These genes have been referred to as the UHB or erp gene family. We previously demonstrated that among a limited number of B. burgdorferi isolates, the UHB gene family is variable in composition and organization. Prior to this report the UHB gene family in other species of the B. burgdorferi sensu lato complex had not been studied, and if this family is important in the pathogenesis or biology of the Lyme disease spirochetes, then a wide distribution among species and isolates of the B. burgdorferi sensu lato complex would be expected. To assess this, we screened for the UHB element by Southern hybridization and determined its restriction fragment length polymorphism (RFLP) patterns. The UHB element was found to be carried by all B. burgdorferi sensu lato complex species tested (B. burgdorferi, B. garinii, B. afzelii, B. japonica, B. valaisiana sp. nov., and B. andersonii), but the RFLP patterns varied widely at both the inter- and intraspecies levels. Variation in both the number and size of the hybridizing restriction fragments was evident. PCR analyses also revealed the presence of polymorphic, ospE-related alleles in many isolates. Sequence analyses identified the molecular basis of the polymorphisms as being primarily insertions and deletions. Sequence variation and the insertions and deletions were found to be clustered in two distinct domains (variable domains 1 and 2). In many isolates variable domain 1 is flanked by direct repeat elements, some as long as 38 bp. Computer analyses of the deduced amino acid sequences encoded within variable domain 1 predict them to be hydrophilic, surface exposed, and antigenic. The analyses conducted here suggest that the UHB gene family, as evidenced by the variable UHB RFLP patterns, is not evolutionarily stable and that the polymorphic ospE alleles are derived from a common ancestral gene which has been modified through mutation or recombination events. The characterization of ospE-related genes of the UHB gene family among B. burgdorferi sensu lato species will prove important in attempts to construct a model for UHB gene family organization and in deciphering the role of the UHB gene family in the biology and pathogenesis of the Lyme disease spirochetes.

Amino Acid Sequence↗

Identification and characterization of a yeast gene encoding the U2 small nuclear ribonucleoprotein particle B" protein.

The inessential yeast gene MUD2 encodes a protein factor that contributes to U1 small nuclear ribonucleoprotein particle (snRNP)-pre-mRNA complex (commitment complex) formation. To identify other genes that contribute to this early splicing step, we performed a synthetic lethal screen with a MUD2 deletion strain. The first characterized gene from this screen, MSL1 (MUD synthetic lethal 1), encodes the yeast homolog of the well studied mammalian snRNP protein U2B". The yeast protein (YU2B") is a component of yeast U2 snRNP, and it is related to other members of the UIA-U2B" family, the human U2B" protein, the human U1A protein, and the yeast U1A protein. It binds in vitro to its RNA target, U2 snRNA stem-loop IV, without a protein cofactor, and the target resembles more closely the U1 snRNA binding site of the human U1A protein than it does the U2 snRNA binding site of human U2B". Surprisingly, the YU2B" protein lacks a C-terminal RNA binding domain, which is conserved in all other family members. Possible functional and evolutionary relationships among these proteins are discussed.

Amino Acid Sequence↗