PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Binary toxins from Bacillus thuringiensis active against the western corn rootworm, Diabrotica virgifera virgifera LeConte.

The western corn rootworm, Diabrotica virgifera virgifera LeConte, is a significant pest of corn in the United States. The development of transgenic corn hybrids resistant to rootworm feeding damage depends on the identification of genes encoding insecticidal proteins toxic to rootworm larvae. In this study, a bioassay screen was used to identify several isolates of the bacterium Bacillus thuringiensis active against rootworm. These bacterial isolates each produce distinct crystal proteins with approximate molecular masses of 13 to 15 kDa and 44 kDa. Insect bioassays demonstrated that both protein classes are required for insecticidal activity against this rootworm species. The genes encoding these proteins are organized in apparent operons and are associated with other genes encoding crystal proteins of unknown function. The antirootworm proteins produced by B. thuringiensis strains EG5899 and EG9444 closely resemble previously described crystal proteins of the Cry34A and Cry35A classes. The antirootworm proteins produced by strain EG4851, designated Cry34Ba1 and Cry35Ba1, represent a new binary toxin. Genes encoding these proteins could become an important component of a sustainable resistance management strategy against this insect pest.

Amino Acid Sequence↗

The tcmVI region of the tetracenomycin C biosynthetic gene cluster of Streptomyces glaucescens encodes the tetracenomycin F1 monooxygenase, tetracenomycin F2 cyclase, and, most likely, a second cyclase.

Certain mutations in the tcmVI region of the Streptomyces glaucescens chromosome affect formation of the D ring of the polyketide antibiotic tetracenomycin C (TCM C). This region lies immediately upstream from the TCM C polyketide synthase genes (tcmKLM), and the nucleotide sequence reveals the presence of three small genes, tcmH, tcmI, and tcmJ. On the basis of the phenotypes of mutants and the effects of these genes, when coupled on a plasmid with the tcmKLMN177 genes (tcmN177 is a 3'-truncated version of tcmN), on the production of TCM intermediates in a TCM- mutant, the tcmH gene encodes the C-5 monooxygenase that converts TCM F1 to TCM D3, the tcmI gene encodes the D-ring cyclase that converts TCM F2 to TCM F1 (mutations in this gene are responsible for the type VI phenotype), and the tcmJ gene most likely encodes the B-ring cyclase that acts in the biosynthesis of TCM F2. Furthermore, it appears that the N-terminal domain of the tcmN gene product (encoded by the tcmN177 gene) acts later in the biosynthesis of TCM F2 than the product of tcmJ, suggesting that the N-terminal domain of the TcmN protein is the C-ring cyclase.

Aldehyde-Lyases↗

A tutorial on Markov models based on Mendel's classical experiments.

Hidden Markov Models (HMM) can be extremely useful tools for the analysis of data from biological sequences, and provide a probabilistic model of protein families. Most reviews and general introductions follow the excellent tutorial by Rabiner, where the focus is outside biology. Mendel's famous experiments in plant hybridisation were published in 1866 and are often considered the icebreaking work of modern genetics. He had no prior knowledge of the dual nature of genes, but through a series of experiments he was able to anticipate the hidden concept and name it "Elemente". In this paper we present the background, theory and algorithms of HMM based on examples from Mendel's experiments, and introduce the toolbox "mendelHMM". This approach is considered to have some intuitive advantages in a biological and bioinformatical setting. Applications to analysing bio-sequences like nucleic acids and proteins are also discussed.

Animals↗

Micro-Mar: a database for dynamic representation of marine microbial biodiversity.

BACKGROUND: The cataloging of marine prokaryotic DNA sequences is a fundamental aspect for bioprospecting and also for the development of evolutionary and speciation models. However, large amount of DNA sequences used to quantify prokaryotic biodiversity requires proper tools for storing, managing and analyzing these data for research purposes. DESCRIPTION: The Micro-Mar database has been created to collect DNA diversity information from marine prokaryotes for biogeographical and ecological analyses. The database currently includes 11874 sequences corresponding to high resolution taxonomic genes (16S rRNA, ITS and 23S rRNA) and many other genes including CDS of marine prokaryotes together with available biogeographical and ecological information. CONCLUSION: The database aims to integrate molecular data and taxonomic affiliation with biogeographical and ecological features that will allow to have a dynamic representation of the marine microbial diversity embedded in a user friendly web interface. It is available online at http://egg.umh.es/micromar/.

Base Sequence↗

Using feature generation and feature selection for accurate prediction of translation initiation sites.

Correct prediction of the translation initiation site (TIS) is an important issue in genomic research. We show that feature generation together with correlation based feature selection can be used with a variety of machine learning algorithms to give highly accurate translation initiation site prediction. Only very few features are needed and the results achieve comparable accuracy to the best existing approaches. Our approach has the advantage that it does not require one to devise a special prediction method; rather standard machine learning classifiers are shown to give very good performance on the selected features. The raw and generated features which we have found to be important are the following: positions -3 and -1 in the sequence; upstream k-grams for k=3, 4, and 5; stop-codon frequency; downstream in-frame 3-gram; and the distance of ATG to the beginning of the sequence. The best result, with an overall accuracy of 90%, is obtained by selecting only seven features from this set. The same features retrained with the use of a scanning model achieves an overall accuracy of 94% on this dataset.

Codon, Initiator↗

Building large knowledge bases in molecular biology.

Large scale genome sequencing projects are now producing hugh amounts of data which can be readily stored and managed within data base management systems, and analyzed using dedicated software packages. The results of these analyzes should also be stored with the input DNA sequences. The increasing complexity and size of the objects to be described and managed have led biologists to rely on advanced data models such as the object-oriented model. As a joint effort between our computer science and molecular biology research projects, the knowledge bases we have developed in molecular genetics have shown however that the basic object-oriented model is not fully adapted to the complexity of some biological situations encountered. Advanced descriptive capabilities, provided only by knowledge models originated from the AI field, are required. Composite or evolving objects, multiple viewpoints, constraints, tasks and methods, textual annotations are some examples of such capabilities. They are illustrated by biological situations for which they appeared to be necessary. Supporting powerful reasoning mechanisms (e.g. object classification, constraint propagation or qualitative simulators), they allow the development of large knowledge bases in molecular biology. These knowledge bases are expected to become the adequate support for co-operative distributed research efforts.

Artificial Intelligence↗

Molecular medicine: a primer for clinicians. Part II: Recombinant DNA molecules.

This is the second paper in our continuing series on the impact of molecular medicine on clinical practice. Discussed are some of the basic methods used to produce recombinant DNA molecules and how these molecules are characterized. Special emphasis is placed on the use of these methods to isolate and characterize human genes.

Clinical Medicine↗

[Contribution of molecular cytogenetics to the diagnosis of chromosome anomalies].

UNLABELLED: MOLECULAR CYTOGENETICS: New fluorescent in situ hybridization (FISH) techniques have been developed using fluorescent non-radioactive DNA probes. FISH: Based on the complementary of nucleotides FISH enables visualization and localization of a DNA fragment on chromosomes by hybridizing the complementary DNA sequence, the probe. Many types of tissues can be analyzed, for example hematopoietic cells in blood or bone marrow, amniotic cells, trophoblasts, fibroblasts, gamete or tumoral cells. APPLICATIONS: Molecular cytogenetics can be used to characterize chromosome anomalies in many fields of cytogenetics (constitutional studies, prenatal diagnosis, hematology, oncology).

Chromosome Aberrations↗

Systematics of basidiomycetous yeasts: a comparison of large subunit D1/D2 and internal transcribed spacer rDNA regions.

Basidiomycetous yeasts in the Urediniomycetes and Hymenomycetes were examined by sequence analysis in two ribosomal DNA regions: the D1/D2 variable domains at the 5' end of the large subunit rRNA gene (D1/D2) and the internal transcribed spacers (ITS) 1 and 2. Four major lineages were recognized in each class: Microbotryum, Sporidiobolus, Erythrobasidium and Agaricostilbum in the Urediniomycetes; Tremellales, Trichosporonales, Filobasidiales and Cystofilobasidiales in the Hymenomycetes. Bootstrap support for many of the clades within those lineages is weak; however, phylogenetic analysis provides a focal point for in-depth study of biological relationships. Combined sequence analysis of the D1/D2 and ITS regions is recommended for species identification, while species definition requires classical biological information such as life cycles and phenotypic characterization.

Basidiomycota↗

Mapping structural determinants of biological activities in snake venom phospholipases A2 by sequence analysis and site directed mutagenesis.

In addition to their catalytic activity, snake venom phospholipases A2 (vPLA2) present remarkable diversity in their biological effects. Sequence alignment analyses of functionally related PLA2 are frequently used to predict the structural determinants of these effects, and the predictions are subsequently evaluated by site directed mutagenesis experiments and functional assays. In order to improve the predictive potential of computer-based analysis, a simple method for scanning amino acid variation analysis (SAVANA) has been developed and included in the analysis of the lysine 49 PLA2 myotoxins (Lys49-PLA2). The SAVANA analysis identified positions in the C-terminal loop region of the protein, which were not identified using previously available sequence analysis tools. Site directed mutagenesis experiments of bothropstoxin-I, a Lys49-PLA2 isolated from the venom of Bothrops jararacussu, reveals that these residues are exactly those involved in the determination of myotoxic and membrane damaging activities. The SAVANA method has been used to analyse presynaptic neurotoxic and anti-coagulant vPLA2s, and the predicted structural determinants of these activities are in excellent agreement with the available results of site directed mutagenesis experiments. The positions of residues involved in the myotoxic and neurotoxic determinants demonstrate significant overlap, suggesting that the multiple biological effects observed in many snake vPLA2s are a consequence of superposed structural determinants on the protein surface.

Amino Acid Sequence↗

Novel transfer RNAs that are active in Escherichia coli.

Many of the mammalian mitochondrial tRNAs contain significant nucleotide deletions in the dihydrouridine (D) stem or T psi C stem, so that they cannot fold into the canonical cloverleaf structure. This suggests that alternative forms and shapes are possible for a mitochondrial tRNA that functions in the specialized translational apparatus of the mammalian mitochondria. The question of whether significant structural alterations may be accommodated by a bacterial protein synthesis machinery, such as in Escherichia coli, is unanswered. In this work, all but ten positions in the gene for the 76-nucleotide coding sequence of an E. coli amber suppressor tRNA were permuted and screened for biological activity in vivo. Sequence analysis of a collection of biologically active variants established that many have unusual structures that include base-pair mismatches in helical stems, substitutions of normally conserved bases, and deletions. Independent mutations were obtained that weaken base pairs or tertiary interactions that normally stabilize the coaxial stacking of the D and anticodon stems, suggesting that the translational apparatus can accommodate considerable flexibility in this part of the molecule. The results demonstrate the capacity of the bacterial protein synthetic apparatus to accommodate altered tRNA structures that are not represented by any naturally occurring tRNAs.

Base Sequence↗

Cooperative computer system for genome sequence analysis.

Analysis of the huge volumes of data generated by large scale sequencing projects clearly requires the construction of new sophisticated computer systems. These systems should be able to handle the biological data as well as the results of the analysis of this data. They should also help the user to choose the most appropriate method for a simple task and to string together the methods needed to solve a global analysis task. In this paper we present the prototype of a software system that provides an environment for the analysis of large-scale sequence data. In a first approach this environment has been put to the test within the B. subtilis sequencing project. This system integrates both a descriptive knowledge of the entities involved (genes, regulatory signals etc.) and the methodological knowledge concerning an extendable set of analytical methods (i.e. how to solve a sequence analysis problem through task decomposition and method selection). A knowledge representation based on two existing object-oriented models, named Shirka and SCARP, is used to implement this integrated system. In addition, the present prototype provides a suitable user interface for both displaying the results generated by several methods and interacting with the objects. We present in this paper an overview of the knowledge-based models used to build this integrated system, and a description of the way in which biological entities and sequence analysis tasks are represented. We give illustrations of the co-operation between user and system during the problem solving process. Such a system constitutes a computer workbench for molecular biologists studying the genetic programs of living organisms.

Bacillus subtilis↗