PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Segmentation of long genomic sequences into domains with homogeneous composition with BASIO software.

UNLABELLED: We present a software system BASIO that allows one to segment a sequence into regions with homogeneous nucleotide composition at a desired length scale. The system can work with arbitrary alphabet and therefore can be applied to various (e.g. protein) sequences. Several sequences of complete genomes of eukaryotes are used to demonstrate the efficiency of the software. AVAILABILITY: The BASIO suite is available for non-commercial users free of charge as a set of executables and accompanying segmentation scenarios from http://www.imb.ac.ru/compbio/basio. To obtain the source code, contact the authors.

Algorithms↗

The Pathway Tools software.

MOTIVATION: Bioinformatics requires reusable software tools for creating model-organism databases (MODs). RESULTS: The Pathway Tools is a reusable, production-quality software environment for creating a type of MOD called a Pathway/Genome Database (PGDB). A PGDB such as EcoCyc (see http://ecocyc.org) integrates our evolving understanding of the genes, proteins, metabolic network, and genetic network of an organism. This paper provides an overview of the four main components of the Pathway Tools: The PathoLogic component supports creation of new PGDBs from the annotated genome of an organism. The Pathway/Genome Navigator provides query, visualization, and Web-publishing services for PGDBs. The Pathway/Genome Editors support interactive updating of PGDBs. The Pathway Tools ontology defines the schema of PGDBs. The Pathway Tools makes use of the Ocelot object database system for data management services for PGDBs. The Pathway Tools has been used to build PGDBs for 13 organisms within SRI and by external users.

Abstracting and Indexing↗

A software tool for finding locally optimal alignments in protein and nucleic acid sequences.

We describe software for aligning protein or nucleic acid sequences based on the concept of match density. This method is especially useful for locating regions of short similarity between two longer sequences which may be largely dissimilar (e.g. locating active site regions in distantly related proteins). Our software is able to identify biologically interesting similarities between two sub-regions because it allows the user to control the matching parameters and the manner in which local alignments are selected for display. Furthermore, the collection and ranking of alignments for display uses a novel, highly efficient algorithm. We illustrate these features with several examples. In addition, we show that this tool can be used to find a new conserved sequence in several viral DNA polymerases, which, we suggest, occurs at a functionally important enzymatic site.

Algorithms↗

Redesigning, implementing and integrating Escherichia coli genome software tools with an object-oriented database system.

This paper reports our exploratory work to redesign, implement and integrate a collection of genome software tools with an object-oriented database system. Our software tools deal with genome data from Escherichia coli K-12, a bacterium that has been studied intensively and provides richer data sets than any other living organism. The object-oriented DBMS used for the integration is ONTOS, a commercial object-oriented system from Ontologic Inc. This redesign and implementation task was performed in two steps. First, C programs were converted into C++, and then the C++ version programs were modified and integrated with an object-oriented modeling of the data to form an ONTOS database application. The first step helps us develop a conceptual view for a DBMS-independent object-oriented construct. The second step elucidates what additional DBMS-dependent modification steps are needed to provide persistency to the objects. Examples are included to illustrate steps of the redesign and implementation. Overall, the outcome of this project demonstrates that programs and data can be successfully integrated with an object-oriented database, while providing the objects with persistency and shareability. This paper includes discussions using concrete examples on what advantage the object-oriented database approach provides over the relational database approach.

Base Sequence↗

CURVATURE: software for the analysis of curved DNA.

Software is presented to plot the sequence-dependent spatial trajectory of the DNA double helix and/or distribution of curvature along the DNA molecule. The nearest-neighbor wedge model is implemented to calculate overall DNA path using local helix parameters: helix twist angle, wedge (deflection) angle and direction (of deflection) angle. The procedures described proved to be very convenient as tools for investigation of a relationship between overall DNA curvature and its gel electrophoretic mobility. All parameters of the model had been estimated from experimental data. Using these wedge parameters the program takes, as input, any DNA sequence and calculates the likely degree of curvature at each point along the molecule. This information is displayed both graphically and in the form of simplified representations of curved double helices. The Software, CURVATURE, can thus be used to investigate possible roles of curvature in modulation of gene expression and for location of curved portions of DNA, which may play an important role in sequence-specific protein--DNA interactions.

Algorithms↗

Multifactor dimensionality reduction software for detecting gene-gene and gene-environment interactions.

MOTIVATION: Polymorphisms in human genes are being described in remarkable numbers. Determining which polymorphisms and which environmental factors are associated with common, complex diseases has become a daunting task. This is partly because the effect of any single genetic variation will likely be dependent on other genetic variations (gene-gene interaction or epistasis) and environmental factors (gene-environment interaction). Detecting and characterizing interactions among multiple factors is both a statistical and a computational challenge. To address this problem, we have developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe the MDR approach and an MDR software package. RESULTS: We developed a program that integrates MDR with a cross-validation strategy for estimating the classification and prediction error of multifactor models. The software can be used to analyze interactions among 2-15 genetic and/or environmental factors. The dataset may contain up to 500 total variables and a maximum of 4000 study subjects. AVAILABILITY: Information on obtaining the executable code, example data, example analysis, and documentation is available upon request. SUPPLEMENTARY INFORMATION: All supplementary information can be found at http://phg.mc.vanderbilt.edu/Software/MDR.

Algorithms↗

HIVbase: a PC/Windows-based software offering storage and querying power for locally held HIV-1 genetic, experimental and clinical data.

BACKGROUND: Human immunodeficiency virus (HIV) research involves ongoing, repetitious sequencing of the HIV genome and the massive accumulation of associated investigational data. As a result, the storage of annotated DNA and/or protein sequences, as well as information retrieval, have become increasingly difficult tasks, with scientists extracting less information from their collected data than they should. OBJECTIVES: Our objective was to design and develop a software package to aid researchers in the storage, analysis and exploration of their HIV-associated data. RESULTS: HIVbase contains familiar, easy-to-use interfaces and functionality for integrating many types of disparate data. The software contains tools that allow for the mass import of raw genetic data, eliminate repetitious sequence translations, have the ability to identify automatically and store HIV regions of interest from nucleic acid or protein sequences, allow for the export of data in commonly used analysis-ready formats, and for unique querying approaches.

Algorithms↗

CisML: an XML-based format for sequence motif detection software.

SUMMARY: CisML is an XML-based format for sequence motif detection software. This proposed standard is applicable to many types of sequence motif detection programs. It is intended to facilitate the integration of data and the comparison of results from different software packages, and to simplify the development of downstream tools. XSL stylesheets are provided for easy generation of text, html and graphical reports from CisML-formatted data. AVAILABILITY: http://zlab.bu.edu/CisML/ SUPPLEMENTARY INFORMATION: Example CisML-formatted data and XSL stylesheets for report generation are available along with the sample output.

Amino Acid Motifs↗

DNAFSMiner: a web-based software toolbox to recognize two types of functional sites in DNA sequences.

UNLABELLED: DNAFSMiner (DNA Functional Sites Miner) is a web-based software toolbox to recognize functional sites in nucleic acid sequences. Currently in this toolbox, we provide two software: TIS Miner and Poly(A) Signal Miner. The TIS Miner can be used to predict translation initiation sites in vertebrate DNA/mRNA/cDNA sequences, and the Poly(A) Signal Miner can be used to predict polyadenylation [poly(A)] signals in human DNA sequences. The prediction results are better than those by literature methods on two benchmark applications. This good performance is mainly attributable to our unique learning method. DNAFSMiner is available free of charge for academic and non-profit organizations. AVAILABILITY: http://research.i2r.a-star.edu.sg/DNAFSMiner/ CONTACT: huiqing@i2r.a-star.edu.sg.

Algorithms↗

The ArrayExpress gene expression database: a software engineering and implementation perspective.

MOTIVATION: The lack of microarray data management systems and databases is still one of the major problems faced by many life sciences laboratories. While developing the public repository for microarray data ArrayExpress we had to find novel solutions to many non-trivial software engineering problems. Our experience will be both relevant and useful for most bioinformaticians involved in developing information systems for a wide range of high-throughput technologies. RESULTS: ArrayExpress has been online since February 2002, growing exponentially to well over 10,000 hybridizations (as of September 2004). It has been demonstrated that our chosen design and implementation works for databases aimed at storage, access and sharing of high-throughput data. AVAILABILITY: The ArrayExpress database is available at http://www.ebi.ac.uk/arrayexpress/. The software is open source. CONTACT: ugis@ebi.ac.uk.

Algorithms↗

BIAS: Bioinformatics Integrated Application Software.

MOTIVATION: We introduce a development platform especially tailored to Bioinformatics research and software development. BIAS (Bioinformatics Integrated Application Software) provides the tools necessary for carrying out integrative Bioinformatics research requiring multiple datasets and analysis tools. It follows an object-relational strategy for providing persistent objects, allows third-party tools to be easily incorporated within the system and supports standards and data-exchange protocols common to Bioinformatics. AVAILABILITY: BIAS is an OpenSource project and is freely available to all interested users at http://www.mcb.mcgill.ca/~bias/. This website also contains a paper containing a more detailed description of BIAS and a sample implementation of a Bayesian network approach for the simultaneous prediction of gene regulation events and of mRNA expression from combinations of gene regulation events. CONTACT: hallett@mcb.mcgill.ca.

Computational Biology↗

A framework for scientific data modeling and automated software development.

MOTIVATION: The lack of standards for storage and exchange of data is a serious hindrance for the large-scale data deposition, data mining and program interoperability that is becoming increasingly important in bioinformatics. The problem lies not only in defining and maintaining the standards, but also in convincing scientists and application programmers with a wide variety of backgrounds and interests to adhere to them. RESULTS: We present a UML-based programming framework for the modeling of data and the automated production of software to manipulate that data. Our approach allows one to make an abstract description of the structure of the data used in a particular scientific field and then use it to generate fully functional computer code for data access and input/output routines for data storage, together with accompanying documentation. This code can be generated simultaneously for different programming languages from a single model, together with, for example for format descriptions and I/O libraries XML and various relational databases. The framework is entirely general and could be applied in any subject area. We have used this approach to generate a data exchange standard for structural biology and analysis software for macromolecular NMR spectroscopy. AVAILABILITY: The framework is available under the GPL license, the data exchange standard with generated subroutine libraries under the LGPL license. Both may be found at http://www.ccpn.ac.uk; http://sourceforge.net/projects/ccpn CONTACT: ccpn@mole.bio.cam.ac.uk.

Biopolymers↗

HTS-Corrector: software for the statistical analysis and correction of experimental high-throughput screening data.

MOTIVATION: High-throughput screening (HTS) plays a central role in modern drug discovery, allowing for testing of >100,000 compounds per screen. The aim of our work was to develop and implement methods for minimizing the impact of systematic error in the analysis of HTS data. To the best of our knowledge, two new data correction methods included in HTS-Corrector are not available in any existing commercial software or freeware. RESULTS: This paper describes HTS-Corrector, a software application for the analysis of HTS data, detection and visualization of systematic error, and corresponding correction of HTS signals. Three new methods for the statistical analysis and correction of raw HTS data are included in HTS-Corrector: background evaluation, well correction and hit-sigma distribution procedures intended to minimize the impact of systematic errors. We discuss the main features of HTS-Corrector and demonstrate the benefits of the algorithms.

Algorithms↗

PROBER: oligonucleotide FISH probe design software.

UNLABELLED: PROBER is an oligonucleotide primer design software application that designs multiple primer pairs for generating PCR probes useful for fluorescence in situ hybridization (FISH). PROBER generates Tiling Oligonucleotide Probes (TOPs) by masking repetitive genomic sequences and delineating essentially unique regions that can be amplified to yield small (100-2000 bp) DNA probes that in aggregate will generate a single, strong fluorescent signal for regions as small as a single gene. TOPs are an alternative to bacterial artificial chromosomes (BACs) that are commonly used for FISH but may be unstable, unavailable, chimeric, or non-specific to small (10-100 kb) genomic regions. PROBER can be applied to any genomic locus, with the limitation that the locus must contain at least 10 kb of essentially unique blocks. To test the software, we designed a number of probes for genomic amplifications and hemizygous deletions that were initially detected by Representational Oligonucleotide Microarray Analysis of breast cancer tumors. AVAILABILITY: http://prober.cshl.edu

Algorithms↗

SEBINI: Software Environment for BIological Network Inference.

UNLABELLED: The Software Environment for BIological Network Inference (SEBINI) has been created to provide an interactive environment for the deployment and evaluation of algorithms used to reconstruct the structure of biological regulatory and interaction networks. SEBINI can be used to compare and train network inference methods on artificial networks and simulated gene expression perturbation data. It also allows the analysis within the same framework of experimental high-throughput expression data using the suite of (trained) inference methods; hence SEBINI should be useful to software developers wishing to evaluate, compare, refine or combine inference techniques, and to bioinformaticians analyzing experimental data. SEBINI provides a platform that aids in more accurate reconstruction of biological networks, with less effort, in less time. AVAILABILITY: A demonstration website is located at https://www.emsl.pnl.gov/NIT/NIT.html. The Java source code and PostgreSQL database schema are available freely for non-commercial use.

Algorithms↗

Software for dynamic analysis of tracer-based metabolomic data: estimation of metabolic fluxes and their statistical analysis.

MOTIVATION: Metabolic flux analysis of biochemical reaction networks using isotope tracers requires software tools that can analyze the dynamics of isotopic isomer (isotopomer) accumulation in metabolites and reveal the underlying kinetic mechanisms of metabolism regulation. Since existing tools are restricted by the isotopic steady state and remain disconnected from the underlying kinetic mechanisms, we have recently developed a novel approach for the analysis of tracer-based metabolomic data that meets these requirements. The present contribution describes the last step of this development: implementation of (i) the algorithms for the determination of the kinetic parameters and respective metabolic fluxes consistent with the experimental data and (ii) statistical analysis of both fluxes and parameters, thereby lending it a practical application. RESULTS: The C++ applications package for dynamic isotopomer distribution data analysis was supplemented by (i) five distinct methods for resolving a large system of differential equations; (ii) the 'simulated annealing' algorithm adopted to estimate the set of parameters and metabolic fluxes, which corresponds to the global minimum of the difference between the computed and measured isotopomer distributions; and (iii) the algorithms for statistical analysis of the estimated parameters and fluxes, which use the covariance matrix evaluation, as well as Monte Carlo simulations. An example of using this tool for the analysis of (13)C distribution in the metabolites of glucose degradation pathways has demonstrated the evaluation of optimal set of parameters and fluxes consistent with the experimental pattern, their range and statistical significance, and also the advantages of using dynamic rather than the usual steady-state method of analysis. AVAILABILITY: Software is available free from http://www.bq.ub.es/bioqint/selivanov.htm

Algorithms↗

Volumetry of hippocampus and amygdala with high-resolution MRI and three-dimensional analysis software: minimizing the discrepancies between laboratories.

Within the medial temporal lobe, both the hippocampus and amygdala are frequently targeted by researchers and clinicians for volumetric analysis based on magnetic resonance imaging (MRI). However, different data acquisition techniques, analysis software and anatomical boundaries have in the past made it difficult to compare results of MRI studies from different laboratories. In order to reduce these differences, a segmentation protocol was established with 40 healthy normal control subjects recently scanned in our laboratory. Data acquisition was performed with a three-dimensional gradient echo technique, and scans were corrected for non-uniformity and registered into standard stereotaxic space prior to segmentation. Volumetric analysis was performed manually using three-dimensional software that allows simultaneous analysis of sagittal, coronal and horizontal images. Intra- and inter-rater coefficients yielded correlation coefficients comparable with other protocols. The hippocampal volume was larger in the right hemisphere (3324 versus 3208 mm(3)), while no interhemispheric differences for the amygdala (1154 versus 1160 mm(3)) could be observed. Most importantly, results from recent segmentation protocols for hippocampus and amygdala seem to approach each other with regard to mean volumes and interhemispheric differences. This indicates that the advances in scanning technique, volume preparation and segmentation protocols allow a more precise definition of medial temporal lobe structures with MRI, and that results for mean volumes for hippocampus and amygdala from different laboratories will eventually become comparable.

Adolescent↗

MyESL: A Software for Evolutionary Sparse Learning in Molecular Phylogenetics and Genomics.

Evolutionary sparse learning uses supervised machine learning to build evolutionary models where genomic sites loci are parameters. It uses the Least Absolute Shrinkage and Selection Operator with bi-level sparsity to connect a specific phylogenetic hypothesis with sequence variation across genomic loci. The MyESL software addresses the need for open-source tools to perform evolutionary sparse learning analyses, offering features to preprocess input phylogenomic alignments, post-process output models to generate molecular evolutionary metrics, and make Least Absolute Shrinkage and Selection Operator regression adaptable and efficient for phylogenetic trees and alignments. The core of MyESL, which constructs models with logistic regressions using bi-level sparsity, is written in C++. Its input data preprocessing and result post-processing tools are developed in Python. Compared to other tools, MyESL is more computationally efficient and provides evolution-friendly inputs and outputs. These features have already enabled the use of MyESL in two phylogenomic applications, one to identify outlier sequences and fragile clades in inferred phylogenies and another to build genetic models of convergent traits. In addition to the use in a Python environment, MyESL is available as a standalone executable compatible across multiple platforms, which can be directly integrated into scripts and third-party software. The source code, executable, and documentation for MyESL are openly accessible at https://github.com/kumarlabgit/MyESL.

Phylogeny↗