PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Programming Languages”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

BUILD: a program generator for modelling experimental biological data.

BUILD is a program generator acting at source code level. The generated code corresponds to a whole application in order to model a biological process of interest using an iterative adjustment of experimental data. The program is designed to be executed in command line mode for processing of multiple data files with an individual execution control for each file. The results are completed by modular statistical and graphical functions. This approach has been shown to reduce the time and the amount of work needed for program development, debugging and maintenance. To date, BUILD has been successfully used in mathematical analysis of phenomenological approaches, but other fields of activity, such as educational software, are also conceivable.

Algorithms↗

On global sequence alignment.

We present a dynamic programming algorithm for computing a best global alignment of two sequences. The proposed algorithm is robust in identifying any of several global relationships between two sequences. The algorithm delivers a best alignment of two sequences in linear space and quadratic time. We also describe a multiple alignment algorithm based on the pairwise algorithm. Both algorithms have been implemented as portable C programs. Experimental results indicate that for a commonly used set of gap penalties, the new programs produce more satisfactory alignments on sequences of various lengths than some existing pairwise and multiple programs based on the dynamic programming algorithm of Needleman and Wunsch.

Algorithms↗

An algorithm based on graph theory for the assembly of contigs in physical mapping of DNA.

An algorithm is described for mapping DNA contigs based on an interval graph (IG) representation. In general terms, the input to the algorithm is a set of binary overlapping relations among finite intervals spread along a real line, from which the algorithm generates sets of ordered overlapping fragments spanning that line. The implications of a more general case of the IG, called a probe interval graph (PIG), in which only a subset of cosmids are used as probes, are also discussed. In the specific case of cosmids hybridizing to regions of a YAC, the algorithm takes cross-hybridization information using the cosmids as probes, and orders them along the YAC; if gaps exist due to insufficient coverage of cosmid contigs along the length of the YAC, repetitive use of the algorithm generates sets of ordered overlapping fragments. Both the IG and the PIG can expose problems caused by false overlaps, such as hybridizations due to repetitive elements. The algorithm, has been coded in C; CPU time is essentially linear with respect to the number of cosmids analyzed. Results are presented for the application of a PIG to cosmid contig assembly along a human chromosome 13-specific YAC. An alignment of 67 cosmids spanning a YAC took 0.28 seconds of CPU time on a Convex 220 computer.

Algorithms↗

Rapid numerical integration algorithm for finding the equilibrium state of a system of coupled binding reactions.

We have adapted a simple method of numerical integration to predict the equilibrium state of a population of components undergoing reversible association according to the Law of Mass Action. Its particular application is to populations of protein molecules in aqueous solution. The method is based on Euler integration but employs an adaptive step size: the time increment being reduced if it would make the concentration of any component negative and increased while the concentration of any component changes at greater than a specified rate. Parameters of the algorithm have been optimized empirically using a model set of binding equilibria with dissociation constants ranging from 10(-5) M to 10(-9) M. The method obtains the solution to a set of binding equilibria more rapidly than the conventional initial value methods (simple Euler, 4th order Runge-Kutta and variable-step Runge-Kutta methods were tested) for the same accuracy. A computer code in standard C is presented.

Algorithms↗

A consensus procedure for predicting the location of alpha-helical transmembrane segments in proteins.

To aid in the development of three-dimensional models of membrane-bound proteins, a consensus procedure for predicting alpha-helical transmembrane segments from amino acid sequence is presented. The algorithm combines the results of six individual prediction methods and some basic properties of membrane-spanning helices to obtain a final consensus prediction. Comparison with experiment and several other recently developed methods shows that the consensus procedure performs quite well in comparison to other recent methods. A FORTRAN program has been developed which takes an input file containing an amino acid sequence in one-letter code and outputs a list of the alpha-helical transmembrane segments predicted by the consensus algorithm.

Algorithms↗

Design and application of PDBlib, a C++ macromolecular class library.

PDBlib is an extensible object-oriented class library written in C++ for representing the three-dimensional structure of biological macromolecules. The software design strategy, features of many of the 129 classes currently distributed with the library, and two sample applications which use the library are described. Version 1.0 of the library represents the structural features of proteins, DNA, RNA and complexes thereof, at a level of detail on a par with that which can be parsed from a Protein Data Bank (PDB) entry. However, the memory-resident representation of the macromolecule is independent of the PDB entry and can be obtained from other sources, e.g. relational and object-oriented databases. PDBlib classes are organized into four categories: (i) classes that model the macromolecule; (ii) classes that enhance the extensibility of the library; (iii) classes that provide navigation facilities of the object-oriented macromolecular structure representation; and (iv) a class that loads a PDB file into the memory-resident object-oriented representation. A number of general-purpose procedures that return features of this representation and that are relevant to all biological disciplines are included in (i). The library has been used to develop PDBtool, a prototype structure verification tool, and PDBview, a structure rendering tool that requires no specialized graphics hardware and software. Current work centers on making the macromolecular structures represented by PDBlib persistent using a commercial object-oriented database and providing an additional class library, MMQLlib, to query those structures.

Databases, Factual↗

DNASTAT: a Pascal unit for the statistical analysis of DNA and protein sequences.

DNASTAT is a collection of Pascal routines for researchers who develop their own application programs for statistical analysis of DNA and protein sequences. Dynamic and file-based data structures allow users to process sets of sequences by simple loop control without limitations on the number of sequences and their individual sizes. This frees the programmer from potentially error-prone tasks like dynamic memory allocation and controlling array sizes. Sequences can be stored in databases along with biological and statistical attributes. Individual sequences can be accessed by column name and row number as with spread-sheets. DNASTAT allows large sets of sequences to be processed using a PC with standard configuration. Its small size, simplicity and free availability make it attractive to students of mathematical biology. Use of DNASTAT is illustrated by two sample programs that generate a database of coding regions from the GenBank entry of the tobacco chloroplast genome. A version of DNASTAT written in ANSI-C for PCs and Unix workstations is also available.

Base Sequence↗

IBIS version 3: an OSF/Motif-based interface for IBIS--integrated biological imaging system.

IBIS is a set of computer programs dedicated to the processing of electron micrographs, mainly for structural analysis of biological macromolecules. We present IBIS version 3, a UNIX/OSF/Motif 1.2-based package which carries out and provides visual display of the many operations essential to image analysis. To ensure portability, the software is written in FORTRAN 77 for computing mathematical functions and in C for display routines. A description of the IBIS OSF/Motif interface is given with the new functions added to the original version. IBIS v.3 is available free of charge to other laboratories on the internet via anonymous ftp (URL: ftp://ftp.univ-rennes1.fr/pub/ incoming/IBIS.tar.Z).

Computer Communication Networks↗

Syntactic recognition of regulatory regions in Escherichia coli.

MOTIVATION: One of the most common methodologies to identify cis-regulatory sites in regulatory regions in the DNA is that of weight matrices, as testified by several articles in this issue. An alternative to strengthen the computational predictions in regulatory regions is to develop methods that incorporate more biological properties present in such DNA regions. The grammatical implementation presented in this paper provides a concrete example in this direction. RESULTS: On the basis of the analysis of an exhaustive collection of regulatory regions in Escherichia coli, a grammatical model for the regulatory regions of sigma 70 promoters has been developed. The terminal symbols of the grammar represent individual sites for the binding of activator and repressor proteins, and include the precise position of sites in relation to transcription initiation. Combining these symbols, the grammar generates a large number of different sentences, each of which can be searched for matching against a collection of regulatory regions by means of weight matrices specific for each set of sites for individual proteins. On the basis of this grammatical model, a Prolog syntactic recognizer is presented here. Specific subgrammars for ArgR, LexA and TyrR were implemented. When parsing a collection of 128 sigma 70 promoter regions, the syntactic recognizer produces a much lower number of false-positive sites than the standard search using weight matrices.

Algorithms↗

Protein data representation and query using optimized data decomposition.

MOTIVATION: To provide data management tools to maintain and query efficiently experimental and derived protein data with the goal of providing new insights into structure-function relationships. The tools should be portable, extensible, and accessible locally, or via the World Wide Web, providing data that would not otherwise be available. RESULTS: The initial phase of the work, the data representation and query of all available macromolecular structure data, including real-time access to complex property patterns based on the amino acid sequence, is reported. protein structure data taken from the Protein Data Bank (PDB) are decomposed into native and derived elementary properties, and represented as compact indexed objects minimizing storage requirements and query time for select types of query. In addition, collections of indices representing a particular property are maintained and can be queried for specific property patterns found across the whole database. The approach is proving applicable to a wide variety of data available on specific protein families.

Algorithms↗

Bi-dimensional scaling map (BDS-Map): an approach for building large genetic maps.

MOTIVATION: The approaches usually used for building large genetic maps consist of dividing the marker set into linkage groups and provide local orders that can be tested by multi-point linkage analysis. To deal with the limitations of these approaches, a strategy taking the marker set into account globally is defined. RESULTS: The paper presents a new approach called 'Bi-Dimensional Scaling Map (BDS-Map) for inferring marker orders and distances in genetic maps based on the use of an additional dimension orthogonal to the map into which markers are projected. Dynamical forces based on a two-point analysis are applied to tend to optimize the marker locations in space. The efficiency of the approach is exemplified on real data (16 and 70 markers on chromosomes 6 and 2, respectively) and simulated data (50 maps of 70 markers).

Algorithms↗

Graphical interface to the genetic network database GeNet.

UNLABELLED: We designed a Java applet which enables the visualization of genetic networks and can be used as a Web publishing tool by molecular biologists studying the mechanisms of gene interactions. AVAILABILITY: http://www. csa.ru/Inst/gorb_dep/inbios/ genet/Graph/Genes_Graph.html CONTACT: samson@fn.csa.ru

Computational Biology↗

RNA movies: visualizing RNA secondary structure spaces.

MOTIVATION: RNA Movies is a system for the visualization of RNA secondary structure spaces. Its input is a script consisting of primary and secondary structure information. From this script, the system fully automatically generates animated graphical structure representations. In this way, it creates the impression of an RNA molecule exploring its own two-dimensional structure space. RESULTS: RNA Movies has been used to generate animations of a switching structure in the spliced leader RNA of Leptomonas collosoma and sequential foldings of potato spindle tuber viroid transcripts. AVAILABILITY: Demonstrations of the animations mentioned in this paper can be viewed on our Bioinformatics web server under the following address: http://BiBiServ.TechFak.Uni-Bielefeld. DE/rnamovies/. The RNA Movies software is available upon request from the authors.

Algorithms↗

WebPHYLIP: a web interface to PHYLIP.

A web interface to PHYLIP (version 3.57 C) is implemented using CGI/Perl programming. It enables users to do phylogenetic analysis through the Internet.

Data Display↗

A proposal for a standard CORBA interface for genome maps.

MOTIVATION: The scientific community urgently needs to standardize the exchange of biological data. This is helped by the use of a common protocol and the definition of shared data structures. We have based our standardization work on CORBA, a technology that has become a standard in the past years and allows interoperability between distributed objects. RESULTS: We have defined an IDL specification for genome maps and present it to the scientific community. We have implemented CORBA servers based on this IDL to distribute RHdb and HuGeMap maps. The IDL will co-evolve with the needs of the mapping community. AVAILABILITY: The standard IDL for genome maps is available at http:// corba.ebi.ac.uk/RHdb/EUCORBA/MapIDL.htm l. The IORs to browse maps from Infobiogen and EBI are at http://www.infobiogen.fr/services/Hugemap/IOR and http://corba.ebi.ac.uk/RHdb/EUCORBA/IOR CONTACT: manu@infobiogen.fr, tome@ebi.ac.uk

Animals↗

Object-oriented parsing of biological databases with Python.

MOTIVATION: While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way. RESULTS: We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.

Databases, Factual↗

A Bayesian framework for the analysis of microarray expression data: regularized t -test and statistical inferences of gene changes.

MOTIVATION: DNA microarrays are now capable of providing genome-wide patterns of gene expression across many different conditions. The first level of analysis of these patterns requires determining whether observed differences in expression are significant or not. Current methods are unsatisfactory due to the lack of a systematic framework that can accommodate noise, variability, and low replication often typical of microarray data. RESULTS: We develop a Bayesian probabilistic framework for microarray data analysis. At the simplest level, we model log-expression values by independent normal distributions, parameterized by corresponding means and variances with hierarchical prior distributions. We derive point estimates for both parameters and hyperparameters, and regularized expressions for the variance of each gene by combining the empirical variance with a local background variance associated with neighboring genes. An additional hyperparameter, inversely related to the number of empirical observations, determines the strength of the background variance. Simulations show that these point estimates, combined with a t -test, provide a systematic inference approach that compares favorably with simple t -test or fold methods, and partly compensate for the lack of replication.

Bayes Theorem↗