PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Programming Languages”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Web-based exchange of biochemical information.

In this paper we present a web-based system for the semantic integration of biological data using a database model and an XML exchange format. The prototype developed integrates data related to enzymes.

Databases, Factual↗

PDB file parser and structure class implemented in Python.

UNLABELLED: The biopython project provides a set of bioinformatics tools implemented in Python. Recently, biopython was extended with a set of modules that deal with macromolecular structure. Biopython now contains a parser for PDB files that makes the atomic information available in an easy-to-use but powerful data structure. The parser and data structure deal with features that are often left out or handled inadequately by other packages, e.g. atom and residue disorder (if point mutants are present in the crystal), anisotropic B factors, multiple models and insertion codes. In addition, the parser performs some sanity checking to detect obvious errors. AVAILABILITY: The Biopython distribution (including source code and documentation) is freely available (under the Biopython license) from http://www.biopython.org

Computer Simulation↗

Libsequence: a C++ class library for evolutionary genetic analysis.

UNLABELLED: A C++ class library is available to facilitate the implementation of software for genomics and sequence polymorphism analysis. The library implements methods for data manipulation and the calculation of several statistics commonly used to analyze SNP data. The object-oriented design of the library is intended to be extensible, allowing users to design custom classes for their own needs. In addition, routines are provided to process samples generated by a widely used coalescent simulation. AVAILABILITY: The source code (in C++) is available from http://www.molpopgen.org

Algorithms↗

PreP: gene expression data pre-processing.

UNLABELLED: PreP is a versatile, powerful, standalone application that aims at pre-processing gene expression data. AVAILABILITY: Documentation and executable file for MS-Windows are available at http://chirimoyo.ac.uma.es/bitlab/services/index.htm

Algorithms↗

NCL: a C++ class library for interpreting data files in NEXUS format.

UNLABELLED: The NEXUS Class Library (NCL) is a collection of C++ classes designed to simplify interpreting data files written in the NEXUS format used by many computer programs for phylogenetic analyses. The NEXUS format allows different programs to share the same data files, even though none of the programs can interpret all of the data stored therein. Because users are not required to reformat the data file for each program, use of the NEXUS format prevents cut-and-paste errors as well as the proliferation of copies of the original data file. The purpose of making the NCL available is to encourage the use of the NEXUS format by making it relatively easy for programmers to add the ability to interpret NEXUS files in newly developed software. AVAILABILITY: The NCL is freely available under the GNU General Public License from http://hydrodictyon.eeb.uconn.edu/ncl/ SUPPLEMENTARY INFORMATION: Documentation for the NCL (general information and source code documentation) is available in HTML format at http://hydrodictyon.eeb.uconn.edu/ncl/

Databases, Bibliographic↗

LDDist: a Perl module for calculating LogDet pair-wise distances for protein and nucleotide sequences.

LDDist is a Perl module implemented in C++ that allows the user to calculate LogDet pair-wise genetic distances for amino acid as well as nucleotide sequence data. It can handle site-to-site rate variation by treating a proportion of the sites as invariant and/or by assigning sites to different, presumably homogenous, rate categories. The rate-class assignments and invariant proportion can be set explicitly, or estimated by the program; the latter using either of two different capture-recapture methods. The assignment to rate categories in lieu of a phylogeny can be done using Shannon-Wiener index as a crude token for relative rate.

Algorithms↗

HIVbase: a PC/Windows-based software offering storage and querying power for locally held HIV-1 genetic, experimental and clinical data.

BACKGROUND: Human immunodeficiency virus (HIV) research involves ongoing, repetitious sequencing of the HIV genome and the massive accumulation of associated investigational data. As a result, the storage of annotated DNA and/or protein sequences, as well as information retrieval, have become increasingly difficult tasks, with scientists extracting less information from their collected data than they should. OBJECTIVES: Our objective was to design and develop a software package to aid researchers in the storage, analysis and exploration of their HIV-associated data. RESULTS: HIVbase contains familiar, easy-to-use interfaces and functionality for integrating many types of disparate data. The software contains tools that allow for the mass import of raw genetic data, eliminate repetitious sequence translations, have the ability to identify automatically and store HIV regions of interest from nucleic acid or protein sequences, allow for the export of data in commonly used analysis-ready formats, and for unique querying approaches.

Algorithms↗

Java-based application framework for visualization of gene regulatory region annotations.

MOTIVATION: The genome sequences of several organisms are either complete, or being sequenced. Each genome needs to be integrated with various types of annotations, e.g. locations of genes, promoters and other functional elements such as transcriptional regulatory elements. A robust application framework will be useful for developing web-based applications to visualize various genome annotations. RESULTS: We developed genome data visualization toolkit (GDVTK) as an application framework that consists of a set of data structures and core classes, using Java technology. GDVTK is a sound framework for developing web-based applications to present the gene regulatory region annotations in visual form. The current version of GDVTK consists of eight packages and 38 Java classes that are portable, reusable and extensible for plugging in new data sources and models. We implemented GDVTK for visualization of promoter annotations in Mammalian Promoter Database (MPromDb), a web-based gene-regulatory information server. AVAILABILITY: GDVTK is available under GNU general public license. Source code and software documentation can be found at the URL http://bioinformatics.med.ohio-state.edu/GDVTK.

Computer Graphics↗

Cluster Analyzer for Transcription Sites (CATS): a C++-based program for identifying clustered transcription factor binding sites.

SUMMARY: We have developed a program, Cluster Analyzer for Transcription Sites (CATS), which identifies clusters of transcription factor binding sites in any genome sequence. The program searches for clusters of the consensus sequence for DNA binding within a window (length of DNA). The window size and the cluster size (number of consensus sequences within a given window) can be varied. CATS can be used for single or multiple transcription factors for which consensus sequences have been deduced based on biochemical and mutational analysis, or by comparative genomics. The use of CATS for clusters of different transcription factor binding sites may facilitate the identification of genes that are co-regulated in a cell type-specific or developmental stage-specific manner. CATS is simple to install and use on computers running any Windows NT-platforms. AVAILABILITY: http://www.healthsciences.columbia.edu/dept/greenwaldlab/links.html

Binding Sites↗

CNplot: visualizing pre-clustered networks.

SUMMARY: CNplot is a simple technique for the visualization of global connectivity within pre-clustered network data. CNplot is easy to implement and in most cases produces informative and satisfactory summary of the data. AVAILABILITY: A Java implementation is available that allows users to modify graphics parameters and produces a LaTeX output. This software is free and is available at http://csb.stanford.edu/nbatada/VCN.html

Cluster Analysis↗

ParSeq: searching motifs with structural and biochemical properties.

SUMMARY: Searches for variable motifs such as protein-binding sites or promoter regions are more complex than the search for casual motifs. For example, in amino acid sequences comparing motifs alone mostly proves to be insufficient to detect regions that represent proteins with a special function, because the function depends on biochemical properties of individual amino acids (such as polarity or hydrophobicity). Pure string matching programs are not able to find these motifs; hence, we developed ParSeq, a program that combines the search for motifs with certain structural properties, the verification of biochemical properties, an approximate search mechanism and a stepwise creation of the motif description by allowing to search on previously obtained results. AVAILABILITY: http://www-pr.informatik.uni-tuebingen.de/parseq

Algorithms↗

galaxie--CGI scripts for sequence identification through automated phylogenetic analysis.

MOTIVATION: The prevalent use of similarity searches like BLAST to identify sequences and species implicitly assumes the reference database to be of extensive sequence sampling. This is often not the case, restraining the correctness of the outcome as a basis for sequence identification. Phylogenetic inference outperforms similarity searches in retrieving correct phylogenies and consequently sequence identities, and a project was initiated to design a freely available script package for sequence identification through automated Web-based phylogenetic analysis. RESULTS: Three CGI scripts were designed to facilitate qualified sequence identification from a Web interface. Query sequences are aligned to pre-made alignments or to alignments made by ClustalW with entries retrieved from a BLAST search. The subsequent phylogenetic analysis is based on the PHYLIP package for inferring neighbor-joining and parsimony trees. The scripts are highly configurable. AVAILABILITY: A service installation and a version for local use are found at http://andromeda.botany.gu.se/galaxiewelcome.html and http://galaxie.cgb.ki.se

Algorithms↗

Sight: automating genomic data-mining without programming skills.

SUMMARY: We created and tested Sight, a Java-based package that provides a user-friendly interface to generate and connect agents for automatic genomic data-mining for individual requirements without requiring programming skills from the user. AVAILABILITY: http://physiologie.uni-ulm.de//Seiten/Arbeitsgruppe/Jurkat-Rott/Jurkat-Rott.htm. The system does not require additional components and runs on IBM PCs under Windows (NT 4.0, 2000 and XP) or Linux (Phat 4.0 and Mandrake 9.0).

Algorithms↗

TOPALi: software for automatic identification of recombinant sequences within DNA multiple alignments.

SUMMARY: TOPALi is a new Java graphical analysis application that allows the user to identify recombinant sequences within a DNA multiple alignment (either automatically or via manual investigation). TOPALi allows a choice of three statistical methods to predict the positions of breakpoints due to past recombination. The breakpoint predictions are then used to identify putative recombinant sequences and their relationships to other sequences. In addition to its sophisticated interface, TOPALi can import many sequence formats, estimate and display phylogenetic trees and allow interactive analysis and/or automatic HTML report generation. AVAILABILITY: TOPALi is freely available from http://www.bioss.ac.uk/software.html

Computer Graphics↗

CisML: an XML-based format for sequence motif detection software.

SUMMARY: CisML is an XML-based format for sequence motif detection software. This proposed standard is applicable to many types of sequence motif detection programs. It is intended to facilitate the integration of data and the comparison of results from different software packages, and to simplify the development of downstream tools. XSL stylesheets are provided for easy generation of text, html and graphical reports from CisML-formatted data. AVAILABILITY: http://zlab.bu.edu/CisML/ SUPPLEMENTARY INFORMATION: Example CisML-formatted data and XSL stylesheets for report generation are available along with the sample output.

Amino Acid Motifs↗

Designing and executing scientific workflows with a programmable integrator.

MOTIVATION: As in many other fields of science, computational methods in molecular biology need to intersperse information access and algorithm execution in a computational workflow. Users often find difficulties when transferring data between data sources and applications. In most cases there is no standard solution for workflow design and execution and tailored scripting mechanisms are implemented in a case by case basis. RESULTS: In this paper, we present a general purpose 'programmable integrator' that can access information from a variety of sources in a coordinated manner. Its usefulness in complex bioinformatics applications is claimed and supported by some application examples. AVAILABILITY: Tools are freely available to non-profit educations and research institutions. Usage by commercial organizations requires a license agreement. Software requirements: Java v1.3 (http://java.sun.com), Xerces XML Parser (http://xml.apache.org/xerces-j) and Kweelt implementation of XQuery (http://kweelt.sourceforge.net/).

Algorithms↗