PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Programming Languages”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Pvclust: an R package for assessing the uncertainty in hierarchical clustering.

SUMMARY: Pvclust is an add-on package for a statistical software R to assess the uncertainty in hierarchical cluster analysis. Pvclust can be used easily for general statistical problems, such as DNA microarray analysis, to perform the bootstrap analysis of clustering, which has been popular in phylogenetic analysis. Pvclust calculates probability values (p-values) for each cluster using bootstrap resampling techniques. Two types of p-values are available: approximately unbiased (AU) p-value and bootstrap probability (BP) value. Multiscale bootstrap resampling is used for the calculation of AU p-value, which has superiority in bias over BP value calculated by the ordinary bootstrap resampling. In addition the computation time can be enormously decreased with parallel computing option.

Algorithms↗

A fast coarse filtering method for peptide identification by mass spectrometry.

MOTIVATION: We reformulate the problem of comparing mass-spectra by mapping spectra to a vector space model. Our search method leverages a metric space indexing algorithm to produce an initial candidate set, which can be followed by any fine ranking scheme. RESULTS: We consider three distance measures integrated into a multi-vantage point index structure. Of these, a semi-metric fuzzy-cosine distance using peptide precursor mass constraints performs the best. The index acts as a coarse, lossless filter with respect to the SEQUEST and ProFound scoring schemes, reducing the number of distance computations and returned candidates for fine filtering to about 0.5% and 0.02% of the database respectively. The fuzzy cosine distance term improves specificity over a peptide precursor mass filter, reducing the number of returned candidates by an order of magnitude. Run time measurements suggest proportional speedups in overall search times. Using an implementation of ProFound's Bayesian score as an example of a fine filter on a test set of Escherichia coli protein fragmentation spectra, the top results of our sample system are consistent with that of SEQUEST.

Algorithms↗

ArrayCluster: an analytic tool for clustering, data visualization and module finder on gene expression profiles.

SUMMARY: One of the significant challenges in gene expression analysis is to find unknown subtypes of several diseases at the molecular levels. This task can be addressed by grouping gene expression patterns of the collected samples on the basis of a large number of genes. Application of commonly used clustering methods to such a dataset however are likely to fail owing to over-learning, because the number of samples to be grouped is much smaller than the data dimension which is equal to the number of genes involved in the dataset. To overcome such difficulty, we developed a novel model-based clustering method, referred to as the mixed factors analysis. The ArrayCluster is a freely available software to perform the mixed factors analysis. It provides us some analytic tools for clustering DNA microarray experiments, data visualization and an automatic detector for module transcriptional of genes that are relevant to the calibrated molecular subtypes and so on.

Bayes Theorem↗

ALTree: association detection and localization of susceptibility sites using haplotype phylogenetic trees.

Finding the genes involved in complex diseases susceptibility and among those genes, localizing the variant sites explaining this susceptibility is a major goal of genetic epidemiology. In this context, haplotypic methods that use the joint information on several markers may be of particular interest. When the number of haplotypes is large, a grouping may be required. Phylogenetic trees allow such groupings of haplotypes based on their evolutionary history and may help in the detection and localization of disease susceptibility sites. In this paper, we present a new software to perform phylogeny-based association and localization analysis.

Computational Biology↗

SNAP: Combine and Map modules for multilocus population genetic analysis.

We have added two software tools to our Suite of Nucleotide Analysis Programs (SNAP) for working with DNA sequences sampled from populations. SNAP Map collapses DNA sequence data into unique haplotypes, extracts variable sites and manipulates output into multiple formats for input into existing software packages for evolutionary analyses. Map collapses DNA sequence data into unique haplotypes, extracts variable sites and manipulates output into multiple formats for input into existing software packages for evolutionary analyses. Map includes novel features such as recoding insertions or deletions, including or excluding variable sites that violate an infinite-sites model and the option of collapsing sequences with corresponding phenotypic information, important in testing for significant haplotype-phenotype associations. SNAP Combine merges multiple DNA sequence alignments into a single multiple alignment file. The resulting file can be the union or intersection of the input files. SNAP Combine currently reads from and writes to several sequence alignment file formats including both sequential and interleaved formats. Combine also keeps track of the start and end positions of each separate alignment file allowing the user to exclude variable sites or taxa, important in creating input files for multilocus analyses.

Algorithms↗

Two Sample Logo: a graphical representation of the differences between two sets of sequence alignments.

SUMMARY: Two Sample Logo is a web-based tool that detects and displays statistically significant differences in position-specific symbol compositions between two sets of multiple sequence alignments. In a typical scenario, two groups of aligned sequences will share a common motif but will differ in their functional annotation. The inclusion of the background alignment provides an appropriate underlying amino acid or nucleotide distribution and addresses intersite symbol correlations. In addition, the difference detection process is sensitive to the sizes of the aligned groups. Two Sample Logo extends WebLogo, a widely-used sequence logo generator. The source code is distributed under the MIT Open Source license agreement and is available for download free of charge.

Algorithms↗

Query Chem: a Google-powered web search combining text and chemical structures.

Query Chem (www.QueryChem.com) is a Web program that integrates chemical structure and text-based searching using publicly available chemical databases and Google's Web Application Program Interface (API). Query Chem makes it possible to search the Web for information about chemical structures without knowing their common names or identifiers. Furthermore, a structure can be combined with textual query terms to further restrict searches. Query Chem's search results can retrieve many interesting structure-property relationships of biomolecules on the Web.

Algorithms↗

Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.

MOTIVATION: In 2001 and 2002, we published two papers (Bioinformatics, 17, 282-283, Bioinformatics, 18, 77-82) describing an ultrafast protein sequence clustering program called cd-hit. This program can efficiently cluster a huge protein database with millions of sequences. However, the applications of the underlying algorithm are not limited to only protein sequences clustering, here we present several new programs using the same algorithm including cd-hit-2d, cd-hit-est and cd-hit-est-2d. Cd-hit-2d compares two protein datasets and reports similar matches between them; cd-hit-est clusters a DNA/RNA sequence database and cd-hit-est-2d compares two nucleotide datasets. All these programs can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on the popular sequence comparison and database search tools, such as BLAST.

Algorithms↗

BAli-Phy: simultaneous Bayesian inference of alignment and phylogeny.

SUMMARY: BAli-Phy is a Bayesian posterior sampler that employs Markov chain Monte Carlo to explore the joint space of alignment and phylogeny given molecular sequence data. Simultaneous estimation eliminates bias toward inaccurate alignment guide-trees, employs more sophisticated substitution models during alignment and automatically utilizes information in shared insertion/deletions to help infer phylogenies. AVAILABILITY: Software is available for download at http://www.biomath.ucla.edu/msuchard/bali-phy.

Algorithms↗

Metatool 5.0: fast and flexible elementary modes analysis.

SUMMARY: Elementary modes analysis is a powerful tool in the constraint-based modeling of metabolic networks. In recent years, new approaches to calculating elementary modes in biochemical reaction networks have been developed. As a consequence, the program Metatool, which is one of the first programs dedicated to this purpose, has been reimplemented in order to make use of these new approaches. The performance of Metatool has been significantly increased and the new version 5.0 can now be run inside the GNU octave or Matlab environments to allow more flexible usage and integration with other tools. AVAILABILITY: The script files and compiled shared libraries can be downloaded from the Metatool website at http://pinguin.biologie.uni-jena.de/bioinformatik/networks/index.html. Metatool consists of script files (m-files) for GNU octave as well as Matlab and shared libraries. The scripts are licensed under the GNU Public License and the use of the shared libraries is free for academic users and testing purposes. Commercial use of Metatool requires a special contract.

Computer Simulation↗

JADE: a distributed Java application for deleterious genomic mutation (DGM) estimation.

SUMMARY: The characterization of deleterious genomic mutation (DGM) is of central significance for evolutionary biology and genetic studies. Fitness moment method has been developed to efficiently characterize DGM from natural population directly. In order to enable researchers to employ this method for theoretical and empirical research on characterizing DGM, we here present a distributed Java Application for DGM Estimation (JADE). AVAILABILITY: http://orclinux.creighton.edu/DGM/index.htm.

Biological Evolution↗

ET viewer: an application for predicting and visualizing functional sites in protein structures.

SUMMARY: The Evolutionary Trace Viewer (ETV) provides a one-stop environment in which to run, visualize and interpret Evolutionary Trace (ET) predictions of functional sites in protein structures. ETV is implemented using Java to run across different operating systems using Java Web Start technology. AVAILABILITY: The ETV is available for download from our website at http://mammoth.bcm.tmc.edu/traceview/index.html. This webpage also links to sample trace results and a user manual that describes ET Viewer functions in detail.

Amino Acid Sequence↗

Building chromosome-wide LD maps.

SUMMARY: BMapBuilder builds maps of pairwise linkage disequilibrium (LD) in either two or three dimensions. The optimized resolution allows for graphical display of LD for single nucleotide polymorphisms (SNPs) in a whole chromosome. AVAILABILITY: The program is coded in Java, which runs on all relevant operating systems, including Windows, Mac and Unix/Linux, and is available from http://bios.ugr.es/BMapBuilder.

Chromosome Mapping↗

JCell--a Java-based framework for inferring regulatory networks from time series data.

MOTIVATION: JCell is a Java-based application for reconstructing gene regulatory networks from experimental data. The framework provides several algorithms to identify genetic and metabolic dependencies based on experimental data conjoint with mathematical models to describe and simulate regulatory systems. Owing to the modular structure, researchers can easily implement new methods. JCell is a pure Java application with additional scripting capabilities and thus widely usable, e.g. on parallel or cluster computers. AVAILABILITY: The software is freely available for download at http://www-ra.informatik.uni-tuebingen.de/software/JCell.

Algorithms↗

HCNet: a database of heart and calcium functional network.

SUMMARY: The Heart and Calcium functional Network (HCNet) database is a collection of functional gene modules calculated from the microarray data compendium available from the GEO database. It is a specialized database designed to assist experimentalists for cardiac calcium signaling research by providing the pre-calculated gene clusters and their potential correlation network in heart. In the current release of HCNet, 57 functional modules from 786 target genes obtained by a bi-clustering analysis of 381 microarray datasets are available. Detailed information of the clusters such as expression profiles, network diagrams is provided in two categories, heart-specific genes and heart-specific genes along with calcium toolkit genes. Overrepresented gene ontological categories and transcription factors in each cluster are also provided to infer the biological implications of the detected functional modules. AVAILABILITY: HCNet is available at http://sbrg2.gist.ac.kr/hcnet.

Algorithms↗

GEL: a novel genotype calling algorithm using empirical likelihood.

MOTIVATION: Preliminary results on the data produced using the Affymetrix large-scale genotyping platforms show that it is necessary to construct improved genotype calling algorithms. There is evidence that some of the existing algorithms lead to an increased error rate in heterozygous genotypes, and a disproportionately large rate of heterozygotes with missing genotypes. Non-random errors and missing data can lead to an increase in the number of false discoveries in genetic association studies. Therefore, the factors that need to be evaluated in assessing the performance of an algorithm are the missing data (call) and error rates, but also the heterozygous proportions in missing data and errors. RESULTS: We introduce a novel genotype calling algorithm (GEL) for the Affymetrix GeneChip arrays. The algorithm uses likelihood calculations that are based on distributions inferred from the observed data. A key ingredient in accurate genotype calling is weighting the information that comes from each probe quartet according to the quality/reliability of the data in the quartet, and prior information on the performance of the quartet. AVAILABILITY: The GEL software is implemented in R and is available by request from the corresponding author at nicolae@galton.uchicago.edu.

Algorithms↗