PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Database-assisted promoter analysis.

The analysis of regulatory sequences is greatly facilitated by database-assisted bioinformatic approaches. The TRANSFAC database contains information on transcription factors and their origins, functional properties and sequence-specific binding activities. Software tools enable us to screen the database with a given DNA sequence for interacting transcription factors. If a regulatory function is already attributed to this sequence then the database-assisted identification of binding sites for proteins or protein classes and subsequent experimental verification might establish functionally relevant sites within this sequence. The binding transcription factors and interacting factors might already be present in the database.

Binding Sites↗

Application of in silico positional cloning and bioinformatic mutation analysis to the study of eye diseases.

A vast amount of DNA and protein sequence is now available and a plethora of programs have been developed to analyse the data. The bewildering variety of analyses that can be performed via the World-Wide Web can deter researchers from applying bioinformatics to augment their traditional genetic research. Focusing on the inherited eye diseases, this paper provides a guide to the appropriate software required for identification of candidate genes through to the detection and analysis of mutations.

Chromosome Mapping↗

Stochastic pairwise alignments.

MOTIVATION: The level of sequence conservation between related nucleic acids or proteins often varies considerably along the sequence. Both regions with high variability (mutational hot-spots) and regions of almost perfect sequence identity may occur in the same pair of molecules. The reliability of an alignment therefore strongly depends on the level of local sequence similarity. Especially in regions of high variability, many alignments of almost equal quality exist, and the optimal alignment is highly arbitrary. RESULTS: We discuss two approaches which deal with the inherent ambiguity of the alignment problem based on the computation of the partition function over all canonical pairwise alignments. The ensemble of possible alignments can be described by the probabilities P(ij) of a match between position i in the first and position j in the second sequence. Alternatively, we introduce a probabilistic backtracking procedure that generates ensembles of suboptimal alignments with correct statistical weights. A comparison between structure based alignments and large samples of stochastic alignments shows that the ensemble contains correct alignments with significant probabilities even though the optimal alignment deviates significantly from the structural alignment. Ensembles of suboptimal alignments obtained by stochastic backtracking can be used as input to any bioinformatics method based on pairwise alignment in order to gain reliability information not available from a single optimal alignment. AVAILABILITY: The software described in this contribution is available for downloading at http://www.tbi.univie.ac.at/~ulim/probA/

Algorithms↗

Functional genomics via multiscale analysis: application to gene expression and ChIP-on-chip data.

UNLABELLED: We present a fast, versatile and adaptive-multiscale algorithm for analyzing a wide-variety of DNA microarray data. Its primary application is in normalization of array data as well as subsequent identification of 'enriched targets', e.g. differentially expressed genes in expression profiling arrays and enriched sites in ChIP-on-chip experimental data. We show how to accommodate the unique characteristics of ChIP-on-chip data, where the set of 'enriched targets' is large, asymmetric and whose proportion to the whole data varies locally. SUPPLEMENTARY INFORMATION: Supplementary figures, related preprint, free software as well as our raw DNA microarray data with PCR validations are available at http://www.math.umn.edu/~lerman/supp/bioinfo06 as well as Bioinformatics online.

Algorithms↗

SeqVISTA: a new module of integrated computational tools for studying transcriptional regulation.

Transcriptional regulation is one of the most basic regulatory mechanisms in the cell. The accumulation of multiple metazoan genome sequences and the advent of high-throughput experimental techniques have motivated the development of a large number of bioinformatics methods for the detection of regulatory motifs. The regulatory process is extremely complex and individual computational algorithms typically have very limited success in genome-scale studies. Here, we argue the importance of integrating multiple computational algorithms and present an infrastructure that integrates eight web services covering key areas of transcriptional regulation. We have adopted the client-side integration technology and built a consistent input and output environment with a versatile visualization tool named SeqVISTA. The infrastructure will allow for easy integration of gene regulation analysis software that is scattered over the Internet. It will also enable bench biologists to perform an arsenal of analysis using cutting-edge methods in a familiar environment and bioinformatics researchers to focus on developing new algorithms without the need to invest substantial effort on complex pre- or post-processors. SeqVISTA is freely available to academic users and can be launched online at http://zlab.bu.edu/SeqVISTA/web.jnlp, provided that Java Web Start has been installed. In addition, a stand-alone version of the program can be downloaded and run locally. It can be obtained at http://zlab.bu.edu/SeqVISTA.

Algorithms↗

SPOP expression is associated with tumor-infiltrating lymphocytes in pancreatic cancer.

BACKGROUND: Speckle Type POZ Protein (SPOP), despite its tumor type-dependent role in tumorigenesis, primarily as a tumor suppressor gene is associated with a variety of different cancers. However, its function in pancreatic cancer remains uncertain. METHODS: SPOP expression and the association between its expression and patient prognosis and immune function were evaluated using The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), The Tumor Immune Estimation Resource 2.0 (TIMER2.0) database, cBioportal, and various bioinformatic databases. Enrichment analysis of SPOP and the association between SPOP expression with clinical stage and grade were analyzed using the R software package. Then immunohistochemistry (IHC) was used to estimate the correlation between SPOP and tumor-infiltrating lymphocytes (TILs) in patients with pancreatic cancer. RESULTS: As part of our study, we assessed that SPOP was anomalously expressed in kinds of cancers, associated with clinical stage and outcomes. Meanwhile, SPOP also played a crucial role in the tumor microenvironment (TME). The expression level of SPOP was significantly correlated to tumor-infiltrating immune cells (TICs) in pancreatic cancer. CONCLUSIONS: Our study uncovered the potential corrections in SPOP with TICs, suggesting that SPOP may act as a biomarker for immunotherapy in pancreatic cancer.

Humans↗

MetaCyc: a multiorganism database of metabolic pathways and enzymes.

MetaCyc is a database of metabolic pathways and enzymes located at http://MetaCyc.org/. Its goal is to serve as a metabolic encyclopedia, containing a collection of non-redundant pathways central to small molecule metabolism, which have been reported in the experimental literature. Most of the pathways in MetaCyc occur in microorganisms and plants, although animal pathways are also represented. MetaCyc contains metabolic pathways, enzymatic reactions, enzymes, chemical compounds, genes and review-level comments. Enzyme information includes substrate specificity, kinetic properties, activators, inhibitors, cofactor requirements and links to sequence and structure databases. Data are curated from the primary literature by curators with expertise in biochemistry and molecular biology. MetaCyc serves as a readily accessible comprehensive resource on microbial and plant pathways for genome analysis, basic research, education, metabolic engineering and systems biology. Querying, visualization and curation of the database is supported by SRI's Pathway Tools software. The PathoLogic component of Pathway Tools is used in conjunction with MetaCyc to predict the metabolic network of an organism from its annotated genome. SRI and the European Bioinformatics Institute employed this tool to create pathway/genome databases (PGDBs) for 165 organisms, available at the BioCyc.org website. These PGDBs also include predicted operons and pathway hole fillers.

Animals↗

Bioinformatics-assisted anti-HIV therapy.

Highly active antiretroviral therapy (HAART), in which three or more drugs are given in combination, has substantially improved the clinical management of HIV-1 infection. Still, the emergence of drug-resistant variants eventually leads to therapy failure in most patients. In such a scenario, the high diversity of resistance-associated mutational patterns complicates the choice of an optimal follow-up regimen. To support physicians in this task, a range of bioinformatics tools for predicting drug resistance or response to combination therapy from the viral genotype have been developed. With several free and commercial software services available, computational advice is rapidly gaining acceptance as an important element of rational decision-making in the treatment of HIV infection.

Anti-HIV Agents↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

A combination of chemical derivatisation and improved bioinformatic tools optimises protein identification for proteomics.

The identification of individual protein species within an organism's proteome has been optimised by increasing the information produced from mass spectral analysis through the chemical derivatisation of tryptic peptides and the development of new software tools. Peptide fragments are subjected to two forms of derivatisation. First, lysine residues are converted to homoarginine moieties by guanidination. This procedure has two advantages, first, it usually identifies the C-terminal amino acid of the tryptic peptide and also greatly increases the total information content of the mass spectrum by improving the signal response of C-terminal lysine fragments. Second, an Edman-type phenylthiocarbamoyl (PTC) modification is carried out on the N-terminal amino acid. The renders the first peptide bond highly susceptible to cleavage during mass spectrometry (MS) analysis and consequently allows the ready identification of the N-terminal residue. The utility of the procedure has been demonstrated by developing novel bioinformatic tools to exploit the additional mass spectral data in the identification of proteome proteins from the yeast Saccharomyces cerevisiae. With this combination of novel chemistry and bioinformatics, it should be possible to identify unambiguously any yeast protein spot or band from either two-dimensional or one-dimensional electropheretograms.

Databases, Factual↗

Plant-based microarray data at the European Bioinformatics Institute. Introducing AtMIAMExpress, a submission tool for Arabidopsis gene expression data to ArrayExpress.

ArrayExpress is a public microarray repository founded on the Minimum Information About a Microarray Experiment (MIAME) principles that stores MIAME-compliant gene expression data. Plant-based data sets represent approximately one-quarter of the experiments in ArrayExpress. The majority are based on Arabidopsis (Arabidopsis thaliana); however, there are other data sets based on Triticum aestivum, Hordeum vulgare, and Populus subsp. AtMIAMExpress is an open-source Web-based software application for the submission of Arabidopsis-based microarray data to ArrayExpress. AtMIAMExpress exports data in MAGE-ML format for upload to any MAGE-ML-compliant application, such as J-Express and ArrayExpress. It was designed as a tool for users with minimal bioinformatics expertise, has comprehensive help and user support, and represents a simple solution to meeting the MIAME guidelines for the Arabidopsis community. Plant data are queryable both in ArrayExpress and in the Data Warehouse databases, which support queries based on gene-centric and sample-centric annotation. The AtMIAMExpress submission tool is available at http://www.ebi.ac.uk/at-miamexpress/. The software is open source and is available from http://sourceforge.net/projects/miamexpress/. For information, contact miamexpress@ebi.ac.uk.

Academies and Institutes↗

Interpolated variable order motifs for identification of horizontally acquired DNA: revisiting the Salmonella pathogenicity islands.

MOTIVATION: There is a growing literature on the detection of Horizontal Gene Transfer (HGT) events by means of parametric, non-comparative methods. Such approaches rely only on sequence information and utilize different low and high order indices to capture compositional deviation from the genome backbone; the superiority of the latter over the former has been shown elsewhere. However even high order k-mers may be poor estimators of HGT, when insufficient information is available, e.g. in short sliding windows. Most of the current HGT prediction methods require pre-existing annotation, which may restrict their application on newly sequenced genomes. RESULTS: We introduce a novel computational method, Interpolated Variable Order Motifs (IVOMs), which exploits compositional biases using variable order motif distributions and captures more reliably the local composition of a sequence compared with fixed-order methods. For optimal localization of the boundaries of each predicted region, a second order, two-state hidden Markov model (HMM) is implemented in a change-point detection framework. We applied the IVOM approach to the genome of Salmonella enterica serovar Typhi CT18, a well-studied prokaryote in terms of HGT events, and we show that the IVOMs outperform state-of-the-art low and high order motif methods predicting not only the already characterized Salmonella Pathogenicity Islands (SPI-1 to SPI-10) but also three novel SPIs (SPI-15, SPI-16, SPI-17) and other HGT events. AVAILABILITY: The software is available under a GPL license as a standalone application at http://www.sanger.ac.uk/Software/analysis/alien_hunter CONTACT: gsv@sanger.ac.uk SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗