PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Proportion of membrane proteins in proteomes of 15 single-cell organisms analyzed by the SOSUI prediction system.

A software system, SOSUI, was previously developed for discriminating between soluble and membrane proteins and predicting transmembrane regions (Hirokawa et al., Bioinformatics, 14 (1998) 378-379). The performance of the system was 99% for the discrimination between two types of proteins and 96% for the prediction of transmembrane helices. When all of the amino acid sequences from 15 single-cell organisms were analyzed by SOSUI, the proportion of predicted polytopic membrane proteins showed an almost constant value of 15-20%, irrespective of the total genome size. However, single-cell organisms appeared to be categorized in terms of the preference of the number of transmembrane segments: species with small genomes were characterized by a significant peak at a helix number of approximately six or seven; species with large genomes showed a peak at 10 or 11 helices; and species with intermediate genome sizes showed a monotonous decrease of the population of membrane proteins against the number of transmembrane helices.

Archaea↗

Quantitative analysis of highly parallel transfection in cell microarrays.

As more genomes are sequenced, we are facing the challenge of rapidly unraveling the functions of genes. To that end, cell microarrays have recently been described that transfect thousands of nucleic acids in parallel and can be used to analyze the phenotypic consequences of such perturbations. As many parameters can influence the efficacy of transfection in such a format, we describe some important features in manufacturing cell microarrays that may improve reliability and efficiency of both plasmid DNA and siRNA transfection. We have also developed image analysis software that allows automatic detection of cell clusters, quantification of transfection efficiency and levels of expression/extinction of genes. Along with cell microarrays, this bioinformatic tool should expedite functional exploration of the human genome.

Automation↗

ProteinParser--a community based tool for the generation of a detailed protein consensus and FASTA output.

Comparison of bioinformatic data is a common application in the life sciences and beyond. In this communication, a novel Java based software tool, ProteinParser, is outlined. This software tool calculates a detailed consensus, or most common, amino acid at a given position in an aligned protein set, whilst also generating a full consensus protein FASTA output. A second application of this software tool, computing a consensus amino acid given a tolerance threshold, is also demonstrated. The phytase and the common bacterial beta-lactamase proteins are analysed as 'proof of concept' examples. Consensus proteins, as generated by ProteinParser, are regularly utilised in the selection of residues for protein stabilisation mutagenesis; however, this widely applicable software tool will find many alternative applications in areas such as protein homology modelling.

Amino Acid Sequence↗

The HIB database of annotated UniGene clusters.

SUMMARY: The HumanInfoBase (HIB) is a database of putative human gene transcripts. UniGene clusters are assembled, and the resulting consensus sequences are submitted to the PEDANT software system (Frishman,D., Albermann,K., Hani,J., Heumann,K., Metanomski,A., Zollner,A. and Mewes,H.-W., 2001, Bioinformatics, 17, 44--57) for fully automatic sequence analysis and annotation. Predicted transcripts are classified using a variety of functional and structural categories, and hyperlinks to various databases are provided for additional information. A WWW-based graphical user interface represents the assembly process as well as functionally important sites in the putative transcripts.

Data Collection↗

NEOBASE: databasing the neocortical microcircuit.

Mammals adapt to a rapidly changing world because of the sophisticated perceptual and cognitive function enabled by the neocortex. The neocortex, which has expanded to constitute nearly 80% of the human brain seems to have arisen from repeated duplication of a stereotypical template of neurons and synaptic circuits with subtle specializations in different brain regions and species. Determining the design and function of this microcircuitry is therefore of paramount importance to understanding normal and abnormal higher brain function. Recent advances in recording synaptically-coupled neurons has allowed rapid dissection of the neocortical microcircuitry thus yielding a massive amount of quantitative anatomical, electrical and gene expression data on the neurons and the synaptic circuits that connect the neurons. Due to the availability of the above mentioned data, it has now become imperative to database the neurons of the microcircuit and their synaptic connections. The NEOBASE project, aims to archive the neocortical microcircuit data in a manner that facilitates development of advanced data mining applications, statistical and bioinformatics analyses tools, custom microcircuit builders, and visualization and simulation applications. The database architecture is based on ROOT, a software environment that allows the construction of an object oriented database with numerous relational capabilities. The proposed architecture allows construction of a database that closely mimics the architecture of the real microcircuit, which facilitates the interface with virtually any application, allows for data format evolution, and aims for full interoperability with other databases. NEOBASE will provide an important resource and research tool for studying the microcircuit basis of normal and abnormal neocortical function. The database will be available to local as well as remote users using Grid based tools and technologies.

Animals↗

Dynamic molecules: molecular dynamics for everyone. An internet-based access to molecular dynamic simulations: basic concepts.

Molecular dynamics is a rapidly developing field of science and has become an established tool for studying the dynamic behavior of biomolecules. Although several high quality programs for performing molecular dynamic simulations are freely available, only well-trained scientists are currently able to make use of the broad scientific potential that molecular dynamic simulations offer to gain insight into structural questions at an atomic level. The "Dynamic Molecules" approach is the first internet portal that provides an interactive access to set up, perform and analyze molecular dynamic simulations. It is completely based on standard web technologies and uses only publicly available software. The aim is to open molecular dynamics techniques to a broader range of users including undergraduate students, teachers and scientists outside the bioinformatics field. The time-limiting factors are the availability of free capacity on the computing server to run the simulations and the time required to transport the history file through the internet for the animation mode. The interactive access mode of the portal is acceptable for animations of molecules having up to about 500 atoms.

Computer Simulation↗

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein↗

A space-efficient algorithm for aligning large genomic sequences.

SUMMARY: In the segment-by-segment approach to sequence alignment, pairwise and multiple alignments are generated by comparing gap-free segments of the sequences under study. This method is particularly efficient in detecting local homologies, and it has been used to identify functional regions in large genomic sequences. Herein, an algorithm is outlined that calculates optimal pairwise segment-by-segment alignments in essentially linear space. AVAILABILTIY: The program is available at the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak. uni-bielefeld.de/dialign/

Algorithms↗

A novel approach for clustering proteomics data using Bayesian fast Fourier transform.

MOTIVATION: Bioinformatics clustering tools are useful at all levels of proteomic data analysis. Proteomics studies can provide a wealth of information and rapidly generate large quantities of data from the analysis of biological specimens. The high dimensionality of data generated from these studies requires the development of improved bioinformatics tools for efficient and accurate data analyses. For proteome profiling of a particular system or organism, a number of specialized software tools are needed. Indeed, significant advances in the informatics and software tools necessary to support the analysis and management of these massive amounts of data are needed. Clustering algorithms based on probabilistic and Bayesian models provide an alternative to heuristic algorithms. The number of clusters (diseased and non-diseased groups) is reduced to the choice of the number of components of a mixture of underlying probability. The Bayesian approach is a tool for including information from the data to the analysis. It offers an estimation of the uncertainties of the data and the parameters involved. RESULTS: We present novel algorithms that can organize, cluster and derive meaningful patterns of expression from large-scaled proteomics experiments. We processed raw data using a graphical-based algorithm by transforming it from a real space data-expression to a complex space data-expression using discrete Fourier transformation; then we used a thresholding approach to denoise and reduce the length of each spectrum. Bayesian clustering was applied to the reconstructed data. In comparison with several other algorithms used in this study including K-means, (Kohonen self-organizing map (SOM), and linear discriminant analysis, the Bayesian-Fourier model-based approach displayed superior performances consistently, in selecting the correct model and the number of clusters, thus providing a novel approach for accurate diagnosis of the disease. Using this approach, we were able to successfully denoise proteomic spectra and reach up to a 99% total reduction of the number of peaks compared to the original data. In addition, the Bayesian-based approach generated a better classification rate in comparison with other classification algorithms. This new finding will allow us to apply the Fourier transformation for the selection of the protein profile for each sample, and to develop a novel bioinformatic strategy based on Bayesian clustering for biomarker discovery and optimal diagnosis.

Algorithms↗

Nuclear Receptor Signaling Atlas (www.nursa.org): hyperlinking the nuclear receptor signaling community.

The nuclear receptor signaling (NRS) field has generated a substantial body of information on nuclear receptors, their ligands and coregulators, with the ultimate goal of constructing coherent models of the biological and clinical significance of these molecules. As a component of the Nuclear Receptor Signaling Atlas (NURSA)--the development of a functional atlas of nuclear receptor biology--the NURSA Bioinformatics Resource is developing a strategy to organize and integrate legacy and future information on these molecules in a single web-based resource (www.nursa.org). This entails parallel efforts of (i) developing an appropriate software framework for handling datasets from NURSA laboratories and (ii) designing strategies for the curation and presentation of public data relevant to NRS. To illustrate our approach, we have described here in detail the development of a web-based interface for the NURSA quantitative PCR nuclear receptor expression dataset, incorporating bioinformatics analysis which provides novel perspectives on functional relationships between these molecules. We anticipate that the free and open access of the community to a platform for data mining and hypothesis generation strategies will be a significant contribution to the progress of research in this field.

Animals↗

EST-PAC a web package for EST annotation and protein sequence prediction.

With the decreasing cost of DNA sequencing technology and the vast diversity of biological resources, researchers increasingly face the basic challenge of annotating a larger number of expressed sequences tags (EST) from a variety of species. This typically consists of a series of repetitive tasks, which should be automated and easy to use. The results of these annotation tasks need to be stored and organized in a consistent way. All these operations should be self-installing, platform independent, easy to customize and amenable to using distributed bioinformatics resources available on the Internet. In order to address these issues, we present EST-PAC a web oriented multi-platform software package for expressed sequences tag (EST) annotation. EST-PAC provides a solution for the administration of EST and protein sequence annotations accessible through a web interface. Three aspects of EST annotation are automated: 1) searching local or remote biological databases for sequence similarities using Blast services, 2) predicting protein coding sequence from EST data and, 3) annotating predicted protein sequences with functional domain predictions. In practice, EST-PAC integrates the BLASTALL suite, EST-Scan2 and HMMER in a relational database system accessible through a simple web interface. EST-PAC also takes advantage of the relational database to allow consistent storage, powerful queries of results and, management of the annotation process. The system allows users to customize annotation strategies and provides an open-source data-management environment for research and education in bioinformatics.

Journal Article↗

The DNA sequence quality machine at IFOM: a simple Web-based tool for quantitative assessment of sequencing reactions.

DNA sequence quality is a factor of paramount importance in the world of modern genetic and genomics. Both the sequencing of Human Genome in the "post-draft" era [NHGRI Standard for quality of Human Genomic Sequences, Rev. 7 July (2002) where http://www.nhgri.nih.gov/Grant_info/Funding/ Statements/RFA/quality_standard.html is the HTTP address] and recent "high-throughput" approaches to genetic investigation such as SAGE [Velculescu, V.E., Zhang, L., Vogelstein, B. et al. (1995) "Serial analysis of gene expression", Science 270, 484-487] need a reliable, standardized measure of the quality of a sequencing reaction. The increasing importance of SNP studies also requires a stronger quality control on sequencing reactions by the final user. We propose here a simple, web-based tool for integrated sequence quality evaluation, high quality region quantitative value calculation and chromatogram display. This software is aimed at the small to medium DNA sequence laboratory or to the single researcher, interested in getting a quantitative measure of the sequence quality, browsing the chromatogram and checking the quality values base by base. The program is freely available from the IFOM bioinformatics web Server at http://bio.ifom-firc.it/Phred20/index.html.

Algorithms↗

A multi-step approach to time series analysis and gene expression clustering.

MOTIVATION: The huge growth in gene expression data calls for the implementation of automatic tools for data processing and interpretation. RESULTS: We present a new and comprehensive machine learning data mining framework consisting in a non-linear PCA neural network for feature extraction, and probabilistic principal surfaces combined with an agglomerative approach based on Negentropy aimed at clustering gene microarray data. The method, which provides a user-friendly visualization interface, can work on noisy data with missing points and represents an automatic procedure to get, with no a priori assumptions, the number of clusters present in the data. Cell-cycle dataset and a detailed analysis confirm the biological nature of the most significant clusters. AVAILABILITY: The software described here is a subpackage part of the ASTRONEURAL package and is available upon request from the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Artificial Intelligence↗