PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

PET-Tool: a software suite for comprehensive processing and managing of Paired-End diTag (PET) sequence data.

BACKGROUND: We recently developed the Paired End diTag (PET) strategy for efficient characterization of mammalian transcriptomes and genomes. The paired end nature of short PET sequences derived from long DNA fragments raised a new set of bioinformatics challenges, including how to extract PETs from raw sequence reads, and correctly yet efficiently map PETs to reference genome sequences. To accommodate and streamline data analysis of the large volume PET sequences generated from each PET experiment, an automated PET data process pipeline is desirable. RESULTS: We designed an integrated computation program package, PET-Tool, to automatically process PET sequences and map them to the genome sequences. The Tool was implemented as a web-based application composed of four modules: the Extractor module for PET extraction; the Examiner module for analytic evaluation of PET sequence quality; the Mapper module for locating PET sequences in the genome sequences; and the Project Manager module for data organization. The performance of PET-Tool was evaluated through the analyses of 2.7 million PET sequences. It was demonstrated that PET-Tool is accurate and efficient in extracting PET sequences and removing artifacts from large volume dataset. Using optimized mapping criteria, over 70% of quality PET sequences were mapped specifically to the genome sequences. With a 2.4 GHz LINUX machine, it takes approximately six hours to process one million PETs from extraction to mapping. CONCLUSION: The speed, accuracy, and comprehensiveness have proved that PET-Tool is an important and useful component in PET experiments, and can be extended to accommodate other related analyses of paired-end sequences. The Tool also provides user-friendly functions for data quality check and system for multi-layer data management.

Animals↗

Ets1 was significantly activated by ERK1/2 in mutant K-ras stably transfected human adrenocortical cells.

In our previous study on the tumorigenesis of human functional adrenal tumors, we observed a high frequency of point mutation in the K-ras gene in clinical adrenal tumors. Therefore, we analyzed gene profiles of mutant K-ras transfected adrenocortical cells by DNA microarray to determine the expression pattern of genes related to cell cycle, signal transduction, apoptosis, tumorigenesis, steroidogenesis, and other expressed sequence tags (ESTs). Then we analyzed all of the significant differentially expressed genes by bioinformatics tools, "Matchminer" and "Gominer." The results revealed that expression of mutant K-ras gene induced by IPTG upregulated Ets1, which was mainly related to cell proliferation. After carefully being analyzed by software "DAVID" and "Pathart," Ets1 was found to be activated by being phosphorylated at theronine 38 by ERK1/2, and in turn, to regulate the following genes: uPA, MMP-3, and prolactin (Ling et al., 2003; Duffy and Daggan, 2004; Maupas-Schwalm et al., 2004; van Themsche et al., 2004). The result of Western blotting analysis confirmed that Ets1 was really phosphorylated when mutant K-ras was activated. On the other hand, the membrane blotting analyses indicated that the expression levels of uPA, MMP-3, and prolactin in human adrenocortical cells stably transfected with the mutant K-ras gene were significantly higher than those in normal control cells. Compared to control cells, the level of prolactin raised 1.4-fold, the level of MMP-3 raised 1.8-fold, and the level of uPA raised 2.1-fold in the transfected cells. From the results of this study, we proposed a mechanism of Ets1 in human adrenocortical cells expressing a mutated K-ras gene.

Adrenal Cortex↗

Comparison of Affymetrix GeneChip expression measures.

MOTIVATION: In the Affymetrix GeneChip system, preprocessing occurs before one obtains expression level measurements. Because the number of competing preprocessing methods was large and growing we developed a benchmark to help users identify the best method for their application. A webtool was made available for developers to benchmark their procedures. At the time of writing over 50 methods had been submitted. RESULTS: We benchmarked 31 probe set algorithms using a U95A dataset of spike in controls. Using this dataset, we found that background correction, one of the main steps in preprocessing, has the largest effect on performance. In particular, background correction appears to improve accuracy but, in general, worsen precision. The benchmark results put this balance in perspective. Furthermore, we have improved some of the original benchmark metrics to provide more detailed information regarding precision and accuracy. A handful of methods stand out as providing the best balance using spike-in data with the older U95A array, although different experiments on more current arrays may benchmark differently. AVAILABILITY: The affycomp package, now version 1.5.2, continues to be available as part of the Bioconductor project (http://www.bioconductor.org). The webtool continues to be available at http://affycomp.biostat.jhsph.edu CONTACT: rafa@jhu.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

[Cloning of an expressed sequence tag with restriction display polymerase chain reaction].

OBJECTIVE: To isolate gene fragments from SH-SY5Y cells by way of restriction display polymerase chain reaction (RD-PCR). METHODS: Total mRNA was extracted from SH-SY5Y cells followed by synthesis of the single-strand cDNA with Oligo (dT18) as the anchored primer, and the second strand was synthesized by nick translation. The double strands were cleft with restriction enzyme Sau3A I and the fragments ligated with a universal adapter to be amplified with the universal primers and selected primers. The products were then ligated into the pMD18-T vector and sequenced. RESULTS: One of the sequenced clones was retrieved in the National Center for Biotechnology Information (NCBI) databases with Blast program. The results showed that the sequence possessed great similarity to one fragment of the 17th chromosome in the genome. Sequence analysis with GenScan software indicated that the EST might be one section of an unknown gene. CONCLUSION: RD-PCR provides simple and efficient approach for isolating EST from cells, and cDNA clone sequencing combined with bioinformatics analysis may be helpful in identifying new genes.

Base Sequence↗

Application of machine learning in SNP discovery.

BACKGROUND: Single nucleotide polymorphisms (SNP) constitute more than 90% of the genetic variation, and hence can account for most trait differences among individuals in a given species. Polymorphism detection software PolyBayes and PolyPhred give high false positive SNP predictions even with stringent parameter values. We developed a machine learning (ML) method to augment PolyBayes to improve its prediction accuracy. ML methods have also been successfully applied to other bioinformatics problems in predicting genes, promoters, transcription factor binding sites and protein structures. RESULTS: The ML program C4.5 was applied to a set of features in order to build a SNP classifier from training data based on human expert decisions (True/False). The training data were 27,275 candidate SNP generated by sequencing 1973 STS (sequence tag sites) (12 Mb) in both directions from 6 diverse homozygous soybean cultivars and PolyBayes analysis. Test data of 18,390 candidate SNP were generated similarly from 1359 additional STS (8 Mb). SNP from both sets were classified by experts. After training the ML classifier, it agreed with the experts on 97.3% of test data compared with 7.8% agreement between PolyBayes and experts. The PolyBayes positive predictive values (PPV) (i.e., fraction of candidate SNP being real) were 7.8% for all predictions and 16.7% for those with 100% posterior probability of being real. Using ML improved the PPV to 84.8%, a 5- to 10-fold increase. While both ML and PolyBayes produced a similar number of true positives, the ML program generated only 249 false positives as compared to 16,955 for PolyBayes. The complexity of the soybean genome may have contributed to high false SNP predictions by PolyBayes and hence results may differ for other genomes. CONCLUSION: A machine learning (ML) method was developed as a supplementary feature to the polymorphism detection software for improving prediction accuracies. The results from this study indicate that a trained ML classifier can significantly reduce human intervention and in this case achieved a 5-10 fold enhanced productivity. The optimized feature set and ML framework can also be applied to all polymorphism discovery software. ML support software is written in Perl and can be easily integrated into an existing SNP discovery pipeline.

Algorithms↗

PROTEOME-3D: an interactive bioinformatics tool for large-scale data exploration and knowledge discovery.

Comprehensive understanding of biological systems requires efficient and systematic assimilation of high-throughput datasets in the context of the existing knowledge base. A major limitation in the field of proteomics is the lack of an appropriate software platform that can synthesize a large number of experimental datasets in the context of the existing knowledge base. Here, we describe a software platform, termed PROTEOME-3D, that utilizes three essential features for systematic analysis of proteomics data: creation of a scalable, queryable, customized database for identified proteins from published literature; graphical tools for displaying proteome landscapes and trends from multiple large-scale experiments; and interactive data analysis that facilitates identification of crucial networks and pathways. Thus, PROTEOME-3D offers a standardized platform to analyze high-throughput experimental datasets for the identification of crucial players in co-regulated pathways and cellular processes.

Computational Biology↗

Building an asynchronous web-based tool for machine learning classification.

Various unsupervised and supervised learning methods including support vector machines, classification trees, linear discriminant analysis and nearest neighbor classifiers have been used to classify high-throughput gene expression data. Simpler and more widely accepted statistical tools have not yet been used for this purpose, hence proper comparisons between classification methods have not been conducted. We developed free software that implements logistic regression with stepwise variable selection as a quick and simple method for initial exploration of important genetic markers in disease classification. To implement the algorithm and allow our collaborators in remote locations to evaluate and compare its results against those of other methods, we developed a user-friendly asynchronous web-based application with a minimal amount of programming using free, downloadable software tools. With this program, we show that classification using logistic regression can perform as well as other more sophisticated algorithms, and it has the advantages of being easy to interpret and reproduce. By making the tool freely and easily available, we hope to promote the comparison of classification methods. In addition, we believe our web application can be used as a model for other bioinformatics laboratories that need to develop web-based analysis tools in a short amount of time and on a limited budget.

Algorithms↗

Can we integrate bioinformatics data on the Internet?

The NETTAB (Network Tools and Applications in Biology) 2001 Workshop entitled 'CORBA and XML: towards a bioinformatics-integrated network environment' was held at the Advanced Biotechnology Centre, Genoa, Italy, 17-18 May 2001.

Computational Biology↗

DIALIGN 2: improvement of the segment-to-segment approach to multiple sequence alignment.

MOTIVATION: The performance and time complexity of an improved version of the segment-to-segment approach to multiple sequence alignment is discussed. In this approach, alignments are composed from gap-free segment pairs, and the score of an alignment is defined as the sum of so-called weights of these segment pairs. RESULTS: A modification of the weight function used in the original version of the alignment program DIALIGN has two important advantages: it can be applied to both globally and locally related sequence sets, and the running time of the program is considerably improved. The time complexity of the algorithm is discussed theoretically, and the program running time is reported for various test examples. AVAILABILITY: The program is available on-line at the Bielefeld University Bioinformatics Server (BiBiServ) http://bibiserv.TechFak.Uni-Bielefeld.DE/dial ign/

Algorithms↗

Evaluation of ontology merging tools in bioinformatics.

Ontologies are being used nowadays in many areas, including bioinformatics. One of the issues in ontology research is the aligning and merging of ontologies. Tools have been developed for ontology merging, but they have not been evaluated for their use in bioinformatics. In this paper we evaluate two of the most well-known ontology merging tools with a bioinformatics perspective. As test ontologies we have used Gene Ontology and Signal-Ontology.

Computational Biology↗

Sigma: multiple alignment of weakly-conserved non-coding DNA sequence.

BACKGROUND: Existing tools for multiple-sequence alignment focus on aligning protein sequence or protein-coding DNA sequence, and are often based on extensions to Needleman-Wunsch-like pairwise alignment methods. We introduce a new tool, Sigma, with a new algorithm and scoring scheme designed specifically for non-coding DNA sequence. This problem acquires importance with the increasing number of published sequences of closely-related species. In particular, studies of gene regulation seek to take advantage of comparative genomics, and recent algorithms for finding regulatory sites in phylogenetically-related intergenic sequence require alignment as a preprocessing step. Much can also be learned about evolution from intergenic DNA, which tends to evolve faster than coding DNA. Sigma uses a strategy of seeking the best possible gapless local alignments (a strategy earlier used by DiAlign), at each step making the best possible alignment consistent with existing alignments, and scores the significance of the alignment based on the lengths of the aligned fragments and a background model which may be supplied or estimated from an auxiliary file of intergenic DNA. RESULTS: Comparative tests of sigma with five earlier algorithms on synthetic data generated to mimic real data show excellent performance, with Sigma balancing high "sensitivity" (more bases aligned) with effective filtering of "incorrect" alignments. With real data, while "correctness" can't be directly quantified for the alignment, running the PhyloGibbs motif finder on pre-aligned sequence suggests that Sigma's alignments are superior. CONCLUSION: By taking into account the peculiarities of non-coding DNA, Sigma fills a gap in the toolbox of bioinformatics.

Algorithms↗

caGEDA: a web application for the integrated analysis of global gene expression patterns in cancer.

The explosion of microarray data from pilot studies, basic research and large-scale clinical trials requires the development of integrative computational tools that can not only analyse gene expression patterns but that can also evaluate the methods of analysis adopted and then provide a boost to post-analysis translational interpretation of those patterns. We have developed a web application called caGEDA (cancer gene expression data analyzer) that can: (1) upload gene expression profiles from cDNA or oligonucleotide microarrays; (2) conduct a diverse range of serial linear normalisations; (3) identify differentially expressed genes using a variety of tests - either threshold or permutation tests; (4) produce tables of literature references to papers reporting that specific genes (identified by accession numbers) are up- or down-regulated in specific cancers; (5) estimate the error of sample class prediction using the significant gene set for features; (6) perform low-bias and accurate validated learning using three computational validation techniques (leave-one out validation, k-fold validation, random re-sampling validation); and (7) validate a classifier with a randomly selected or user-defined validation set. Significant genes are reported in a table of links to entries in the following databases: Locus Link, Genome View, UCSC, Ensembl, UniGene, dbSNP, AmiGO and OMIM. caGEDA is seamlessly integrated via embedded forms with UCSD's (University of California at San Diego) 2HAPI server (for medical subject heading (MeSH) term exploration) and EZ-Retrieve (to identify common transcription factors located upstream of sets of genes that exhibit similar modes of differential expression). caGEDA offers a variety of previously described and novel tests for differentially expressed genes, most notably the permutation percentile separability test, which is most appropriate for identifying genes that are significantly differentially expressed in a subset of patients. caGEDA, which is open source and free to academic users, will soon be greatly enhanced by operating with the components of the National Cancer Institute's new cancer bioinformatics grid (caBIG).

Biomarkers, Tumor↗

Prediction and visualization of structural switches in RNA.

There are various cases where the biological function of an RNA molecule involves a reversible change of conformation. paRNAss is a software approach to the prediction of such structural switching in RNA. It is based on three hypotheses about the secondary structure space of a switching RNA molecule, which can be evaluated by RNA folding and structure comparison. In the positive case, the predicted structures must be verified experimentally. Additionally, we give an animated visualization of an energetically favourable transition between the predicted structures. paRNAss is available via the Bielefeld Bioinformatics Server. This paper explains the underlying model and shows that the approach performs well in a variety of applications.

Base Sequence↗

AMDA: an R package for the automated microarray data analysis.

BACKGROUND: Microarrays are routinely used to assess mRNA transcript levels on a genome-wide scale. Large amount of microarray datasets are now available in several databases, and new experiments are constantly being performed. In spite of this fact, few and limited tools exist for quickly and easily analyzing the results. Microarray analysis can be challenging for researchers without the necessary training and it can be time-consuming for service providers with many users. RESULTS: To address these problems we have developed an automated microarray data analysis (AMDA) software, which provides scientists with an easy and integrated system for the analysis of Affymetrix microarray experiments. AMDA is free and it is available as an R package. It is based on the Bioconductor project that provides a number of powerful bioinformatics and microarray analysis tools. This automated pipeline integrates different functions available in the R and Bioconductor projects with newly developed functions. AMDA covers all of the steps, performing a full data analysis, including image analysis, quality controls, normalization, selection of differentially expressed genes, clustering, correspondence analysis and functional evaluation. Finally a LaTEX document is dynamically generated depending on the performed analysis steps. The generated report contains comments and analysis results as well as the references to several files for a deeper investigation. CONCLUSION: AMDA is freely available as an R package under the GPL license. The package as well as an example analysis report can be downloaded in the Services/Bioinformatics section of the Genopolis http://www.genopolis.it/.

Algorithms↗

[Cloning and subcellular localization of apr-1--a new gene of tumor specific antigen family].

BACKGROUND & OBJECTIVE: apr-1 was cloned by improved polymerase chain reaction (PCR)-based subtractive hybridization from all-trans retinoic acid (ATRA)-induced apoptotic leukemia HL-60 cells in 1999. Preliminary results showed that apr-1 might be an apoptosis-related gene (GenBank ID: NM_014061). This study was to explore the background of apr-1 through gene cloning, bioinformatic analysis, and subcellular locating. METHODS: The cDNA encoding Apr-1 was amplified by reverse transcription-PCR (RT-PCR), and sequenced. Open reading frame (ORF) of apr-1 was analyzed with ORF finder software. Chromosome locus was defined by genome blast software. Conserved domains of amino acids were analyzed by protein blast software. Align (Cluster W) software in Vector NTI software package was used to analyze homogeneous genes (or proteins), and to draw the Phylogenetic Tree. Subcellular localization of apr-1 was performed. RESULTS: apr-1 was mapped to chromosome Xp11.22 with the ORF locating in 1 exon. Two MAGE conserved domains were found in Apr-1. Apr-1 shared homology with MAGE-A1, MAGE-B1, MAGE-C1, MAGE-D1, and Necdin. Phylogenetic analysis showed that Apr-1 was more closely related to MAGE-D1 and Necdin. Gene products of apr-1 were located in the nuclei of eukaryocytes. CONCLUSIONS: apr-1 is a member of MAGE family, and might belong to type II MAGE genes.

Amino Acid Sequence↗