PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Augmented assessment as a means to augmented reality.

Rigorous scientific assessment of educational technologies typically lags behind the availability of the technologies by years because of the lack of validated instruments and benchmarks. Even when the appropriate assessment instruments are available, they may not be applied because of time and monetary constraints. Work in augmented reality, instrumented mannequins, serious gaming, and similar promising educational technologies that haven't undergone timely, rigorous evaluation, highlights the need for assessment methodologies that address the limitations of traditional approaches. The most promising augmented assessment solutions incorporate elements of rapid prototyping used in the software industry, simulation-based assessment techniques modeled after methods used in bioinformatics, and object-oriented analysis methods borrowed from object oriented programming.

Computer Simulation↗

Novel retinal genes discovered by mining the mouse embryonic RetinalExpress database.

PURPOSE: Bioinformatics has emerged as a powerful tool for identifying novel genes and pathways associated with retinal biology and disease. The developing mouse retina expresses an exceedingly large and complex variety of genes. Many of these genes have not been characterized but nevertheless are likely to have important developmental or physiological functions. The purpose of this study was to use an in silico approach with a mouse embryonic retinal database of cDNAs/expressed sequence tags (ESTs) named RetinalExpress to identify previously uncharacterized genes that are represented in the developing retina. METHODS: cDNA clones unique to the RetinalExpress database were identified by comparing clones in the RetinalExpress database with those in other cDNA/EST databases. We used a hierarchical filtering procedure with high stringency criteria that included sequence quality, colinearity with hypothetical gene sequences, and absence of any substantial existing annotation to select clones that were likely to represent novel genes. Selected clones were located on mouse chromosomes using National Center for Biotechnology Informatics Map Viewer software and the database from the University of California at Santa Cruz Genome Bioinformatics Web browser. The expression of selected retinal transcripts was determined using reverse transcriptase (RT)-PCR. In situ hybridization of sectioned embryonic and postnatal retinas was performed to determine spatial expression patterns of selected transcripts. RESULTS: Of the 27,765 cDNA clones from RetinalExpress that we filtered through several public cDNA/EST databases, 26 cDNA/EST sequences were identified that, at the time of the analysis, were unique to RetinalExpress. Seventeen clones were selected for RT-PCR analysis, and retinal transcripts corresponding to previously uncharacterized genes were unambiguously detected for six clones. Three genes encoded open reading frames containing putative functional domains; one sequence contained an HMG DNA binding domain, another, an RFX DNA binding domain, and another, a phospholipase C catalytic domain X. Transcripts from the genes encoding DNA binding domains were expressed in embryonic and postnatal retinas with distinct spatial patterns. CONCLUSIONS: The characterization of 26 mouse genes whose partial nucleotide sequences were uniquely represented in the RetinalExpress cDNA/EST database demonstrated the feasibility of retinal gene discovery using in silico analysis. Two of these genes had distinctive spatial expression patterns in the retina and one was likely to function as a DNA binding protein in embryonic and postnatal retinas. The gene identification approach described here demonstrates the usefulness of establishing large cDNA/EST databases from highly specialized neuronal tissues such as the retina to find novel genes.

Animals↗

Bioinformatics in mass spectrometry data analysis for proteomics studies.

Mass spectrometry is a technique widely employed for the identification and characterization of proteins. The role of bioinformatics is fundamental for the elaboration of mass spectrometry data due to the amount of data that this technique can produce. To process data efficiently, new software packages and algorithms are continuously being developed to improve protein identification and characterization in terms of high-throughput and statistical accuracy. However, many limitations exist concerning bioinformatics spectral data elaboration. This review aims to critically cover the recent and future developments of new bioinformatics approaches in mass spectrometry data analysis for proteomics studies.

Computational Biology↗

Design and implementation of a library-based information service in molecular biology and genetics at the University of Pittsburgh.

SETTING: In summer 2002, the Health Sciences Library System (HSLS) at the University of Pittsburgh initiated an information service in molecular biology and genetics to assist researchers with identifying and utilizing bioinformatics tools. PROGRAM COMPONENTS: This novel information service comprises hands-on training workshops and consultation on the use of bioinformatics tools. The HSLS also provides an electronic portal and networked access to public and commercial molecular biology databases and software packages. EVALUATION MECHANISMS: Researcher feedback gathered during the first three years of workshops and individual consultation indicate that the information service is meeting user needs. NEXT STEPS/FUTURE DIRECTIONS: The service's workshop offerings will expand to include emerging bioinformatics topics. A frequently asked questions database is also being developed to reuse advice on complex bioinformatics questions.

Computational Biology↗

RNA movies: visualizing RNA secondary structure spaces.

MOTIVATION: RNA Movies is a system for the visualization of RNA secondary structure spaces. Its input is a script consisting of primary and secondary structure information. From this script, the system fully automatically generates animated graphical structure representations. In this way, it creates the impression of an RNA molecule exploring its own two-dimensional structure space. RESULTS: RNA Movies has been used to generate animations of a switching structure in the spliced leader RNA of Leptomonas collosoma and sequential foldings of potato spindle tuber viroid transcripts. AVAILABILITY: Demonstrations of the animations mentioned in this paper can be viewed on our Bioinformatics web server under the following address: http://BiBiServ.TechFak.Uni-Bielefeld. DE/rnamovies/. The RNA Movies software is available upon request from the authors.

Algorithms↗

JDotter: a Java interface to multiple dotplots generated by dotter.

UNLABELLED: Java-Dotter (JDotter) is a platform-independent Java interactive interface for the Linux version of Dotter, a widely used program for generating dotplots of large DNA or protein sequences. JDotter runs as a client-server application and can send new sequences to the Dotter program for alignment as well as rapidly access a repository of preprocessed dotplots. JDotter also interfaces with a sequence database or file system to display supplementary feature data. Thus, JDotter greatly simplifies access to dotplot data in laboratories that deal with large numbers of genomes and have a multi-platform organization. AVAILABILITY: Currently, JDotter is used via Java Web Start by the Poxvirus Bioinformatics Resource for examining dotplots of complete poxvirus genomes; http://athena.bioc.uvic.ca/pbr/jdotter/. The software is available for download from the same location. SUPPLEMENTARY INFORMATION: Installation instructions, the User's Manual, screenshots and examples are available at the JDotter home page http://athena.bioc.uvic.ca/pbr/jdotter/. The software and source code is free for non-commercial applications.

Computer Graphics↗

Prospective use of DNA microarrays for evaluating renal function and disease.

At the forefront of the revolution in human genomics is DNA microarray technology, which evaluates expression levels or genotypes of thousands of genes simultaneously, by means of miniaturization and parallel processing. Furthermore, advances in bioinformatics will result in the creation of large databases, which will require complex software programming for structural analysis. Over the next decade, DNA microarrays, combined with sophisticated informatics and genomic databases, will provide molecular fingerprints of disease processes and prognoses. This review provides an update on DNA microarray technology and its application to renal diseases.

DNA↗

Discovery 99: accelerate and improve the drug discovery process. 26-29 April 1999, San Diego, CA, USA.

The stated purpose of this conference was the examination of emerging technologies for drug discovery. The meeting was divided into numerous pre- and post-conferences and tracks. A pre-conference workshop examined enabling software and business strategies. Nine tracks covered natural products, high-throughput screening, bioinformatics, proteomics/functional genomics, pharmacogenomics, combichem, chemoinformatics and assay methods. A post-conference symposium covered ADME-toxicology screening. In short, the gamut of drug discovery.

Journal Article↗

Support vector machine classification on the web.

The support vector machine (SVM) learning algorithm has been widely applied in bioinformatics. We have developed a simple web interface to our implementation of the SVM algorithm, called Gist. This interface allows novice or occasional users to apply a sophisticated machine learning algorithm easily to their data. More advanced users can download the software and source code for local installation. The availability of these tools will permit more widespread application of this powerful learning algorithm in bioinformatics.

Algorithms↗

caCORE: a common infrastructure for cancer informatics.

MOTIVATION: Sites with substantive bioinformatics operations are challenged to build data processing and delivery infrastructure that provides reliable access and enables data integration. Locally generated data must be processed and stored such that relationships to external data sources can be presented. Consistency and comparability across data sets requires annotation with controlled vocabularies and, further, metadata standards for data representation. Programmatic access to the processed data should be supported to ensure the maximum possible value is extracted. Confronted with these challenges at the National Cancer Institute Center for Bioinformatics, we decided to develop a robust infrastructure for data management and integration that supports advanced biomedical applications. RESULTS: We have developed an interconnected set of software and services called caCORE. Enterprise Vocabulary Services (EVS) provide controlled vocabulary, dictionary and thesaurus services. The Cancer Data Standards Repository (caDSR) provides a metadata registry for common data elements. Cancer Bioinformatics Infrastructure Objects (caBIO) implements an object-oriented model of the biomedical domain and provides Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. caCORE has been used to develop scientific applications that bring together data from distinct genomic and clinical science sources. AVAILABILITY: caCORE downloads and web interfaces can be accessed from links on the caCORE web site (http://ncicb.nci.nih.gov/core). caBIO software is distributed under an open source license that permits unrestricted academic and commercial use. Vocabulary and metadata content in the EVS and caDSR, respectively, is similarly unrestricted, and is available through web applications and FTP downloads. SUPPLEMENTARY INFORMATION: http://ncicb.nci.nih.gov/core/publications contains links to the caBIO 1.0 class diagram and the caCORE 1.0 Technical Guide, which provide detailed information on the present caCORE architecture, data sources and APIs. Updated information appears on a regular basis on the caCORE web site (http://ncicb.nci.nih.gov/core).

Animals↗

Open Source Software meets gene expression.

Use of the Open Source Software (OSS) development model has been crucial in a number of recent technological areas, including operating systems, applications and bioinformatics. The rationale for why OSS is often a better development model than proprietary development and some of the results of this model in the field of Gene Expression are reviewed. The paper concludes with a discussion of why funding agencies should endorse OSS and require funded software projects to be released Open Source.

Access to Information↗

Predicting the nuclear localization signals of 107 types of HPV L1 proteins by bioinformatic analysis.

In this study, 107 types of human papillomavirus (HPV) L1 protein sequences were obtained from available databases, and the nuclear localization signals (NLSs) of these HPV L1 proteins were analyzed and predicted by bioinformatic analysis. Out of the 107 types, the NLSs of 39 types were predicted by PredictNLS software (35 types of bipartite NLSs and 4 types of monopartite NLSs). The NLSs of the remaining HPV types were predicted according to the characteristics and the homology of the already predicted NLSs as well as the general rule of NLSs. According to the result, the NLSs of 107 types of HPV L1 proteins were classified into 15 categories. The different types of HPV L1 proteins in the same NLS category could share the similar or the same nucleocytoplasmic transport pathway. They might be used as the same target to prevent and treat different types of HPV infection. The results also showed that bioinformatic technology could be used to analyze and predict NLSs of proteins.

Amino Acid Sequence↗

Bioinformatic and experimental tools for identification of single-nucleotide polymorphisms in genes with a potential role for the development of the insulin resistance syndrome.

OBJECTIVES: Genes with a possible role for the development of the insulin resistance syndrome (IRS) were scanned for novel single-nucleotide polymorphisms (SNPs) using bioinformatics. METHODS: GenBank mRNA sequences were compared to the human EST database using gapped BLAST, software that is available on the internet. Mismatches between the search and the EST sequences indicated potential SNPs. Thirty-two SNPs in 13 genes were randomly chosen for experimental verification. PCR and direct sequencing were used to determine the 'true' SNPs. A random sample of 30 Swedish men with slightly elevated diastolic blood pressure (85-94 mmHg) obtained from a population-based study was selected for the sequencing. After completion of these stages, the potential SNPs were checked against the large and rapidly expanding SNP databases HGBASE and NCBI. RESULTS: EST searches of 146 genes revealed 106 potential SNPs in 44 genes. Experimental analysis of 32 of these potential SNPs verified two SNPs; endothelin receptor A 1471 G/C (3' UTR) and PAI-1 Trp514Arg from a T/C exchange. These two SNPs were also identified in the NCBI and HGBASE databases together with two polymorphisms that were not experimentally identified in our homogeneous Swedish population. Overall, the HGBASE and NCBI databases contained entries of 22% (23 out of 106) of the SNPs identified through our EST searches. CONCLUSIONS: In the search for genetic variations causing complex diseases like IRS in homogeneous populations (such as the Swedish one used here), important information can be obtained through bioinformatic searches of human genome databases and experimental verification.

Adult↗

Identification and genetic mapping of highly polymorphic microsatellite loci from an EST database of the septoria tritici blotch pathogen Mycosphaerella graminicola.

A database of 30,137 EST sequences from Mycosphaerella graminicola, the septoria tritici blotch fungus of wheat, was scanned with a custom software pipeline for di- and trinucleotide units repeated tandemly six or more times. The bioinformatics analysis identified 109 putative SSR loci, and for 99 of them, flanking primers were developed successfully and tested for amplification and polymorphism by PCR on five field isolates of diverse origin, including the parents of the standard M. graminicola mapping population. Seventy-seven of the 99 primer pairs generated an easily scored banding pattern and 51 were polymorphic, with up to four alleles per locus, among the isolates tested. Among these 51 loci, 23 were polymorphic between the parents of the mapping population. Twenty-one of these as well as two previously published microsatellite loci were positioned on the existing genetic linkage map of M. graminicola on 13 of the 24 linkage groups. Most (66%) of the primer pairs also amplified bands in the closely related barley pathogen Septoria passerinii, but only six were polymorphic among four isolates tested. A subset of the primer pairs also revealed polymorphisms when tested with DNA from the related banana black leaf streak (Black Sigatoka) pathogen, M. fijiensis. The EST database provided an excellent source of new, highly polymorphic microsatellite markers that can be multiplexed for high-throughput genetic analyses of M. graminicola and related species.

Alleles↗

Bioinformatic identification and characterization of new members of short-chain dehydrogenase/reductase superfamily.

With about 60 genes known in the human genome, short-chain dehydrogenases/reductases (SDRs) form a large gene family with important implications for medicine. They are known to be involved in carcinogenesis (e.g. breast and prostate cancer) as well as in metabolic and degenerative defects such as the pathogenesis of Alzheimer's disease, osteoporosis and diabetes. Uncharacterized SDRs are thus potential candidates for many monogenic and multifactorial human diseases. The identification and functional analysis of such SDR enzymes is therefore the primary goal of the study leading to new targets for drug development. In all taxa (bacteria, plants, insects, vertebrates), members of SDR superfamily are known. Up to now, there are several thousand members annotated many of which have not been characterized biochemically with regard to enzymatic activity, substrate specificity, or subcellular localization. We bioinformatically identified 250 vertebrate candidate genes belonging to the SDR superfamily using the BioNetWorks software SDR finder. The number was reduced to 95 after continuative analysis, including manual SDR motif verification and focus on human, rat and murine enzymes. Here, we present several new mammalian SDRs that were clustered into several enzymatically different groups by detailed phylogenetic analyses. Furthermore, characteristic mRNA expression patterns were identified for some of these genes by a recently developed in silico Northern blot method supporting their putative functions in retinoid, steroid, sugar and other metabolic pathways.

Amino Acid Sequence↗

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software↗

Workflows in bioinformatics: meta-analysis and prototype implementation of a workflow generator.

BACKGROUND: Computational methods for problem solving need to interleave information access and algorithm execution in a problem-specific workflow. The structures of these workflows are defined by a scaffold of syntactic, semantic and algebraic objects capable of representing them. Despite the proliferation of GUIs (Graphic User Interfaces) in bioinformatics, only some of them provide workflow capabilities; surprisingly, no meta-analysis of workflow operators and components in bioinformatics has been reported. RESULTS: We present a set of syntactic components and algebraic operators capable of representing analytical workflows in bioinformatics. Iteration, recursion, the use of conditional statements, and management of suspend/resume tasks have traditionally been implemented on an ad hoc basis and hard-coded; by having these operators properly defined it is possible to use and parameterize them as generic re-usable components. To illustrate how these operations can be orchestrated, we present GPIPE, a prototype graphic pipeline generator for PISE that allows the definition of a pipeline, parameterization of its component methods, and storage of metadata in XML formats. This implementation goes beyond the macro capacities currently in PISE. As the entire analysis protocol is defined in XML, a complete bioinformatic experiment (linked sets of methods, parameters and results) can be reproduced or shared among users. AVAILABILITY: http://if-web1.imb.uq.edu.au/Pise/5.a/gpipe.html (interactive), ftp://ftp.pasteur.fr/pub/GenSoft/unix/misc/Pise/ (download). CONCLUSION: From our meta-analysis we have identified syntactic structures and algebraic operators common to many workflows in bioinformatics. The workflow components and algebraic operators can be assimilated into re-usable software components. GPIPE, a prototype implementation of this framework, provides a GUI builder to facilitate the generation of workflows and integration of heterogeneous analytical tools.

Algorithms↗