PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Development of public toxicogenomics software for microarray data management and analysis.

A robust bioinformatics capability is widely acknowledged as central to realizing the promises of toxicogenomics. Successful application of toxicogenomic approaches, such as DNA microarray, inextricably relies on appropriate data management, the ability to extract knowledge from massive amounts of data and the availability of functional information for data interpretation. At the FDA's National Center for Toxicological Research (NCTR), we are developing a public microarray data management and analysis software, called ArrayTrack. ArrayTrack is Minimum Information About a Microarray Experiment (MIAME) supportive for storing both microarray data and experiment parameters associated with a toxicogenomics study. A quality control mechanism is implemented to assure the fidelity of entered expression data. ArrayTrack also provides a rich collection of functional information about genes, proteins and pathways drawn from various public biological databases for facilitating data interpretation. In addition, several data analysis and visualization tools are available with ArrayTrack, and more tools will be available in the next released version. Importantly, gene expression data, functional information and analysis methods are fully integrated so that the data analysis and interpretation process is simplified and enhanced. ArrayTrack is publicly available online and the prospective user can also request a local installation version by contacting the authors.

Databases, Genetic↗

BioWareDB: the biomedical software and database search engine.

UNLABELLED: A wealth of bioinformatics tools and databases has been created over the last decade and most are freely available to the general public. However, these valuable resources live a shadow existence compared to experimental results and methods that are widely published in journals and relatively easily found through publication databases such as PubMed. For the general scientist as well as bioinformaticists, these tools can deliver great value to the design and analysis of biological and medical experiments, but there is no inventory presenting an up-to-date and easily searchable index of all these resources. To remedy this, the BioWareDB search engine has been created. BioWareDB is an extensive and current catalog of software and databases of relevance to researchers in the fields of biology and medicine, and presently consists of 2800 validated entries. AVAILABILITY: BioWareDB is freely available over the Internet at http://www.biowaredb.org/

Abstracting and Indexing↗

map3C: a computational tool for processing multiomic single-cell Hi-C data.

SUMMARY: The emergence of multiomic single-cell Hi-C (scHi-C) methods, which simultaneously profile chromatin conformation and other modalities such as gene expression or DNA methylation, creates tremendous opportunities for studying the genome's structure-function relationships. Existing tools for processing multiomic scHi-C datasets lack certain key functions for downstream bioinformatics analysis. We present map3C, a software tool that incorporates additional key functions. Specifically, we demonstrate that map3C facilitates multiomic scHi-C processing, quality control, and identification of structural variant locations in the genome. AVAILABILITY AND IMPLEMENTATION: map3C is available at https://github.com/luogenomics/map3C and is archived at https://doi.org/10.5281/zenodo.20724719.

Software↗

Molecular biology and reproduction.

Modern molecular biology has provided unique insights into the fundamental understanding of reproductive disorders and the detection of microorganisms. The remarkable advances in DNA diagnostics have been expedited by the development of polymerase chain reaction (PCR) and the ability to isolate DNA and RNA from many different sources such as blood, saliva, hair roots, microscopic slides, paraffin-embedded tissue sections, clinical swabs, and even cancellous bone. These technical advances have been bolstered by the development of an increasing number of effective screening techniques to scan genomic DNA for unknown point mutations. The continued development of technology will ultimately result in automated DNA (desoxyribonucleic acid) diagnosis for the practicing clinician. The continuing expansion of information concerning the human genome will place an increasing emphasis on bioinformatics and the use of computer software for analyzing DNA sequences. With the automation of DNA diagnosis and the use of small samples (500 nanograms), the direct examination of the DNA of a patient, fetus, or microorganism will emerge as a definitive means of establishing the presence of the specific genetic change that causes disease. A knowledge of the precise pathology at the molecular level has and will provide important insights into the biochemical basis for many human diseases. A firm knowledge of the DNA alterations in disease and expression patterns of specific genes will provide for more directed therapeutic strategies. The refinement of vector technology and nuclear transplantion techniques will provide the opportunity for directed gene therapy to the early human embryo. This presentation is designed to acquaint the reader with current techniques of testing at the DNA level, prototype mutations in the reproductive sciences, new concepts in the molecular mechanisms of disease that affect reproduction, and therapeutic opportunities for the future. It is hoped that future refinement of these techniques combined with the ability to maintain genetic modification of these cells with recombinant vector technology will provide a definitive therapy for many single gene disorders, such as sickle cell anemia and thalassaemia. It is truly the challenge of the next century to decipher how these legions of newly discovered genes work, and to create a molecular language that can extend across all organisms.

DNA↗

Current Comparative Table (CCT) automates customized searches of dynamic biological databases.

The Current Comparative Table (CCT) software program enables working biologists to automate customized bioinformatics searches, typically of remote sequence or HMM (hidden Markov model) databases. CCT currently supports BLAST, hmmpfam and other programs useful for gene and ortholog identification. The software is web based, has a BioPerl core and can be used remotely via a browser or locally on Mac OS X or Linux machines. CCT is particularly useful to scientists who study large sets of molecules in today's evolving information landscape because it color-codes all result files by age and highlights even tiny changes in sequence or annotation. By empowering non-bioinformaticians to automate custom searches and examine current results in context at a glance, CCT allows a remote database submission in the evening to influence the next morning's bench experiment. A demonstration of CCT is available at http://orb.public.stolaf.edu/CCTdemo and the open source software is freely available from http://sourceforge.net/projects/orb-cct.

Computational Biology↗

Chemistry in bioinformatics.

Chemical information is now seen as critical for most areas of life sciences. But unlike Bioinformatics, where data is openly available and freely re-usable, most chemical information is closed and cannot be re-distributed without permission. This has led to a failure to adopt modern informatics and software techniques and therefore paucity of chemistry in bioinformatics. New technology, however, offers the hope of making chemical data (compounds and properties) free during the authoring process. We argue that the technology is already available; we require a collective agreement to enhance publication protocols.

Access to Information↗

The BioTools Suite. A comprehensive suite of platform-independent bioinformatics tools.

The BioTools Suite is a set of three comprehensive, platform-independent software packages (PepTool, GeneTool, and ChromaTool) developed for sequence assembly and analysis. In addition to supporting a large number of standard bioinformatics functions, these programs also incorporate a number of useful innovations including uniform graphical-user interface (GUI) design, direct internet connectivity, a novel approach to feature annotation, and a variety of enhanced algorithms for large scale proteome and genome analysis. This article describes the key features, recent changes, and general operation of all three programs.

Computational Biology↗

WU-Blast2 server at the European Bioinformatics Institute.

Since 1995, the WU-BLAST programs (http://blast.wustl.edu) have provided a fast, flexible and reliable method for similarity searching of biological sequence databases. The software is in use at many locales and web sites. The European Bioinformatics Institute's WU-Blast2 (http://www.ebi.ac.uk/blast2/) server has been providing free access to these search services since 1997 and today supports many features that both enhance the usability and expand on the scope of the software.

Computational Biology↗

Evolution of matrix and bone gamma-carboxyglutamic acid proteins in vertebrates.

The evolution of calcified tissues is a defining feature in vertebrate evolution. Investigating the evolution of proteins involved in tissue calcification should help elucidate how calcified tissues have evolved. The purpose of this study was to collect and compare sequences of matrix and bone gamma-carboxyglutamic acid proteins (MGP and BGP, respectively) to identify common features and determine the evolutionary relationship between MGP and BGP. Thirteen cDNAs and genes were cloned using standard methods or reconstructed through the use of comparative genomics and data mining. These sequences were compared with available annotated sequences (a total of 48 complete or nearly complete sequences, 28 BGPs and 20 MGPs) have been identified across 32 different species (representing most classes of vertebrates), and evolutionarily conserved features in both MGP and BGP were analyzed using bioinformatic tools and the Tree-Puzzle software. We propose that: 1) MGP and BGP genes originated from two genome duplications that occurred around 500 and 400 million years ago before jawless and jawed fish evolved, respectively; 2) MGP appeared first concomitantly with the emergence of cartilaginous structures, and BGP appeared thereafter along with bony structures; and 3) BGP derives from MGP. We also propose a highly specific pattern definition for the Gla domain of BGP and MGP.

Amino Acid Sequence↗

Identification of NLRP3 and TIPE2 as asthma biomarkers via integrative bioinformatics and Mendelian randomization.

Asthma is a chronic inflammatory airway disease imposing a substantial global health burden. NLRP3 is an immune sensor involved in infection and cellular stress responses. Recent studies suggest that NLRP3 may be involved in the pathogenesis of asthma. We hypothesized that genetic variation in NLRP3 may contribute to asthma susceptibility. However, the causal relationship between NLRP3 and asthma still remains unclear. In this study, bioinformatics analysis using asthma data and R software was performed to identify NLRP3-related genes. We performed weighted gene co-expression network analysis to identify co-expressed genes, resulting in 12 candidate genes. Kyoto Encyclopedia of Genes and Genomes and Gene Ontology enrichment analyses were used to identify the functions of these candidate genes, revealing their involvement in cellular metabolism. Mendelian randomization analysis of the 12 candidate genes identified 2 biomarkers: NLRP3 and TNFAIP8L2 (TIPE2). We validated their diagnostic value for asthma using the GSE182503 dataset, with area under the curve values of 0.83 and 0.66 for NLRP3 and TIPE2, respectively. This project discusses how NLRP3 promotes asthma pathogenesis, whereas TIPE2 may alleviate it, and explores the potential interplay between them. NLRP3 and TIPE2 may serve as diagnostic biomarkers for asthma: NLRP3 may promote, whereas TIPE2 may alleviate asthma development. Both genes represent potential diagnostic biomarkers and therapeutic targets that warrant further functional investigation.

Asthma↗

Proteomic fingerprints for potential application to early diagnosis of severe acute respiratory syndrome.

BACKGROUND: Definitive early-stage diagnosis of severe acute respiratory syndrome (SARS) is important despite the number of laboratory tests that have been developed to complement clinical features and epidemiologic data in case definition. Pathologic changes in response to viral infection might be reflected in proteomic patterns in sera of SARS patients. METHODS: We developed a mass spectrometric decision tree classification algorithm using surface-enhanced laser desorption/ionization time-of-flight mass spectrometry. Serum samples were grouped into acute SARS (n = 74; <7 days after onset of fever) and non-SARS [n = 1067; fever and influenza A (n = 203), pneumonia (n = 176); lung cancer (n = 29); and healthy controls (n = 659)] cohorts. Diluted samples were applied to WCX-2 ProteinChip arrays (Ciphergen), and the bound proteins were assessed on a ProteinChip Reader (Model PBS II). Bioinformatic calculations were performed with Biomarker Wizard software 3.1.1 (Ciphergen). RESULTS: The discriminatory classifier with a panel of four biomarkers determined in the training set could precisely detect 36 of 37 (sensitivity, 97.3%) acute SARS and 987 of 993 (specificity, 99.4%) non-SARS samples. More importantly, this classifier accurately distinguished acute SARS from fever and influenza with 100% specificity (187 of 187). CONCLUSIONS: This method is suitable for preliminary assessment of SARS and could potentially serve as a useful tool for early diagnosis.

Adolescent↗

Open Source software in medical informatics--why, how and what.

'Open Source' is a 20-40 year old approach to licensing and distributing software that has recently burst into public view. Against conventional wisdom this approach has been wildly successful in the general software market--probably because the openness lets programmers the world over obtain, critique, use, and build upon the source code without licensing fees. Linux, a UNIX-like operating system, is the best known success. But computer scientists at the University of California, Berkeley began the tradition of software sharing in the mid 1970s with BSD UNIX and distributed the major internet network protocols as source code without a fee. Medical informatics has its own history of Open Source distribution: Massachusetts General's COSTAR and the Veterans Administration's VISTA software have been distributed as source code at no cost for decades. Bioinformatics, our sister field, has embraced the Open Source movement and developed rich libraries of open-source software. Open Source has now gained a tiny foothold in health care (OSCAR GEHR, OpenEMed). Medical informatics researchers and funding agencies should support and nurture this movement. In a world where open-source modules were integrated into operational health care systems, informatics researchers would have real world niches into which they could engraft and test their software inventions. This could produce a burst of innovation that would help solve the many problems of the health care system. We at the Regenstrief Institute are doing our part by moving all of our development to the open-source model.

Database Management Systems↗

ESTWeb: bioinformatics services for EST sequencing projects.

ESTWeb is an internet based software package designed for uniform data processing and storage for large-scale EST sequencing projects. The package provides for: (a) reception of sequencing chromatograms; (b) sequence processing such as base-calling, vector screening, comparison with public databases; (c) storage of data and analysis in a relational database, (d) generation of a graphical report of individual sequence quality; and (e) issuing of reports with statistics of productivity and redundancy. The software facilitates real-time monitoring and evaluation of EST sequence acquisition progress along an EST sequencing project.

Computational Biology↗

GLYCOSCIENCES.de: an Internet portal to support glycomics and glycobiology research.

The development of glycan-related databases and bioinformatics applications is considerably lagging behind compared with the wealth of available data and software tools in genomics and proteomics. Because the encoding of glycan structures is more complex, most of the bioinformatics approaches cannot be applied to glycan structures. No standard procedures exist where glycan structures found in various species, organs, tissues or cells can be routinely deposited. In this article the concepts of the GLYCOSCIENCES.de portal are described. It is demonstrated how an efficient structure-based cross-linking of various glycan-related data originating from different resources can be accomplished using a single user interface. The structure oriented retrieval options-exact structure, substructure, motif, composition and sugar components-are discussed. The types of available data-references, composition, spatial structures, nuclear magnetic resonance (NMR) shifts (experimental and estimated), theoretically calculated fragments and Protein Database (PDB) entries-are exemplified for Man(3.) The free availability and unrestricted use of glycan-related data is an absolute prerequisite to efficiently share distributed resources. Additionally, there is an urgent need to agree to a generally accepted exchange format as well as to a common software interface. An open access repository for glyco-related experimental data will secure that the loss of primary data will be considerably reduced.

Computational Biology↗

Models@Home: distributed computing in bioinformatics using a screensaver based approach.

MOTIVATION: Due to the steadily growing computational demands in bioinformatics and related scientific disciplines, one is forced to make optimal use of the available resources. A straightforward solution is to build a network of idle computers and let each of them work on a small piece of a scientific challenge, as done by Seti@Home (http://setiathome.berkeley.edu), the world's largest distributed computing project. RESULTS: We developed a generally applicable distributed computing solution that uses a screensaver system similar to Seti@Home. The software exploits the coarse-grained nature of typical bioinformatics projects. Three major considerations for the design were: (1) often, many different programs are needed, while the time is lacking to parallelize them. Models@Home can run any program in parallel without modifications to the source code; (2) in contrast to the Seti project, bioinformatics applications are normally more sensitive to lost jobs. Models@Home therefore includes stringent control over job scheduling; (3) to allow use in heterogeneous environments, Linux and Windows based workstations can be combined with dedicated PCs to build a homogeneous cluster. We present three practical applications of Models@Home, running the modeling programs WHAT IF and YASARA on 30 PCs: force field parameterization, molecular dynamics docking, and database maintenance.

Computational Biology↗

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine↗

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics↗