PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Earliest pages of bioinformatics.

This review is a brief outline of the chronology and essence of early events in bioinformatics, covering the period from 1869 (discovery of DNA by Miescher) to 1980-1981 (beginning of massive sequencing). For the purpose of this review, bioinformatics is understood as a chapter of molecular biology dealing with the amino acid and nucleotide sequences and with the information they carry.

Animals↗

XML, bioinformatics and data integration.

MOTIVATION: The eXtensible Markup Language (XML) is an emerging standard for structuring documents, notably for the World Wide Web. In this paper, the authors present XML and examine its use as a data language for bioinformatics. In particular, XML is compared to other languages, and some of the potential uses of XML in bioinformatics applications are presented. The authors propose to adopt XML for data interchange between databases and other sources of data. Finally the discussion is illustrated by a test case of a pedigree data model in XML. CONTACT: Emmanuel.Barillot@infobiogen.fr

Computational Biology↗

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics↗

Secure bioinformatics: privacy-preserving federated analytics using homomorphic encryption.

MOTIVATION: Large-scale bioinformatics analyses increasingly require collaboration across multiple cohorts and institutions, yet existing workflows often rely on data co-localization, which is slow, difficult to scale, and raises privacy concerns. We present a privacy-preserving federated analytics framework that enables secure statistical analysis across distributed datasets without transferring raw data, by performing all computations on encrypted data via cryptographic methods. RESULTS: We evaluate the framework by validating polygenic risk scores and conducting meta-analyses on two real-world cohorts. The proposed solution achieves over 99.9% accuracy relative to plaintext analyses, while maintaining scalable runtime performance with increasing data size and number of participating sites. These results demonstrate the feasibility of secure federated analytics for practical bioinformatics applications involving sensitive data.

Computational Biology↗

The discovery net system for high throughput bioinformatics.

MOTIVATION: Bioinformatics requires Grid technologies and protocols to build high performance applications without focusing on the low level detail of how the individual Grid components operate. RESULTS: The Discovery Net system is a middleware that allows service developers to integrate tools based on existing and emerging Grid standards such as web services. Once integrated, these tools can be used to compose reusable workflows using these services that can later be deployed as new services for others to use. Using the Discovery Net system and a range of different bioinformatics tools, we built a Grid based application for Genome Annotation. This includes workflows for automatic nucleotide annotation, annotation of predicted proteins and text analysis based on metabolic profiles and text analysis.

Algorithms↗

Early bioinformatics: the birth of a discipline--a personal view.

MOTIVATION: The field of bioinformatics has experienced an explosive growth in the last decade, yet this 'new' field has a long history. Some historical perspectives have been previously provided by the founders of this field. Here, we take the opportunity to review the early stages and follow developments of this discipline from a personal perspective. RESULTS: We review the early days of algorithmic questions and answers in biology, the theoretical foundations of bioinformatics, the development of algorithms and database resources and finally provide a realistic picture of what the field looked like from a resources and finally provide a realistic picture of what the field looked like from a practitioner's viewpoint 10 years ago, with a perspective for future developments.

Algorithms↗

Automatic discovery and classification of bioinformatics Web sources.

MOTIVATION: The World Wide Web provides an incredible resource to genomics researchers in the form of query access to distributed data sources--e.g. BLAST sequence homology search interfaces. The number of these autonomous sources and their rate of change outpaces the speed at which they can be manually classified, meaning that the available data is not being utilized to its full potential. Manually maintaining a wrapper library will not scale to accommodate the growth of genomics data sources on the Web, challenging us to produce an automated system that can find, classify and wrap new sources without constant human intervention. Previous research has not addressed the problem of automatically locating, classifying and integrating classes of bioinformatics data sources. RESULTS: This paper presents an overview of a system for finding classes of bioinformatics data sources and integrating them behind a unified interface. We describe our approach for automatic classification of new Web sources into relevance categories that eliminates the human effort required to maintain a current repository of sources. Our approach is based on a meta-data description of classes of interesting sources that describes the important features of an entire class of services without tying that description to any particular Web source. We examine the features of this format in the context of BLAST sources to show how it relates to Web sources that are being described. We then show how a description can be used to determine if an arbitrary Web source is an instance of the described service. To validate the effectiveness of this approach, we have constructed a prototype that correctly classifies approximately two-thirds of the BLAST sources we tested. We conclude with a discussion of these results, the factors that affect correct automatic classification and areas for future study.

Algorithms↗

A criticality-based framework for task composition in multi-agent bioinformatics integration systems.

MOTIVATION: During task composition, such as can be found in distributed query processing, workflow systems and AI planning, decisions have to be made by the system and possibly by users with respect to how a given problem should be solved. Although there is often more than one correct way of solving a given problem, these multiple solutions do not necessarily lead to the same result. Some researchers are addressing this problem by providing data provenance information. Others use expert advice encoded in a supporting knowledge-base. In this paper, we propose an approach that assesses the importance of such decisions with respect to the overall result. We present a way of measuring decision criticality and describe its potential use. RESULTS: A multi-agent bioinformatics integration system is used as the basis of a framework that facilitates such functionality. We propose an agent architecture, and a concrete bioinformatics example (prototype) is used to show how certain decisions may not be critical in the context of more complex tasks.

Algorithms↗

Capturing expert knowledge with argumentation: a case study in bioinformatics.

MOTIVATION: The output of a bioinformatic tool such as BLAST must usually be interpreted by an expert before reliable conclusions can be drawn. This may be based upon the expert's experience, additional data and statistical analysis. Often the process is laborious, goes unrecorded and may be biased. Argumentation is an established technique for reasoning about situations where absolute truth or precise probability is impossible to determine. RESULTS: We demonstrate the application of argumentation to 3D-PSSM, a protein structure prediction tool. The expert's interpretation of results is represented as an argumentation framework. Given a 3D-PSSM result, an automated procedure constructs arguments for and against the conclusion that the result is a good predictor of protein structure. In addition to capturing the unique expertise of the author of 3D-PSSM for distribution to users, an improvement in recall of 5-10 percentage points is achieved. This technique can be applied to a wide range of bioinformatic tools. AVAILABILITY: Example public server and benchmarking data are available at http://www.sbg.bio.ic.ac.uk/~brj03/argumentation/paper/. Source code available on request.

Amino Acid Sequence↗

springScape: visualisation of microarray and contextual bioinformatic data using spring embedding and an 'information landscape'.

The interpretation of microarray and other high-throughput data is highly dependent on the biological context of experiments. However, standard analysis packages are poor at simultaneously presenting both the array and related bioinformatic data. We have addressed this challenge by developing a system springScape based on 'spring embedding' and an 'information landscape' allowing several related data sources to be dynamically combined while highlighting one particular feature. Each data source is represented as a network of nodes connected by weighted edges. The networks are combined and embedded in the 2-D plane by spring embedding such that nodes with a high similarity are drawn close together. Complex relationships can be discovered by varying the weight of each data source and observing the dynamic response of the spring network. By modifying Procrustes analysis, we find that the visualizations have an acceptable degree of reproducibility. The 'information landscape' highlights one particular data source, displaying it as a smooth surface whose height is proportional to both the information being viewed and the density of nodes. The algorithm is demonstrated using several microarray data sets in combination with protein-protein interaction data and GO annotations. Among the features revealed are the spatio-temporal profile of gene expression and the identification of GO terms correlated with gene expression and protein interactions. The power of this combined display lies in its interactive feedback and exploitation of human visual pattern recognition. Overall, springScape shows promise as a tool for the interpretation of microarray data in the context of relevant bioinformatic information.

Algorithms↗

The bioinformatics resource for oral pathogens.

Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.

Bacteria↗

CryptoDB: a Cryptosporidium bioinformatics resource update.

The database, CryptoDB (http://CryptoDB.org), is a community bioinformatics resource for the AIDS-related apicomplexan-parasite, Cryptosporidium. CryptoDB integrates whole genome sequence and annotation with expressed sequence tag and genome survey sequence data and provides supplemental bioinformatics analyses and data-mining tools. A simple, yet comprehensive web interface is available for mining and visualizing the data. CryptoDB is allied with the databases PlasmoDB and ToxoDB via ApiDB, an NIH/NIAID-fundedBioinformatics Resource Center. Recent updates to CryptoDB include the deposition of annotated genome sequences for Cryptosporidium parvum and Cryptosporidium hominis, migration to a relational database (GUS), a new query and visualization interface and the introduction of Web services.

Animals↗

The MPI Bioinformatics Toolkit for protein sequence analysis.

The MPI Bioinformatics Toolkit is an interactive web service which offers access to a great variety of public and in-house bioinformatics tools. They are grouped into different sections that support sequence searches, multiple alignment, secondary and tertiary structure prediction and classification. Several public tools are offered in customized versions that extend their functionality. For example, PSI-BLAST can be run against regularly updated standard databases, customized user databases or selectable sets of genomes. Another tool, Quick2D, integrates the results of various secondary structure, transmembrane and disorder prediction programs into one view. The Toolkit provides a friendly and intuitive user interface with an online help facility. As a key feature, various tools are interconnected so that the results of one tool can be forwarded to other tools. One could run PSI-BLAST, parse out a multiple alignment of selected hits and send the results to a cluster analysis tool. The Toolkit framework and the tools developed in-house will be packaged and freely available under the GNU Lesser General Public Licence (LGPL). The Toolkit can be accessed at http://toolkit.tuebingen.mpg.de.

Computational Biology↗

The MIGenAS integrated bioinformatics toolkit for web-based sequence analysis.

We describe a versatile and extensible integrated bioinformatics toolkit for the analysis of biological sequences over the Internet. The web portal offers convenient interactive access to a growing pool of chainable bioinformatics software tools and databases that are centrally installed and maintained by the RZG. Currently, supported tasks comprise sequence similarity searches in public or user-supplied databases, computation and validation of multiple sequence alignments, phylogenetic analysis and protein-structure prediction. Individual tools can be seamlessly chained into pipelines allowing the user to conveniently process complex workflows without the necessity to take care of any format conversions or tedious parsing of intermediate results. The toolkit is part of the Max-Planck Integrated Gene Analysis System (MIGenAS) of the Max Planck Society available at www.migenas.org (click 'Start Toolkit').

Animals↗

A comprehensive expression analysis of the Arabidopsis proline-rich extensin-like receptor kinase gene family using bioinformatic and experimental approaches.

The Arabidopsis proline-rich extensin-like receptor kinase (PERK) family consists of 15 predicted receptor kinases. A comprehensive expression analysis was undertaken to identify overlapping and unique expression patterns within this family relative to their phylogeny. Three different approaches were used to study AtPERK gene family expression, and included analyses of the EST, MPSS and NASCArrays databases as well as experimental RNA blot analyses. Some of the AtPERK members were identified as tissue-specific genes while others were more broadly expressed. While in some cases there was a good association between these different expression patterns and the position of the AtPERK members in the kinase phylogeny, in other cases divergence of expression patterns was seen. The PERK expression data identified by the bioinformatics and experimental approaches were found generally to show similar trends and supported the use of data from large-scale expression studies for obtaining preliminary expression data. Thus, the bioinformatics survey for ESTs and microarrays is a powerful comprehensive approach for obtaining a genome-wide view of genes in a multigene family.

Arabidopsis↗

Bioinformatics-driven, rational engineering of protein thermostability.

A longstanding goal in protein engineering is to identify specific sequence changes that endow proteins with desired functional properties. As opposed to traditional rational and random protein engineering techniques, we have employed a bioinformatic approach to identify specific sequence changes that influence key functional properties of a protein within a defined superfamily. Specifically, we have used the Bayesian sequence-based algorithms PROBE and Classifier to identify a strand-turn-strand motif that contributes to thermophilicity among members of the serine protease subtilase superfamily. By replacing a 16 amino acid sequence in the mesophilic subtilisin E (from Bacillus subtilis) with a bioinformatics-generated thermophilic model sequence, the melting temperature of subtilisin E was increased by 13 degrees C. While wild-type subtilisin E was inactive at 90 degrees C, the mutant retained a substantial fraction of its function, with ca. one-third of the activity that it has at 45 degrees C.

Algorithms↗

Bioinformatics and genomic medicine.

Bioinformatics is a rapidly emerging field of biomedical research. A flood of large-scale genomic and postgenomic data means that many of the challenges in biomedical research are now challenges in computational science. Clinical informatics has long developed methodologies to improve biomedical research and clinical care by integrating experimental and clinical information systems. The informatics revolution in both bioinformatics and clinical informatics will eventually change the current practice of medicine, including diagnostics, therapeutics, and prognostics. Postgenome informatics, powered by high-throughput technologies and genomic-scale databases, is likely to transform our biomedical understanding forever, in much the same way that biochemistry did a generation ago. This paper describes how these technologies will impact biomedical research and clinical care, emphasizing recent advances in biochip-based functional genomics and proteomics. Basic data preprocessing with normalization and filtering, primary pattern analysis, and machine-learning algorithms are discussed. Use of integrative biochip informatics technologies, including multivariate data projection, gene-metabolic pathway mapping, automated biomolecular annotation, text mining of factual and literature databases, and the integrated management of biomolecular databases, are also discussed.

Computational Biology↗

Exploring the mechanism of Shengmai San in treating lung adenocarcinoma based on bioinformatics and molecular dynamics simulation.

To investigate the mechanism of Shengmai San (SMS) in the treatment of lung adenocarcinoma (LUAD) based on an integrated strategy combining "network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulation," aiming to provide a precise combination therapy strategy and identify potential bioactive compounds. Differentially expressed genes in LUAD were identified from the Gene Expression Omnibus database using R (originally developed at Bell Laboratories and currently managed by Lucent Technologies). SMS components (ginseng, Ophiopogon japonicus, and Schisandra chinensis) were retrieved from encyclopaedia of traditional Chinese medicine, with Lipinski-compliant compounds selected. Compound targets were predicted via SwissTargetPrediction and Similarity Ensemble Approach. Intersecting targets between differentially expressed genes and compound targets were identified for "herbs-compounds-targets-disease" network construction. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses were performed. Hub targets were identified by analyzing the protein-protein interaction network. High-prognostic relevance targets were screened from The Cancer Genome Atlas. Compounds targeting these were identified through the herbs-compounds-targets-disease network, and absorption, distribution, metabolism, excretion, and toxicity-compliant compounds were selected using SwissADME (a web-based tool provided by the Molecular Modeling Group of the Swiss Institute of Bioinformatics). Core regulatory targets were identified through molecular docking, with complex stability assessed by molecular dynamics simulations. The key bioactive compounds of SMS for treating LUAD were identified as 7-hydroxy-2,5-dimethyl-4H-1-benzopyran-4-one, N-trans-feruloyltyramine, paprazine, and (E)-N-[(2S)-2-hydroxy-2-(4-hydroxyphenyl)ethyl]-3-(4-hydroxyphenyl)prop-2-enamide. Hub targets included AURKA, CCNA2, CCNB1, CDK1, CHEK1, KIF11, NEK2, PLK1, TTK, and TYMS. Among these, CDK1, CHEK1, and PLK1 demonstrated both high-prognostic relevance and strong binding affinity with SMS, emerging as core regulatory targets for SMS in LUAD treatment. Mechanistically, SMS exerts its anticancer effects primarily by modulating the tumor necrosis factor, interleukin-17, cell cycle, and Lipid and atherosclerosis signaling pathways. The active components of SMS, such as paprazine, may exert antitumor effects partly through downregulating CDK1, CHEK1, and PLK1 expression. Although the present study did not examine drug-resistance models or combination regimens, our findings raise the possibility that, in patients with high expression of these genes, combining SMS with standard chemotherapy or targeted therapy could potentially enhance chemosensitivity and mitigate the development of resistance. This hypothesis, however, requires formal testing in appropriate preclinical models and functional validation studies.

Molecular Dynamics Simulation↗