PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

Gene expression profiles of breast cancer obtained from core cut biopsies before neoadjuvant docetaxel, adriamycin, and cyclophoshamide chemotherapy correlate with routine prognostic markers and could be used to identify predictive signatures.

BACKGROUND: Neoadjuvant administration of chemotherapy provides a unique opportunity to monitor response to treatment in breast cancer and assesses response exactly. Global gene expression profiling by microarrays has been used as a valuable tool for the identification of prognostic and predictive marker genes. Even though this technology is now wide spread and relatively standardized, there are only few data available which compare established parameters with expression values to determine reliability of this method. Therefore we analyzed gene expression data of pretreatment biopsies of breast cancer patients and compared them with the results of the immunohistochemical receptor expression for ER/ PR and Her-2, as well as FISH testing for HER-2 amplification. We analyzed the change of expression of these markers before and after neoadjuvant chemotherapy. Furthermore we evaluated the predictive significance of prognostic gene signatures as described by Sorlie, van't Veer and Ahr for response to neoadjuvant chemotherapy. METHODS: Pretherapeutic core biopsies were obtained from 70 patients undergoing neoadjuvant TAC chemotherapy within the GEPARTRIO-trial. Samples were characterized according to standard pathology including ER, PR and HER2 IHC and amount of cancer cells. Only biopsies with more than 80 % tumor cells were considered for further examination. RNA was isolated and expression profiling performed using Affymetrix Hg U133 Arrays (22 500 genes). GeneData's Expressionist software was used for bioinformatic analyses. RESULTS: More than two thirds of the biopsies yielded sufficient amounts (> 5 microg) of RNA for expression profiling and high quality data were obtained for 50 samples. Unsupervised clustering broadly revealed a correlation with hormone receptor status. When ER-alpha, PR and HER2 as analyzed by immunohistochemistry were compared to the corresponding mRNA data from gene chips more than 90 % concordance was observed. We could observe a switch of receptor expression for ER, PR or HER-2 from positive to negative and vice versa in 16/35 cases (45.7 %) and 5/22 cases (22.7 %) respectively. The prognostic marker sets of Sorlie, van't Veer and Ahr could not discriminate responders from non-responders in our patient group. CONCLUSIONS: Our results demonstrate that reliable expression profiles can be achieved by using limited amounts of tissue obtained during neoadjuvant chemotherapy. Microarray data capture conventional prognostic markers but might contain additional informative gene sets correlated with treatment outcome. Prognostic marker sets are not suitable to predict tumor response in the neoadjuvant setting, suggesting the necessity of class prediction methods to identify marker sets predictive for the type of therapy used.

Adult↗

Plant genome databases: from references to inference tools.

Plant genome databases play an important role in the archiving and dissemination of data arising from the international genome projects. Recent developments in bioinformatics, such as new software tools, programming languages and standards, have produced better access across the Internet to the data held within them. An increasing emphasis is placed on data analysis and indeed many resources now provide tools allied to the databases, to aid in the analysis and interpretation of the data. However, a considerable wealth of information lies untapped by considering the databases as single entities and will only be exploited by linking them with a wide range of data sources. Data from research programs such as comparative mapping and germplasm studies may be used as tools, to gain additional knowledge but without additional experimentation. To date, the current plant genome databases are not yet linked comprehensively with each other or with these additional resources, although they are clearly moving toward this. Here, the current wealth of public plant genome databases is reviewed, together with an overview of initiatives underway to bind them to form a single plant genome infrastructure.

Computational Biology↗

Base-By-Base: single nucleotide-level analysis of whole viral genome alignments.

BACKGROUND: With ever increasing numbers of closely related virus genomes being sequenced, it has become desirable to be able to compare two genomes at a level more detailed than gene content because two strains of an organism may share the same set of predicted genes but still differ in their pathogenicity profiles. For example, detailed comparison of multiple isolates of the smallpox virus genome (each approximately 200 kb, with 200 genes) is not feasible without new bioinformatics tools. RESULTS: A software package, Base-By-Base, has been developed that provides visualization tools to enable researchers to 1) rapidly identify and correct alignment errors in large, multiple genome alignments; and 2) generate tabular and graphical output of differences between the genomes at the nucleotide level. Base-By-Base uses detailed annotation information about the aligned genomes and can list each predicted gene with nucleotide differences, display whether variations occur within promoter regions or coding regions and whether these changes result in amino acid substitutions. Base-By-Base can connect to our mySQL database (Virus Orthologous Clusters; VOCs) to retrieve detailed annotation information about the aligned genomes or use information from text files. CONCLUSION: Base-By-Base enables users to quickly and easily compare large viral genomes; it highlights small differences that may be responsible for important phenotypic differences such as virulence. It is available via the Internet using Java Web Start and runs on Macintosh, PC and Linux operating systems with the Java 1.4 virtual machine.

Base Composition↗

An interactive visualization tool to explore the biophysical properties of amino acids and their contribution to substitution matrices.

BACKGROUND: Quantitative descriptions of amino acid similarity, expressed as probabilistic models of evolutionary interchangeability, are central to many mainstream bioinformatic procedures such as sequence alignment, homology searching, and protein structural prediction. Here we present a web-based, user-friendly analysis tool that allows any researcher to quickly and easily visualize relationships between these bioinformatic metrics and to explore their relationships to underlying indices of amino acid molecular descriptors. RESULTS: We demonstrate the three fundamental types of question that our software can address by taking as a specific example the connections between 49 measures of amino acid biophysical properties (e.g., size, charge and hydrophobicity), a generalized model of amino acid substitution (as represented by the PAM74-100 matrix), and the mutational distance that separates amino acids within the standard genetic code (i.e., the number of point mutations required for interconversion during protein evolution). We show that our software allows a user to recapture the insights from several key publications on these topics in just a few minutes. CONCLUSION: Our software facilitates rapid, interactive exploration of three interconnected topics: (i) the multidimensional molecular descriptors of the twenty proteinaceous amino acids, (ii) the correlation of these biophysical measurements with observed patterns of amino acid substitution, and (iii) the causal basis for differences between any two observed patterns of amino acid substitution. This software acts as an intuitive bioinformatic exploration tool that can guide more comprehensive statistical analyses relating to a diverse array of specific research questions.

Amino Acid Sequence↗

Multiple alignment of genomic sequences using CHAOS, DIALIGN and ABC.

Comparative analysis of genomic sequences is a powerful approach to discover functional sites in these sequences. Herein, we present a WWW-based software system for multiple alignment of genomic sequences. We use the local alignment tool CHAOS to rapidly identify chains of pairwise similarities. These similarities are used as anchor points to speed up the DIALIGN multiple-alignment program. Finally, the visualization tool ABC is used for interactive graphical representation of the resulting multiple alignments. Our software is available at Göttingen Bioinformatics Compute Server (GOBICS) at http://dialign.gobics.de/chaos-dialign-submission.

Computer Graphics↗

Poxvirus bioinformatics.

Biochemical and functional analysis of poxvirus genomes, genes, and proteins has entered a new era with the recent sequencing of more than 30 poxvirus genomes. The management and analysis of this volume of sequence data in an efficient and effective manner requires specialized computer software. This chapter describes a number of bioinformatics techniques useful for analyzing poxvirus genomes. Some of the software discussed here have been developed by members of the Poxvirus Bioinformatics Resource Center (PBRC; funded by National Institutes of Health [NIH]) specifically for use with poxvirus genomes. These programs or, more accurately, suites of programs have many functions dedicated to poxvirus genome characterization. Significantly, this software has been designed with ease of use at a single location as the major goal.

Computational Biology↗

G-language Genome Analysis Environment: a workbench for nucleotide sequence data mining.

SUMMARY: G-language Genome Analysis Environment (G-language GAE) is an open source generic software package aimed for higher efficiency in bioinformatics analysis. G-language GAE has an interface as a set of Perl libraries for software development, and a graphical user interface for easy manipulation. Both Windows and Linux versions are available. AVAILABILITY: From http://www.g-language.org/ under GNU General Public License. CD-ROMs are distributed freely in major conferences.

Database Management Systems↗

ISYS: a decentralized, component-based approach to the integration of heterogeneous bioinformatics resources.

MOTIVATION: Heterogeneity of databases and software resources continues to hamper the integration of biological information. Top-down solutions are not feasible for the full-scale problem of integration across biological species and data types. Bottom-up solutions so far have not integrated, in a maximally flexible way, dynamic and interactive graphical-user-interface components with data repositories and analysis tools. RESULTS: We present a component-based approach that relies on a generalized platform for component integration. The platform enables independently-developed components to synchronize their behavior and exchange services, without direct knowledge of one another. An interface-based data model allows the exchange of information with minimal component interdependency. From these interactions an integrated system results, which we call ISYSf1.gif" BORDER="0">. By allowing services to be discovered dynamically based on selected objects, ISYS encourages a kind of exploratory navigation that we believe to be well-suited for applications in genomic research.

Arabidopsis↗

Bioinformatics and its applications in plant biology.

Bioinformatics plays an essential role in today's plant science. As the amount of data grows exponentially, there is a parallel growth in the demand for tools and methods in data management, visualization, integration, analysis, modeling, and prediction. At the same time, many researchers in biology are unfamiliar with available bioinformatics methods, tools, and databases, which could lead to missed opportunities or misinterpretation of the information. In this review, we describe some of the key concepts, methods, software packages, and databases used in bioinformatics, with an emphasis on those relevant to plant science. We also cover some fundamental issues related to biological sequence analyses, transcriptome analyses, computational proteomics, computational metabolomics, bio-ontologies, and biological databases. Finally, we explore a few emerging research topics in bioinformatics.

Computational Biology↗

Bioinformatics in drug development and assessment.

Bioinformatics is playing an increasingly important role in nearly all aspects of drug discovery, drug assessment, and drug development. This growing importance lies not only in the role that bioinformatics plays in handling large volumes of data, but also in the utility of bioinformatics tools to predict, analyze, or help interpret clinical and preclinical findings. This review focuses on describing and evaluating some of the newer or more important bioinformatics resources (i.e., databases and software) that are of growing importance to understanding or predicting drug metabolism, especially with respect to the absorption, distribution, metabolism, excretion, (ADME), and toxicity (T) of both existing drugs and potential drug leads. Detailed descriptions and critical assessments of a number of potentially useful bioinformatics/cheminformatics databases and predictive ADMET software tools are provided. Additionally, several pharmaceutically important applications of both the databases and software are highlighted. Given the rapid growth in this area and the rapid changes that are taking place, a special emphasis is placed on freely available or Web-accessible resources.

Animals↗

Methods for gene expression profiling in dermatology research using DermArray nylon filter DNA microarrays.

Here we present methods of gene expression profiling using nylon filter deoxyribonucleic acid (DNA) microarrays and radiolabeled and nonradiolabeled hybridization probes. DermArray(R) nylon filter DNA microarrays were designed specifically for use in dermatology research. A patent-pending method was used to select approx 4400 highly informative, sequence-verified human cDNA clones for this DNA micro array. Using DermArray(R) filters, biomarkers have been discovered for normal and pathologic cells from skin, and for responses to dermatologic drugs. As an example, gene expression profiling was performed with hydroquinone-treated SKMel-28 cells, a melanoma cell line. Also included are the methods for bioinformatic analysis using Pathwaystrade mark software.

Computational Biology↗

KDE Bioscience: platform for bioinformatics analysis workflows.

Bioinformatics is a dynamic research area in which a large number of algorithms and programs have been developed rapidly and independently without much consideration so far of the need for standardization. The lack of such common standards combined with unfriendly interfaces make it difficult for biologists to learn how to use these tools and to translate the data formats from one to another. Consequently, the construction of an integrative bioinformatics platform to facilitate biologists' research is an urgent and challenging task. KDE Bioscience is a java-based software platform that collects a variety of bioinformatics tools and provides a workflow mechanism to integrate them. Nucleotide and protein sequences from local flat files, web sites, and relational databases can be entered, annotated, and aligned. Several home-made or 3rd-party viewers are built-in to provide visualization of annotations or alignments. KDE Bioscience can also be deployed in client-server mode where simultaneous execution of the same workflow is supported for multiple users. Moreover, workflows can be published as web pages that can be executed from a web browser. The power of KDE Bioscience comes from the integrated algorithms and data sources. With its generic workflow mechanism other novel calculations and simulations can be integrated to augment the current sequence analysis functions. Because of this flexible and extensible architecture, KDE Bioscience makes an ideal integrated informatics environment for future bioinformatics or systems biology research.

Biological Science Disciplines↗

[From bioinformatics to systems biology: account of the 12th international conference on intelligent systems in molecular biology].

The paper reviews the 12th International Conference on Intelligent Systems for Molecular Biology/Third European Conference on Computational Biology 2004 that was held in Glasgow, UK, during July 31-August 4. A number of talks, papers and software demos from the conference in bioinformatics, genomics, proteomics, transcriptomics and systems biology are described. Recent applications of liquid chromatography - tandem mass spectrometry, comparative genomics and DNA microarrays are given along with the discussion of bioinformatics curricular in higher education.

Computational Biology↗

[Bioinformatic analysis of adenoma-normal mucosa SSH library of colon].

We established a colonic adenoma-normal mucosa suppressive subtraction hybridization (SSH) library in 1999. In this study, we wanted to explore the expression profile of all candidate genes in this library. We developed an EST pipeline which contained two in-house software packages, nucleic acid analytical software and GetUni. The nucleic acid analytical software, an integrator of the universal bioinformatics tools including phred, phd2fasta, cross_match, repeatmasker and blast2.0, can blast sequences of differential clones with the downloaded non-redundant nucleotide (NR) database. GetUni can cluster these NR sequences into Unigene via matching with the downloaded Homo Sapiens UniGene database. Sixty-two candidate genes in A-N library were obtained via the high throughput automatic gene expression bioinformatics pipeline. Gene Ontology online analysis revealed that ribosome genes and immunity-regulating genes were the two most common categories in the KEGG or Biocarta Pathway. We also detected the expression of 2 genes with highest hits, Reg4 and FAM46A, by semi-quantitative RT-PCR. Both genes were up-regulated in 10 or 9 out of 10 adenomas in comparison with the paired normal mucosa, respectively. The candidate genes in A-N library would be of great significance in disclosing the molecular mechanism underlying in colonic adenoma initiation and progression.

Adenoma↗

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry↗

DivergentSet, a tool for picking non-redundant sequences from large sequence collections.

DivergentSet addresses the important but so far neglected bioinformatics task of choosing a representative set of sequences from a larger collection. We found that using a phylogenetic tree to guide the construction of divergent sets of sequences can be up to 2 orders of magnitude faster than the naive method of using a full distance matrix. By providing a user-friendly interface (available online) that integrates the tasks of finding additional sequences, building and refining the divergent set, producing random divergent sets from the same sequences, and exporting identifiers, this software facilitates a wide range of bioinformatics analyses including finding significant motifs and covariations. As an example application of DivergentSet, we demonstrate that the motifs identified by the motif-finding package MEME (Motif Elicitation by Maximum Entropy) are highly unstable with respect to the specific choice of sequences. This instability suggests that the types of sensitivity analysis enabled by DivergentSet may be widely useful for identifying the motifs of biological significance.

Amino Acid Sequence↗

Spatial Genomic Approaches to Investigate HOX Genes in Mouse Brain Tissues.

Spatial transcriptomic tools are an upcoming and powerful way to investigate targeted gene expression patterns within tissues. These tools offer the unique advantage of visualizing and understanding gene expression while preserving tissue integrity, thereby maintaining the spatial context of genes. Curio is a robust spatial transcriptomic tool that facilitates high throughput comprehensive spatial gene expression analysis across the entir e transcriptome with high efficiency. Here, we present a bioinformatics protocol for performing whole transcriptome gene expression analysis of mouse brain tissue using Curio. Specifically, we demonstrate using computational techniques to visualize expression patterns of various HOX genes in the mouse brain.

Animals↗