PubMed HealthSearch

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines

A distributed environment for physical map construction.

MOTIVATION: With the main focus of the Human Genome Project shifting to sequencing, bioinformatics support for constructing large-scale genomic maps of other organisms is still required. We attempt to provide for this with our work, aimed at the delivery of robust and user-friendly contig-building software on the WWW. RESULTS: We present a prototype distributed analytical environment for molecular biologists working in the area of genomic mapping. It consists of the WWW server for constructing contigs from users' data with a hypertext output connected to Java-based map visualization software. AVAILABILITY: Freely available on http://www.mpimg-berlin-dahlem.mpg. de/ approximately andy/server/ CONTACT: andy@rag3.rz-berlin.mpg.de

Algorithms

Ubiquitous distributed objects with CORBA.

Database interoperation is becoming a bottleneck for the research community in biology. In this paper, we first discuss the question of interoperability and give a brief overview of CORBA. Then, an example is explained in some detail: a simple but realistic data bank of STSs is implemented. The Object Request Broker is the media for communication between an object server (the data bank) and a client (possibly a genome center). Since CORBA enables easy development of networked applications, we meant this paper to provide an incentive for the bioinformatics community to develop distributed objects.

Base Sequence

Expression and prognosis of CXCL13 in uterine corpus endometrial carcinoma based on bioinformatics analysis.

OBJECTIVE: The biological significance of the chemokine ligand C-X-C motif chemokine ligand 13 (CXCL13) may play a significant role in the pathogenesis of uterine corpus endometrial carcinoma (UCEC). This study aims to identify and verify CXCL13 with predictive value for prognosis in UCEC. METHODS: CXCL13 mRNA expression differences were analyzed using R software in three independent datasets: one each from The Cancer Genome Atlas (TCGA) and two from the Gene Expression Omnibus (GEO), namely GSE17025 and GSE106191. The correlation between CXCL13 expression and prognosis was evaluated by Kaplan-Meier analysis. Univariate and multivariate Cox analyses were utilized to construct a prognostic nomogram. Tumor Immune Estimation Resource (TIMER) and the Tumor and Immune System Interaction Database (TISIDB) were employed to assess the relationship between CXCL13 and tumor immune infiltration. Coexpressed genes with CXCL13 were identified by the Spearman correlation analysis. A CXCL13 protein-protein interaction (PPI) network was constructed with the STRING website tool and hub genes were screened out. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genome (KEGG) analyses were performed with the "clusterProfiler" R package. Gene set enrichment analysis (GSEA) was used to identify underlying biological mechanisms. A drug-gene interaction network was constructed in the Comparative Toxicogenomics Database (CTD). RESULTS: High CXCL13 mRNA expression were validated in UCEC in the above three independent datasets. High CXCL13 expression was associated with favorable prognosis in UCEC. A nomogram for predicting the 1-, 3-, and 5-year survival probability in UCEC was construct based on CXCL13 expression and other clinical parameters. The use of Spearman correlation indicated certain correlation between CXCL13 and immune cells and immune checkpoint (ICP) genes. Seven hub genes were upregulated in UCEC, namely CXCL9, IFNG, CXCL10, CXCL11, GBP5, CCL18, and GZMB. The expression and prognostic relevance of CXCL9, IFNG, GBP5, and GZMB were in accordance with CXCL13. The main biological processes enriched were cytokine-cytokine receptor interaction and chemokine signaling pathway. CONCLUSIONS: The above comprehensive analyses suggest that CXCL13 may serve as a potential prognostic biomarker for UCEC, specifically for early-stage UCEC.

CXCL13

Bayesian inference on biopolymer models.

MOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/

Bayes Theorem

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome

AutoPVPrimer: A comprehensive AI-Enhanced pipeline for efficient plant virus primer design and assessment.

Plant viruses pose a significant threat to global agriculture and require efficient tools for their timely detection. We present AutoPVPrimer, an innovative pipeline that integrates artificial intelligence (AI) and machine learning to accelerate the development of plant virus primers. The pipeline uses Biopython to automatically retrieve different genomic sequences from the NCBI database to increase the robustness of the subsequent primer design. The design_primers_with_tuning module uses a random forest classifier that optimizes parameters and provides flexibility for different experimental conditions. Quality control measures, including the evaluation of poly-X content and melting temperature, increase primer reliability. Unique to AutoPVPrimer is the visualize_primer_dimer module, which supports the visual evaluation of primer dimers-a feature missing in other tools. Primer specificity is validated via primer BLAST, which contributes to the overall efficiency of the pipeline. AutoPVPrimer has been successfully applied to the tomato mosaic virus, proving its adaptability and efficiency. The modular design allows customization by the user and extends the applicability to different plant viruses and experimental scenarios. The pipeline represents a significant advance in primer design and provides researchers with an effective tool to accelerate molecular biology experiments. Future developments aim to extend compatibility and incorporate user feedback to consolidate AutoPVPrimer as an innovative contribution to the bioinformatics toolbox and a promising resource for the advancement of plant virology research.

DNA Primers

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence

Whole-genome automated assembly pipeline for Chlamydia trachomatis strains from reference, in vitro and clinical samples using the integrated CtGAP pipeline.

Whole genome sequencing (WGS) is pivotal for the molecular characterization of Chlamydia trachomatis (Ct)-the leading bacterial cause of sexually transmitted infections and infectious blindness worldwide. Ct WGS can inform epidemiologic, public health and outbreak investigations of these human-restricted pathogens. However, challenges persist in generating high-quality genomes for downstream analyses given its obligate intracellular nature and difficulty with in vitro propagation. No single tool exists for the entirety of Ct genome assembly, necessitating the adaptation of multiple programs with varying success. Compounding this issue is the absence of reliable Ct reference strain genomes. We, therefore, developed CtGAP-Chlamydia trachomatisGenome Assembly Pipeline-as an integrated 'one-stop-shop' pipeline for assembly and characterization of Ct genome sequencing data from various sources including isolates, in vitro samples, clinical swabs and urine. CtGAP, written in Snakemake, enables read quality statistics output, adapter and quality trimming, host read removal, de novo and reference-guided assembly, contig scaffolding, selective ompA, multi-locus-sequence and plasmid typing, phylogenetic tree construction, and recombinant genome identification. Twenty Ct reference genomes were also generated. Successfully validated on a diverse collection of 363 samples containing Ct, CtGAP represents a novel pipeline requiring minimal bioinformatics expertise with easy adaptation for use with other bacterial species.

Chlamydia trachomatis

TFinder: A Python Web Tool for Predicting Transcription Factor Binding Sites.

Transcription is a key cell process that consists of synthesizing several copies of RNA from a gene DNA sequence. This process is highly regulated and closely linked to the ability of transcription factors to bind specifically to DNA. TFinder is an easy-to-use Python web portal allowing the identification of Individual Motifs (IM) such as Transcription Factor Binding Sites (TFBS). Using the NCBI API, TFinder extracts either promoter or gene terminal regulatory regions, through a simple query of NCBI gene name or ID. It enables simultaneous analysis across five different species for an unlimited number of genes. TFinder searches for Individual Motifs in different formats, including IUPAC codes and JASPAR entries. Moreover, TFinder also allows de novo generations of a Position Weight Matrix (PWM) and the use of already established PWM. Finally, the data are provided in a tabular and a graph format showing the relevance and the P-value of the Individual Motifs found as well as their location relative to the Transcription Start Site (TSS) or the terminal region of the gene. The results are then sent by email to users facilitating the subsequent data analysis and sharing. TFinder is written in Python and freely available on GitHub under the MIT license: https://github.com/Jumitti/TFinder. It can be accessed as a web application implemented in Streamlit at https://tfinder-ipmc.streamlit.app. Resources are available on Streamlit "Resources" tab. TFINDER strength is that it relies on an all-in-one intuitive tool allowing users inexperienced with bioinformatics tools to retrieve gene regulatory regions sequences in multiple species and to search for individual motifs in a huge number of genes.

Transcription Factors

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny

Comprehensive circRNA expression profile and hub genes screening during human liver development.

BACKGROUND: Understanding the expression of non-coding RNA in the liver during embryonic development provides important insights into liver diseases. Therefore, we investigated circular RNA (circRNA) roles in human liver development, an unexplored research domain. METHODS: Using high-throughput sequencing and bioinformatics, we analysed foetal liver samples across developmental stages (7-20 weeks post-conception). Differentially expressed (DE) genes were identified and subjected to enrichment analysis using Gene Ontology (GO), Kyoto Encyclopaedia of Genes and Genomes (KEGG), and Disease Ontology (DO). Modular analysis was performed using the Search Tool for Retrieval of Interacting Genes (STRING), followed by construction of a protein-protein interaction (PPI) network using Cytoscape software. The key genes were screened using Molecular Complex Detection (MCODE). The mRNA levels of hub genes were validated using quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: There were 645 DE circRNAs and 5,145 DE mRNAs between human livers at the three growth stages (HB, EH, and LH). It was found that the activity of circRNAs was boosted remarkably in the hepatoblastic stage. Enrichment analysis found they mainly involved in nervous system regulation of liver function, embryonic organ development and digestive system development. In addition, DE circRNAs were primarily involved in the PI3K-AKT, MAPK and calcium pathways, potentially contributing to adult liver diseases. Notably, only hsa_circ_001471 and novel_circ_017382 were simultaneously identified at all stages and were persistently downregulated. A co-expression regulatory network involving these circRNAs was established. Three hub genes (LGR5, FOXL1 and RSPO3) were identified from the PPI network of 167 genes and may play key roles in human liver development. The RT-qPCR validation results were in agreement with the sequencing data. CONCLUSIONS: Our findings provide the first insights into the roles and regulatory networks of circRNAs in human liver development, laying the groundwork for further investigations of molecular and signalling networks.

Humans

Information services of the European Bioinformatics Institute.

The scope of the EBI is focused on providing better services to the scientific community. Technological advancements in the hardware area provide EBI with means of producing data much faster than before, and with greater accuracy since there is now a better technical ability to produce more exhaustive searches through larger indices. Hand in hand with the technological developments, research and development work is continuing on better indexing systems and more efficient ways of establishing and maintaining the future databases. The existing links of communication between EBI and the user community are exploited to study the needs of the scientific community, to provide better services, and to enhance the quality of databases by interpreting user feedback and updates. A very important goal is to enhance the awareness of the scientific (and, maybe even more, the nonscientific) public of the importance of the modern field of bioinformatics and to introduce special meetings and courses, in which more specific subjects will be studied in depth. Another aspect of this goal is to help in constructing special bioinformatics programs in university faculties. In such programs, in contrast to the existing layout, students will pursue studies in a combined environment that provides basic training in biology and in computation. Currently, one of the main problems in the field is that scientists are either biologists, who are self-educated in the field of computers and programming, or computer scientists without sufficient knowledge of biology. It is hoped that a combined program will provide a high level of education in both fields of interest at the appropriate ratios. Building an efficient and friendly interface between the EBI and the user community is the basis for any future development. This aim is achieved by using the most modern server systems while continuously researching newer and better systems and interfaces. This task can never be complete without involvement of the user community by providing feedback to any of EBI's services. A better bioinformatics community is a necessity for any future development of the biological research aiming at a better society.

Amino Acid Sequence

Identification of autophagy-related genes as potential biomarkers correlated with immune infiltration in bipolar disorder: a bioinformatics analysis.

BACKGROUND: Bipolar disorder (BPD) is a kind of manic and depressive phase alternate episodes of serious mental illness, and it is correlated with well-documented cortical brain abnormalities. Emerging evidence supports that autophagy dysfunction in neuronal system contributes to pathophysiological changes in neurological disease. However, the role of autophagy in bipolar disorder has rarely been elucidated. This study aimed to identify the autophagy-related gene as a potential biomarker Correlated to immune infiltration in BPD. METHODS: The microarray dataset GSE23848 and autophagy-related genes (ARGs) were downloaded. Differentially expressed genes (DEGs) between normal and BPD samples were screened using the R software. Machine learning algorithms were performed to screen the significant candidate biomarker from autophagy-related differentially expressed genes (ARDEGs). The correlation between the screened ARDEGs and infiltrating immune cells was explored through correlation analysis. RESULTS: In this study, the autophagy pathway was abundantly enriched and activated in BPD, as indicated by Pathway enrichment analysis. We identified 16 ARDEGs in BPD compared to the normal group. A signature of 4 ARDEGs (ERN1, ATG3, CTSB, and EIF2AK3) was screened. ROC analysis showed that the above genes have good diagnostic performance. In addition, immune correlation analysis considered that the above four genes significantly correlated with immune cells in BPD. CONCLUSIONS: Autophagy - immune cell axis mediates pathophysiological changes in BPD. Four important ARDEGs are prospective to be potential biomarkers associated with immune infiltration in BPD and helpful for the prediction or diagnosis of BPD.

Bipolar Disorder

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation