PubMed HealthSearch

SEARCH · PubMed Health

Results for “Web server”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins.

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

Archaea

CancerOmicsStudio (CoS): a web server for integrative and interpretable analysis of multi-omics cancer data.

MOTIVATION: Large-scale omics resources, including The Cancer Genome Atlas, Genomics of Drug Sensitivity in Cancer, and the Cancer Dependency Map, have become essential for cancer research. However, these datasets are distributed across different platforms, formats and analysis frameworks, which limits their practical use by researchers without extensive computational expertise. RESULTS: We developed CancerOmicsStudio (CoS), a web server for integrative and interpretable analysis of multi-omics cancer data across 33 cancer types. CoS provides five major modules: CosAI, Traditional Analysis, Drug Sensitivity, CRISPR Dependency and Single-Cell Tumor Microenvironment. The Traditional Analysis module supports expression comparison, diagnostic evaluation, survival analysis, enrichment analysis and gene correlation. The Drug Sensitivity and CRISPR Dependency modules enable systematic evaluation of gene-drug response associations and gene essentiality in cancer cell lines. The Single-Cell Tumor Microenvironment module supports tumor microenvironment analysis at single-cell resolution. In total, approximately 1.23 million results have been precomputed to enable rapid retrieval. CosAI further allows users to submit natural-language queries and obtain results through a Real-time Analysis as Retrieval framework, with responses summarized by a lightweight language model. AVAILABILITY AND IMPLEMENTATION: CancerOmicsStudio is freely available at Zenodo (doi: 10.5281/zenodo.18744990) and https://cos.wanglab.bio.

Humans

Effectidor II: a pan-genomic AI-based algorithm for the prediction of type III secretion system effectors.

MOTIVATION: Type III secretion systems are used by many Gram-negative bacteria to inject type 3 effectors (T3Es) directly into eukaryotic cells, promoting disease or provoking immune response. Because of these opposing evolutionary forces, T3E repertoires often vary within taxonomic groups. Identifying the full effector gene repertoire in genomes of related individuals is crucial for determining core and specialized effectors, understanding the disease dynamics, and developing appropriate management strategies against pathogens. It can also help uncover novel T3Es that have recently emerged in a population. Our previously published Effectidor web server successfully addressed the challenge of identifying T3Es in a single bacterial genome. Here, we enriched the web server with various novel capabilities, including the identification of T3Es from multiple genome sequences simultaneously. RESULTS: We present Effectidor II, a web server that relies on machine learning to predict T3E-encoding genes within bacterial pan-genomes. We demonstrate the benefit of learning based on features extracted from the entire sequences comprising the pan-genome and report a novel T3E discovered by it in Xanthomonas euroxanthea. AVAILABILITY AND IMPLEMENTATION: Effectidor II is available at: https://effectidor.tau.ac.il and the source code is available at: https://github.com/naamawagner/Effectidor. A stand-alone version of Effectidor II is available at: https://github.com/naamawagner/Effectidor/tree/StandAlone. The source code for the standalone version and the data used in this work are also provided in https://doi.org/10.5281/zenodo.15081636.

Type III Secretion Systems

GRNContext: an interactive web platform for contextualized gene regulatory networks visualization across human cancers.

SUMMARY: While current Gene Regulatory Network (GRN) databases provide comprehensive reference maps of potential interactions between transcription factors and target genes, they do not specify which regulatory interactions are active within specific biological contexts. This limitation is particularly critical in cancer, where transcriptional programs are inherently tissue-specific. To address this gap, we developed GRNContext, an interactive web platform designed for the visualization, exploration, and comparative analysis of gene regulatory networks contextualized across 33 cancer types from The Cancer Genome Atlas (TCGA). Our approach uses the TFLink human reference GRN as a starting point and integrates TCGA transcriptomic profiles to infer cancer-specific regulatory activity. Regulatory relevance was assessed using complementary machine learning and statistical methods, which were unified into a consensus score to prioritize and filter the most relevant candidate regulators for each target gene. By providing both curated context-specific GRNs and a user-friendly platform, GRNContext constitutes a comprehensive and accessible resource that supports mechanistic investigations, hypothesis generation, and translational research focused on transcriptional regulation in cancer. AVAILABILITY AND IMPLEMENTATION: GRNContext is supported by all major browsers and freely available on the web at https://apps.cienciavida.org/grncontext. It is implemented as a client-server web application featuring a FastAPI backend and a React frontend utilizing Cytoscape.js for interactive network visualization, all containerized via Docker for cross-platform compatibility.

Humans

Functional Prediction of Epitranscriptome.

N6-methyladenosine (m6A) is one of the most prevalent and well-studied RNA modifications, playing a pivotal role in many biological processes. With the recent advances in high-throughput sequencing technologies, tens of thousands of m6A sites have been reported. However, not all m6A sites are important or functionally significant, highlighting the need to distinguish biologically relevant m6As from non-functional or technically artefactual ones. Here, we describe ConsRM, which is a web-based resource that was designed to evaluate the importance of m6As from an evolutionary perspective. It introduced a novel scoring framework for quantifying the conservation degree of m6As in humans. Its web interface includes a database of 177998 distinct human m6A sites along with their calculated conservation score, and allows users to analyze their own data via the web server. ConsRM is freely accessible at: http://180.208.58.19/conservation/browser.html .

Humans

One thousand SARS-CoV-2 antibody structures reveal convergent binding and near-universal immune escape.

Understanding antibody recognition and adaptation to viral evolution is central to vaccine and therapeutic development. Over 1,100 SARS-CoV-2 antibody structures have been resolved, marking the largest structural biology effort for a single pathogen. We present a comprehensive analysis of this landmark dataset to investigate the principles of antibody recognition and immune escape. Human immunoglobulins and camelid single-chain antibodies dominate, collectively mapping 99% of the receptor-binding domain. Despite remarkable sequence and conformational diversity, antibodies exhibit convergence in their paratope structures, revealing evolutionary constraints in epitope selection. Analyses reveal near-universal immune escape of antibodies, including all clinical monoclonals, by advanced variants such as KP3.1.1. On average, over one-third of antibody epitope residues are mutated. These findings support pervasive immune escape, underscoring the need to effectively leverage multi-epitope-targeting strategies to achieve durable immunity. To support community accessibility, we developed an interactive web server for visualization and analysis of antibody-antigen complexes and mutational data.

SARS-CoV-2

DeepMASS v.2: An enhanced deep learning platform for large-scale discovery and structural annotation of unknown plant metabolites.

Determining the structures of unknown metabolites remains a fundamental bottleneck in plant metabolomics, as the vast chemical diversity of plant secondary metabolites far exceeds the coverage of existing spectral libraries. Here, we present DeepMASS v.2, a substantially enhanced platform for annotating unknown metabolites from liquid chromatography-tandem mass spectrometry data, designed to address this challenge at scale. DeepMASS v.2 leverages a semantic spectral representation model trained on millions of spectra from GNPS, NIST, and in-house resources. By integrating Spec2Vec-based embeddings with HNSW (hierarchical navigable small world) graph retrieval and a unified chemical space defined by molecular fingerprints, DeepMASS v.2 identifies structurally related neighbors of unknown spectra and ranks candidate structures according to their proximity to the predicted structural neighborhoods within chemical space. Benchmarking against Critical Assessment of Small Molecule Identification datasets and a curated natural product collection demonstrated that DeepMASS v.2 outperforms state-of-the-art in silico annotation tools, including SIRIUS, CFM-ID, MetFrag, and MS-Finder. Importantly, DeepMASS v.2 maintains strong performance for metabolites absent from spectral libraries, highlighting its capacity to annotate genuinely unknown compounds. Application of DeepMASS v.2 to large-scale plant metabolomics datasets demonstrated its ability to expand accessible metabolome coverage. Implemented as an intuitive web platform, DeepMASS v.2 provides the community with a scalable, interpretable, and high-throughput solution for structural annotation, enabling more comprehensive characterization of plant chemical diversity and accelerating natural product discovery in molecular plant science. The DeepMASS v.2 web server is publicly available at http://deepmass.cn.

Metabolomics

Towards efficient perturbation for the noncoding genome.

Deciphering the functionality of the noncoding genome, which includes important cis-regulatory elements (CREs) and transcribed noncoding RNA genes, remains technically challenging. Here, using massively parallel genetic screening, we systematically benchmark the performance of five representative loss-of-function perturbation tools, including single-guide RNA (gRNA) mediated SpCas9 cleavage or CRISPR interference, and paired gRNA (pgRNA) involved dual-SpCas9, Big Papi (paired SpCas9 and SaCas9) or dual-enAsCas12a fragment deletion methods, in decoding the roles of the noncoding genome. For targeting CREs such as enhancers, dual-SpCas9 outperforms other methods with superior efficiency in destroying functional genomic regions. For perturbing noncoding RNA genes, in addition to dual-SpCas9, other RNA-targeting methods such as RNA interference are recommended to discriminate transcript-dependent or -independent roles. A deep learning model, DeepDC, with an associated web server, is built to facilitate optimal dual-SpCas9 pgRNA design for efficiently deleting a genomic fragment. Together, our work provides practical guidance on selecting appropriate loss-of-function tools to resolve the functional complexity of the noncoding genome.

CRISPR-Cas Systems

Missense variants pathogenicity annotation from homologous proteins.

MOTIVATION: High-throughput DNA sequencing has revealed millions of single nucleotide variants (SNVs) in the human genome, with a small fraction linked to disease. The effect of missense variants, which alter the protein sequence, is particularly challenging to interpret due to the scarcity of clinical annotations and experimental information. While using conservation and structural information, current prediction tools still struggle to predict variant pathogenicity. In this study, we explored the pathogenicity of homologous missense variants-variants in equivalent positions across homologous proteins-focusing on proteins involved in autosomal dominant diseases. RESULTS: Our analysis of 2976 pathogenic and 17 555 non-pathogenic homologous variants demonstrated that pathogenicity can be extrapolated with 95% accuracy within a family, or up to 98% for closer homologs. Remarkably, the evaluation of 27 commonly used mutation predictor methods revealed that they were not fully capturing this biological feature. To facilitate the exploration of homologous variants, we created HomolVar, a web server that computationally predicts the pathogenesis of missense variants using annotations from homologous variants, freely available at https://rarevariants.org/HomolVar. Overall, these findings and the accompanying tool offer a robust method for predicting the pathogenicity of unannotated variants, enhancing genotype-phenotype correlations, and contributing to diagnosing rare genetic disorders. AVAILABILITY AND IMPLEMENTATION: HomolVar is freely available at https://rarevariants.org/HomolVar.

Mutation, Missense

CCRR: a user-friendly platform for analyzing complex chromosomal rearrangements in tumors.

SUMMARY: Complex chromosomal rearrangements in tumors involve intricate genomic alterations that significantly affect gene function and contribute to cancer development. Identifying these events is crucial for cancer research but is often challenging due to the complexity and limitations of existing tools. We developed the Complex Chromosomal Rearrangements Resolver (CCRR), a comprehensive, reproducible, and user-friendly platform for analyzing complex rearrangements in tumors. CCRR integrates multiple SV and CNV detection tools within a Docker container environment, simplifying installation and configuration. It can be easily deployed, automating the execution and merging of results, providing high-confidence consensus SV and CNV calls, allowing researchers to efficiently analyze complex chromosomal rearrangements in tumors without extensive bioinformatics expertise. CCRR also includes a web server for one-click analysis and customized visualization. AVAILABILITY AND IMPLEMENTATION: The CCRR platform is freely available at https://www.ccrr.life. Source code and executables can be accessed at https://github.com/laslk/CCRR. An archived version is available at Zenodo: https://doi.org/10.5281/zenodo.15386513.

Software

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors

PYRAMA: an open-source tool for advanced meta-analysis of genome wide association studies.

MOTIVATION: Genome-wide association study (GWAS) meta-analysis tools are essential for integrating summary statistics across multiple cohorts, thereby increasing statistical power and validating genetic associations. Widely cited tools, such as METAL, PLINK, and GWAMA, have facilitated numerous significant discoveries in the field of GWAS. Nevertheless, these tools offer a limited set of meta-analysis methods and typically require users to have prior experience with command-line tools to be executed. RESULTS: We present here PYRAMA, an open-source tool which is designed for meta-analysis of genome wide association studies. This work introduces an easy-to-use software package that includes several meta-analysis methods that are absent in similar software packages. PYRAMA is faster compared to other tools, supports robust methods for analysis and meta-analysis, fixed-effects, random-effects and Bayesian meta-analysis and it is currently the only tool that supports meta-analysis with imputation of summary statistics. It is available both as a standalone tool and as a freely available web server. AVAILABILITY AND IMPLEMENTATION: https://github.com/pbagos/PYRAMA, https://doi.org/10.5281/zenodo.17830449.

Genome-Wide Association Study

SynFlow: an interactive online genome structural variant viewer.

MOTIVATION: Structural variations (SVs), including inversions, translocations (TRAs), duplications, and large insertions or deletions, are key drivers of genome evolution and phenotypic diversity. With the increasing number of high-quality, chromosome-scale genome assemblies, the ability to detect and interpret SVs has become a crucial aspect of modern genomics. While SV detection has advanced, most visualization methods produce static plots that fall short when researchers, particularly in comparative genomics, need to interactively explore large datasets, zoom into specific genomic regions, or dynamically filter structural events in real time. RESULTS: To address this gap, we introduce SynFlow, a lightweight, web-based interactive application specifically designed for exploring and visualizing SVs identified by SyRI. We demonstrate that SynFlow can reproduce complex static synteny plots published in literature, but transforms them into dynamic, shareable visualizations that support real-time filtering, reordering, and deep exploration of specific SVs, including TRAs. SynFlow is available as a web server and offers multiple entry points: browsing precomputed datasets (e.g. banana and grapevine genomes), uploading user-provided SyRI outputs, or running an integrated workflow to produce and visualize SVs on the fly. AVAILABILITY AND IMPLEMENTATION: https://synflow.southgreen.fr; source code https://github.com/SouthGreenPlatform/synflow; preprocessing Snakemake workflow https://gitlab.cirad.fr/agap/cluster/snakemake/synflow.

Software

pLAST-a tool for rapid comparison and classification of bacterial plasmid sequences.

MOTIVATION: The increasing number of fully sequenced bacterial plasmids being annotated and catalogued has prompted the development of computational tools for comparing and classifying them. Existing approaches typically compare full-length DNA sequences (e.g. Mash, BLASTn, and ANI-based methods) or translated open reading frames (ORFs) (e.g. DIAMOND), with plasmid-level scores obtained by aggregating ORF-to-ORF similarities; however, they are either restricted to closely related plasmids or become computationally demanding in large-scale analyses. RESULTS: We describe pLAST (plasmid Language Analysis and Search Tool), a plasmid-search tool built using word2vec representations of protein-family content informed by local genomic context. Benchmarks indicate that pLAST outperforms nucleotide-based methods and performs comparably to DIAMOND in identifying functionally similar plasmids and compared with the widely used Mash, it achieves 26% and 24% improvements in detecting shared mating-pair formation system type and relaxase type, respectively. This performance scales to database searches across hundreds of thousands of sequences, as demonstrated using the precomputed PlasmidScope collection of ∼750 000 plasmids. Beyond global similarity, pLAST also returns per-ORF plasmid-plasmid alignments, enabling detection of shared functional modules. AVAILABILITY AND IMPLEMENTATION: pLAST is freely accessible as a web server at https://plast.lbs.cent.uw.edu.pl/ or https://plast.lbs.biol.uw.edu.pl/ and available as a Python module along with a precomputed database at https://github.com/labstructbioinf/pLAST for customized analysis.

Plasmids

MultiDMPcaller: a one-stop software for detection and visualization of differentially methylated positions and regions.

MOTIVATION: Whole-genome bisulfite sequencing (WGBS/BS-Seq) is the gold standard for single-base resolution DNA methylome profiling. However, the diverse statistical models of existing computational methods lead to limited overlap between their results, highlighting the need for novel methods to detect differentially methylated positions (DMPs) and differentially methylated regions (DMRs). RESULTS: We developed MultiDMPcaller, an automated downstream methylome analysis software. It processes upstream outputs to profile DMPs, non-DMPs, DMRs, and context-specific (CpG/CHG/CHH) methylation status, alongside visualizing their chromosomal distribution and enrichment. The software features two key innovations: (i) an adaptive two-step P-value adjustment strategy based on organism-specific methylation patterns, with raw P-value ≤0.05 pre-filtering followed by false discovery rate (FDR) correction, to recover potential DMPs usually missed by standard FDR correction in plant CHG/CHH and animal CpG contexts; and (ii) a multiple pairwise comparison approach, which performs m × n pairwise comparisons for m control and n experimental replicates, followed by a voting system supporting both user-defined majority thresholds and model-based adaptive thresholds, to identify robust and reliable DMPs (with a stricter voting threshold exclusively for loci with low methylation differences) and DMRs. On real datasets from Arabidopsis, apple, and mouse, as well as simulated human datasets, MultiDMPcaller's results showed good agreement with those of other software, exhibiting high conservativeness and superior precision, which suggested a low false discovery proportion. AVAILABILITY AND IMPLEMENTATION: MultiDMPcaller is available at GitHub (https://github.com/jiantaoyuNWAFU/MultiDMPcaller) and via a web server (https://ciebioinfo.nwafu.edu.cn).

Software

UFold-X: an enhanced Dual & Dynamic U-Mamba model for long-range RNA secondary structure prediction.

RNA secondary structure is essential for understanding the functions of non-coding RNAs, ribosomal RNAs, and viral genomes. However, accurate prediction of long RNA structures remains challenging due to complex long-range interactions and the limited availability of long-RNA training data. We present UFold-X, a dual-branch deep learning framework that combines a convolutional encoder for local structure modeling with a Mamba-based Visual State Space Module for capturing long-range dependencies. A dynamic gating mechanism adaptively integrates the two branches according to sequence length. UFold-X was evaluated on multiple benchmark datasets containing RNAs up to 5000 nucleotides. To rigorously assess generalization, we introduced a cross-clan benchmark for long RNAs. Under this stringent setting, UFold-X achieved performance comparable to state-of-the-art classical approaches while achieving the best performance among deep learning-based methods. Additional cross-family and within-family evaluations further demonstrated robust transferability and competitive predictive performance. UFold-X also maintained excellent computational efficiency, requiring only 0.08 s per sequence on average. To assess biological consistency, we developed a SHAPE-based reactivity prediction variant (UFold-X-R) and an integrated metric, the Hybrid Reactivity-Pairing Score (HRPS). UFold-X-R showed strong agreement with experimental icSHAPE data and achieved the highest HRPS among all evaluated methods. A user-friendly web server is available at https://ufold-x.ai4bread.com.

Nucleic Acid Conformation

CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity.

Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network to learn more robust protein representations alongside a standard graph neural network to learn molecular representations to predict drug-target affinity. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting drug-target affinity and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every FDA-approved drug against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.

deep learning