PubMed HealthSearch

SEARCH · PubMed Health

Results for “Computational Biology”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Catalyzing computational biology research at an academic institute through an interest network.

Biology has been transformed by the rapid development of computing and the concurrent rise of data-rich approaches such as, omics or high-resolution imaging. However, there is a persistent computational skills gap in the biomedical research workforce. Inherent limitations of classroom teaching and institutional core support highlight the need for accessible ways for researchers to explore developments in computational biology. An analysis of the Scripps Research Genomics Core revealed increases in the total number and diversity of experiments: the share of experiments other than bulk RNA- or DNA-sequencing increased from 34% to 60% within 10 years, requiring more tailored computational analyses. These challenges were tackled by forming a volunteer-led affinity group of approximately 300 academic biomedical researchers interested in computational biology, referred to as the Computational Biology and Bioinformatics (CBB) affinity group. This adaptive group has provided continuing education and networking opportunities through seminars, workshops, and coding sessions while evolving along with the needs of its members. A survey of CBB's impact confirmed the group's events increased the members' exposure to computational biology educational and research events (79% respondents) and networking opportunities (61% respondents). Thus, volunteer-led affinity groups may be a viable complement to traditional institutional resources for enhancing the application of computing in biomedical research.

Computational Biology

An educator framework for organizing Wikipedia editathons for computational biology.

MOTIVATION: Wikipedia is a vital open educational resource in computational biology; however, a significant knowledge gap exists between English and non-English Wikipedias. Reducing this knowledge gap via intensive editing events, or "editathons," would be beneficial in reducing language barriers that disadvantage learners whose native language is not English. Results: We present a framework to guide educators in organizing editathons for learners to improve and create relevant Wikipedia articles. As a case study, we present the results of an editathon held at the 2024 ISCB Latin America conference, in which ten new articles were created for the Spanish-language edition of Wikipedia. We also present a web tool, "compbio-on-wiki," which identifies relevant English Wikipedia articles missing in other languages. We demonstrate the value of editathons to expand the accessibility and visibility of computational biology content in multiple languages. AVAILABILITY AND IMPLEMENTATION: Source code for the compbio-on-wiki Toolforge site is available at: https://github.com/lubianat/compbio-on-wiki.

Computational Biology

Computational network biology analysis revealed COVID-19 severity markers: Molecular interplay between HLA-II with CIITA.

COVID-19, severe acute respiratory syndrome coronavirus 2, rapidly spread worldwide. Severe and critical patients are expected to rapidly deteriorate. Although several studies have attempted to uncover the mechanisms underlying COVID-19 severity, most have focused on the perturbations of single genes. However, the complex mechanism of COVID-19 involves numerous perturbed genes in a molecular network rather than a single abnormal gene. Thus, we aimed to identify COVID-19 severity-specific markers in the Japanese population using gene network analysis. In order to reveal the severity-specific molecular interplays, we developed a novel computational network biology strategy that measures dissimilarity between networks based on the comprehensive information of gene network (i.e., expression levels of genes and network structure) by using Kullback-Leibler divergence. Monte Carlo simulations demonstrated the effectiveness of our strategy for differential gene network analysis. We applied this method to publicly available whole blood RNA-seq data from the Japan coronavirus disease 2019 Task Force and identified differentially regulated molecular interplays between 368 severe and 105 non-severe samples. Our analysis suggests the gene network between HLA class II, CIITA, and CD74 as a COVID-19 severity specific molecular marker. Although the association between HLA class II and COVID-19 has been demonstrated, our data analysis revealed that the molecular interplay of HLA class II with its target and/or regulator is a crucial marker for COVID-19 severity. Our findings from computational network biology analysis suggest that suppression and activation of the molecular interplay between HLA class II, CIITA, and CD74 provide crucial clues to uncover the mechanisms of COVID-19 severity.

Humans

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6 mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6 mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6 mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6 mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6 mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning

Integrated experimental and bioinformatics analysis reveals ECM-integrin and redox signaling associated with PMMA/NiO nanocomposites for craniofacial applications.

BACKGROUND: Poly(methyl methacrylate) (PMMA) is widely used in dental and craniofacial applications; however, its clinical performance is limited by poor surface wettability, moderate mechanical strength, and restricted biological activity. Integrating nanomaterial engineering with computational biology offers an opportunity to better understand biomaterial-cell interactions and support the rational design of functional biomaterials. METHODS: Nickel oxide (NiO) nanoparticles were synthesized via chemical precipitation and incorporated into PMMA to fabricate nanocomposites. Physicochemical characterization included contact angle measurements, Fourier-transform infrared spectroscopy (FTIR), scanning electron microscopy (SEM), energy-dispersive X-ray spectroscopy (EDX), and Vickers hardness testing. Biocompatibility was evaluated using zebrafish embryo developmental assays. To explore biological processes potentially associated with biomaterial-cell interactions, bioinformatics analyses including Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and STRING protein-protein interaction (PPI) network analyses were performed. RESULTS: Incorporation of NiO nanoparticles improved the surface and mechanical properties of PMMA, reducing the contact angle from 105.35° to 90.46° and increasing Vickers hardness compared with unmodified PMMA. Structural and morphological analyses confirmed successful synthesis and homogeneous nanoparticle incorporation. Zebrafish embryo studies demonstrated minimal developmental toxicity, supporting the biocompatibility of the nanocomposite. Bioinformatics analyses identified significant enrichment of pathways related to extracellular matrix organization, cell adhesion, focal adhesion, PI3K-Akt signaling, and oxidative stress regulation. Protein-protein interaction analysis revealed highly interconnected networks associated with ECM-integrin signaling and redox homeostasis, highlighting biological processes potentially associated with biomaterial-cell communication. CONCLUSIONS: PMMA/NiO nanocomposites exhibited improved physicochemical performance and favorable biocompatibility characteristics. The integration of experimental characterization with bioinformatics and network-based analyses provides a systems-level perspective on biomaterial-associated cellular processes and identifies ECM-integrin signaling and oxidative stress-related pathways as candidate biological processes for future experimental validation. These findings support the continued development of PMMA/NiO nanocomposites for oral and craniofacial biomedical applications.

Nanocomposites

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides insight into the importance of training data selection relative to model size, the need for domain-specific training, and the capabilities and limitations of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

Prophages

Assessment of Gene Set Enrichment Analysis using curated RNA-seq-based benchmarks.

Pathway enrichment analysis is a ubiquitous computational biology method to interpret a list of genes (typically derived from the association of large-scale omics data with phenotypes of interest) in terms of higher-level, predefined gene sets that share biological function, chromosomal location, or other common features. Among many tools developed so far, Gene Set Enrichment Analysis (GSEA) stands out as one of the pioneering and most widely used methods. Although originally developed for microarray data, GSEA is nowadays extensively utilized for RNA-seq data analysis. Here, we quantitatively assessed the performance of a variety of GSEA modalities and provide guidance in the practical use of GSEA in RNA-seq experiments. We leveraged harmonized RNA-seq datasets available from The Cancer Genome Atlas (TCGA) in combination with large, curated pathway collections from the Molecular Signatures Database to obtain cancer-type-specific target pathway lists across multiple cancer types. We carried out a detailed analysis of GSEA performance using both gene-set and phenotype permutations combined with four different choices for the Kolmogorov-Smirnov enrichment statistic. Based on our benchmarks, we conclude that the classic/unweighted gene-set permutation approach offered comparable or better sensitivity-vs-specificity tradeoffs across cancer types compared with other, more complex and computationally intensive permutation methods. Finally, we analyzed other large cohorts for thyroid cancer and hepatocellular carcinoma. We utilized a new consensus metric, the Enrichment Evidence Score (EES), which showed a remarkable agreement between pathways identified in TCGA and those from other sources, despite differences in cancer etiology. This finding suggests an EES-based strategy to identify a core set of pathways that may be complemented by an expanded set of pathways for downstream exploratory analysis. This work fills the existing gap in current guidelines and benchmarks for the use of GSEA with RNA-seq data and provides a framework to enable detailed benchmarking of other RNA-seq-based pathway analysis tools.

Humans

The molecular landscape of chordoma: Current frontiers from multi-omics to artificial intelligence.

Chordoma is a rare and aggressive malignant bone tumor of the axial skeleton that has historically challenged clinicians due to its complex anatomical locations and a high recurrence rate of up to 85%. This review synthesizes the most recent advances in chordoma research and offers an overview of how multi-omics, advanced immunology, and artificial intelligence are reshaping the treatment paradigm. Central to its pathogenesis is the T-box transcription factor Brachyury, which this review highlights as both the pathognomonic diagnostic marker and the primary therapeutic vulnerability. Cutting-edge innovations targeting this driver include covalent small-molecule binders, targeted protein degradation, and peptide-centric CAR-T cells designed to attack the intracellular oncoprotein. The tumor immune microenvironment is functionally dynamic, and new dimensions in cellular therapy, such as dual-specific CAR constructs and NK-cell platforms, are being engineered to neutralize immunosuppressive factors. Beyond biological insights, the review emphasizes the role of computational biology, specifically how deep-learning and machine-learning models achieve expert-level precision in tumor segmentation and personalized survival forecasting. By integrating genomic, transcriptomic, epigenomic, and proteomic data, multiomics approaches can fully elucidate chordoma subtypes and underlying resistance mechanisms, ultimately paving the way for more precise and personalized therapeutic strategies.

Humans

Pedigree Painter (pepa): a tool for the visualization of genetic inheritance in chromosomal context.

MOTIVATION: Data visualization is increasingly important in genomics, enabling researchers to uncover inheritance and recombination patterns across generations. While most existing tools focus on ancestry prediction, they lack functionality for analyzing known ancestries in controlled settings, such as determining parental contributions to offspring genomes. To address this gap, I developed pepa, a lightweight, deterministic, modular tool that visualizes and quantifies genomic inheritance, designed for beginner and advanced users. RESULTS: pepa is a program for processing VCF files, assigning ancestries to homozygous SNPs, and clustering them into biologically meaningful regions. It generates human-readable comparison tables and visualizes inheritance patterns with chromosome paintings through R. Tested on fission yeast, pepa revealed non-uniform recombination patterns, with chromosomes largely inherited from one parent and seemingly random recombination. Quantitative analyses showed differences in parental contributions at the nucleotide and gene levels, with some offspring inheriting similar percentages from parents. However, the painted chromosomes revealed that even offspring with similar percentages from one parent rarely inherit the same genomic region, highlighting the importance of this tool in drawing biologically meaningful insights. pepa provides an accessible and powerful solution for analyzing genomic inheritance, bridging experimental and computational biology. Its modular design and minimal dependencies allow adaptation to diverse organisms, facilitating intuitive visualization and quantitative insights into recombination dynamics.

Pedigree

Multi-scale phylodynamic modelling of rapid punctuated pathogen evolution.

Computational multi-scale pandemic modelling remains a major and timely challenge. Here we identify specific requirements for a new class of models simulating pandemics across three scales: (1) pathogen evolution, often punctuated by the rapid emergence of new variants, (2) human interactions within a heterogeneous population, and (3) public health responses which constrain individual actions to control the disease transmission. We then present a pandemic modelling framework satisfying these requirements and capable of simulating feedback loops between dynamics unfolding at these different scales. The developed framework comprises a stochastic agent-based model of pandemic spread, coupled with a phylodynamic model that incorporates within-host pathogen evolution. It is validated with a case study, modelling the punctuated evolution of SARS-CoV-2, based on global and contemporary genomic surveillance data, which captures a large heterogeneous population. We demonstrate that the model replicates the essential features of the COVID-19 pandemic and virus evolution, while retaining computational tractability and scalability.

SARS-CoV-2

Influence of Major Histocompatibility Complex (MHC) Diversity on Immune Modulation, Pathogenesis, and Control of Lumpy Skin Disease Virus.

INTRODUCTION: Lumpy Skin Disease Virus (LSDV), a member of the genus Capripoxvirus within the family Poxviridae, is an economically important transboundary viral pathogen affecting cattle and water buffalo. The disease causes severe production losses through decreased milk yield, infertility, hide damage, reduced growth performance, and occasional mortality. The rapid geographic spread of LSDV, together with its vectorborne transmission and emerging recombinant strains, has intensified the need for improved understanding of viral pathogenesis, host immune responses, and effective prevention strategies. In particular, the role of the bovine Major Histocompatibility Complex (BoLA/MHC) in regulating antiviral immunity, disease susceptibility, and vaccine responsiveness has gained increasing scientific attention. METHODS: This review summarises the published literature related to the epidemiology, transmission, structure, pathogenesis, diagnosis, prevention, and control of LSDV, with special emphasis on the immunological and molecular role of bovine MHC molecules. Relevant studies concerning BoLA-mediated antigen presentation, immunoinformaticsbased epitope prediction, vaccine development, antiviral drug repurposing, molecular docking, genomic surveillance, and diagnostic approaches, including PCR- and ELISAbased assays, were critically evaluated. Recent advances in computational biology, molecular virology, and host-pathogen interaction studies were also reviewed. RESULTS: The reviewed studies demonstrate that Lumpy Skin Disease Virus (LSDV) possesses a complex double-stranded DNA genome enabling immune modulation and efficient transmission through arthropod vectors such as mosquitoes, ticks, and biting flies. Disease progression involves systemic viral replication, vascular injury, dermal necrosis, and inflammatory skin lesions. Real-time PCR remains the most sensitive diagnostic method for early detection, while ELISA supports surveillance. Evidence highlights the central role of bovine Major Histocompatibility Complex (BoLA) molecules in antigen presentation and T-cell activation. Computational studies identified promising BoLA-binding epitopes and repurposed antiviral candidates, including ivermectin, theaflavin, canagliflozin, and tepotinib, for future therapeutic development. DISCUSSION: Current evidence indicates that effective LSDV control requires integration of molecular diagnostics, vector management, vaccination, and host immunogenetics. BoLAguided immunoinformatics provides promising opportunities for developing multi-epitope vaccines, although experimental validation remains essential. Similarly, repurposed antiviral candidates require comprehensive in vivo and pharmacological evaluation before clinical application. Future research should focus on elucidating viral immune-evasion mechanisms, validating predicted epitopes, and translating computational findings into practical vaccines and therapeutics for sustainable disease control. CONCLUSION: Lumpy Skin Disease continues to pose a major threat to global cattle health and livestock economies. Advances in molecular diagnostics, genomic surveillance, antiviral drug discovery, and BoLA-guided vaccine design provide promising opportunities for improved disease control. Understanding the interaction between LSDV and the bovine MHC system is essential for developing next-generation vaccines, immunotherapeutics, and precision disease-management strategies. Future research should prioritise experimental validation of predicted epitopes, large-scale vaccine trials, and mechanistic studies on host-virus immune interactions to establish effective and sustainable global control programs for LSDV.

BoLA

LAMBDA: A Prophage Detection Benchmark for Genomic Language Models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, indicating a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides novel insights into the importance of training data quality relative to model size, the need for domain-specific training, and the application of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

DNA language model

Evolution, Mechanisms, and Therapeutic Implications of Mobile Tetracycline Destructases.

Tetracycline destructases (TDases) pose an emerging global threat by enzymatically inactivating all generations of tetracycline (Tet) antibiotics, including last-resort agents such as tigecycline. Despite their recent identification, TDases have rapidly disseminated worldwide, largely driven by mobile genetic elements and environmental reservoirs. This review synthesizes current knowledge on TDase genomics, structural and catalytic mechanisms, ecological niches, and clinical impacts. We detail the mechanistic distinctions between type 1 and type 2 TDases, emphasizing their divergent structural configurations and substrate specificity profiles. Additionally, we examine strategies for therapeutic intervention, highlighting progress in structure-guided inhibitor development. Key gaps remain in understanding ancestral reservoirs, evolutionary trajectories, and effective surveillance strategies. Addressing these areas through integrative evolutionary, biochemical, and ecological studies is critical for mitigating the clinical spread and therapeutic impact of TDases globally.

Humans

Interrogating functional connectivity of in vitro neural glia tissue model modulated through integrative control of matrix stiffness and a neurotrophic factor.

Brain function emerges from intricate cellular communication within neural networks. Both In silico neuronal models and primary neuron cells have revealed that the branching architecture of individual neurons determines the bioelectrical signal propagation pattern and dynamics. However, whether stem cell-differentiated neurons can build functional connectivity regulated by neuronal morphology has yet to be determined. Here, we hypothesized that neurite length, branching, or both factors would regulate the functional connectivity of the stem cell-differentiated neural network. We examined this hypothesis by differentiating mouse cortical neural stem cells (NSCs) on Matrigel substrates with varying storage moduli, both with and without basic fibroblast growth factor (bFGF). Interestingly, with bFGF, Matrigel with a storage modulus (G') of 100 Pa drives NSCs to differentiate into neurons with more dendritic branches, while the gel with G' of 50 Pa led to the development of longer neurites with fewer branches. Notably, branch-rich neural networks exhibited an increased frequency of calcium transients. Using a MATLAB-based analysis pipeline incorporating graph theory, we constructed spatial and temporal calcium activity maps, revealing that branching complexity, more than neurite length, correlates with the density and strength of functional neural circuits. Overall, this study demonstrates that the dendritic branching of neurons, modulated with matrix stiffness and neurotrophic factors, is a key element in enhancing the electrophysiological functionality of the stem cell-differentiated neural network. This finding will have a significant impact on efforts to reconstruct functional neural tissue models, advancing both regenerative therapies and unexplored applications, including biological computing.

Animals

Secure bioinformatics: privacy-preserving federated analytics using homomorphic encryption.

MOTIVATION: Large-scale bioinformatics analyses increasingly require collaboration across multiple cohorts and institutions, yet existing workflows often rely on data co-localization, which is slow, difficult to scale, and raises privacy concerns. We present a privacy-preserving federated analytics framework that enables secure statistical analysis across distributed datasets without transferring raw data, by performing all computations on encrypted data via cryptographic methods. RESULTS: We evaluate the framework by validating polygenic risk scores and conducting meta-analyses on two real-world cohorts. The proposed solution achieves over 99.9% accuracy relative to plaintext analyses, while maintaining scalable runtime performance with increasing data size and number of participating sites. These results demonstrate the feasibility of secure federated analytics for practical bioinformatics applications involving sensitive data.

Computational Biology

Genome-wide identification and characterization of 1-amino-cyclopropane-1- carboxylate synthase (ACS) gene family in Carica papaya and expression insights in response to hormone stress.

ACC-synthase (1-aminocyclopropane-1-carboxylate synthase), also known as the ACS gene, plays a pivotal role in ethylene production, which is of great importance in the fruit ripening process for producing saleable yield (marketable fruit). The ACS gene family presumably controls stress responses, plant growth and development, and particularly fruit ripening. Computational biology was used as an essential tool to identify seven ACS genes in Carica papaya (red hermaphrodite) using an RNA-seq database (NCBI GEO). Further, the phylogenetic relationships of ACS genes determined gene family resemblance in the genomes of Hordeum vulgare, Musa acuminata, C. papaya, and Arabidopsis thaliana; therefore, the identified gene families were further classified into four distinct clades (Type-I, Type-II, Type-III, and Type-IV) in alignment with the well-established Arabidopsis classification. Moreover, encompassing gene structure, domain motifs, cis-element phylogenetic profiling, synteny, and transcriptomic profiling unveiled latent structural and functional attributes within CpACS genes. Through segmental duplication of CpACS, insights into evolutionary duplication events were predicted. The paralogous behavior of ACS genes in C. papaya and a comprehensive transcriptomic analysis demonstrated both up- and down-regulation patterns in response to ethylene treatment at different time points during the fruit ripening process, using the papaya manual handbook V2 (2021). Gene expression showed upregulation of two essential CpACS genes, CpACS5 and CpACS6. RT-qPCR validates the expression of these important genes during fruit ripening. However, one gene, CpACS7, is expressed in the later stages of fruit development. Our results demonstrated novel avenues for understanding the expression pathways of the ACS gene family in red hermaphrodite papaya, and most of these genes were linked to regulating various abiotic stresses, plant growth, and fruit development.

Carica