PubMed HealthSearch

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

From the establishment of a national bioinformatics society to the development of a national bioinformatics infrastructure.

We describe the evolution of a bioinformatics national capacity from scattered professionals into a collaborative organisation, and advancements in the adoption of the bioinformatics infrastructure philosophy by the national community. The Romanian Society of Bioinformatics (RSBI), a national professional society, was founded in 2019 to accelerate the development of Romanian bioinformatics. Incrementally, RSBI expanded its role to include: i) developing a community and engaging the public and stakeholders, ii) a national training approach, including through increased interactions with European training resources, and iii) advocating national participation in European bioinformatics infrastructures. In a next step RSBI led the development of the national bioinformatics infrastructure, the Romanian Bioinformatics Cluster (CRB) with the mission to act as an ELIXIR National Node. In this paper we report both the successful projects in training, public engagement, and policy projects, as well as initiatives related to data federation that, while not successful, can serve as valuable learning experiences for future implementations. We explain CRB's structure and the role such an entity can play in the national bioinformatics infrastructure for data, tools, and training. Finally, we offer insights into the evolving role of the bioinformatics professional society and the synergies and interactions with the forthcoming National ELIXIR Node.

Computational Biology

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

Integrated experimental and bioinformatics analysis reveals ECM-integrin and redox signaling associated with PMMA/NiO nanocomposites for craniofacial applications.

BACKGROUND: Poly(methyl methacrylate) (PMMA) is widely used in dental and craniofacial applications; however, its clinical performance is limited by poor surface wettability, moderate mechanical strength, and restricted biological activity. Integrating nanomaterial engineering with computational biology offers an opportunity to better understand biomaterial-cell interactions and support the rational design of functional biomaterials. METHODS: Nickel oxide (NiO) nanoparticles were synthesized via chemical precipitation and incorporated into PMMA to fabricate nanocomposites. Physicochemical characterization included contact angle measurements, Fourier-transform infrared spectroscopy (FTIR), scanning electron microscopy (SEM), energy-dispersive X-ray spectroscopy (EDX), and Vickers hardness testing. Biocompatibility was evaluated using zebrafish embryo developmental assays. To explore biological processes potentially associated with biomaterial-cell interactions, bioinformatics analyses including Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and STRING protein-protein interaction (PPI) network analyses were performed. RESULTS: Incorporation of NiO nanoparticles improved the surface and mechanical properties of PMMA, reducing the contact angle from 105.35° to 90.46° and increasing Vickers hardness compared with unmodified PMMA. Structural and morphological analyses confirmed successful synthesis and homogeneous nanoparticle incorporation. Zebrafish embryo studies demonstrated minimal developmental toxicity, supporting the biocompatibility of the nanocomposite. Bioinformatics analyses identified significant enrichment of pathways related to extracellular matrix organization, cell adhesion, focal adhesion, PI3K-Akt signaling, and oxidative stress regulation. Protein-protein interaction analysis revealed highly interconnected networks associated with ECM-integrin signaling and redox homeostasis, highlighting biological processes potentially associated with biomaterial-cell communication. CONCLUSIONS: PMMA/NiO nanocomposites exhibited improved physicochemical performance and favorable biocompatibility characteristics. The integration of experimental characterization with bioinformatics and network-based analyses provides a systems-level perspective on biomaterial-associated cellular processes and identifies ECM-integrin signaling and oxidative stress-related pathways as candidate biological processes for future experimental validation. These findings support the continued development of PMMA/NiO nanocomposites for oral and craniofacial biomedical applications.

Nanocomposites

Bioinformatics in crop research: using genomic data for crop improvement.

Sustainable crop development aims to maintain or increase yields while reducing environmental impact and managing the challenges imposed by climate change. As the global population grows and arable land becomes scarcer, the integration of molecular breeding with bioinformatics has emerged as an effective strategy for long-term crop improvement. Bioinformatics enables researchers to analyze and interpret the vast quantities of genetic data generated by high-throughput sequencing, making it possible to identify molecular markers, candidate genes, and regulatory networks linked to specific agronomic traits, which breeders then translate into focused, ecologically sustainable breeding programs. This approach has enabled major progress across several fronts: the identification of genes conferring resistance to biotic stressors (pests, pathogens) and abiotic stressors (drought, salinity, heat); the development of nutrient-efficient, low-input crop varieties; the improvement of agronomic performance and nutritional quality through identification of yield- and quality-related genes; and the conservation and deployment of genetic diversity to safeguard long-term breeding sustainability. By combining genomic data with precision breeding techniques, researchers are developing crops that are better adapted to a growing population and a changing climate, positioning the integration of molecular breeding and bioinformatics as a central pillar of future global food security.

bioinformatics

VirDetector: a bioinformatic pipeline for virus surveillance using nanopore sequencing.

SUMMARY: Virus surveillance programmes are designed to counter the growing threat of viral outbreaks to human health. Nanopore sequencing, in particular, has proven to be suitable for this purpose, as it is readily available and provides rapid results. However, as special bioinformatic programs are required to extract the relevant information from the sequencing data, applications are needed that allow users without extensive bioinformatics knowledge to carry out the relevant analysis steps. We present VirDetector, a bioinformatic pipeline for virus surveillance using nanopore sequencing. The pipeline automatically installs all required programs and databases and allows all its steps to be executed with a single console command. After preprocessing the samples, including the possibility for basecalling, the pipeline classifies each sample taxonomically and reconstructs the viral consensus genomes, which are then used in phylogenetic analyses. This streamlined workflow provides a user-friendly and efficient solution for monitoring viral pathogens. AVAILABILITY AND IMPLEMENTATION: VirDetector is freely available at https://github.com/NLKaiser/VirDetector and https://zenodo.org/records/14637302 (10.5281/zenodo.14637302).

Nanopore Sequencing

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology

Unlocking microbial potential: advances in omics and bioinformatics for aromatic hydrocarbon degradation.

Aromatic hydrocarbons (AHs) are persistent environmental pollutants with high toxicity. Bacterial degradation of AHs provides a sustainable and cost-effective approach for the remediation of sites contaminated with both mono- and polycyclic aromatic hydrocarbons. Aerobic degradation of AHs typically involves oxygenases-mediated hydroxylation followed by aromatic ring cleavage. In contrast, anaerobic degradation relies on diverse activation mechanisms that ultimately converge on the central intermediate benzoyl-CoA. Over the past decades, research on bacterial degradation of AHs has grown steadily, supported by advances in omics and bioinformatics. In this review, we summarize the current knowledge on the pathways, enzymes, and microbial diversity involved in AH degradation, highlighting how omics and bioinformatic approaches are advancing our understanding of this process. However, to improve our knowledge of microbial AHs catabolism, it is crucial to prioritize the characterization of novel enzymes and pathways, especially those mediating anaerobic and hybrid degradation strategies. Addressing this gap requires the development of specialized resources that incorporate a broader taxonomic diversity and an expanded inventory of anaerobic genes and enzymes supported by experimental evidence. Equally important is the integration of multi-omics technologies, artificial intelligence, and ecological modeling into unified analytical pipelines. These efforts will be key to fully unlocking microbial metabolic potential and guiding more effective bioremediation and monitoring strategies for AHs.

Biodegradation, Environmental

Screening and identification of key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods.

OBJECTIVE: Radiotherapy (RT) plays a crucial role in the comprehensive treatment of rectal cancer. However, the impact of radiotherapy on the tumor microenvironment (TME), especially its effect on immune cell infiltration and immune-related gene expression, has not been fully studied. This study aims to screen and analyze key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods for the purpose of identifying potential biomarkers and providing new insights for the personalized therapy of rectal cancer. METHODS: Using data from the Public Gene Expression Database (GEO) and the Cancer Genomics Database (TCGA), the impact of radiotherapy on the immune microenvironment of rectal cancer was explored using bioinformatics tools. Through screening differentially expressed genes (DEGs), correlation analysis, TIMER database analysis, immune infiltration score, and correlation analysis between key genes and prognosis, the effects of radiotherapy on the immune microenvironment of rectal cancer were investigated. RESULTS: Totally 7 upregulated and 4 downregulated differentially expressed genes were identified, among which MASP1, LTK, SLC9A3R2 were negatively correlated with myeloid suppressor cell infiltration (MDSCs), while ZP2 was positively correlated. The expression of MASP1 and SLC9A3R2 was closely related to the level of immune cell infiltration and played significant roles in the immune microenvironment. High expression of MASP1 was significantly correlated with survival benefits from immune checkpoint inhibitor therapy, while SLC9A3R2 was closely related to the efficacy of PD-L1 inhibitors and CTLA4 inhibitors. CONCLUSIONS: MASP1 and SLC9A3R2, as two key genes that may be related to the immune microenvironment of rectal cancer radiotherapy, deserve further exploration of their roles in the mechanism. The combination of radiotherapy and immunotherapy holds promising prospects in the treatment of rectal cancer, and exploration of related mechanisms will provide new strategies and targets for the treatment of various tumors and rectal cancer.

Bioinformatics

Genetic and clinical insights into the coexistence of multiple myeloma and diffuse large B cell lymphoma from a case report and systematic review with bioinformatics analysis.

BACKGROUND: Multiple myeloma (MM) and diffuse large B-cell lymphoma (DLBCL) are B-cell malignancies that rarely coexist in a single patient, presenting significant diagnostic and therapeutic challenges. While MM primarily involves clonal plasma cells, DLBCL is an aggressive lymphoid neoplasm. Investigating shared genetic mutations and understanding their clinical relevance in both cancers could provide novel insights into their pathogenesis and underlying molecular mechanisms, thereby informing future translational research. MATERIALS AND METHODS: A case report was conducted on a 52-year-old male who presented with abdominal pain and anemia. Imaging revealed lymphadenopathy, and biopsy confirmed high-grade DLBCL with concurrent bone marrow involvement suggestive of MM. Laboratory tests identified monoclonal IgM gammopathy, and the patient was treated with R-CHOP (Rituximab, Cyclophosphamide, Doxorubicin, Vincristine, and Prednisone) chemotherapy for DLBCL followed by autologous stem cell transplantation (ASCT) for MM relapse. A systematic review of the literature was performed using PubMed, Scopus, and Web of Science databases to identify cases of patients diagnosed with both MM and DLBCL. Data on patient demographics, clinical features, treatment regimens, and outcomes were extracted. Additionally, bioinformatics analysis was conducted using publicly available genomic data from cBioPortal and IntOGen to identify driver gene mutations in MM and DLBCL. Functional and pathway enrichment analysis was performed with KEGG and Gene Ontology (GO) databases. RESULTS: The case report highlighted a complex clinical course where the patient initially responded well to R-CHOP chemotherapy for DLBCL, achieving remission, but later relapsed with MM, treated with ASCT and lenalidomide. The systematic review revealed 14 eligible studies in which MM and DLBCL often occur in older patients, either simultaneously or sequentially, with variable treatment responses, including complete remission, partial remission, or relapse. The bioinformatics analysis identified several shared function and cancer-related pathways between two cancers including interleukin and cytokine-mediated signaling pathways, regulation of cell cycle, neurotrophin signaling pathway, FOXO signaling pathway, Epstein Barr virus infection, and viral carcinogenesis. CONCLUSION: This study provides valuable insights into the dual occurrence of MM and DLBCL, emphasizing the importance of tailored treatment approaches. The driver mutations identified highlight overlapping oncogenic pathways rather than implying a shared clonal origin, and may inform future studies exploring their biological and clinical implications. Further research into these shared molecular mechanisms could lead to more effective treatments for patients with coexisting MM and DLBCL.

Bioinformatics analysis

Exploring shared biomarkers and their mechanisms in thyroid cancer and systemic lupus erythematosus via bioinformatics analysis.

BACKGROUND: Systemic lupus erythematosus (SLE), an autoimmune disorder, is linked to a heightened risk of multiple malignancies, including thyroid cancer. Thyroid cancer is the most prevalent malignancy of the endocrine system, and its autoimmune-related pathological features render it an optimal subject for investigating the mechanisms of their comorbidity. The molecular mechanisms underlying this comorbidity are still ambiguous. The accurate diagnosis and treatment of thyroid cancer urgently necessitate innovative molecular targets that extend beyond conventional pathological characteristics. This study seeks to employ integrated bioinformatics approaches to elucidate potential shared molecular mechanisms and immunological features between thyroid cancer and systemic lupus erythematosus (SLE), aiming to enhance understanding of their comorbidity and identify novel intervention targets. METHODS: This study initially acquired gene expression data for TC and SLE from the GEO database and subsequently screened and identified differentially expressed genes (DEGs) shared by both diseases. Subsequently, we conducted Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Reactome functional enrichment analyses on these 46 shared differentially expressed genes (DEGs) and further assessed the activation status of pertinent pathways using Gene Set Enrichment Analysis (GSEA). Subsequently, we employed CIBERSORTx to examine immune infiltration patterns and developed protein-protein interaction networks utilising the STRING database. We identified hub genes utilising the MCODE and cytoHubba plugins and visualised the findings with Cytoscape software. We additionally assessed the diagnostic efficacy of these core hub genes in an independent dataset utilising ROC curves and investigated their prognostic relevance in thyroid cancer through Kaplan-Meier survival analysis and multivariate Cox proportional hazards regression. Ultimately, we employed the Network Analyst platform to forecast transcription factor-gene and miRNA-gene regulatory networks and identified potential targeted therapeutic compounds utilising the DSigDB database. RESULTS: This study identified 46 differentially expressed genes (DEGs) commonly linked to thyroid cancer and systemic lupus erythematosus (SLE), which were significantly enriched in signalling pathways associated with immune-inflammatory activation, type I interferon responses, and complement pathway activation. Moreover, GSEA findings validated that immune-inflammatory and autoimmune-related pathways are markedly activated in both conditions. Twelve hub genes were discerned through protein-protein interaction networks. Analysis of immune infiltration indicated that thyroid cancer and systemic lupus erythematosus exhibit a shared characteristic of innate immune dysregulation, marked by the infiltration of myeloid cells (neutrophils, M0/M2 macrophages). Receiver operating characteristic (ROC) curve analysis identified six significant core hub genes with substantial diagnostic value: C1QB, LCN2, C1QC, LTF, VSIG4, and C3AR1. Univariate survival analysis indicated that elevated expression of C1QC and C3AR1 significantly enhances overall survival in thyroid cancer patients; however, multivariate COX regression analysis revealed that their independent prognostic significance necessitates further validation. This study predicted the interaction networks of transcription factors and miRNAs regulating key genes, with LCN2 demonstrating the highest connectivity to miRNAs, and identified candidate therapeutic compounds linked to it. CONCLUSION: This study employed bioinformatics analysis to identify critical shared hub genes and molecular pathways connecting thyroid cancer and systemic lupus erythematosus, offering novel insights into their shared pathogenesis and the advancement of targeted biomarkers and therapeutic strategies.

Bioinformatics analysis

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

Identification of key genes related to bone metastasis of breast cancer using bioinformatics methods and construction of a prognostic model.

Breast cancer (BC) ranks among the most prevalent cancers in females, with bone metastasis significantly compromising patients' quality of life and survival rates. Enhancing our comprehension of BC bone metastasis mechanisms at the molecular level holds promise for improving BC treatment and prognosis. Leveraging bioinformatics tools, we integrated multiple datasets, conducted comprehensive analyses across various databases, identified biomarkers associated with BC bone metastasis, and constructed a prognostic model. Firstly, 3 BC bone metastasis-related datasets were downloaded from gene expression omnibus, the data were merged, and batch effects were removed, followed by identification of differentially expressed genes (DEGs). Gene ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the DEGs. A protein-protein interaction network was constructed using the STRING database to screen hub genes. Then, survival analysis of hub genes was performed using the Cancer Genome Atlas (TCGA) database. A prognostic model was constructed using key genes with survival differences, and the model was evaluated. Two hundred ninety-two DEGs were identified. Gene ontology and KEGG pathway enrichment analysis yielded 769 biological processes (BPs), 78 cellular components, 43 molecular functions, and 50 KEGG pathways. Fifteen hub genes were selected from the protein-protein interaction network. Survival analysis revealed 6 genes related to BC survival. The prognostic model identified 4 genes with important predictive value for BC prognosis. Our study utilized bioinformatics analysis to identify a series of DEGs related to BC bone metastasis. Based on further selection of hub genes, we constructed a relatively ideal prognostic model for BC, and identified 4 genes (DLGAP5, TPX2, PLK1, and CENPN) with valuable predictive value for BC prognosis.

Humans

Integrative analyses of mendelian randomization and bioinformatics reveal casual relationship and genetic links between COVID-19 and knee osteoarthritis.

BACKGROUND: Clinical and epidemiological analyses have found an association between coronavirus disease 2019 (COVID-19) and knee osteoarthritis (KOA). Infection with COVID-19 may increase the risk of developing KOA. OBJECTIVES: This study aimed to investigate the potential causal relationship between COVID-19 and KOA using Mendelian randomization (MR) and to explore the underlying mechanisms through a systematic bioinformatics approach. METHODS: Our investigation focused on exploring the potential causal relationship between COVID-19, acute upper respiratory tract infection (URTI) and KOA utilizing a bidirectional MR approach. Additionally, we conducted differential gene expression analysis using public datasets related to these three conditions. Subsequent analyses, including transcriptional regulation analysis, immune cell infiltration analysis, single-cell analysis, and druggability evaluation, were performed to explore potential mechanisms and prioritize therapeutic targets. RESULTS: The results indicate that COVID-19 has a one-way impact on KOA, while URTI does not play a causal role in this association. Ribosomal dysfunction may serve as an intermediate factor connecting COVID-19 with KOA. Specifically, COVID-19 has the potential to influence the metabolic processes of the extracellular matrix, potentially impacting the joint homeostasis. A specific group of genes (COL10A1, BGN, COL3A1, COMP, ACAN, THBS2, COL5A1, COL16A1, COL5A2) has been identified as a shared transcriptomic signature in response to KOA with COVID-19. Imatinib, Adiponectin, Myricetin, Tranexamic acid, and Chenodeoxycholic acid are potential drugs for the treatment of KOA patients with COVID-19. CONCLUSIONS: This study uniquely combines Mendelian randomization and bioinformatics tools to explore the possibility of a causal relationship and genetic association between COVID-19 and KOA. These findings are expected to provide novel perspectives on the underlying biological mechanisms that link COVID-19 and KOA.

Humans

Investigating the molecular mechanisms, drug prediction, and validation of CCNA2 and MAD2L1 in esophageal squamous cell carcinoma based on bioinformatics.

OBJECTIVE: Aims to comprehensively investigate the expression patterns of CCNA2 and MAD2L1 in esophageal squamous cell carcinoma using bioinformatics methods. METHODS: Based on WGCNA analysis of gene mutation expression, methylation level distribution, mRNA expression and ESCC-related genes in public databases, were employed for investigating potential biomarkers for prognosis of esophageal squamous cell carcinoma(ESCC).Finally,. performing qRT-PCR and immunohistochemistry to validate. RESULTS: Ultimately identified 4 hub genes: CDK1, CCNA2,TOP2A and MAD2L1. Bioinformatics analysis showed high expression of these four genes in ESCC (P&#x2009;<&#x2009;0.05). CCNA2 and MAD2L1 were selected for subsequent analysis based on literature.3.Single gene enrichment analysis revealed significant enrichment of CCNA2 and MAD2L1 in pathways related to splicing, bladder cancer, non-homologous end joining and homologous recombination, glycosaminoglycan biosynthesis chondroitin sulfate, progesterone-mediated oocyte maturation and mismatch repair. PASTAA database indicated the involvement of transcription factors such as Roralpha1, Pou6f1, Roralpha2, Atf-1, Pax-3, C/ebpalpha, Nkx2-1 in the regulation of CCNA2, while no transcription factors were predicted for MAD2L1..Immune infiltration analysis revealed a close association between ESCC and plasma cells, CD8&#x2009;+&#x2009;T cells, monocytes, M0 macrophages, M1 macrophages, dendritic cells, and resting mast cells.Drug prediction for CCNA2 included 7 drugs such as ETHINYL ESTRADIOL, Seliciclib and TAMOXIFEN, while no drugs were predicted for MAD2L1.qRT-PCR and immunohistochemistry demonstrated high expression of CCNA2 in ESCC, while MAD2L1 showed no significant difference between ESCC and normal esophageal squamous epithelial tissues. CONCLUSION: CCNA2 and MAD2L1 may be potential biomarkers for ESCC, providing a novel basis for understanding the molecular mechanisms underlying ESCC pathogenesis.Additionally, the potential drugs predicted for CCNA2 may emerge as a new hope for ESCC patients in the future.

Humans

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics