PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy.

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Databases, Protein↗

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans↗

Unveiling the power of TIIC: A prognostic tool for esophageal adenocarcinoma.

BACKGROUND: Esophageal adenocarcinoma (EAC) remains a lethal malignancy with limited prognostic tools for guiding immunotherapy. Tumor-infiltrating immune cells (TIICs) play a critical role in EAC prognosis and treatment response. METHODS: We integrated single-cell RNA sequencing and bulk transcriptome data from TCGA and GEO databases. TIIC-specific RNAs were identified via tissue specificity index calculation combined with machine learning feature selection. Twenty machine learning algorithms were benchmarked to construct an optimal TIIC signature score (TIIC-Score) based on the comprehensive C-index. Immunotherapy response, genomic mutation, and copy number variation were analyzed. Summary-data-based Mendelian randomization (SMR) and two-sample Mendelian randomization (MR) were performed to explore genetic associations. Core prognostic TIIC-related genes were functionally validated in esophageal cancer cell lines through loss-of-function assays. RESULTS: The TIIC-Score demonstrated robust prognostic value for 1-, 2-, and 3-year overall survival across multiple cohorts, outperforming 22 published models. High TIIC-Score was associated with poor survival and increased chromosomal instability. Mutation profiling revealed high frequencies of TP53 (78.2%), TTN (48.7%), and SYNE1 (30.8%). MR analysis identified a significant association between gastro-oesophageal reflux and EAC risk at SNP rs8130507. Functionally, CCNI was upregulated in esophageal cancer cells, and its knockdown suppressed malignant phenotypes while promoting apoptosis, supporting its pro-tumorigenic role. CONCLUSION: The TIIC-Score provides a novel prognostic framework for EAC that effectively stratifies patient risk and may help identify individuals most likely to benefit from immunotherapy.

Esophageal adenocarcinoma↗

RAS signaling in lung adenocarcinoma is defined by lineage context and DUSP4 loss.

BACKGROUNDThe molecular landscape of lung adenocarcinoma (LUAD) is often illustrated as a driver-oncogene pie chart, but identical mutations exhibit heterogeneous signaling shaped by comutations, transcriptional programs, and lineage context. We propose a lineage-integrated signaling framework using an EGFR mutation signature (mSig).METHODSWe defined EGFR mSig using differentially expressed genes in EGFR-mutant (EGFR-mt) LUADs. Semisupervised clustering and machine learning models were used to test reproducibility in different combinations of datasets. We analyzed molecular subtypes, lineage markers, co-occurring mutations, and EGFR copy number alterations in EGFR mSig-defined subtypes of LUAD.RESULTSEGFR mSig showed robust classification performance (area under receiver operating characteristic curve = 0.83-0.95; mean negative predictive value = 96.3%). Validated gene expression subtypes and lung lineage markers were closely aligned with EGFR mSig status. Most EGFR mSig+ tumors, including many without EGFR mutations, belonged to the bronchioid subtype. A subset of canonical RAS mutations were mSig+ and mirrored the EGFR mutation pattern. EGFR WT/mSig- tumors were enriched for nonbronchioid subtypes and had comutations in TP53 or RAS/RAF/RTKs. We highlight a parsimonious collection of coordinated mutations, including RAS, KEAP1, STK11, TP53, and CDKN2A, that taken together suggest coordination of tumor signaling previously suggested but now reproduced and expanded.CONCLUSIONA potentially novel EGFR mSig that captures the transcriptional footprint of EGFR activation revealed a subset of EGFR WT LUADs with mt-like features. mSig refines LUAD taxonomy beyond mutation-only pie-chart models by incorporating lineage and comutation context. Lineage-directed stratification with coalteration identifies clinically relevant groups across EGFR and RAS states and highlights treatment opportunities for patients currently considered oncogene-negative.FUNDINGNational Cancer Institute (NCI) U01CA272541, R01CA262296, U24CA264021, UG1CA233333, R01CA211939.

Humans↗

LINC01871-Mediated Sensitivity to Cyclin-Dependent Kinase 4/6 Inhibitors in Human Breast Cancer.

Breast cancer remains the most frequently diagnosed malignancy in women, and resistance to cyclin-dependent kinase 4 and 6 (CDK4/6) inhibitors limits long-term treatment efficacy. This study aimed to identify long non-coding RNAs (lncRNAs) associated with predicted sensitivity to CDK4/6 inhibitors and to investigate their biological functions in breast cancer. Transcriptomic data from The Cancer Genome Atlas (TCGA) and drug sensitivity data from the Genomics of Drug Sensitivity in Cancer 2 (GDSC2) database were integrated, and drug sensitivity was predicted using the oncoPredict algorithm. Candidate lncRNAs were identified through differential expression analysis, weighted gene co-expression network analysis, prognostic analysis, and machine learning. The biological functions of LINC01871 were subsequently evaluated using in vitro and in vivo experiments. Sixty-two lncRNAs associated with predicted sensitivity to ribociclib and palbociclib were identified, and six core lncRNAs were selected. LINC01871 showed the highest discriminatory performance for predicted drug sensitivity. Overexpression of LINC01871 was associated with increased sensitivity of breast cancer cells to ribociclib and palbociclib, inhibition of cell proliferation, promotion of apoptosis, and suppression of nuclear factor kappa B (NF-&#x3ba;B) signaling. Single-cell transcriptomic analysis demonstrated high LINC01871 expression in T cells and natural killer (NK) cells, while transcriptome-based immune infiltration analyses showed that high LINC01871 expression was associated with increased immune infiltration. These findings identify LINC01871 as a candidate biomarker of sensitivity to CDK4/6 inhibitors and demonstrate its tumor-suppressive effects in breast cancer. Further clinical and mechanistic studies are required to validate its predictive value and therapeutic relevance.

Humans↗

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

Construction of a molecular diagnostic system for neurogenic rosacea by combining transcriptome sequencing and machine learning.

Patients with neurogenic rosacea (NR) frequently demonstrate pronounced neurological manifestations, often unresponsive to conventional therapeutic approaches. A molecular-level understanding and diagnosis of this patient cohort could significantly guide clinical interventions. In this study, we amalgamated our sequencing data (n&#x2009;=&#x2009;46) with a publicly accessible database (n&#x2009;=&#x2009;38) to perform an unsupervised cluster analysis of the integrated dataset. The eighty-four rosacea patients were partitioned into two distinct clusters. Neurovascular biomarkers were found to be elevated in cluster 1 compared to cluster 2. Pathways in cluster 1 were predominantly involved in neurotransmitter synthesis, transmission, and functionality, whereas cluster 2 pathways were centered on inflammation-related processes. Differential gene expression analysis and WGCNA were employed to delineate the characteristic gene sets of the two clusters. Subsequently, a diagnostic model was constructed from the identified gene sets using linear regression methodologies. The model's C index, comprising genes PNPLA3, CUX2, PLIN2, and HMGCR, achieved a remarkable value of 0.9683, with an area under the curve (AUC) for the training cohort's nomogram of 0.9376. Clinical characteristics from our dataset (n&#x2009;=&#x2009;46) were assessed by three seasoned dermatologists, forming the NR validation cohort (NR, n&#x2009;=&#x2009;18; non-neurogenic rosacea, n&#x2009;=&#x2009;28). Upon application of our model to NR diagnosis, the model's AUC value reached 0.9023. Finally, potential therapeutic candidates for both patient groups were predicted via the Connectivity Map. In summation, this study unveiled two clusters with unique molecular phenotypes within rosacea, leading to the development of a precise diagnostic model instrumental in NR diagnosis.

Humans↗

The role of circulating tumor DNA (ctDNA) to detect minimal residual disease in locally advanced gastroesophageal carcinoma: the BUTTERFLY study.

BACKGROUND: Despite advances in perioperative and neoadjuvant strategies, patients with locally advanced gastroesophageal cancers remain at high risk of recurrence after curative intent treatment. No validated biomarkers are available to detect minimal residual disease (MRD) or to guide post-operative risk-adapted management. Circulating tumor DNA (ctDNA) has emerged as a noninvasive tool for disease monitoring; single-parameter or tumor-informed assays, however, may lack sensitivity in low-tumor burden settings. Multimodal, tumor-agnostic approaches may overcome these limitations. METHODS: The BUTTERFLY study is a prospective, multicenter observational study enrolling patients with stage II-III gastric, gastroesophageal junction, or esophageal cancer treated with perioperative chemotherapy or neoadjuvant chemoradiotherapy followed by surgery. It evaluates the diagnostic performance and prognostic value of an academic, tumor-agnostic, multimodal ctDNA assay for MRD detection and prognostic stratification. Serial plasma samples are collected from baseline through post-operative follow-up and at relapse. Cell-free DNA is analyzed using the Agnostic Liquid Biopsy Multimodal Advancement (ALMA) platform, integrating tumor fraction estimation, somatic copy number alterations, fragmentomic features, single-nucleotide variants, and whole-genome methylation profiling. Multimodal features are combined with clinical variables using machine learning-based models to enhance MRD detection and relapse risk stratification. The primary endpoint includes sensitivity and specificity of ALMA-defined ctDNA/MRD status at the 4-8 weeks after surgery landmark, whereas secondary endpoints assess diagnostic performance at other time points and associations between ctDNA status and dynamics with disease-free survival, overall survival, treatment response, and lead time to recurrence. FUTURE PERSPECTIVES: If validated, this tumor-agnostic, multimodal ctDNA approach may enable earlier molecular relapse detection and support personalized post-operative management strategies.

circulating tumor DNA (ctDNA)↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33&#x202f;have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(&#x3b5;)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from &#x2248;32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

Integration of single-cell transcriptomics and genomic mutation analysis identifies an immunotherapy-resistant tumor subcluster and validates ARNTL2 as a malignant driver in lung adenocarcinoma.

BACKGROUND: Immunotherapy resistance in lung adenocarcinoma (LUAD) remains a critical clinical challenge, and the mechanisms underlying resistance-associated intratumoral heterogeneity are poorly characterized. METHODS: We performed single-cell RNA sequencing of LUAD patients receiving neoadjuvant immunotherapy (responders vs. non-responders), integrating inferCNV, GSVA, and differential expression analyses. Cluster-specific genes were validated across seven independent cohorts (TCGA-LUAD, GSE13213, GSE26939, GSE29016, GSE30219, GSE31210, GSE42127). A multi-algorithm machine learning framework was used to construct a prognostic model, and the immune microenvironment was characterized using TCIA scoring, seven infiltration algorithms, and ESTIMATE. ARNTL2 function was assessed by CCK-8 and Transwell assays in A549 and H1299 cells. RESULTS: Non-responders showed significant enrichment of epithelial cells, depletion of cytotoxic T/NK cells, and elevated copy number variation burden versus responders (p < 0.0001). A resistance-enriched malignant subcluster (Cluster 2) exhibited hyperproliferative and metabolic reprogramming signatures with upregulated KRT17, S100A2, and CST6, which showed tumor-specific overexpression, adverse prognostic value, and genomic amplification across cohorts. CoxBoost combined with survivalSVM achieved optimal predictive performance (C-index = 0.686), yielding robust risk stratification (HR: 2.54-10.51, all p < 0.05). Low-risk patients showed greater immune infiltration and higher TCIA immunophenoscores. ARNTL2 was an independent prognostic factor (HR: 2.07-4.64) strongly correlated with risk score (r = 0.69), and its knockdown suppressed proliferation and invasion in both LUAD cell lines (all p < 0.05). CONCLUSION: This study identifies a resistance-associated malignant subcluster in LUAD, constructs a validated CoxBoost + survivalSVM prognostic model with robust immune stratification, and establishes ARNTL2 as a core oncogenic driver and therapeutic target.

ARNTL2↗

Accelerate Your Science: Direct-to-Biology Strategies in Medicinal Chemistry.

Direct-to-biology (D2B) is a powerful strategy that accelerates early drug discovery. It enables compounds to be synthesized in miniaturized formats and evaluated directly as crude reaction mixtures. This bypasses the need for purification during the initial design-make-test cycle. Advances in robust synthetic methodologies, automation, reaction miniaturization, and biological screening have transformed D2B from a proof-of-concept approach into a versatile medicinal chemistry platform. This platform is applicable to fragment optimization, covalent ligands, macrocycles, proteolysis-targeting chimeras (PROTACs), molecular glues, and cellular phenotypic screening. This perspective focuses on the synthetic transformations, assay technologies, and platform implementations that drive modern D2B workflows. It emphasizes reaction robustness, assay compatibility, and practical implementation. Analysis of the current literature revealed that D2B is more governed by reaction reliability than synthetic diversity. Amide coupling and click chemistry dominate reported workflows, while more complex transformations remain underexplored. We discuss the complementary strengths and limitations of biochemical, biophysical, and cellular readouts, identify current bottlenecks in reaction scope and data management, and highlight emerging opportunities arising from reaction miniaturization, machine learning, automated experimentation, and advanced synthetic methodologies. Rather than replacing conventional medicinal chemistry, D2B fundamentally shifts experimental effort from purification toward early biological validation and is poised to become an integral component of future medicinal chemistry workflows.

Humans↗

Bioactive peptides for meat quality and preservation: Integrating peptidomics and computational screening.

Bioactive peptides generated from meat proteins, fermented meat products, and slaughter by-products have attracted increasing attention as functional molecules for improving meat quality and preservation. In meat systems, peptides can be produced through endogenous postmortem proteolysis, microbial fermentation, gastrointestinal digestion, or controlled enzymatic hydrolysis of underutilized animal by-products. These peptides are closely associated with key meat science endpoints, including postmortem tenderization, oxidative stability, color retention, flavor development, microbial inhibition, and the valorization of processing by-products. However, although high-resolution peptidomics has greatly expanded the identification of meat-derived peptide sequences, their translation into practical meat applications remains limited by matrix interactions, processing stability, sensory constraints, safety concerns, and insufficient validation in real meat systems. This review synthesizes recent advances in meat-related peptidomics and computational screening, including sequence-based prediction, machine learning, molecular docking, molecular dynamics, stability assessment, and safety-oriented filtering. Particular attention is given to how these approaches can prioritize peptides with antioxidant, antimicrobial, flavor-modulating, and preservation-related functions under meat-specific technological constraints. By integrating peptide generation pathways, mass spectrometry-based identification, in silico prioritization, and meat quality endpoints, this review proposes a stage-gated framework for translating meat-derived bioactive peptides from discovery to application. Future research should strengthen matrix-specific validation, standardized peptidomic reporting, and safety assessment to support the use of bioactive peptides in meat quality improvement, clean-label preservation, and circular utilization of meat industry by-products.

Animals↗