PubMed HealthSearch

SEARCH · PubMed Health

Results for “algorithmic awareness”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Navigating Social Media: Balancing Connectivity With Media Literacy to Combat Misinformation and Protect Mental Well-Being.

BACKGROUND: The pervasive use of social media has created a complex digital ecosystem where high connectivity coexists with significant challenges, including the rapid spread of misinformation, particularly regarding mental health, and documented negative impacts on psychological well-being. Platform architectures designed for engagement maximization have been identified as central factors in both issues. OBJECTIVE: This paper critically analyzes the interconnected relationships between social media use, misinformation dissemination, and mental health impacts, with particular attention to psychiatric misinformation across diagnostic categories (e.g., depression, anxiety, ADHD). A primary objective is to evaluate the potential of advanced critical digital literacy frameworks to serve as protective mechanisms against these dual threats. METHODS: A systematic search was conducted following PRISMA 2020 guidelines across APA PsycInfo, PubMed, JSTOR, and Google Scholar for literature published between January 2018 and March 2026 (updated from the original 2023 search). The search yielded 2672 records. After removing 624 duplicates, 2048 records underwent title and abstract screening, with 1802 excluded. The remaining 246 full-text articles were assessed for eligibility, resulting in 86 studies included in the final qualitative synthesis. Inter-rater reliability was established (Cohen's κ = 0.82). Quality assessment was conducted using the Joanna Briggs Institute Checklist, AXIS, and CASP tools, with findings weighted by methodological quality. A thematic analysis was undertaken to synthesize findings. RESULTS: The analysis reveals that core architectural features of social media platforms, algorithmic curation and engagement-based metrics, simultaneously foster environments ripe for misinformation spread and contribute to psychological distress, including anxiety, depression, and harmful social comparison. Psychiatric misinformation specifically (e.g., inaccurate claims about treatment effectiveness, diagnostic criteria, and medication side effects) represents a growing concern, particularly on image- and video-based platforms. The findings indicate that conventional media literacy approaches focused solely on fact-checking are insufficient. Instead, a critical digital literacy framework encompassing algorithmic awareness, data literacy, and emotional awareness is essential for building user resilience, with evidence from high-quality systematic reviews supporting this approach. CONCLUSIONS: Navigating the complexities of modern social media requires an integrated approach combining "pedagogies of play" for experiential skill development with advocacy for structural change (e.g., algorithmic transparency, well being by design principles). This dual strategy empowers individual users to critically engage with digital content while advocating for ethical platform design, thereby safeguarding both mental well-being and democratic discourse. Implications for educators, mental health professionals (including competencies for addressing patient encounters with psychiatric misinformation), policymakers, and platform designers are discussed.

Humans

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Reliability-aware hierarchical learning for Chagas disease screening from 12-lead ECGs: tackling label uncertainty and class imbalance.

Objective.Chagas disease, a neglected tropical disease (NTD) with significant cardiovascular impact, remains underdiagnosed in resource-limited regions. Electrocardiogram (ECG) screening offers a low-cost tool for detecting cardiac involvement, yet algorithm development is challenged by label noise, data scarcity, and the latent nature of infection. This study proposes a robust ECG-based screening framework that explicitly addresses these constraints.Approach.We introduce aReliability-Aware Hierarchical Learningstrategy that calibrates supervision according to data provenance, prioritizing serology-confirmed labels over noisy self-reports. To mitigate data scarcity, we compare a specialized convolutional neural network (CNN) trained from scratch with a transfer learning approach based on a Spatio-Temporal ECG foundation Model (FM). Performance is evaluated across varying data scales, and the representation structure is analyzed to interpret model behavior.Main results.On the official hidden test set of the George B. Moody PhysioNet/Computing in Cardiology Challenge 2025, our approach achieved a Challenge Score of 0.163. We observe that while the specialized CNN performs competitively in data-rich regimes, the FM exhibits superior robustness in extreme low-resource settings. Furthermore, performance reaches a plateau imposed by underlying disease physiology. Bimodal score distributions suggest that models distinguish established cardiomyopathy from indeterminate infection, which remains electrophysiologically indistinguishable from healthy controls.Significance.These findings clarify both the potential and intrinsic limits of ECG-based AI screening for NTD-associated cardiac involvement. Reliability-aware supervision and data-efficient transfer learning provide a practical framework toward scalable and clinically meaningful ECG screening systems in resource-constrained environments.

Humans

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis.

MOTIVATION: Sexual dimorphism is a fundamental biological determinant driving systematic differences in disease susceptibility, progression, and clinical outcomes. However, current sex-combined AI-based genomic models often exhibit algorithmic bias and fail to capture these sex-specific mechanisms, creating a critical barrier to unbiased precision medicine. Ensuring fairness in the context of sexual dimorphism requires understanding and addressing the distinct biological mechanisms functioning in each sex, rather than focusing solely on equalizing predictive performance. RESULTS: We propose a fairness-aware supervised hierarchical contrastive learning approach, called FairHICON, to discover unbiased sex-common and sex-specific predictive features. Evaluations on cancer and asthma transcriptomic datasets demonstrate that FairHICON significantly outperforms state-of-the-art benchmarks, improving predictive performance by up to 9% while effectively reducing the performance gap between male and female sexes. Furthermore, prognostic validation confirms that the identified sex-specific pathways stratify patient survival significantly better within their corresponding sex groups. This validates FairHICON to elucidate the molecular heterogeneity of sexual dimorphism, advancing inclusive precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data is available at https://github.com/datax-lab/FairHICON.

Sex Characteristics

Psychological consequences of AI-assisted training and the buffering role of mindfulness.

The integration of artificial intelligence (AI) into athletic training is accelerating, yet its psychological implications for athletes remain insufficiently understood. Drawing on the transactional model of stress and the stress-buffering framework of mindfulness, this study examined whether mindfulness training can mitigate adverse psychological responses associated with AI-assisted training. Using a randomized controlled factorial design, 160 collegiate athletes were assigned to AI-assisted training or standard training, with or without concurrent mindfulness intervention, and assessed at baseline, week 4, and week 8. Athletes exposed to AI-assisted training without psychological support exhibited increases in perceived stress and AI dependence over time. In contrast, these stress increases were substantially attenuated when mindfulness training was implemented alongside AI-assisted training. A significant AI × Mindfulness × Time interaction emerged for perceived stress at post-intervention, and difference-in-differences analyses corroborated a robust buffering effect. Mediation analyses further indicated that mindfulness training reduced stress partially through enhancing mindful awareness; a three-wave cross-lagged analysis showed that mindful awareness and stress were reciprocally related over time, with the hypothesized awareness-to-stress pathway remaining robust. Together, these findings suggest that AI-assisted training introduces a distinct form of evaluative pressure, and that mindfulness training may serve as an effective psychological buffer during the adoption of continuous algorithmic performance evaluation systems.

Humans

SpacerScope: binary-vectorized, genome-wide off-target profiling for RNA-guided nucleases without prior candidate-site bias.

The precision of CRISPR/Cas systems is fundamental to their application in plant and animal biotechnology. However, comprehensive sequence-based off-target candidate discovery remains a computational bottleneck, particularly in large and complex genomes. Here we developed SpacerScope, an off-target candidate discovery framework that enables unbiased, genome-wide discovery by leveraging binary vectorization, bitwise filtering, and right-end-anchored alignment. Benchmarking against human CIRCLE-seq data demonstrated that SpacerScope recovered 100% of validated off-target sites (6142/6142), matching the sensitivity of exhaustive algorithms. Crucially, SpacerScope achieved this maximum candidate recovery while substantially reducing computational overhead. In large-genome evaluations, SpacerScope maintained low peak memory usage of 2.20 GiB and achieved substantial runtime improvements over indel-aware comparator tools, including more than 50-fold speedup relative to Cas-OFFinder 3 (544 s versus 29 185 s). Furthermore, comparative analyses in polyploid species, such as the octoploid strawberry, revealed that SpacerScope identified larger sequence-compatible candidate burdens than standard web-based design platforms. Our results establish SpacerScope as a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes. The source code and program was publicly available at https://github.com/charlesqu666/SpacerScope. Short Abstract CRISPR/Cas sequence-based off-target candidate discovery remains computationally challenging in large, repetitive, and polyploid genomes. Existing tools either miss indel-containing candidate sites or incur prohibitive runtime and memory costs. We developed SpacerScope, a binary-vectorized framework that enables unbiased, genome-wide off-target candidate discovery without pre-selected candidate sites. By integrating bitwise filtering with right-end-anchored alignment, SpacerScope recovered 100% of validated off-target sites in human CIRCLE-seq data while using only 2.20 GiB of memory and achieving more than 10-fold speedup over indel-aware alternatives. Evaluation in plant genomes, including rice and octoploid strawberry, further demonstrated SpacerScope's capacity to identify larger sequence-compatible candidate burdens overlooked by standard tools. SpacerScope thus provides a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes, supporting downstream prioritization.

CRISPR-Cas Systems

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products

Toward Class Imbalance and Uncertainty in Powder XRD Analysis: A Dual-Channel Fusion Network for Space Group Classification.

Accurate identification of space groups from powder X-ray diffraction (pXRD) is essential for understanding crystal structures and accelerating materials discovery. However, this task remains highly challenging due to inherent peak overlap, experimental noise, and the complexity of the 230-class classification problem. To address the critical issues of class imbalance and data scarcity, we first design a general physics-informed data augmentation pipeline. We then propose a dual-channel fusion uncertainty-aware network (DFUN) for automated space group classification. The DFUN architecture integrates two complementary feature representations: convolutional features extracted directly from raw diffraction profiles and domain-specific peak descriptors. These distinct representations are adaptively fused through a gating mechanism. Furthermore, to mitigate the inherent long-tailed distribution of crystallographic data, we employ a hybrid loss function that combines Focal Loss with Label Smoothing. Finally, we incorporate Monte Carlo Dropout to provide predictive uncertainty estimation, thereby enabling not only accurate classification but also a crucial assessment of the model's reliability. Evaluated on large-scale simulated data and two public data sets (opXRD and RRUFF), DFUN outperforms the evaluated baseline methods across the reported metrics. The framework also provides uncertainty-aware predictions, establishing DFUN as a robust and interpretable solution for high-throughput automated crystallographic analysis from powder diffraction.

Uncertainty

EPIC: Event Prototyping via Information Constrained graph learning for personalized cancer driver gene prediction.

MOTIVATION: Precision oncology relies on accurately distinguishing patient-specific driver mutations from the vast background of passenger alterations. While graph-based computational methods have emerged as powerful tools for this task, they often struggle to preserve the distinct genomic context of individual mutations within complex biological networks. Consequently, subtle patient-specific driver signals are frequently obscured by dominant topological patterns, critically impeding the identification of individualized oncogenic events essential for personalized cancer therapy. RESULTS: To address this, we propose EPIC, a novel framework for Event Prototyping via Information Constrained Graph Learning. Unlike traditional node-centric approaches, EPIC redefines driver prediction as a metric learning task in an event embedding space. We introduce an information-constrained learning strategy that imposes explicit geometric constraints on feature variance, effectively preventing feature collapse and ensuring that low-frequency driver signals are distinctively preserved. Experiments on large-scale cancer cohorts demonstrate that EPIC significantly outperforms established baselines. Notably, the model prioritizes low-frequency driver variants typically overlooked by population-based methods, mapping them to critical oncogenic mechanisms associated with drug resistance and metastasis. Furthermore, clinical actionability analysis confirms that EPIC substantially expands the patient population eligible for targeted therapies. EPIC provides a robust and context-aware solution for personalized cancer driver discovery, bridging the gap between genomic data and actionable therapeutic insights. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/spcho-dev/EPIC.

Humans

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

Genetically distinct within-host subpopulations of hepatitis C virus persist after Direct-Acting Antiviral treatment failure.

Analysis of viral genetic data has previously revealed distinct within-host population structures in both untreated and interferon-treated chronic hepatitis C virus (HCV) infections. While multiple subpopulations persisted during the infection, each subpopulation was observed only intermittently. However, it was unknown whether similar patterns were also present after Direct-Acting Antiviral (DAA) treatment, where viral populations were often assumed to go through narrow bottlenecks. Here we tested for the maintenance of population structure after DAA treatment failure, and whether there were different evolutionary rates along distinct lineages where they were observed. We analysed whole-genome next-generation sequencing data generated from a randomised study using DAAs (the BOSON study). We focused on samples collected from patients (N=84) who did not achieve sustained virological response (i.e., treatment failure) and had sequenced virus from multiple timepoints. Given the short-read nature of the data, we used a number of methods to identify distinct within-host lineages including tracking concordance in intra-host nucleotide variant (iSNV) frequencies, applying sequenced-based and tree-based clustering algorithms to sliding windows along the genome, and haplotype reconstruction. Distinct viral subpopulations were maintained among a high proportion of individuals post DAA treatment failure. Using maximum likelihood modelling and model comparison, we found an overdispersion of viral evolutionary rates among individuals, and significant differences in evolutionary rates between lineages within individuals. These results suggest the virus is compartmentalised within individuals, with the varying evolutionary rates due to different viral replication rates and/or different selection pressures. We endorse lineage awareness in future analyses of HCV evolution and infections to avoid conflating patterns from distinct lineages, and to recognise the likely existence of unsampled subpopulations.

Humans

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing

CaLMPhosKAN: prediction of general phosphorylation sites in proteins via fusion of codon aware embeddings with amino acid aware embeddings and wavelet-based Kolmogorov-Arnold network.

MOTIVATION: The mapping from codon to amino acid is surjective due to codon degeneracy, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various protein downstream tasks. However, predictive models for residue-level tasks such as phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites prediction in general, have predominantly relied on representations in amino acid space. RESULTS: We introduce a novel approach for predicting phosphorylation sites by utilizing codon-level information through embeddings from the codon adaptation language model (CaLM), trained on protein-coding DNA sequences. Protein sequences are first reverse-translated into reliable coding sequences by mapping UniProt sequences to their corresponding NCBI reference sequences and extracting the exact coding sequences from their GenBank format using a dynamic programming-based global pairwise alignment. The resulting coding sequences are encoded using the CaLM encoder to generate codon-aware embeddings, which are subsequently integrated with amino acid-aware embeddings obtained from a protein language model, through an early fusion strategy. Next, a window-level representation of the site of interest, retaining the full sequence context, is constructed from the fused embeddings. A ConvBiGRU network extracts feature maps that capture spatiotemporal correlations between proximal residues within the window. This is followed by a prediction head based on a Kolmogorov-Arnold network (KAN) using the derivative of gaussian wavelet transform to generate the inference for the site. The overall model, dubbed CaLMPhosKAN, performs better than the existing approaches across multiple datasets. AVAILABILITY AND IMPLEMENTATION: CaLMPhosKAN is publicly available at https://github.com/KCLabMTU/CaLMPhosKAN.

Codon