PubMed HealthSearch

SEARCH · PubMed Health

Results for “Global research representation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The value of international collaborations for supporting neuroanesthesia practice, education, and research in resource-constrained settings.

PURPOSE OF REVIEW: Neuroanesthesia practice in low- and middle-income countries is constrained by workforce shortages, limited infrastructure, and variability in clinical practice. Growing global interest in collaboration makes it timely to evaluate how international partnerships can address these gaps and improve equity in care, education, and research. RECENT FINDINGS: Recent literature highlights substantial variability in neuroanesthesia practice and limited access to context-appropriate guidelines and advanced technologies. International collaborations, including training partnerships, scholarship programs, and research networks, have improved knowledge exchange, workforce development, and the adoption of standardized practices. Evidence suggests that specialized training is associated with improved clinical outcomes. However, persistent inequities in research participation, authorship, and leadership, as well as concerns regarding sustainability and 'parachute research', remain. SUMMARY: International collaboration is a key strategy for advancing neuroanesthesia in resource-constrained settings. Sustainable, equitable partnerships that prioritize local ownership, capacity building, and contextual adaptation are essential to improving clinical practice, strengthening education, and enhancing global research representation.

Humans

What's the meta now? More updates on the problems with systematic reviews.

BACKGROUND: Systematic reviews are intended to provide trustworthy evidence synthesis, yet previous iterations of this living review have identified numerous recurring problems in their conduct and reporting. This article presents the third version and second update of the living systematic review examining issues raised across the academic literature. METHODS: Using consistent eligibility criteria and methods from earlier versions, literature searches were updated to May 2025. Eligible meta-research and editorial articles describing problems with systematic reviews were analyzed to identify emerging themes. Additionally, four basic indicators of methodological quality of the included meta-research were presented across review versions. RESULTS: The update included 209 additional articles. Critically low methodological quality and absence of protocols remained among the most frequently reported issues in systematic reviews across disciplines and journals but notably in evidence underpinning clinical practice guidelines. Spin in abstracts and conflicts of interest continued to be common. Apparent improvements in reporting quality were inconsistent, with modest gains in some full-text reporting but persistent deficiencies in abstracts. Authorship diversity of systematic reviews improved in gender representation but remained geographically concentrated in high-income countries, and primary research included in reviews similarly lacked global representativeness. The issue of misalignment between systematic review evidence bases and global burden of disease bring the total number of problems with systematic reviews to 69. Emerging use of automation and artificial intelligence was variably reported. Descriptive comparison of meta-research articles over the three versions of this living review suggests a greater proportion meeting basic quality indicators in more recent updates. CONCLUSION: Across successive updates, problems with systematic reviews remain widespread and consistent rather than isolated. Incremental reporting improvements coexist with persistent concerns about transparency, bias, and representativeness. Future efforts should prioritize evaluating interventions and aligning research incentives to support genuinely trustworthy evidence synthesis.

Humans

Next-generation phenotyping: introducing phecodeX for enhanced discovery research in medical phenomics.

MOTIVATION: Phecodes are widely used and easily adapted phenotypes based on International Classification of Diseases codes. The current version of phecodes (v1.2) was designed primarily to study common/complex diseases diagnosed in adults; however, there are numerous limitations in the codes and their structure. RESULTS: Here, we present phecodeX, an expanded version of phecodes with a revised structure and 1,761 new codes. PhecodeX adds granularity to phenotypes in key disease domains that are under-represented in the current phecode structure-including infectious disease, pregnancy, congenital anomalies, and neonatology-and is a more robust representation of the medical phenome for global use in discovery research. AVAILABILITY AND IMPLEMENTATION: phecodeX is available at https://github.com/PheWAS/phecodeX.

Phenomics

Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers.

MOTIVATION: Machine-generated or synthetic data is a valuable resource for training artificial intelligence algorithms, evaluating rare workflows, and sharing data under stricter data legislations. However, current statistical and deep learning methods struggle with large data volumes, are prone to hallucinating scenarios incompatible with reality, and seldom quantify privacy meaningfully. RESULTS: Here, we introduce Genomator, a logic solving approach (SAT solving), which efficiently produces private and realistic representations of the original data. We demonstrate the method on genomic data, which arguably is the most complex and private information. We benchmark Genomator against state-of-the-art methodologies (Markov generation, Wasserstein Generative Adversarial Network and Conditional Restricted Boltzmann Machines), demonstrating a 40%-530% accuracy improvement and 57%-172% higher privacy. Genomator is also 3-100 times more efficient, making it the only tested method that scales to whole genomes. We show the universal trade-off between privacy and accuracy, and use Genomator's tuning capability to cater to all applications along the spectrum, from provable private representations of sensitive cohorts, to datasets with indistinguishable pharmacogenomic profiles. Demonstrating the production-scale generation of tuneable synthetic genomes hold great potential for balancing underrepresented populations in medical research and advancing global data exchange. AVAILABILITY AND IMPLEMENTATION: Genomator is available at https://github.com/csiro/genomator.

Algorithms

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S

The landscape of suicide risk factors in Latin America.

Suicide is a leading cause of death globally, with over 75 % of cases occurring in low- and middle-income countries (LMICs). Latin America, a severely underrepresented region in suicide research, has experienced a 6 % increase in its suicide rate from 2010 to 2016, contrasting with a global decline. In this commentary, we provide an overview of the epidemiological, social, and clinical factors influencing suicide in Latin America. We highlight the distinct trends in Latin America compared to other world regions, including the significant disparities in mental health care accessibility. Our review also considers the role of sociocultural factors such as strong family ties, traditional gender norms, and the impact of religiosity on stigma and care-seeking behaviors. Additionally, we address the lack of representation of Latin American populations in genetic studies on suicide, which hinders the identification of relevant biological markers and the development of tailored interventions. We call for increased research on the clinical, genetic, and social determinants of suicide in this region to better understand risk factors and ultimately inform suicide prevention strategies. Addressing this gap is critical for reducing suicide rates in Latin America.

Humans

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans

Triphenyl Phosphate Alters Methyltransferase Expression and Induces Genome-Wide Aberrant DNA Methylation in Zebrafish Larvae.

Emerging environmental contaminants, organophosphate flame retardants (OPFRs), pose significant threats to ecosystems and human health. Despite numerous studies reporting the toxic effects of OPFRs, research on their epigenetic alterations remains limited. In this study, we investigated the effects of exposure to 2-ethylhexyl diphenyl phosphate (EHDPP), tricresyl phosphate (TMPP), and triphenyl phosphate (TPHP) on DNA methylation patterns during zebrafish embryonic development. We assessed general toxicity and morphological changes, measured global DNA methylation and hydroxymethylation levels, and evaluated DNA methyltransferase (DNMT) enzyme activity, as well as mRNA expression of DNMTs and ten-eleven translocation (TET) methylcytosine dioxygenase genes. Additionally, we analyzed genome-wide methylation patterns in zebrafish larvae using reduced-representation bisulfite sequencing. Our morphological assessment revealed no general toxicity, but a statistically significant yet subtle decrease in body length following exposure to TMPP and EHDPP, along with a reduction in head height after TPHP exposure, was observed. Eye diameter and head width were unaffected by any of the OPFRs. There were no significant changes in global DNA methylation levels in any exposure group, and TMPP showed no clear effect on DNMT expression. However, EHDPP significantly decreased only DNMT1 expression, while TPHP exposure reduced the expression of several DNMT orthologues and TETs in zebrafish larvae, leading to genome-wide aberrant DNA methylation. Differential methylation occurred primarily in introns (43%) and intergenic regions (37%), with 9% and 10% occurring in exons and promoter regions, respectively. Pathway enrichment analysis of differentially methylated region-associated genes indicated that TPHP exposure enhanced several biological and molecular functions corresponding to metabolism and neurological development. KEGG enrichment analysis further revealed TPHP-mediated potential effects on several signaling pathways including TGFβ, cytokine, and insulin signaling. This study identifies specific changes in DNA methylation in zebrafish larvae after TPHP exposure and brings novel insights into the epigenetic mode of action of TPHP.

Animals

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral

S-GMAS: Genome-Wide Mediation Analysis With Brain Subcortical Shape Mediators.

Mediation analysis is widely utilized in neuroscience to investigate the role of brain image phenotypes in the neurological pathways from genetic exposures to clinical outcomes. However, it is still difficult to conduct mediation analyses with whole genome-wide exposures and brain subcortical shape mediators due to several challenges including (i) large-scale genetic exposures, that is, millions of single-nucleotide polymorphisms (SNPs); (ii) nonlinear Hilbert space for shape mediators; and (iii) statistical inference on the direct and indirect effects. To tackle these challenges, this paper proposes a genome-wide mediation analysis framework with brain subcortical shape mediators. First, to address the issue caused by the high dimensionality in genetic exposures, a fast genome-wide association analysis is conducted to discover potential genetic variants with significant genetic effects on the clinical outcome. Second, the square-root velocity function representations are extracted from the brain subcortical shapes, which fall in an unconstrained linear Hilbert subspace. Third, to identify the underlying causal pathways from the detected SNPs to the clinical outcome implicitly through the shape mediators, we utilize a shape mediation analysis framework consisting of a shape-on-scalar model and a scalar-on-shape model. Furthermore, the bootstrap resampling approach is adopted to investigate both global and spatial significant mediation effects. Finally, our framework is applied to the corpus callosum shape data from the Alzheimer's Disease Neuroimaging Initiative.

Humans

Selecting candidate Neisseria gonorrhoeae strains for oropharyngeal gonorrhoea human challenge: a genomics-based analysis of clinical isolates.

BACKGROUND: Neisseria gonorrhoeae is a human pathogen of major public health importance due to its increasing global prevalence and antimicrobial resistance (AMR). Evidence suggests that oropharyngeal infection plays a key role in N gonorrhoeae transmission and AMR; however, our understanding of oropharyngeal gonorrhoea pathogenesis is poor. A controlled human infection model (CHIM) for oropharyngeal gonorrhoea will improve understanding of infection and accelerate urgently needed novel gonorrhoea prevention and therapeutic strategies. As the first step in the development of this CHIM, we describe a systematic approach to CHIM strain selection that leverages genomics and clinical data. METHODS: In this genomics-based analysis, we applied a systematic N gonorrhoeae challenge strain selection strategy incorporating genomic and clinical data to a primary dataset of clinical isolates of N gonorrhoeae collected from adult patients in Victoria, Australia, between Jan 1 and Dec 31, 2017, and July 1, 2019, and June 30, 2021. This selection strategy used clinical, phenotypic, and genomic characteristics to define a set of eight criteria that aimed to ensure the contemporary global clinical relevance of the candidate strains; select strains that would be applicable for the assessment of current and future gonorrhoea vaccines; and maximise participant safety by reducing the risk of disseminated gonococcal infection and clinically significant AMR. We applied these criteria to our primary dataset to generate a panel of potential challenge strains. From this final dataset of potential challenge strains, we predetermined that we would select up to ten isolates to proceed to the next stage of detailed phenotypic characterisation for final N gonorrhoeae CHIM strain selection. FINDINGS: 5881 isolates comprised the primary dataset. After application of the selection criteria, most of the isolates (5795 [98·6%] of 5881) were excluded, mostly due to having clinically significant AMR and poor contemporary global clinical relevance. The remaining 86 N gonorrhoeae challenge strain candidates comprised five multilocus sequence types and six N gonorrhoeae multiantigen sequence types, many of which were represented by a single isolate. Of these 86 strains, five isolates were selected to maximise coverage of the phylogenetically distinct groups within the 86 candidate challenge strains and ensure representation of strains collected from various anatomical sites. INTERPRETATION: We transparently describe a novel, systematic, and rational genomics-based strategy for oropharyngeal gonorrhoea CHIM strain selection that improves the efficiency and transparency of CHIM strain selection and enables identification of contemporary and clinically relevant potential challenge strains. A final N gonorrhoeae challenge strain will be selected from the subset of five shortlisted candidates after detailed phenotypic assessment. FUNDING: Medical Research Future Fund, Australian National Health and Medical Research Council and Australian Government Research Training Program.

Humans

Characterization of DPYD pharmacogenetic variation in Mexican patients with gastrointestinal malignancies.

PURPOSE: Fluoropyrimidines are among the most widely used chemotherapeutic agents for gastrointestinal malignancies, but interindividual variability in dihydropyrimidine dehydrogenase (DPD) activity, encoded by DPYD, can lead to severe or lethal toxicities. Most pharmacogenetic data on DPYD originates from European populations, limiting the applicability of current guidelines in admixed groups. METHODS: We evaluated DPYD pharmacogenetic variation and its association with fluoropyrimidine-related adverse events in Mexican patients with gastrointestinal cancers. Adverse events were prospectively assessed using CTCAE v5.0. Genotyping was performed with the Illumina Global Screening Array and analyzed using PLINK and R. RESULTS: A total of 208 patients were enrolled, and 192 samples passed genotyping quality control; 156 patients received fluoropyrimidines. Only three patients (1.5%) carried actionable DPYD variants (rs3918290, rs67376798 and rs75017182), yielding allele frequencies of 0.26%, approximately ten-fold lower than those reported in European cohorts. Genome-wide analyses did not reveal significant genotype-phenotype associations, though suggestive variants in SDK1, ZPBP, and FGF12 were observed. Pharmacodynamic analyses identified frequent variation in TYMS rs2847153 and MTHFR rs1801133, both previously associated with fluoropyrimidine toxicity. Overall, patients exhibited a predominantly Native Mexican ancestry (56.5%), which may explain the markedly low frequency of actionable DPYD alleles commonly found in European populations. CONCLUSIONS: These findings highlight the limited representation of admixed populations in pharmacogenetic research and underscore the need for population-specific data to inform safe and equitable fluoropyrimidine dosing.

Humans

Ten-Year Update of Nurse Practitioner Service Impact on Patient and Health Service Outcomes in Emergency Care Settings-A Systematic Review.

AIMS: To provide a 10-year update on the best available evidence evaluating the impact of nurse practitioner services on cost, waiting times, patient satisfaction, representation rates, and length of stay in emergency and urgent care settings. DESIGN: Systematic review. DATA SOURCES: The search was completed on January 28, 2025, in Embase (Elsevier), Medline (EBSCOhost), CINAHL (EBSCOhost), Cochrane Library (Wiley), Emcare (Ovid), Web of Science Core Collection (Clarivate) and Scopus (Elsevier). The data range (2014-2024) was used to limit the search. METHODS: The search was conducted with results imported into Covidence. In Covidence, two reviewers conducted screening, data extraction, and quality appraisal of articles, and findings were analysed using a narrative synthesis approach. Eligible studies examined nurse practitioner services in emergency or urgent care settings, reporting outcomes of cost, waiting times, patient satisfaction, representation rates, and length of stay. RESULTS: Title and abstract screening were performed on 2329 records. Of these, 236 full-text articles were reviewed, and 17 underwent critical appraisal and data extraction. Narrative analysis of outcome measures yielded mixed results, with both favourable and unfavourable findings reported regarding nurse practitioner services. CONCLUSIONS: Global evaluation of nurse practitioner services in emergency care remains inconsistent. Nevertheless, emerging evidence supports their positive impact, particularly in improving patient outcomes. To effectively inform policy, workforce planning and clinical integration, there is a need for professional benchmarks that provide clear frameworks for the evaluation of patient-centred outcomes and operational impacts in emergency departments. IMPLICATIONS: Evidence related to nurse practitioner services in emergency and urgent care clinics highlights the positive impact of nurse practitioner services on patient wait times and satisfaction; however, there is limited and variable evidence of impact on health care costs and outcomes. IMPACT: This paper recommends that evaluating emergency nurse practitioner services requires homogeneous research using consistent professional benchmarks and evaluation frameworks. REPORTING METHOD: This systematic review follows the Preferred Reporting Items for Systematic Review and Meta-Analysis (PRISMA) guidelines. PATIENT OR PUBLIC CONTRIBUTION: This study did not include patient or public involvement in its design, conduct, or reporting. TRAIL REGISTRATION: PROSPERO 2025 CRD420250645148.

Humans

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA

From transcriptomic profiling to precision oncology: a bibliometric analysis of RNA sequencing in acute myeloid leukemia.

BACKGROUND: RNA sequencing (RNA-seq) has become an important tool for investigating the molecular heterogeneity of acute myeloid leukemia (AML); however, the global development and thematic evolution of this field remain inadequately characterized. OBJECTIVE: To map the global landscape of AML RNA-seq research and identify major knowledge domains, emerging themes, and temporal changes in research priorities. METHODS: Publications indexed in the Web of Science Core Collection and Scopus between January 1, 2007, and August 18, 2025, were retrieved. After database filtering, merging, and deduplication, 3,460 articles and reviews were included. CiteSpace, VOSviewer, the bibliometrix R package, and Microsoft Excel were used to analyze publication trends, collaboration networks, co-citation structures, keyword evolution, and citation bursts. RESULTS: Publication output increased steadily, accelerating after 2014. China contributed the largest number of publications (n = 547, 15.8%), whereas the United States had the highest total citation count. Major publication outlets spanned hematology, oncology, genomics, and molecular biology. Co-citation analysis identified prominent themes involving next-generation sequencing, gene mutations, KMT2A rearrangements, epigenetic dysregulation, leukemia-initiating cells, drug resistance, biomarkers, T-cell biology, and single-cell sequencing. Earlier literature emphasized sequencing technologies, gene expression profiling, and molecular alterations, whereas recent publications show increasing representation of cellular heterogeneity, single-cell transcriptomics, drug resistance, biomarker applications, immune-related research, and computational interpretation. CONCLUSION: While molecular characterization remains foundational, AML RNA-seq research has broadened to encompass increasingly prominent cellular, functional, computational, and translational dimensions. This study provides a structured overview of the field; nevertheless, bibliometric prominence should not be interpreted as direct evidence of clinical utility.

RNA sequencing

Structural modelling and preventive strategy targeting of WSSV hub proteins to combat viral infection in shrimp Penaeus monodon.

White spot syndrome virus (WSSV) presents a considerable peril to the aquaculture sector, leading to notable financial consequences on a global scale. Previous studies have identified hub proteins, including WSSV051 and WSSV517, as essential binding elements in the protein interaction network of WSSV. This work further investigates the functional structures and potential applications of WSSV hub complexes in managing WSSV infection. Using computational methodologies, we have successfully generated comprehensive three-dimensional (3D) representations of hub proteins along with their three mutual binding counterparts, elucidating crucial interaction locations. The results of our study indicate that the WSSV051 hub protein demonstrates higher binding energy than WSSV517. Moreover, a unique motif, denoted as "S-S-x(5)-S-x(2)-P," was discovered among the binding proteins. This pattern perhaps contributes to the detection of partners by the hub proteins of WSSV. An antiviral strategy targeting WSSV hub proteins was demonstrated through the oral administration of dual hub double-stranded RNAs to the black tiger shrimp, Penaeus monodon, followed by a challenge assay. The findings demonstrate a decrease in shrimp mortality and a cessation of WSSV multiplication. In conclusion, our research unveils the structural features and dynamic interactions of hub complexes, shedding light on their significance in the WSSV protein network. This highlights the potential of hub protein-based interventions to mitigate the impact of WSSV infection in aquaculture.

Animals

pLAST-a tool for rapid comparison and classification of bacterial plasmid sequences.

MOTIVATION: The increasing number of fully sequenced bacterial plasmids being annotated and catalogued has prompted the development of computational tools for comparing and classifying them. Existing approaches typically compare full-length DNA sequences (e.g. Mash, BLASTn, and ANI-based methods) or translated open reading frames (ORFs) (e.g. DIAMOND), with plasmid-level scores obtained by aggregating ORF-to-ORF similarities; however, they are either restricted to closely related plasmids or become computationally demanding in large-scale analyses. RESULTS: We describe pLAST (plasmid Language Analysis and Search Tool), a plasmid-search tool built using word2vec representations of protein-family content informed by local genomic context. Benchmarks indicate that pLAST outperforms nucleotide-based methods and performs comparably to DIAMOND in identifying functionally similar plasmids and compared with the widely used Mash, it achieves 26% and 24% improvements in detecting shared mating-pair formation system type and relaxase type, respectively. This performance scales to database searches across hundreds of thousands of sequences, as demonstrated using the precomputed PlasmidScope collection of ∼750 000 plasmids. Beyond global similarity, pLAST also returns per-ORF plasmid-plasmid alignments, enabling detection of shared functional modules. AVAILABILITY AND IMPLEMENTATION: pLAST is freely accessible as a web server at https://plast.lbs.cent.uw.edu.pl/ or https://plast.lbs.biol.uw.edu.pl/ and available as a Python module along with a precomputed database at https://github.com/labstructbioinf/pLAST for customized analysis.

Plasmids

CaLMPhosKAN: prediction of general phosphorylation sites in proteins via fusion of codon aware embeddings with amino acid aware embeddings and wavelet-based Kolmogorov-Arnold network.

MOTIVATION: The mapping from codon to amino acid is surjective due to codon degeneracy, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various protein downstream tasks. However, predictive models for residue-level tasks such as phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites prediction in general, have predominantly relied on representations in amino acid space. RESULTS: We introduce a novel approach for predicting phosphorylation sites by utilizing codon-level information through embeddings from the codon adaptation language model (CaLM), trained on protein-coding DNA sequences. Protein sequences are first reverse-translated into reliable coding sequences by mapping UniProt sequences to their corresponding NCBI reference sequences and extracting the exact coding sequences from their GenBank format using a dynamic programming-based global pairwise alignment. The resulting coding sequences are encoded using the CaLM encoder to generate codon-aware embeddings, which are subsequently integrated with amino acid-aware embeddings obtained from a protein language model, through an early fusion strategy. Next, a window-level representation of the site of interest, retaining the full sequence context, is constructed from the fused embeddings. A ConvBiGRU network extracts feature maps that capture spatiotemporal correlations between proximal residues within the window. This is followed by a prediction head based on a Kolmogorov-Arnold network (KAN) using the derivative of gaussian wavelet transform to generate the inference for the site. The overall model, dubbed CaLMPhosKAN, performs better than the existing approaches across multiple datasets. AVAILABILITY AND IMPLEMENTATION: CaLMPhosKAN is publicly available at https://github.com/KCLabMTU/CaLMPhosKAN.

Codon