PubMed HealthSearch

SEARCH · PubMed Health

Results for “language”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Effectiveness of Caregiver-Mediated Spoken Language Interventions for Children Under Five at Risk of Developmental Language Disorder: A Systematic Review and Meta-Analysis.

BACKGROUND AND AIMS: Caregiver-mediated interventions are commonly used by Speech and Language Therapists to support early language development. Developmental Language Disorder (DLD) is associated with reduced quality of life throughout the lifespan. Understanding factors that predict intervention success is essential for developing appropriate, cost-effective therapy provision for the approximately 12% of preschool children who present with early markers for Developmental Language Disorder (DLD). This systematic review and meta-analysis examined the effectiveness of caregiver-mediated spoken language interventions for under-fives at risk of DLD, and factors influencing intervention effectiveness. METHODS: A systematic review following PRISMA guidelines was conducted. Five electronic databases were searched to identify experimental studies comparing caregiver-mediated spoken language interventions to control conditions in under-fives presenting with risk factors for DLD. Risk factors included prematurity, socioeconomic factors, caregiver language development concerns, and formal or informal language screening or assessment scores. Twenty-six experimental studies with 1407 child participants were included in qualitative synthesis. Meta-analysis was performed on nine Randomised Controlled Trials involving 947 children. RESULTS: Effectiveness was examined for outcomes including child language gains, child wellbeing, inclusion and attainment. Meta-analysis indicated a significant effect of caregiver-mediated spoken language interventions on language outcomes compared to treatment-as-usual, non-language intervention or waitlist control conditions. Non-language outcomes were evaluated via qualitative synthesis. Interventions significantly improved language development trajectories for under-fives presenting with risk factors or early markers for DLD. CONCLUSION AND IMPLICATIONS: This review contributes to the growing evidence base demonstrating that caregiver-mediated interventions can positively impact language development and wellbeing outcomes for children under five at risk of DLD. These findings support the implementation of caregiver-mediated environmental language interventions in clinical practice to maximise accessibility and cost-effectiveness while delivering optimal outcomes for vulnerable populations. WHAT THIS PAPER ADDS: What is already known on this subject Previous research on caregiver-mediated spoken language interventions has highlighted gaps in the evidence regarding the impact of risk factors, demographic characteristics, dosage and intervention components on child language outcomes. Developmental Language Disorder has relatively high population prevalence, estimated at 7%. Prevalence is associated with risk factors including low household socioeconomic status (SES), prematurity and late language emergence. In contrast to its prevalence, there is low public and professional awareness of DLD and a low diagnostic rate. Therefore, a strengthened evidence base and additional insights into the factors affecting success of family-based interventions is important in order to increase the effectiveness of service provision and care planning for this underserved population. Timely and effective intervention with young children presenting with early markers for DLD has the potential to offer lifelong improvement to their wellbeing, inclusion and attainment outcomes. Recent systematic reviews of the effectiveness of caregiver-mediated language interventions had differences in population age range and diagnostic inclusion criteria. What this paper adds to existing knowledge Our review examines the effectiveness of caregiver-mediated early spoken language interventions on child language, attainment and wellbeing, and on caregiver self-efficacy and adherence to language support strategies. Our population was children under five presenting with risk factors for Developmental Language Disorder, in the absence of other neurodevelopmental or genetic conditions such as intellectual disability or autism. This review adds depth and detail to the evidence base supporting the effectiveness of caregiver-mediated spoken language interventions in improving outcomes for this population of young children, and factors that influence their success. What are the potential or actual clinical implications of this work? The high prevalence of Developmental Language Disorder, estimated at around 7% of the population, and the strong association with risk factors including low SES, prematurity and late language emergence, coupled with the low awareness of DLD and low diagnostic rate, mean that a strengthened evidence base and additional insights into the factors affecting success of family-based interventions can increase the effectiveness of service provision and care planning for this population. Timely and effective intervention in this group of young children has the potential to improve wellbeing and attainment outcomes across the lifespan. This review contributes to our understanding of how to implement cost-effective, socially valid and maximally engaging partnership working with families of young children at risk for DLD.

Humans

The impact of tokenizer selection in genomic language models.

MOTIVATION: Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and nonoverlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make subword tokenization in genomic language models significantly different from both traditional language models and protein language models. RESULTS: This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on 44 classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms subword tokenization methods on tasks that rely on nucleotide-level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks. AVAILABILITY AND IMPLEMENTATION: Detailed results of all benchmarking experiments are available in https://github.com/leannmlindsey/DNAtokenization. Training datasets and pretrained models are available at https://huggingface.co/datasets/leannmlindsey. Datasets and processing scripts are available at doi: 10.5281/zenodo.16287401 and doi: 10.5281/zenodo.16287130.

Natural Language Processing

Autism ableism seen through research abstract contents: A mixed-methods analysis of language in NIH-funded genetic and genomic autism research.

In recent years, genetic and genomic autism research has come under increasing scrutiny, moving to the center of debates about ableism, neurodiversity, autism acceptance, and the future of research and care. At the same time, both autism research and genetics and genomics research have, as fields, begun to reckon with the significance of the language researchers use in the course of their work and the harmful ideas that may thereby be reinforced. Although the language of research cannot be assumed to straightforwardly correspond to individual researchers' beliefs, the presence of widespread ableist language may indicate structural and institutionalized ableism, including ableist assumptions at the foundations of research. We conducted a mixed-methods analysis of 166 genetic and genomic autism research projects funded by the US National Institutes of Health, in order to understand the prevalence of potentially ableist discourse, language, and stigmatizing language about autistic people. We found that such discourse and language was ubiquitous across our sample, including a discourse of prevention. This study lends empirical evidence to current debates about language in autism research. Evaluating language can prompt researchers and institutions to reflect on how they conceptualize, design, discuss, and pursue their work.Lay abstractGenetic research about autism is controversial. Researchers are starting to think more carefully about the words they use to talk about autism and the way they do their research. Past research has found that researchers sometimes write about autism in ableist ways. This means that they write about autistic people as though they are less important than nonautistic people. We looked at the way genetics researchers have written about autism in the paperwork for their research. We found that they often write about autistic people in an ableist way. We think that researchers should think carefully about the way they write about autistic people, and how they plan and do their research.

Humans

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides insight into the importance of training data selection relative to model size, the need for domain-specific training, and the capabilities and limitations of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

Prophages

LAMBDA: A Prophage Detection Benchmark for Genomic Language Models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, indicating a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides novel insights into the importance of training data quality relative to model size, the need for domain-specific training, and the application of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

DNA language model

Novel Proactive Speech-Language Intervention Is More Effective Than Usual Care: Randomized Controlled Trial of Babble Boot Camp for Infants With Classic Galactosemia.

PURPOSE: Speech and language disorders cannot be diagnosed and treated until children are approximately 2-4 years old. To investigate whether these disorders can be prevented, we developed and trialed Babble Boot Camp (BBC), the first proactive sustained intervention starting with precursor skills including cooing and babbling. METHOD: Participants were two randomly assigned groups of 22 infants with classic galactosemia, a metabolic disease with known risks for severe speech and language disorders. One group started BBC at under 6 months of age, and the other started at 15 months of age, both completing BBC at 24 months of age. Coached by a speech-language pathologist in weekly telehealth sessions, caregivers implemented BBC activities and routines daily at home. A typical control group and a group of children with classic galactosemia who received usual care participated as well. All children completed standardized assessments of speech and language at postintervention. RESULTS: Assessment scores showed that BBC was more effective than usual care for both intervention groups. Greatest benefits were seen in the group that started at or before 6 months of age, with a proportion of clinically concerning scores equal to that in the typically developing peers. No effects of sex, genotype, or milk consumption were evident in the outcomes. CONCLUSIONS: Findings motivate a paradigm shift from deficit-based to proactive approaches for infants with classic galactosemia. BBC is extensible to many other disorders, with trials currently underway for infants with Down syndrome and infants born preterm.

Humans

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Searching the druggable genome using large language models.

SUMMARY: The druggable genome encompasses the genes that are known or predicted to interact with drugs. The Drug-Gene Interaction Database (DGIdb) provides an integrated resource for discovering and contextualizing these interactions, supporting a broad range of research and clinical applications. DGIdb is currently accessed through structured web interfaces and API calls, requiring users to translate natural-language questions into database-specific query patterns. To allow for the use of DGIdb through natural language, we developed the DGIdb Model Context Protocol (MCP) server, which allows large language models (LLMs) access to up-to-date information through the DGIdb API. We demonstrate that the MCP server improves an LLM's ability to answer questions requiring accurate, up-to-date biomedical knowledge drawn from structured external resources. AVAILABILITY AND IMPLEMENTATION: The DGIdb MCP server is detailed at https://github.com/dgidb/dgidb-mcp-server and includes instructions for accessing the server through the Claude desktop app.

Large Language Models

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

Large Language Model and Knowledge Graph-Driven AJCC Staging of Prostate Cancer Using Pathology Reports.

Background/Objectives: To develop an automated American Joint Committee on Cancer (AJCC) staging system for radical prostatectomy pathology reports using large language model-based information extraction and knowledge graph validation. Methods: Pathology reports from 152 radical prostatectomy patients were used. Five additional parameters (Prostate-specific antigen (PSA) level, metastasis stage (M-stage), extraprostatic extension, seminal vesicle invasion, and perineural invasion) were extracted using GPT-4.1 with zero-shot prompting. A knowledge graph was constructed to model pathological relationships and implement rule-based AJCC staging with consistency validation. Information extraction performance was evaluated using a local open-source large language model (LLM) (Mistral-Small-3.2-24B-Instruct) across 16 parameters. The LLM-extracted information was integrated into the knowledge graph for automated AJCC staging classification and data consistency validation. The developed system was further validated using pathology reports from 88 radical prostatectomy patients in The Cancer Genome Atlas (TCGA) dataset. Results: Information extraction achieved an accuracy of 0.973 and an F1-score of 0.986 on the internal dataset, and 0.938 and 0.968, respectively, on external validation. AJCC staging classification showed macro-averaged F1-scores of 0.930 and 0.833 for the internal and external datasets, respectively. Knowledge graph-based validation detected data inconsistencies in 5 of 150 cases (3.3%). Conclusions: This study demonstrates the feasibility of automated AJCC staging through the integration of large language model information extraction and knowledge graph-based validation. The resulting system enables privacy-protected clinical decision support for cancer staging applications with extensibility to broader oncologic domains.

artificial intelligence

CodonMoE: DNA language models for codon-dependent mRNA prediction.

MOTIVATION: Genomic language models (gLMs) face a fundamental efficiency challenge: one must either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multimodal architectures. Both approaches impose significant computational burdens-modality-specific models require redundant infrastructure despite inherent biological connections, while multi-modal architectures demand increased parameter counts and extensive cross-modality pretraining. RESULTS: To address this limitation, we introduce CodonMoE (Adaptive Mixture of Codon Reformative Experts), a lightweight adapter that transforms DNA language models into effective RNA analyzers without RNA-specific pretraining. Our theoretical analysis establishes CodonMoE as a universal approximator at the codon level, capable of mapping arbitrary functions from codon sequences to codon-dependent RNA properties given sufficient expert capacity. Across four RNA prediction tasks spanning stability, expression, and regulation, DNA models augmented with CodonMoE significantly outperform their unmodified counterparts, with the HyenaDNA+CodonMoE series achieving state-of-the-art results using 80% fewer parameters than specialized RNA models. By maintaining sub-quadratic complexity while achieving superior performance, our approach provides a principled path toward unifying genomic language modeling, leveraging more abundant DNA data and reducing computational overhead while preserving modality-specific performance advantages. AVAILABILITY AND IMPLEMENTATION: Source code for the method and to reproduce the results is available at https://github.com/Kingsford-Group/CodonMoE.

Codon

Searching the Druggable Genome using Large Language Models.

SUMMARY: The druggable genome encompasses the genes that are known or predicted to interact with drugs. The Drug-Gene Interaction Database (DGIdb) provides an integrated resource for discovering and contextualizing these interactions, supporting a broad range of research and clinical applications. DGIdb is currently accessed through structured web interfaces and API calls, requiring users to translate natural-language questions into database-specific query patterns. To allow for the use of DGIdb through natural language, we developed the DGIdb Model Context Protocol (MCP) server, which allows large language models (LLMs) access to up-to-date information through the DGIdb API. We demonstrate that the MCP server greatly enhances an LLM's ability to answer questions requiring accurate, up-to-date biomedical knowledge drawn from structured external resources. AVAILABILITY AND IMPLEMENTATION: The DGIdb MCP server is detailed at https://github.com/griffithlab/dgidb-mcp-server and includes instructions for accessing the server through the Claude desktop app.

Journal Article

Distinct cellular phenotypes of language and executive decline in amyotrophic lateral sclerosis.

Cognitive manifestations, including impairments in language and executive functions, are seen in amyotrophic lateral sclerosis (ALS), but the underlying mechanisms remain unclear. We mapped prefrontal cortex regions from ALS patients by integrating spatial and single-nucleus transcriptomics in a cognitively stratified patient cohort. We uncover that cognitive impairment in ALS is associated with distinct patterns of neuronal dysfunction and glial-vascular dysregulation that vary by region and cognitive subtype. Executive dysfunction is linked to reduced mitochondrial and synaptic activity in deep-layer dorsolateral prefrontal cortex neurons, whereas language-related deficits track with a diffuse pan-regional response involving glial and vascular abnormalities. Our analyses, validated by multiplexed imaging, further identify signatures in the prefrontal cortex that span both motor and cognitive phenotypes, including a multicellular gliosis response. The findings reveal that clinical heterogeneity in ALS is driven by phenotype-specific cellular interactions in motor and non-motor regions of the brain.

Amyotrophic Lateral Sclerosis

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models

BMT4me En Espa&#xf1;ol: Multisite Feasibility and Usability Testing of a Spanish-Language mHealth Adherence Support App for Spanish-Speaking Caregivers of Children After Hematopoietic Stem Cell Transplantation and Cancer Treatment.

BACKGROUND: Medication nonadherence during the first 100 days after pediatric hematopoietic stem cell transplantation (HSCT) and during oncology treatment increases risk for complications. BMT4me is a caregiver-facing mobile health (mHealth) application providing medication reminders, symptom tracking, and note-taking features to support medication management. Spanish-speaking caregivers are frequently excluded from digital adherence interventions due to the lack of language-accessible tools. PROCEDURE: We conducted a multisite, mixed-methods usability testing of a Spanish-language version of BMT4me ("BMT4me en Espa&#xf1;ol") with Spanish-speaking caregivers of children (ages 2-17 years) post-HSCT or with an oncology diagnosis on active treatment. Caregivers completed a facilitated, three-step usability session (unobtrusive observation, interactive observation, and debriefing), followed by a semi-structured interview, and then completed the system usability scale (SUS). Quantitative outcomes were summarized descriptively; qualitative data were analyzed using content analysis with constant comparison. RESULTS: Fifteen participants enrolled at each site for a total of 30 participants. Across both sites, the recruitment rate was 91%. All participants completed all parts of the study. The SUS score (M&#xa0;=&#xa0;80.09; SD&#xa0;=&#xa0;17.35) was above average (>68). Two key qualitative themes emerged: (1) the perceived positive impact of BMT4me on managing a serious illness and (2) the acceptance and sociocultural relevance of BMT4me for Spanish-speaking families. Caregivers also shared suggestions to add educational content and multiuser functionalities to BMT4me. CONCLUSIONS: The acceptance and perceived positive impact of the Spanish BMT4me app indicates that socioculturally relevant, Spanish mHealth interventions have strong potential to support Spanish-speaking caregivers in pediatric oncology and HSCT settings. CLINICAL TRIALS NCT: NCT06361173.

Adolescent

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). &#xa9; 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models