PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Large Language Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Variation in the application of natural processes: language-dependent constraints in the phonological acquisition of bilingual children.

This paper studies phonological processes and constraints on early phonological and lexical development, as well as the strategies employed by a young Spanish-, Portuguese-, and Hebrew-speaking child-Nurit (the author's niece)-in the construction of her early lexicon. Nurit's linguistic development is compared to that of another Spanish-, Portuguese-, and Hebrew-speaking child-Noam (the author's son). Noam and Nurit's linguistic development is contrasted to that of Berman's (1977) English- and Hebrew-speaking daughter (Shelli). The simultaneous acquisition of similar (closely related languages) such as Spanish and Portuguese versus that of nonrelated languages such as English and Hebrew yields different results: Children acquiring similar languages seem to prefer maintenance as a strategy for the construction of their early lexicon, while children exposed to nonrelated languages appear to prefer reduction to a large extent (Faingold, 1990). The Spanish- and Portuguese-speaking children's high accuracy stems from a wider choice of target words, where the diachronic development of two closely related languages provides a simplified model lexicon to the child.

Child Language↗

Context-free evolutionary grammars and the structural language of nucleic acids.

This paper introduces and investigates a generative mechanism based on some operations inspired by the large-scale mutations in genomes (deletion, inversion, transposition, duplication). Basic questions regarding these devices and their generated languages are investigated: generative capacity, closure properties, decidability. We also briefly discuss a few problems concerning our model with respect to some structural features of the nucleic acids.

Evolution, Molecular↗

Evidence supporting the role of GIGYF2 in synapse development and autism.

Autism spectrum disorder (ASD) is a heterogeneous condition in which genetically defined subtypes offered insights into underlying biological mechanisms and potential targeted treatments. Here, we investigate the clinical and pathogenic significance of GIGYF2 variants in ASD through an integrated approach combining clinical genetics, conditional knockout (cKO) mouse models, neurobiology, and molecular studies. Through targeted sequencing, large-scale genomic data analysis of neurodevelopmental disorder cohorts, and international collaborations, we identified ten affected individuals from eight families harboring de novo or dominantly inherited likely gene-disruptive (LGD) variants and 13 affected individuals from 13 families with de novo missense variants in GIGYF2. Clinical characterization of 16 probands with GIGYF2 variants revealed common features, including ASD, language problems, intellectual disability, and anxiety. In a Gigyf2 cKO mouse model, we observed pronounced autistic-like behaviors, cognitive deficits, and anxiety-like behaviors, mirroring phenotypes observed in affected individuals. Mechanistically, Gigyf2 deficiency disrupted synaptic homeostasis, as evidenced by altered spine density and miniature excitatory postsynaptic currents, and impaired IGF-1R/mTOR signaling, along with dysregulation of synapse-related genes such as Nrp2. Pharmacological inhibition of mTOR with rapamycin or Torin1, as well as Nrp2 knockdown rescued synaptic defects in Gigyf2 KO neurons. These findings define a novel ASD subtype associated with GIGYF2 variants and establish GIGYF2 as a key regulator of synaptic development and function, implicating GIGYF2 dysfunction in ASD pathogenesis and highlighting the IGF-1R/mTOR pathway as a potential therapeutic target for GIGYF2-related ASD subtype.

Journal Article↗

Structural basis of carbohydrate recognition by lectin II from Ulex europaeus, a protein with a promiscuous carbohydrate-binding site.

Protein-carbohydrate interactions are the language of choice for inter- cellular communication. The legume lectins form a large family of homologous proteins that exhibit a wide variety of carbohydrate specificities. The legume lectin family is therefore highly suitable as a model system to study the structural principles of protein-carbohydrate recognition. Until now, structural data are only available for two specificity families: Man/Glc and Gal/GalNAc. No structural data are available for any of the fucose or chitobiose specific lectins. The crystal structure of Ulex europaeus (UEA-II) is the first of a legume lectin belonging to the chitobiose specificity group. The complexes with N-acetylglucosamine, galactose and fucosylgalactose show a promiscuous primary binding site capable of accommodating both N-acetylglucos amine or galactose in the primary binding site. The hydrogen bonding network in these complexes can be considered suboptimal, in agreement with the low affinities of these sugars. In the complexes with chitobiose, lactose and fucosyllactose this suboptimal hydrogen bonding network is compensated by extensive hydrophobic interactions in a Glc/GlcNAc binding subsite. UEA-II thus forms the first example of a legume lectin with a promiscuous binding site and illustrates the importance of hydrophobic interactions in protein-carbohydrate complexes. Together with other known legume lectin crystal structures, it shows how different specificities can be grafted upon a conserved structural framework.

Amino Acid Sequence↗

Methodology for using the UMLS as a background knowledge for the description of surgical procedures.

The Unified Medical Language System (UMLS) contains and organizes a large number of terms from a variety of biomedical terminology systems. This study examines the relevance of the UMLS content and structures to the specific purpose of the conceptual representation of medical procedures. The MAOUSSC modelling is a compositional formalism with a description of elementary procedures in terms of elementary concept entities and combinations of such descriptions into more complex ones. The UMLS knowledge base is expected to provide semantically categorized medical concepts and interconcept relations. A method to reuse the UMLS has been developed. Quantitative and qualitative results are presented. Some difficulties in reusing the UMLS as a background knowledge are related to the preeminence of some terminology sources and to the instanciation of interconcept links. Other ones suggest that purpose-independence in categorization cannot be achieved.

Artificial Intelligence↗

Novel approaches and applications in identifying DNA methylation markers of cardio-kidney-metabolic disease.

Cardio-kidney-metabolic (CKM) diseases represent a major public health challenge, accounting for a large proportion of global burden of morbidity and mortality. These conditions share risk factors, including genetic predisposition, environmental exposures, and lifestyle influences, which collectively drive disease development and progression. Epigenetic modifications, particularly DNA methylation (DNAm), serve as key mediators and biomarkers between these risk factors and disease phenotypes by regulating gene expression without altering the DNA sequence. Epigenome-wide association studies have identified DNAm markers associated with CKM diseases and related phenotypes, highlighting both shared pathways and disease-specific epigenetic signatures in inflammation, metabolic dysfunction, and aging-related processes. Longitudinal studies further demonstrate the dynamic nature of DNAm changes over time, offering insights into disease trajectories. Additionally, methylation risk scores integrating multiple epigenetic markers show promise in improving disease prediction and risk stratification beyond traditional clinical factors. To synthesize the current evidence, we conducted a targeted literature search in PubMed for English-language, peer-reviewed articles published between 2014 and the present. Future research leveraging large, well-phenotyped cohorts, advanced statistical methods, and innovative study designs will be critical for uncovering novel biomarkers, refining risk prediction models, and developing targeted epigenetic therapies to mitigate the global burden.

Humans↗

The development of language-like communication without a language model.

Deaf children who are unable to acquire oral language naturally and who are not exposed to a standard manual language can spontaneously develop a structured sign system that has many of the properties of natural spoken language. This communication system appears to be largely the invention of the child himself rather than of the caretakers.

Child, Preschool↗

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing↗

How knowledge drives understanding--matching medical ontologies with the needs of medical language processing.

In this article, we introduce a knowledge-based approach to medical text understanding. From an in-depth consideration of deep sentence and text understanding we distill basic requirements for an adequate knowledge representation framework. These requirements are then matched with currently available medical ontologies (thesauri, terminologies, etc.). A fundamental trade-off is recognized between large-scale conceptual coverage on the one hand, and formal mechanisms for integrity preservation and conceptual expressiveness on the other hand. We discuss various shortcomings of the most wide-spread ontologies to capture medical knowledge in-the-large. As a result, we argue for the need of a formally sound and expressive model along the lines of KL-ONE-style terminological representation systems in the format of description logics. These provide an adequate methodology for designing more sophisticated, flexible medical ontologies serving the needs of 'deep' knowledge applications which are by no means restricted to medical language processing.

Artificial Intelligence↗

A system that facilitates the orientation within procedure nomenclatures through a semantic approach.

The representation of medical concepts should provide the flexibility required to support several purposes. We have implemented a model in which medical terms are represented in a standard format based on a semantic description of the terms. We have focused on the description of procedures. Underlying this project is the assumption that information about medical procedures is crucial in the healthcare system. A prototype has been developed for urology. Because of the large number of terms in the Unified Medical Language System (UMLS) and the abundance of links between them, we have experimented in the use of the UMLS as the foundation for our concept base. We assess the usefulness of this approach and discuss its improvements.

Algorithms↗

The ERATO Systems Biology Workbench: enabling interaction and exchange between software tools for computational biology.

Researchers in computational biology today make use of a large number of different software packages for modeling, analysis, and data manipulation and visualization. In this paper, we describe the ERATO Systems Biology Workbench (SBW), a software framework that allows these heterogeneous application components--written in diverse programming languages and running on different platforms--to communicate and use each others' data and algorithmic capabilities. Our goal is to create a simple, open-source software infrastructure which is effective, easy to implement and easy to understand. SBW uses a broker-based architecture and enables applications (potentially running on separate, distributed computers) to communicate via a simple network protocol. The interfaces to the system are encapsulated in client-side libraries that we provide for different programming languages. We describe the SBW architecture and the current set of modules, as well as alternative implementation technologies.

Computational Biology↗

[Use of the ACSL simulation language for physiologic toxicokinetic models].

For the description of the processes of absorption, excretion or elimination of chemicals, the open one- or two-compartment models have been used thus far. The latter consist mainly of the fast (central) and slow (peripheral) compartments. The toxicological studies were based on an assumption that the organic processes develop according to is the first order kinetic reaction. However, the absorption, elimination or excretion of toxic chemicals are in fact much more complicated processes that should be explained using, e.g. the physiologically-based toxicokinetic (PBTK) models, covering physiological, biochemical and metabolic parameters, as well as the allometric calibration of selected parameters for interspecies extrapolations, and in vitro/in vivo extrapolations of metabolic parameters. Simulation languages, e.g. ACSL (Advanced Continuous Simulation Language) are indispensable application tools to be operated with PBTK models. They have been developed for modelling systems described by time-dependent non-linear differential equations and/or transfer functions. ACSL with its interfaces (ACSL Builder, ACSL Graphic Modeller, ACSL Math) ensures data input and communication inside the model by the control, transfer and computed parameters. The physiologically-based toxicokinetic models employ a large number of different parameters, which enables, e.g. forecasting the dose/effect or dose/response relationship absorption rate, metabolic pathways, excretion or elimination according to the absorbed dose of xenobiotic; evaluation of risk assessment; extrapolation from high to low doses characteristic of environmental exposure or setting biological exposure limits.

Body Fluid Compartments↗

Knowledge representation of signal transduction pathways.

MOTIVATIONS: Signal transduction is the common term used to define a diverse topic that encompasses a large body of knowledge about the biochemical mechanisms. Since most of the knowledge of signal transduction resides in scientific articles and is represented by texts in natural language or by diagrams, there is the need of a knowledge representation model for signal transduction pathways that can be as readily processed by a computer as it is easily understood by humans. RESULTS: A signal transduction pathway representation model is presented. It is based on a compound graph structure and is designed to handle the diversity and hierarchical structure of pathways. A prototype knowledge base was implemented on a deductive database and a number of biological queries are demonstrated on it.

Amino Acid Motifs↗

Simulators in clinical surgery.

Simulators are no replacement for patients in surgical learning. Live patients are required for teaching clinical signs and skills. Large numbers of students, a relative lack of motivation, a decreasing number of common cases, unwilling patients, differences in language, etc., make clinical teaching in India a bitter problem. Because patient-related problems are important, surgical training using models can help students to gain effective control over surgical signs and skills.

Education, Medical, Undergraduate↗

Trends in computational tools for biomagnetism: from procedural codes to intelligent scientific models.

The nature of the available computing tools strongly influences modern scientific investigations. The sources of well known problems associated with the use of procedural computer languages are traced and their consequences investigated. The likely impact of recent quantitative and qualitative advances in software and hardware is examined with emphasis on its relevance to the biomagnetic inverse problem. Gradual changes in the use of computers, some already employed in a recent study of a specific biomagnetic inverse problem, are outlined which take into account the large investment in conventional codes.

Animals↗

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning↗

Age, consumer direction, and outcomes of supportive services at home.

PURPOSE: Supportive services at home are essential for older people with severe chronic impairments. Newer "consumer-directed" models of organizing home-based services rely heavily on service recipients rather than home care agencies to arrange and direct care at home. This study examined differences in service experience and outcomes between recipients over and under age 65 who direct their own services in one large Medicaid program. DESIGN AND METHODS: A random sample of 1,095 recipients of In-Home Supportive Services in California was selected and interviewed by telephone. Interviews were conducted in English, Spanish, and three Asian languages; those with severe cognitive impairment were excluded from the study. RESULTS: Findings indicate that although younger recipients embrace self-direction more enthusiastically than older ones, age differences are small on a majority of service outcomes. On average, older users embrace this model and manage within it much like younger users. Some differences emerge between the young-old (65-74) and old-old (75+), but these are neither consistent nor determinative. IMPLICATIONS: Old age is far from an inevitable barrier to self-direction. As with other age groups, there are opportunities and obstacles to be addressed as this newer approach to home care is disseminated.

Activities of Daily Living↗

Modeling global and focal hyperarticulation during human-computer error resolution.

When resolving errors with interactive systems, people sometimes hyperarticulate--or adopt a clarified style of speech that has been associated with increased recognition errors. The primary goals of the present study were: (1) to provide a comprehensive analysis of acoustic, prosodic, and phonological adaptations to speech during human-computer error resolution after different types of recognition error; and (2) to examine changes in speech during both global and focal utterance repairs. A semi-automatic simulation method with a novel error-generation capability was used to compare speech immediately before and after system recognition errors. Matched original-repeat utterance pairs then were analyzed for type and magnitude of linguistic adaption during global and focal repairs. Results indicated that the primary hyperarticulate changes in speech following all error types were durational, with increases in number and length of pauses most noteworthy. Speech also was adapted toward a more deliberate and hyperclear articulatory style. During focal error repairs, large durational effects functioned together with pitch and amplitude to provide selective prominence marking of the repair region. These results corroborate and generalize the computer-elicited hyperarticulate adaptation model (CHAM). Implications are discussed for improved error handling in next-generation spoken language and multimodal systems.

Computers↗