PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

SSB deficiency-induced R-loop accumulation triggers podocyte inflammation in DKD.

INTRODUCTION: Diabetic kidney disease (DKD) is fundamentally a podocytopathy in which sterile inflammation plays a central pathogenic role, yet the upstream triggers that initiate inflammatory cascades in podocytes remain elusive. R-loops are critical regulators of genomic stability, and their pathological accumulation triggers DNA damage and innate immune activation. Whether R-loop dysregulation contributes to podocyte-driven inflammation in DKD is unknown. METHODS: We integrated single-cell transcriptomic profiling, dual machine learning algorithms, and functional experiments to dissect the R-loop regulatory network in the diabetic kidney. RESULTS: Integrated analysis of human diabetic kidney single-cell RNA-seq data revealed a globally compromised R-loop regulatory network selectively within podocytes. Intersection of podocyte-specific transcriptomic shifts with validated R-loop regulators identified 93 candidate genes, from which dual machine learning algorithms pinpointed SSB (Sjögren syndrome antigen B) as the principal podocyte-selective R-loop resolver and a superior diagnostic biomarker (AUC = 0.983). SSB expression was selectively downregulated in diabetic podocytes and showed the strongest positive correlation with the R-loop resolution module. Mechanistically, SSB loss impaired RNA splicing and stability pathways, leading to aberrant R-loop accumulation that activated the cGAS-dependent inflammatory signaling in podocytes. In two murine DKD models and high glucose-challenged podocytes, SSB was markedly reduced. Remarkably, SSB knockdown in podocytes alone sufficed to trigger R-loop accumulation and pro-inflammatory cytokine expression, whereas both RNase H1-mediated R-loop removal and cGAS co-depletion blunted this response. DISCUSSION: These findings suggest that an SSB-governed R-loop -cGAS -inflammatory signaling axis may link genomic instability to podocyte inflammation and contribute to DKD progression, nominating R-loop homeostasis as a previously unrecognized potential therapeutic target.

Podocytes

Identification of NR4A2 as a Potential Predictive Biomarker for Atherosclerosis.

INTRODUCTION/OBJECTIVE: Atherosclerosis, a leading cause of death globally, is characterized by the buildup of immune cells and lipids in medium to large-sized arteries. However, its precise mechanism remains unclear. The purpose of this study is to explore innovative and reliable biomarkers as a viable approach for the identification and management of atherosclerosis. METHODS: The atherosclerosis-related datasets GSE100927 and GSE66360 were retrieved from the Gene Expression Omnibus (GEO) database. The Limma package in the R programming language was utilized, applying the criteria of |logFC| > 1 and P < 0.05. Subsequently, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the 127 identified DEGs using R. Machine learning techniques were then applied to these data to explore and pinpoint potential biomarkers. The diagnostic potential of these markers was assessed via Receiver Operating Characteristic (ROC) curve analysis. Finally, western blot, real-time quantitative PCR (qRT-PCR), and immunohistochemistry (IHC) were employed to confirm the key biomarkers. RESULTS: Our research indicated that a total of 127 DEGs linked to atherosclerosis were successfully identified. Through the application of machine learning methods, eight critical genes were highlighted. Among these, Nuclear Receptor Subfamily 4 Group A Member-2 (NR4A2) emerged as the most promising marker for further investigation. CIBERSORT analysis revealed that NR4A2 expression levels were significantly correlated with multiple immune cell types, including B cells, plasma cells, and macrophages. Additional validation experiments confirmed that NR4A2 expression was indeed elevated in atherosclerotic plaques, supporting its potential as a biomarker for atherosclerosis. CONCLUSION: Our study identified NR4A2 as a potential immune-related biomarker for the diagnosis and treatment of atherosclerosis.

Atherosclerosis

CCT2 defines a highly cisplatin-resistant and poor-prognosis subtype of lung adenocarcinoma.

Cisplatin-based chemotherapy is a standard treatment for lung adenocarcinoma (LUAD), yet acquired cisplatin resistance remains a marked cause of treatment failure. The molecular mechanisms driving cisplatin resistance in LUAD have not been fully elucidated. The present study integrated bulk transcriptomic data, genomic mutation profiles and single-cell RNA sequencing data to systematically investigate cisplatin resistance in LUAD. Resistance-associated genes were identified through differential expression, survival analysis and database integration. Unsupervised clustering was used to define cisplatin resistance-associated subtypes. Functional characteristics were explored using pathway enrichment, immune infiltration, tumor mutation burden and weighted gene co-expression network analysis. A machine learning framework incorporating 101 algorithms was applied to identify key genes and construct a prognostic model. Single-cell analyses and in vitro experiments were performed to validate the biological role of the core gene. Molecular docking and molecular dynamics simulations were conducted to identify potential therapeutic compounds. A total of two molecular subtypes with distinct cisplatin resistance levels and prognostic outcomes were identified. The high-resistance subtype exhibited enhanced cell cycle activity, DNA repair signaling and immune heterogeneity. Machine learning analysis revealed a five-gene signature, with chaperonin-containing TCP1 subunit 2 (CCT2) emerging as a key regulator of cisplatin resistance. Single-cell analyses showed that CCT2 was predominantly enriched in resistant epithelial cell subpopulations. Functional experiments demonstrated that CCT2 knockdown significantly inhibited cell proliferation and enhanced cisplatin sensitivity in LUAD cell lines. A number of candidate compounds targeting CCT2 exhibited stable binding in silico. The present findings identified CCT2 as a key mediator of cisplatin resistance in LUAD and provided potential therapeutic strategies to overcome chemotherapy resistance.

chaperonin-containing TCP-1 subunit 2

In silico prediction method for plant Nucleotide-binding leucine-rich repeat- and pathogen effector interactions.

Plant Nucleotide-binding leucine-rich repeat (NLR) proteins play a crucial role in effector recognition and activation of Effector triggered immunity following pathogen infection. Genome sequencing advancements have led to the identification of a myriad of NLRs in numerous agriculturally important plant species. However, deciphering which NLRs recognize specific pathogen effectors remains challenging. Predicting NLR-effector interactions in silico will provide a more targeted approach for experimental validation, critical for elucidating function, and advancing our understanding of NLR-triggered immunity. In this study, NLR-effector protein complex structures were predicted using AlphaFold2-Multimer for all experimentally validated NLR-effector interactions reported in literature. Binding affinities- and energies were predicted using 97 machine learning models from Area-Affinity. We show that AlphaFold2-Multimer predicted structures have acceptable accuracy and can be used to investigate NLR-effector interactions in silico. Binding affinities for 58 NLR-effector complexes ranged between -8.5 and -10.6 log(K), and binding energies between -11.8 and -14.4&#x2009;kcal/mol-1, depending on the Area-Affinity model used. For 2427 "forced" NLR-effector complexes, these estimates showed larger variability, enabling identification of novel NLR-effector interactions with 99% accuracy using an Ensemble machine learning model. The narrow range of binding energies- and affinities for "true" interactions suggest a specific change in Gibbs free energy, and thus conformational change, is required for NLR activation. This is the first study to provide a method for predicting NLR-effector interactions, applicable to all pathosystems. Finally, the NLR-Effector Interaction Classification (NEIC) resource can streamline research efforts by identifying NLRs important for plant-pathogen resistance, advancing our understanding of plant immunity.

Plant Proteins

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin

Subphenogroups of acute heart failure with preserved ejection fraction: comprehensive proteomics and pathway analysis.

BACKGROUND: Heterogeneity of heart failure with preserved ejection fraction (HFpEF) results in significant challenges for treatment development. Identifying and characterising distinct HFpEF phenogroups may aid in tailoring therapeutic strategies for these patients. The objective of this study was to assess proteomic patterns of HFpEF phenogroups identified through a machine-learning-based clustering model, with the aim of uncovering specific biological pathways associated with each phenogroup. METHODS: This study represents a post-hoc analysis of the ongoing Prospective mUlticenteR obServational stUdy of patIenTs with Heart Failure with preserved Ejection Fraction (PURSUIT-HFpEF) study, which is a multicentre prospective observational study of hospitalised patients with acute decompensated HFpEF. Of the overall cohort (N=1238), this study analysed 198 patients with HFpEF with available proteomics data. These patients were classified into four phenogroups using the machine-learning-based clustering model. The SomaScan assay V.4.1 was used to measure levels of >7000 plasma proteins, and subsequent pathway analysis was conducted to determine the biological differences among the phenogroups. RESULTS: We identified four distinct phenogroups: Phenogroup 1 ('rhythm trouble'), Phenogroup 2 ('ventricular-arterial uncoupling'), Phenogroup 3 ('low output and systemic congestion') and Phenogroup 4 ('systemic failure'). The proteomics revealed distinct protein expression profiles among the phenogroups, with ribonuclease 4, tax1-binding protein 1, regenerating islet-derived protein 3-gamma and alpha-1-antichymotrypsin being the most significant markers to specific identified phenogroups. Pathway analysis suggested differences in immune response, autonomic activation, cellular homeostasis and tissue repair mechanisms across the phenogroups. CONCLUSIONS: Using a comprehensive plasma proteomics approach, our study identified distinct proteomic profiles of HFpEF phenogroups, which in turn suggest specific underlying biological processes. These profiles suggest the involvement of inflammatory activation, tissue injury and regenerative responses, immune modulation and systemic stress signalling as key components of HFpEF pathophysiology. TRIAL REGISTRATION NUMBER: UMIN-CTR ID: UMIN000021831.

Humans

Prematurity and Genetic Liability for Autism Spectrum Disorder.

BACKGROUND: Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by diverse presentations and a strong genetic component. Environmental factors, such as prematurity, have also been linked to increased liability for ASD, though the interaction between genetic predisposition and prematurity remains unclear. This study aims to investigate the impact of genetic liability and preterm birth on ASD conditions. METHODS: We analyzed phenotype and genetic data from two large ASD cohorts, the Simons Foundation Powering Autism Research for Knowledge (SPARK) and Simons Simplex Collection (SSC), encompassing 78,559 individuals for phenotype analysis, 12,519 individuals with genome sequencing data, and 8,104 individuals with exome sequencing data. Statistical significance of differences in clinical measures was evaluated between individuals with different ASD and preterm status. We assessed the rare variants burden using generalized estimating equations (GEE) models and polygenic load using ASD-associated polygenic risk score (PRS). Furthermore, we developed a machine learning model to predict ASD in preterm children using phenotype and genetic features available at birth. RESULTS: Individuals with both preterm birth and ASD exhibit more severe phenotypic outcomes despite similar levels of genetic liability for ASD across the term and preterm groups. Notably, preterm ASD individuals showed an elevated rate of de novo variants identified in exome sequencing (GEE model, p=0.005) in comparison to the non-ASD preterm group. Additionally, a GEE model showed that a higher ASD PRS, preterm birth, and male sex were positively associated with a higher predicted probability for ASD, reaching a probability close to 90% in SPARK. Lastly, we developed a machine learning model using phenotype and genetic features available at birth with limited predictive power (AUROC = 0.65). CONCLUSIONS: Preterm birth may exacerbate the multimorbidity present in ASD, which was not due to the ASD genetic factors. However, increased genetic factors may elevate the likelihood of a preterm child being diagnosed with ASD. Additionally, a polygenic load of ASD-associated variants had an additive role with preterm birth in the predicted probability for ASD, especially for boys. We propose that incorporating genetic assessment into neonatal care could benefit early ASD identification and intervention for preterm infants.

Autism Spectrum Disorder

NRG-P0074 Viral Sample RU1 from Unclassified Mosigvirus Genomic Characterization and Host Range Analysis.

BACKGROUND: Machine learning models for phage-host range prediction and design require comprehensive training data on phage genomes and host ranges to predict phage-host interactions effectively. MATERIALS AND METHODS: This study characterizes phage sample NRG-P0074 viral sample RU1 from unclassified Mosigvirus, originally isolated by the Betty Kutter. The complete genome of NRG-P0074 was sequenced, annotated, and analyzed using various bioinformatic tools. Host range analysis was conducted using the Escherichia coli Reference (ECOR) Library and nine Escherichia coli (E. coli) K12 strains (Keio Knockout Collection) with single nonessential gene deletions. RESULTS: The genome of NRG-P0074 spans 168,357 base pairs with a guanine-cytosine (GC) content of 37.5%. NRG-P0074 exhibited permissiveness in 15.28% of the ECOR isolates and all 9 Keio knockout strains. Comparative genomic analysis revealed that NRG-P0074 is closely related to E. coli phage a20. Its genome is comprised of 270 coding sequences, 153 known genes, 16 terminators, 3 ribosomal-binding sites, 0 tRNAs, and 117 hypothetical proteins. CONCLUSIONS: This research provides valuable data for developing machine learning models to predict phage-host interactions, aiding the development of targeted phage therapies against antibiotic-resistant bacteria.

ECOR Library

Multi-criteria decision making and its application to in silico discovery of vaccine candidates for Toxoplasma gondii.

Vaccine discovery against eukaryotic parasites is not trivial and few exist. Reverse vaccinology is an in silico vaccine discovery approach, designed to identify vaccine candidates from the thousands of protein sequences encoded by a target genome. Previously, we produced the Vacceed bioinformatics pipeline for identification of parasite membrane and excreted/secreted proteins that were likely be exposed to the hosts immune system. More recently, we improved upon machine learning as the final decision-making process to identify parasite proteins that induce a protective response in an animal model. Subsequently, we combined Vacceed with metrics on B and T cell epitope types to produce a new in silico discovery workflow. In this study we extend this in silico workflow to the developability of proteins as vaccines by the incorporation of metrics on the physicochemical properties of proteins. To demonstrate this process, every Toxoplasma gondii protein was ranked in its capacity to provide exposure to the immune system (Vacceed exposure score), presence of epitopes and solubility characteristics by several multicriteria decision making (MCDM) tools (such as TOPSIS, VIKOR and MABAC). A consensus rank was subsequently generated from the results of these tools using a variety of aggregate ranking methods. Levels of uncertainty in the aggregate protein rankings was assessed by conformal interval prediction in association with a machine learning model. Several of the top ranked proteins identified by this approach were novel, uncharacterized membrane transporters or proteins associated with RNA metabolism. In conclusion, MCDM automated the decision making using well known algorithms while conformal prediction intervals varied significantly across the 8000+ proteins of T. gondii. Highly ranked proteins (e.g. the top 100) typically generated low prediction intervals, providing high levels of confidence in their ranks.

Toxoplasma

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans

Systemic Proteome Profiling to Differentiate Primary Glomerular Diseases.

KEY POINTS: Plasma proteome profiling identified distinct signatures across biopsy-proven primary glomerular disease subtypes. An elastic net model using 93 proteins classified primary glomerular disease subtypes and controls, with external validation. Integrating proteomics with machine learning yields biologically interpretable insights in primary glomerular diseases. BACKGROUND: Primary GN is a heterogeneous group of kidney disorders where understanding of their pathophysiology remains incomplete. Despite the diagnostic potential of high-throughput proteomics, constrained proteomic depth and a reliance on binary comparisons have left the feasibility of using systemic signatures to differentiate multiple GN subtypes largely unexplored. METHODS: To identify protein signatures that noninvasively differentiate major primary glomerular disease subtypes and provide mechanistic insights, we performed large-scale systemic proteome profiling of 5416 plasma proteins via Olink Explore HT in a discovery cohort ( n =147) and an external validation cohort ( n =85) of Korean participants (mean age, 41&#xb1;13 years; 46% female). The study population included patients with four GN subtypes-focal segmental glomerulosclerosis, IgA nephropathy, minimal change disease, and membranous nephropathy-alongside healthy controls. We developed a machine learning (ML) model using logistic regression with elastic net regularization to classify disease groups based on proteomic profiles and evaluated its performance in the independent validation cohort. RESULTS: Plasma proteome profiles were distinct among disease subtypes, emerging as a significant source of data variation independent of conventional markers such as eGFR or proteinuria levels. The ML model performed robustly in both the discovery and validation cohorts, achieving an area under the receiver operating characteristic curve >0.8 for differentiating minimal change disease, membranous nephropathy, and IgA nephropathy. The model, even without clinical information, correctly identified 93% of minimal change disease cases (14 of 15) and 63% of IgA nephropathy cases (20 of 32), but its performance was limited for focal segmental glomerulosclerosis, with only 21% of cases (three of 14) correctly classified. Functional analysis of key proteins highlighted distinct biologic pathways, such as hemostasis in minimal change disease. CONCLUSIONS: We identified distinct systemic proteome signatures for primary glomerular diseases, where disease subtype served as a major determinant of proteomic variance alongside conventional clinical markers. ML models demonstrated robust discriminatory performance for minimal change disease, membranous nephropathy, and IgA nephropathy, underscoring the potential for proteome-based classification.

Humans

Computer-derived nuclear "grade" and breast cancer prognosis.

Visual assessments of nuclear grade are subjective yet still prognostically important. Now, computer-based analytical techniques can objectively and accurately measure size, shape and texture features, which constitute nuclear grade. The cell samples used in this study were obtained by fine needle aspiration (FNA) during the diagnosis of 187 consecutive patients with invasive breast cancer. Regions of FNA preparations to be analyzed were digitized and displayed on a computer monitor. Nuclei to be analyzed were roughly outlined by an operator using a mouse. Next, the computer generated a "snake" that precisely enclosed each designated nucleus. Ten nuclear features were then calculated for each nucleus based on these snakes. These results were analyzed statistically and by an inductive machine learning technique that we developed and call "recurrence surface approximation" (RSA). Both the statistical and RSA machine learning analyses demonstrated that computer-derived nuclear features are prognostically more important than are the classic prognostic features, tumor size and lymph node status.

Adult

Potential evaluation of SULT1A3 as an early diagnostic marker for nasopharyngeal carcinoma: a study based on serum proteomics screening and ELISA validation.

BACKGROUND: Nasopharyngeal carcinoma (NPC) represents a highly prevalent and aggressive malignancy endemic to Southeast Asia. Early and accurate diagnosis is critical to improving survival outcomes; however, the absence of robust, stage-specific biomarkers remains a key obstacle to clinical implementation of early screening strategies. METHODS: We performed untargeted serum proteomic profiling using mass spectrometry in 15 treatment-na&#xef;ve early-stage NPC patients and 15 VCA-IgA-positive healthy controls. Bioinformatics analyses were conducted to identify differentially expressed proteins (DEPs). Machine learning (random forest combined with recursive feature elimination) was employed to prioritize candidate biomarkers, which were subsequently verified using enzyme-linked immunosorbent assay (ELISA) in independent sample cohorts. RESULTS: In total, 1,428 serum proteins were identified, among which 1,410 were reliably quantified. We observed 31 upregulated and 189 downregulated proteins in NPC patients relative to controls. Spearman correlation analysis revealed significant associations: LTA4H (leukotriene A4 hydrolase) levels correlated with serum cell infiltration (r&#x2009;=&#x2009;0.383, p&#x2009;=&#x2009;0.032) and CD8&#x2009;+&#x2009;T-cell abundance (r&#x2009;=&#x2009;0.408, p&#x2009;=&#x2009;0.021); both SULT1A3 (sulfotransferase family 1&#xa0;A member 3) and FGL1 (fibrinogen-like protein 1) levels were positively associated with M1 macrophage infiltration (r&#x2009;=&#x2009;0.510, p&#x2009;=&#x2009;0.003 and r&#x2009;=&#x2009;0.430, p&#x2009;=&#x2009;0.015, respectively). In a preliminary validation cohort (n&#x2009;=&#x2009;80), ELISA yielded AUC values of 0.631 (95% CI: 0.515-0.736, p&#x2009;=&#x2009;0.04) for LTA4H, 0.787 (95% CI: 0.681-0.871, p&#x2009;<&#x2009;0.001) for SULT1A3, and 0.688 (95% CI: 0.575-0.787, p&#x2009;=&#x2009;0.002) for FGL1. In large-scale independent validation, SULT1A3 achieved an AUC of 0.826 (95% CI: 0.766-0.876; sensitivity&#x2009;=&#x2009;78.89%, specificity&#x2009;=&#x2009;75.47%) in cohort 1 (n&#x2009;=&#x2009;196) and 0.796 (95% CI: 0.723-0.857; sensitivity&#x2009;=&#x2009;76.67%, specificity&#x2009;=&#x2009;76.67%) in cohort 2 (n&#x2009;=&#x2009;150). CONCLUSIONS: Through an integrated workflow combining proteomic screening, machine learning prioritization, and multi-stage ELISA validation, we identified SULT1A3 as a candidate serum-based biomarker for early detection of NPC. Preliminary findings suggest that SULT1A3 may have potential utility in clinical screening, though further validation in independent, multi&#x2011;center cohorts is required.

Humans

Induction of decision trees and Bayesian classification applied to diagnosis of sport injuries.

Machine learning techniques can be used to extract knowledge from data stored in medical databases. In our application, various machine learning algorithms were used to extract diagnostic knowledge which may be used to support the diagnosis of sport injuries. The applied methods include variants of the Assistant algorithm for top-down induction of decision trees, and variants of the Bayesian classifier. The available dataset was insufficient for reliable diagnosis of all sport injuries considered by the system. Consequently, expert-defined diagnostic rules were added and used as pre-classifiers or as generators of additional training instances for diagnoses for which only few training examples were available. Experimental results show that the classification accuracy and the explanation capability of the naive Bayesian classifier with the fuzzy discretization of numerical attributes were superior to other methods and estimated as the most appropriate for practical use.

Artificial Intelligence

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

Application of genetic search in derivation of matrix models of peptide binding to MHC molecules.

T cells of the vertebrate immune system recognise peptides bound by major histocompatibility complex (MHC) molecules on the surface of host cells. Peptide binding to MHC molecules is necessary for immune recognition, but only a subset of peptides are capable of binding to a particular MHC molecule. Common amino acid patterns (binding motifs) have been observed in sets of peptides that bind to specific MHC molecules. Recently, matrix models for peptide/MHC interaction have been reported. These encode the rules of peptide/ MHC interactions for an individual MHC molecule as a 20 x 9 matrix where the contribution to binding of each amino acid at each position within a 9-mer peptide is quantified. The artificial intelligence techniques of genetic search and machine learning have proved to be very useful in the area of biological sequence analysis. The availability of peptide/MHC binding data can facilitate derivation of binding matrices using machine learning techniques. We performed a simulation study to determine the minimum number of peptide samples required to derive matrices, given the pre-defined accuracy of the matrix model. The matrices were derived using a genetic search. In addition, matrices for peptide binding to the human class I MHC molecules, HLA-B35 and -A24, were derived, validated by independent experimental data and compared to previously-reported matrices. The results indicate that at least 150 peptide samples are required to derive matrices of acceptable accuracy. This result is based on a maximum noise content of 5%, the availability of precise affinity measurements and that acceptable accuracy is determined by an area under the Relative Operating Characteristic curve (Aroc) of > 0.8. More than 600 peptide samples are required to derive matrices of excellent accuracy (Aroc > 0.9). Finally, we derived a human HLA-B27 binding matrix using a genetic search and 404 experimentally-tested peptides, and estimated its accuracy at Aroc > 0.88. The results of this study are expected to be of practical interest to immunologists for efficient identification of peptides as candidates for immunotherapy.

Amino Acid Sequence

Computer-derived nuclear features distinguish malignant from benign breast cytology.

This article describes the use of computer-based analytical techniques to define nuclear size, shape, and texture features. These features are then used to distinguish between benign and malignant breast cytology. The benign and malignant cell samples used in this study were obtained by fine needle aspiration (FNA) from a consecutive series of 569 patients: 212 with cancer and 357 with fibrocystic breast masses. Regions of FNA preparations to be analyzed were converted by a video camera to computer files that were displayed on a computer monitor. Nuclei to be analyzed were roughly outlined by an operator using a mouse. Next, the computer generated a "snake" that precisely enclosed each designated nucleus. The computer calculated 10 features for each nucleus. The ability to correctly classify samples as benign or malignant on the basis of these features was determined by inductive machine learning and logistic regression. Cross-validation was used to test the validity of the predicted diagnosis. The logistic regression cross validated classification accuracy was 96.2% and the inductive machine learning cross-validated classification accuracy was 97.5%. Our computerized system provides a probability that a sample is malignant. Should this probability fall between 30% and 70%, the sample is considered "suspicious," in the same way a visually graded FNA may be termed suspicious. All of the 128 consecutive cases obtained since the introduction of this system were correctly diagnosed, but nine benign aspirates fell into the suspicious category.(ABSTRACT TRUNCATED AT 250 WORDS)

Breast