PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Comprehensive analysis of diagnostic biomarkers related to histone acetylation in acute myocardial infarction.

BACKGROUND: Acute myocardial infarction (AMI) has become a serious disease that endangers human health, with high morbidity and mortality. Numerous studies have reported histone acetylation can result in the occurrence of cardiovascular diseases. This article aims to explore the potential biomarkers of histone acetylation regulatory genes (ARGs) in AMI patients. METHODS: Five AMI datasets were downloaded from the Gene Expression Omnibus (GEO) database. Next, ARG-related genes were gathered by gene set variation analysis (GSVA) and Spearman's correlation analysis. Subsequently, weighted gene co-expression network analysis (WGCNA) was performed to identify the module genes related to histone acetylation regulation. In the GSE60993 and GSE48060 datasets, the common differentially expressed genes (DEGs) between AMI and control samples were screened. Importantly, the intersecting genes were obtained by overlapping ARGs-related genes, common DEGs, and module genes. Then, the biomarkers in AMI were determined by machine learning, receiver operating characteristic (ROC) curves, and quantitative PCR (qPCR). In addition, immune analysis, drug prediction, molecular docking, and the lncRNA-miRNA-mRNA regulatory network targeting the biomarkers were analyzed, respectively. RESULTS: Here, a total of 18 intersecting genes were identified by overlapping 7,349 ARGs-related genes, 5,565 module genes, and 25 common DEGs. Further, five biomarkers (AQP9, HLA-DQA1, MCEMP1, NKG7, and S100A12) were obtained, and a nomogram was constructed and verified based on these biomarkers. Notably, the biomarkers were significantly associated with CD8 T cells and neutrophils. In addition, the drugs related to biomarkers were predicted, and ATOGEPANT with the molecular target (S100A12) had a high binding affinity (docking score = -10 kcal/mol). CONCLUSION: AQP9, HLA-DQA1, MCEMP1, NKG7, and S100A12 were identified as biomarkers related to ARGs in AMI, which provides a new perspective to study the relationship between ARGs and AMI.

Humans↗

Predicting DNA-binding sites of proteins from amino acid sequence.

BACKGROUND: Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. RESULTS: We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. CONCLUSION: Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.

Algorithms↗

Paternally Expressed Gene 10 Promoter Methylation Level as a Predictor of HBeAg Seroconversion in Chronic Hepatitis B Patients.

The management of chronic hepatitis B (CHB) encounters challenges like suboptimal antiviral response and the lack of predictive biomarkers. In this study, the role of paternally expressed gene 10 (PEG10) in hepatitis B e antigen (HBeAg) seroconversion (HBeAg SC) was explored to identify a therapeutic target and predictive model. In total, 349 participants were recruited, and 141 HBeAg-positive patients were followed up after 48 weeks of antiviral therapy. Key genes were screened by machine learning algorithms (BORUTA, RF and LASSO). PEG10 mRNA, promoter methylation and plasma levels were examined. The effect of PEG10 was assessed by logistic regression, and HBeAg SC was predicted by nomograms. HBeAg-positive patients showed markedly elevated PEG10 mRNA expression (p&#x2009;<&#x2009;0.001), which correlated strongly with major virological markers such as HBV DNA (r&#x2009;=&#x2009;0.520, p&#x2009;<&#x2009;0.001), HBeAg (r&#x2009;=&#x2009;0.490, p&#x2009;<&#x2009;0.001) and HBsAg (r&#x2009;=&#x2009;0.400, p&#x2009;<&#x2009;0.001). In addition, HBeAg-positive patients exhibited a significant reduction in PEG10 promoter methylation levels compared with controls (p&#x2009;<&#x2009;0.001). According to logistic regression analysis, PEG10 promoter methylation status was an independent predictor of HBeAg SC. The predictive nomogram incorporating PEG10 promoter methylation ratio (PMR), albumin (ALB), aspartate aminotransferase (AST) and HBeAg demonstrated excellent clinical predictive value (area under curve (AUC)&#x2009;=&#x2009;0.895,95% confidence interval (CI): 0.808&#x2009;~&#x2009;0.963). The methylation status of the PEG10 promoter represents a promising biomarker for the prediction of HBeAg SC in patients with CHB. CLINICAL TRIAL REGISTRATION: Not applicable.

Humans↗

Predicting gene function in Saccharomyces cerevisiae.

MOTIVATION: S.cerevisiae is one of the most important model organisms, and has has been the focus of over a century of study. In spite of these efforts, 40% of its open reading frames (ORFs) remain classified as having unknown function (MIPS: Munich Information Center for Protein Sequences). We wished to make predictions for the function of these ORFs using data mining, as we have previously successfully done for the genomes of M.tuberculosis and E.coli. Applying this approach to the larger and eukaryotic S.cerevisiae genome involves modifying the machine learning and data mining algorithms, as this is a larger organism with more data available, and a more challenging functional classification. RESULTS: Novel extensions to the machine learning and data mining algorithms have been devised in order to deal with the challenges. Accurate rules have been learned and predictions have been made for many of the ORFs whose function is currently unknown. The rules are informative, agree with known biology and allow for scientific discovery. AVAILABILITY: All predictions are freely available from http://www.genepredictions.org, all datasets used in this study are freely available from http://www.aber.ac.uk/compsci/Research/bio/dss/yeastdataand software for relational data mining is available from http://www.aber.ac.uk/compsci/Research/bio/dss/polyfarm.

Chromosome Mapping↗

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans↗

Transcriptome-based high-frequency recurrence index predicts frequent recurrence in non-muscle-invasive bladder cancer after Bacillus Calmette-Gu&#xe9;rin therapy.

BACKGROUND: High-frequency recurrence (HfR,&#x2009;&#x2265;&#x2009;2 recurrences) in non-muscle-invasive bladder cancer (NMIBC) poses a significant clinical burden. Current risk models, such as the European Organization for Research and Treatment of Cancer (EORTC), the European Association of Urology (EAU), and the UROMOL classification, offer limited predictive accuracy for identifying patients at risk for frequent recurrence despite appropriate treatment. METHODS: A 75-gene high-frequency recurrence index (HfRI) was constructed by selecting recurrence-associated genes using differential expression and Cox regression analyses. The HfRI was computed as a weighted sum of normalized gene expression values. The model was trained on a discovery cohort and validated in multiple cohorts (n&#x2009;=&#x2009;1379) using machine-learning approaches. Clinical relevance was assessed using recurrence-free survival (RFS) and Cox models, and predictive performance was compared with that of the EORTC, EAU, and UROMOL classifications using the area under the curve (AUC) and the concordance index (c-index). RESULTS: The HfRI robustly stratified patients into high-risk and low-risk groups across six independent NMIBC cohorts. Patients classified as HfRI-high had a significantly greater likelihood of experiencing&#x2009;&#x2265;&#x2009;2 recurrences (&#x3c7;2, p&#x2009;=&#x2009;0.001) and showed markedly reduced RFS (log-rank test, p&#x2009;<&#x2009;0.001). The adverse prognostic effect of the HfRI persisted even among patients treated with BCG therapy (log-rank test, p&#x2009;=&#x2009;0.02). Multivariate analysis revealed that the HfRI was an independent predictor of HfR (HR&#x2009;=&#x2009;2.82, 95% CI&#x2009;=&#x2009;1.89-4.20, p&#x2009;<&#x2009;0.001). Compared with established clinical risk classifiers, the HfRI demonstrated superior predictive performance (AUC&#x2009;=&#x2009;0.736, c-index&#x2009;=&#x2009;0.673) in terms of the EORTC (AUC&#x2009;=&#x2009;0.594), EAU (AUC&#x2009;=&#x2009;0.557) risk groups, and UROMOL2021 (AUC&#x2009;=&#x2009;0.596) classification. Pathway analysis revealed that HfRI-high tumors were characterized by upregulation of cell cycle progression and DNA replication pathways, accompanied by suppression of immune signaling pathways. These biological features provide a mechanistic explanation for the reduced responsiveness to intravesical BCG therapy, underscoring the role of HfRI not only as a predictor of recurrence risk but also as a biomarker capable of identifying patients unlikely to benefit from standard BCG treatment. CONCLUSIONS: HfRI represents a robust, transcriptome-based tool for predicting frequent recurrence in NMIBC patients. The HfRI supports earlier identification of patients at risk of high-frequency recurrence, thereby supporting personalized treatment strategies.

Humans↗

Tumour class prediction and discovery by microarray-based DNA methylation analysis.

Aberrant DNA methylation of CpG sites is among the earliest and most frequent alterations in cancer. Several studies suggest that aberrant methylation occurs in a tumour type-specific manner. However, large-scale analysis of candidate genes has so far been hampered by the lack of high throughput assays for methylation detection. We have developed the first microarray-based technique which allows genome-wide assessment of selected CpG dinucleotides as well as quantification of methylation at each site. Several hundred CpG sites were screened in 76 samples from four different human tumour types and corresponding healthy controls. Discriminative CpG dinucleotides were identified for different tissue type distinctions and used to predict the tumour class of as yet unknown samples with high accuracy using machine learning techniques. Some CpG dinucleotides correlate with progression to malignancy, whereas others are methylated in a tissue-specific manner independent of malignancy. Our results demonstrate that genome-wide analysis of methylation patterns combined with supervised and unsupervised machine learning techniques constitute a powerful novel tool to classify human cancers.

Algorithms↗

Therapeutic melanoma vaccines: Platforms, neoantigen strategies, and emerging combination immunotherapies.

Melanoma has emerged as a major focus of cancer immunotherapy research because of its highly immunogenic nature and responsiveness to immune-based treatments. Therapeutic melanoma vaccines are designed to stimulate tumor-specific immune responses through the delivery of Tumor-Associated Antigens (TAAs), Tumor-Specific Antigens (TSAs), and personalized neoantigens. This narrative review provides an overview of current melanoma vaccine strategies, including peptide-based vaccines, dendritic cell vaccines, nucleic acid-based platforms such as mRNA, DNA, and viral vector vaccines. Recent advances in vaccine engineering and tumor genomics have accelerated the development of personalized neoantigen vaccines capable of targeting mutations unique to individual tumors. In parallel, Artificial Intelligence (AI) and Machine Learning (ML) are increasingly being incorporated into neoantigen identification pipelines to improve epitope prediction and optimize vaccine design. Combination strategies involving Immune Checkpoint Inhibitors (ICIs), particularly anti-PD-1 and anti-CTLA-4 therapies, have further enhanced interest in melanoma vaccines by helping overcome tumor-induced immune suppression and augment T-cell activation. In addition to reviewing vaccine mechanisms and emerging technologies, this manuscript examines the evolving clinical trial landscape through analysis of melanoma vaccine studies registered on ClinicalTrials.gov. Although many studies have reported encouraging safety and immunogenicity findings, challenges related to tumor heterogeneity, immune evasion, biomarker selection, and manufacturing complexity continue to limit widespread clinical implementation. Ongoing advances in computational immunology, biomaterial engineering, and precision oncology are expected to further refine melanoma vaccine development and improve therapeutic efficacy. Collectively, these innovations may help establish melanoma vaccines as an increasingly important component of future personalized cancer immunotherapy strategies.

DNA vaccines↗

Genetic mapping and predictive modeling of paralog synthetic lethality.

Paralogs are abundant in the human genome and thought to be a primary source of synthetic lethality, yet the vast paralogome remains largely uncharacterized. A digenic screen of 36,648 paralogous pairs in the human genome revealed that synthetic lethalities were infrequent and varied in penetrance in different tumor backgrounds. We hypothesized that the variable penetrance of synthetic lethalities resulted from complex polygenic interactions with different cellular contexts. A machine learning classifier of a subset of paralog pairs tested across 49 cancer models revealed that endogenous perturbations in related pathways predicted paralog synthetic lethality. Further, predictive modeling of paralog synthetic lethality showed that the strength of synthetic lethal interactions was largely due to the overlap and essentiality of the protein-protein interaction networks shared by the paralog pairs. Collectively, this study tested 36,648 digenic paralog interactions and delineated the key feature classes that underlie the heterogeneity of paralog synthetic lethalities.

Humans↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

Application of machine learning to structural molecular biology.

A technique of machine learning, inductive logic programming implemented in the program GOLEM, has been applied to three problems in structural molecular biology. These problems are: the prediction of protein secondary structure; the identification of rules governing the arrangement of beta-sheets strands in the tertiary folding of proteins; and the modelling of a quantitative structure activity relationship (QSAR) of a series of drugs. For secondary structure prediction and the QSAR, GOLEM yielded predictions comparable with contemporary approaches including neural networks. Rules for beta-strand arrangement are derived and it is planned to contrast their accuracy with those obtained by human inspection. In all three studies GOLEM discovered rules that provided insight into the stereochemistry of the system. We conclude machine learning used together with human intervention will provide a powerful tool to discover patterns in biological sequences and structures.

Amino Acid Sequence↗

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans↗

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models↗

Divergent microbial preludes to necrotising enterocolitis defined by gut phages and bacterial resistomes.

BACKGROUND: Translating microbiome correlations into robust predictive features for complex gut disorders remains elusive, partly due to oversimplified models of pathogenesis and neglect of the virome, a key player in microbial ecosystems. Necrotising enterocolitis (NEC), a devastating disease of preterm infants with no reliable clinical predictors, exemplifies this challenge. OBJECTIVE: To determine the predictive potential of the gut prophageome and polymicrobial aetiologies for NEC. DESIGN: We applied integrated metagenomic and metatranscriptomic analyses and machine learning to 1825 longitudinal stool samples from 43 preterm infants who later developed NEC and 86 gestational age-matched and birthweight-matched controls across three US hospitals. We characterised gut prophageome acquisitions and their association with clinical exposures, including antibiotics, diet and pharmacotherapies. To predict NEC risk, we integrated pre-onset prophageome, antibacterial resistome and bacteriome profiles with neonatal pathology, stratifying the cohort by disease onset timing (early: &#x2264;40 days; late: >40&#x2009;days) for separate analysis. RESULTS: NEC cases exhibited distinct viral diversity trajectories before disease onset. Early-onset NEC was best predicted by phage-bacterial interaction signatures (75% accuracy, 81% sensitivity). Metatranscriptomics revealed increased phage DNA abundance with low gene expression, suggesting a lysogenic lifestyle that may stabilise pathobionts. These phages encode metabolic genes potentially enhancing pathobiont resilience. Late-onset NEC was best predicted by antibacterial resistome profiles (83% accuracy). CONCLUSION: The gut prophageome serves as both a source of pre-symptomatic predictive signals and an active modulator of NEC pathogenesis, with distinct microbial mechanisms driving early-onset and late-onset disease. These polymicrobial etiologies inform strategies for early detection, risk stratification and the development of microbiome-targeted preventive and therapeutic interventions.

BIOMARKERS↗

Detection of antibiotic heteroresistance in clinical microbiology: current and emerging methodologies.

BACKGROUND: Antibiotic heteroresistance (HR) is characterised by the coexistence of susceptible and resistant subpopulations within an apparently isogenic bacterial isolate. Because routine antimicrobial susceptibility testing (AST) primarily assesses the dominant population, HR may escape detection, potentially leading to discrepancies between laboratory susceptibility categorisation and the underlying bacterial population structure. OBJECTIVES: To provide a critical and practice-oriented evaluation of current and emerging methodologies for HR detection and to discuss their strengths, limitations, and potential for clinical implementation. SOURCES: Narrative review based on PubMed searches, complemented by screening of key reference lists and relevant EUCAST and CLSI documents. Peer-reviewed literature was prioritised. CONTENT: Phenotypic approaches, particularly population analysis profiling, remain the reference method for HR definition, but their labour-intensive workflows, long turnaround times, and limited standardisation restrict routine implementation. Alternative strategies, including modified AST assays, metabolic assays, and single-cell platforms, offer gains in speed or throughput but require broader validation. Molecular approaches such as quantitative PCR, droplet digital PCR, targeted deep sequencing, and whole-genome sequencing improve detection of minority resistance determinants. Emerging computational frameworks, including machine learning models integrating phenotypic and genomic data, represent a promising frontier for scalable HR prediction. IMPLICATIONS: Available evidence supports the clinical relevance of HR, although its association with adverse outcomes varies across bacterial species and antibiotic classes. Harmonised methodologies and clinically validated interpretive criteria are needed to support integration of HR assessment into routine diagnostics. Prospective multicentre studies and further standardisation, including engagement with EUCAST and CLSI, will be important to advance clinical implementation.

Antimicrobial resistance↗

Automated interpretation of subcellular patterns from immunofluorescence microscopy.

Immunofluorescence microscopy is widely used to analyze the subcellular locations of proteins, but current approaches rely on visual interpretation of the resulting patterns. To facilitate more rapid, objective, and sensitive analysis, computer programs have been developed that can identify and compare protein subcellular locations from fluorescence microscope images. The basis of these programs is a set of features that numerically describe the characteristics of protein images. Supervised machine learning methods can be used to learn from the features of training images and make predictions of protein location for images not used for training. Using image databases covering all major organelles in HeLa cells, these programs can achieve over 92% accuracy for two-dimensional (2D) images and over 95% for three-dimensional images. Importantly, the programs can discriminate proteins that could not be distinguished by visual examination. In addition, the features can also be used to rigorously compare two sets of images (e.g., images of a protein in the presence and absence of a drug) and to automatically select the most typical image from a set. The programs described provide an important set of tools for those using fluorescence microscopy to study protein location.

Automation↗

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our&#xa0;prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by&#xa0;the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article↗

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping↗