PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Ribo-ITP enables identification of translons from limited input samples.

In the last decade, an unexpectedly large number of translated regions (translons) have been discovered using ribosome profiling and proteomics. Translons can act as regulatory elements or encode functional micropeptides. However, identification of translons has been limited to cell lines or large organs due to high input requirements for conventional ribosome profiling and mass spectrometry. Here, we address this input limitation using Ribo-ITP on difficult-to-collect samples such as microdissected hippocampal tissues and single preimplantation embryos to identify thousands of translons. To test the translational capacity of the identified translons, we engineer a translon-dependent GFP reporter system and detect expression of translons initiating at ATG and near-cognate start codons in mouse embryonic stem cells (mESCs). We identify distinct expression patterns of translons using a comparative analysis of more than a thousand ribosome profiling datasets across a wide range of cell types. Further, using a machine learning model, we predict that specific upstream translons in synaptically enriched mRNAs regulate translation efficiency of the annotated coding region. Taken together, we present a proof-of-concept study to identify non-canonical translation events from low input samples which can be applied to cell and tissue types inaccessible to conventional methods.

Animals↗

Clinical translation of senescence-related pan-cancer multi-omics: tools for assessment and immunotherapy prediction.

Cellular senescence (CS) exerts dual roles in tumorigenesis, yet its pan-cancer molecular characteristics and clinical value remain unclear, hindering its translation to oncology and personalized therapy. To address the lack of specific and universal tools for senescence assessment and immunotherapy response prediction, this study systematically analyzed 1259 CS-related genes from the CellAge database across 31 cancer types by integrating multi-omics data, including bulk RNA-seq, single-cell/spatial transcriptomics, and CRISPR screening. We developed a rank-based algorithm SenScoreR (publicly available at https://gxhub.shinyapps.io/SenScoreR/ ) for senescence quantification, validated with 10 independent datasets, and constructed a machine learning-based predictive model CS.Sig for immunotherapy response. Results showed that tumors had significantly lower Rank-based Senescence Score (RSS) than normal tissues across 31 cancers (average diagnostic AUC = 0.895), with low RSS linked to poor survival; high RSS correlated with reduced genomic instability, enriched CD8⁺ T/NK cell/macrophage infiltration, upregulated PD-L1 expression, and elevated immune cytolytic activity. CS.Sig demonstrated robust performance in predicting ICI response (AUC = 0.716 across 10 cohorts), outperforming 13 existing signatures, while CRISPR screening identified 17 senescence-related targets (e.g., CEP55, PPP1CC) whose knockout enhanced anti-tumor immunity. Our findings clarify CS's role in maintaining tumor genomic stability and shaping immune microenvironments, and the developed SenScoreR, CS.Sig, and identified targets bridge basic CS research with clinical oncology, providing a translational resource and hypothesis basis for future experimental and clinical validation.

Journal Article↗

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans↗

Support vector machines for prediction of protein subcellular location.

Support Vector Machine (SVM), which is one kind of learning machines, was applied to predict the subcellular location of proteins from their amino acid composition. In this research, the proteins are classified into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracall, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane, and (12) vacuole, which have covered almost all the organelles and subcellular compartments in an animal or plant cell. The examination for the self-consistency and the jackknife test of the SVMs method was tested for the three sets: 2022 proteins, 2161 proteins, and 2319 proteins. As a result, the correct rate of self-consistency and jackknife test reaches 91 and 82% for 2022 proteins, 89 and 75% for 2161 proteins, and 85 and 73% for 2319 proteins, respectively. Furthermore, the predicting rate was tested by the three independent testing datasets containing 2240 proteins, 2513 proteins, and 2591 proteins. The correct prediction rates reach 82, 75, and 73% for 2240 proteins, 2513 proteins, and 2591 proteins, respectively.

Algorithms↗

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals↗

Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis.

PURPOSE: Given the existing uncertainties regarding the link between Di(2-ethylhexyl) phthalate (DEHP) exposure and gastric cancer (GC) progression, this study aimed to clarify their association, identify the toxic targets of DEHP, and elucidate the underlying molecular mechanisms. METHODS: Multiple integrated approaches were employed, including Gene Expression Omnibus (GEO) data analysis, network toxicology, molecular docking, and machine learning. STRING and Cytoscape tools were utilized to identify key targets, while Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the functional enrichment of intersecting targets. Machine learning and SHAP analysis were applied to screen core targets in GC. Molecular docking was performed to evaluate the binding affinity of DEHP toward core targets, and 200 ns molecular dynamics simulations were further conducted for representative complexes to validate their dynamic stability. RESULTS: A total of 18 key targets were identified using STRING and Cytoscape. GO and KEGG enrichment analyses demonstrated that these intersecting targets were primarily enriched in the extracellular region, as well as the Calcium signaling pathway and cAMP signaling pathway. Through machine learning analyses, 7 key genes (ADRB2, ESRRG, GRIA4, IL13RA2, NR3C2, PLA2G1B, and SULT2A1) were identified as core targets in GC through machine learning analyses. Molecular docking simulations revealed strong binding specificity between DEHP and the target proteins. Among them, NR3C2 and ADRB2 exhibited relatively high predictive importance in the machine learning models. DEHP showed favorable binding affinity toward these core targets, and molecular dynamics simulations further confirmed that ADRB2-DEHP and NR3C2-DEHP complexes maintained stable conformations throughout the simulation. CONCLUSIONS: Our findings identified GC associated genes that were computationally predicted as potential targets of DEHP. These results indicated structural compatibility between DEHP and its target proteins but did not prove that DEHP exposure accounts for the gene expression changes in GC.

Molecular Docking Simulation↗

Antimicrobial resistance analysis of Klebsiella pneumoniae bloodstream infections based on a random forest algorithm: a longitudinal study based on data from tertiary hospitals in China from 2012 to 2023.

BACKGROUND: Bloodstream infections (BSIs) caused by Klebsiella pneumoniae pose a significant global health burden, complicated by rising antimicrobial resistance (AMR). This study aimed to characterize resistance patterns, identify predictors of carbapenem resistance, and develop a machine learning model to predict patient outcomes. METHODS: In a retrospective analysis of 109 279 K. pneumoniae BSIs from tertiary hospitals in China (2012-2023), 11&#x2009;000 isolates underwent whole-genome sequencing (WGS) and antimicrobial susceptibility testing. Cox proportional hazards and logistic regression models identified predictors of 30-day mortality and carbapenem-resistant K. pneumoniae (CRKP), respectively. A random forest model predicted AMR trends and outcomes, evaluated by accuracy, precision, recall, and ROC-AUC using R Studio (R Studio, Inc., Boston, MA, USA). RESULTS: Carbapenem resistance occurred in 32.3% of isolates, with rates of 41.9% for third-generation cephalosporins and 41.2% for fluoroquinolones. Among sequenced isolates, ST11 with blaKPC was the dominant CRKP genotype (12.0%). blaKPC (OR 3.97, 95% CI 3.10-5.11) and blaNDM (OR 2.80, 95% CI 2.07-3.71) strongly predicted carbapenem resistance; ICU admission predicted 30-day mortality (HR 2.10, 95% CI 1.80-2.46, p<0.001). Mortality was higher in CRKP (40.2%) vs. susceptible cases (21.5%). The random forest model achieved 89.2% accuracy and 0.92 ROC-AUC, with drug share, age, and CRKP status as top predictors. CONCLUSIONS: CRKP, especially ST11-blaKPC, drives excess mortality. Key predictors highlight the urgency for enhanced AMR surveillance and targeted therapy.

Humans↗

Integration of Genetic Information to Improve Brain Age Gap Estimation Models in the UK Biobank.

Neurodegeneration occurs when the body's central nervous system becomes impaired as a person ages, which can happen at an accelerated pace. Neurodegeneration impairs quality of life, affecting essential functions, including memory and the ability to self-care. Genetics play an important role in neurodegeneration and longevity. Brain age gap estimation (BrainAGE) is a biomarker that quantifies the difference between a machine learning model-predicted biological age of the brain and the true chronological age for healthy subjects; however, a large portion of the variance remains unaccounted for in these models, attributed to individual differences. This study focuses on predicting the BrainAGE more accurately, aided by genetic information associated with neurodegeneration. To achieve this, a BrainAGE model was developed based on MRI measures, and then the associated genes were determined with a Genome-Wide Association Study. Subsequently, genetic information was incorporated into the models. The incorporation of genetic information yielded improvements in the model performances by 7% to 12%, showing that the incorporation of genetic information can notably reduce unexplained variance. This work helps to define new ways of determining persons susceptible to neurological aging decline and reveals genes for targeted precision medicine therapies.

Humans↗

Predicting training outcomes for developmental dyslexia from EEG data.

Developmental dyslexia (DD) is characterised by lower-than-average reading abilities and is diagnosed in approximately 10% of individuals. The societal barriers may limit professional fulfilment and psychological wellbeing of individuals with DD, calling for the development of effective interventions to counteract them. As DD is associated with challenges in both phonological and visuo-attentional domains, different longitudinal training approaches were developed to strengthen them. However, they require a considerable amount of personal, social and economic resources and the outcomes may vary depending on individual differences in behavioural and neurophysiological functionality. Hence, predicting training outcomes might help in developing personalised treatment protocols and optimising the use of resources. In the present work we applied machine learning to resting-state EEG to predict longitudinal training outcomes in adults with DD enrolled in a randomized clinical trial. In particular, one group received a visuo-attentional training combined with transcranial alternating current stimulation (tACS), another group received visuo-attentional training with sham/placebo stimulation, and the third group received a phonological training with sham/placebo stimulation. The improvement in text reading speed was associated with spectral power in low-beta and individual frequencies in the alpha (IAF) and beta (IBF) bands, while the improvement in pseudoword reading was associated with IBF. The findings highlight the potential of capturing neural markers of treatment responsiveness in DD. Future studies should focus on the generalisability of predictive models to real-world settings, while investigating whether specific EEG markers predict responsiveness to distinct remediation protocols, thus supporting the development of personalised interventions.

Humans↗

A comprehensive meta-analysis of tissue resident memory T cells and their roles in shaping immune microenvironment and patient prognosis in non-small cell lung cancer.

Tissue-resident memory T cells (TRM) are a specialized subset of long-lived memory T cells that reside in peripheral tissues. However, the impact of TRM-related immunosurveillance on the tumor-immune microenvironment (TIME) and tumor progression across various non-small-cell lung cancer (NSCLC) patient populations is yet to be elucidated. Our comprehensive analysis of multiple independent single-cell and bulk RNA-seq datasets of patient NSCLC samples generated reliable, unique TRM signatures, through which we inferred the abundance of TRM in NSCLC. We discovered that TRM abundance is consistently positively correlated with CD4+ T helper 1 cells, M1 macrophages, and resting dendritic cells in the TIME. In addition, TRM signatures are strongly associated with immune checkpoint and stimulatory genes and the prognosis of NSCLC patients. A TRM-based machine learning model to predict patient survival was validated and an 18-gene risk score was further developed to effectively stratify patients into low-risk and high-risk categories, wherein patients with high-risk scores had significantly lower overall survival than patients with low-risk. The prognostic value of the risk score was independently validated by the Cancer Genome Atlas Program (TCGA) dataset and multiple independent NSCLC patient datasets. Notably, low-risk NSCLC patients with higher TRM infiltration exhibited enhanced T-cell immunity, nature killer cell activation, and other TIME immune responses related pathways, indicating a more active immune profile benefitting from immunotherapy. However, the TRM signature revealed low TRM abundance and a lack of prognostic association among lung squamous cell carcinoma patients in contrast to adenocarcinoma, indicating that the two NSCLC subtypes are driven by distinct TIMEs. Altogether, this study provides valuable insights into the complex interactions between TRM and TIME and their impact on NSCLC patient prognosis. The development of a simplified 18-gene risk score provides a practical prognostic marker for risk stratification.

Humans↗

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI↗

Plasma Exosome Metabolomics Reveal Stage-Specific Alterations in Elderly Women With Premetabolic and Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) is a chronic disorder that poses a major threat to global health. Exosomes have emerged as promising biomarkers for diagnosing and monitoring chronic diseases. However, stage-specific alterations in the exosomal metabolome during MetS development remain poorly understood. This study aimed to characterize the plasma exosomal metabolome and explore candidate exosomal biomarkers in individuals with MetS. METHODS: This study included 20 patients with MetS, 23 individuals with pre-MetS, and 45 healthy controls. Plasma exosomes were isolated and analyzed using untargeted liquid chromatography-mass spectrometry-based metabolomics. Differential metabolites were defined by a dual-threshold, that is, p&#x2009;<&#x2009;0.05 from t-test and variable importance in projection >&#x2009;1 from partial least squares discriminant analysis, with fold change indicating their expression changes. Further, we employed machine learning algorithms to predict MetS status. RESULTS: We identified 27 differential metabolites between the pre-MetS and control groups, mainly enriched in histidine metabolism and the tricarboxylic acid cycle. Of these, 12 metabolites were upregulated, and 15 were downregulated, with 1-methylhistidine and isocitrate playing central regulatory roles. Comparison between the MetS and control groups revealed 45 differentially expressed metabolites, mainly enriched in thiamine metabolism, including 13 upregulated and 32 downregulated. In the pre-MetS group, cladribine showed the highest area under the curve (AUC) (0.743, p&#x2009;<&#x2009;0.05), whereas 3-methylxanthine yielded the largest AUC (0.714, p&#x2009;<&#x2009;0.05) in the MetS group. CONCLUSION: Our study characterized stage-dependent alterations in the plasma exosome-derived metabolome in MetS and suggests that exosomal metabolomics may provide complementary molecular information on early MetS metabolic perturbations.

exosomal features↗

Engineering Bacillus Subtilis for Efficient Biosynthesis of Riboflavin: Current Knowledge and Future Perspectives.

Riboflavin is an essential water-soluble vitamin that serves as a precursor for the biosynthesis of the flavin cofactors FMN and FAD, which play pivotal roles in numerous redox and energy metabolism reactions. With the growing global demand for sustainable vitamin production, microbial fermentation has become an attractive alternative to chemical synthesis due to its environmental and economic advantages. Among microbial hosts, Bacillus subtilis has emerged as a leading cell factory for riboflavin production owing to its GRAS status, well-characterized genetics, and efficient protein secretion system. This review provides a comprehensive overview of recent advances in metabolic engineering strategies to enhance riboflavin biosynthesis in B. subtilis. Key topics include strengthening biosynthetic and precursor pathways, relieving feedback inhibition, balancing metabolic flux and cell growth, employing adaptive laboratory evolution, and utilizing omics-guided optimization and 13C metabolic flux analysis. Moreover, the integration of synthetic biology tools such as riboswitch engineering, regulatory element design, and high-throughput screening has significantly accelerated strain improvement. Despite remarkable progress, challenges remain in achieving precise regulatory control, optimizing multi-gene expression, and enhancing genome integration efficiency. Future research combining multi-omics data, synthetic regulatory design, and machine learning-driven predictive modeling is expected to further advance the development of intelligent B. subtilis cell factories. However, the practical implementation of these systems remains constrained by the metabolic burden of overproduction and the lack of universal regulatory models that can predict strain performance across varying industrial scales.

Bacillus subtilis↗

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO↗

ERCC2 mutations alter the genomic distribution pattern of somatic mutations and are independently prognostic in bladder cancer.

Excision repair cross-complementation group 2 (ERCC2) encodes the DNA helicase xeroderma pigmentosum group D, which functions in transcription and nucleotide excision repair. Point mutations in ERCC2 are putative drivers in around 10% of bladder cancers (BLCAs) and a potential positive biomarker for cisplatin therapy response. Nevertheless, the prognostic significance directly attributed to ERCC2 mutations and its pathogenic role in genome instability remain poorly understood. We first demonstrated that mutant ERCC2 is an independent predictor of prognosis in BLCA. We then examined its impact on the somatic mutational landscape using a cohort of ERCC2 wild-type (n&#xa0;= 343) and mutant (n&#xa0;= 39) BLCA whole genomes. The genome-wide distribution of somatic mutations is significantly altered in ERCC2 mutants, including T[C>T]N enrichment, altered replication time correlations, and CTCF-cohesin binding site mutation hotspots. We leverage these alterations to develop a machine learning model for predicting pathogenic ERCC2 mutations, which may be useful to inform treatment of patients with BLCA.

Humans↗

Support vector machines for predicting the specificity of GalNAc-transferase.

Support Vector Machines (SVMs) which is one kind of learning machines, was applied to predict the specificity of GalNAc-transferase. The examination for the self-consistency and the jackknife test of the SVMs method were tested for the training dataset (305 oligopeptides), the correct rate of self-consistency and jackknife test reaches 100% and 84.9%, respectively. Furthermore, the prediction of the independent testing dataset (30 oligopeptides) was tested, the rate reaches 76.67%.

Algorithms↗

A genomic catalog of Earth's bacterial and archaeal symbionts.

Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15-23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host-microorganism associations.

Journal Article↗

Proteome Analyst: custom predictions with explanations in a web-based tool for high-throughput proteome annotations.

Proteome Analyst (PA) (http://www.cs.ualberta.ca/~bioinfo/PA/) is a publicly available, high-throughput, web-based system for predicting various properties of each protein in an entire proteome. Using machine-learned classifiers, PA can predict, for example, the GeneQuiz general function and Gene Ontology (GO) molecular function of a protein. In addition, PA is currently the most accurate and most comprehensive system for predicting subcellular localization, the location within a cell where a protein performs its main function. Two other capabilities of PA are notable. First, PA can create a custom classifier to predict a new property, without requiring any programming, based on labeled training data (i.e. a set of examples, each with the correct classification label) provided by a user. PA has been used to create custom classifiers for potassium-ion channel proteins and other general function ontologies. Second, PA provides a sophisticated explanation feature that shows why one prediction is chosen over another. The PA system produces a Naïve Bayes classifier, which is amenable to a graphical and interactive approach to explanations for its predictions; transparent predictions increase the user's confidence in, and understanding of, PA.

Internet↗