PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Radiogenomic MRI biomarkers for noninvasive prediction of GPC3 expression and tumor microenvironment in hepatocellular carcinoma.

BACKGROUND: Glypican-3 (GPC3) is frequently overexpressed in hepatocellular carcinoma (HCC) and plays a key role in immune and metabolic remodeling of the tumor microenvironment. Reliable noninvasive biomarkers for predicting GPC3 status could improve patient stratification and support precision immunotherapy. METHODS: This multicenter retrospective study included 274 patients with pathologically confirmed hepatocellular carcinoma from three institutions, 34 external cases with MRI from The Cancer Imaging Archive, and 363 transcriptomic profiles from The Cancer Genome Atlas. Contrast-enhanced T1-weighted imaging and diffusion-weighted imaging were analyzed. Tumor and peritumoral regions were segmented manually and radiomic features extracted using PyRadiomics. Feature selection was performed with correlation filtering and least absolute shrinkage and selection operator regression. Machine learning classifiers including logistic regression, random forest, support vector machine, k-nearest neighbor, and decision tree were trained with 10-fold cross-validation and tested on independent external cohorts. A radiomics score was calculated for each patient. Radiogenomic analysis correlated radiomics scores with transcriptomic data using weighted gene co-expression network analysis. Hub genes and enriched pathways were identified, and immune infiltration and predicted immunotherapy response were assessed using computational methods. RESULTS: The random forest model using contrast-enhanced T1-weighted imaging achieved an area under the curve of 0.966 in training and 0.935 in internal validation. The integrated contrast-enhanced T1-weighted imaging plus diffusion-weighted imaging model reached an internal validation area under the curve of 0.979. In external testing, the best performance was obtained with a support vector machine model (area under the curve 0.756). Radiomics scores were significantly correlated with GPC3 expression (R&#x2009;=&#x2009;0.78, p&#x2009;<&#x2009;0.05). Transcriptomic analysis identified a 10-gene signature enriched in hypoxia and lipid metabolism pathways that stratified patients into prognostic subgroups (concordance index 0.720, hazard ratio 4.07, p&#x2009;<&#x2009;0.0001). High-risk patients had greater immune infiltration and a lower predicted immune evasion score, suggesting a potential benefit from immunotherapy. CONCLUSIONS: MRI-based radiomics models can noninvasively predict GPC3 expression in hepatocellular carcinoma. Radiomics scores reflect underlying hypoxia and lipid metabolism pathways and stratify patients by prognosis and predicted immunotherapy response. These findings support radiogenomics as a translational approach to imaging-guided precision treatment in hepatocellular carcinoma.

Humans↗

COMPETE: a model for vocational evaluation, training, employment, and community for integration for persons with cognitive impairments.

This paper describes a job skills training program for young adults with cognitive impairments called COMPETE, an acronym for Computer Preparation: Evaluation, Training, and Employment. Participants in COMPETE trained in use of information-age machines, (e.g., computers, fax machines). The 3-year demonstration project resulted in successful placement of 17 of 27 persons who were unemployed before entering the program.

Activities of Daily Living↗

Proteomic signature of human cancer cells.

We assessed proteomic profiles as biomarkers for monitoring cell phenotypes. Protein expression profiles were obtained by fluorescence two-dimensional difference gel electrophoresis (2-D-DIGE), in which quantitative ability is improved by labeling proteins with fluorescent dyes prior to electrophoresis. Integrated protein spot intensities were analyzed by a statistical approach. The proteomic data of two groups of cell lines: (1) adenocarcinoma (AC) cell lines derived from lung, pancreas and colon tissues and (2) lung cancer cell lines with different histological backgrounds, including AC, squamous cell carcinoma and small cell carcinoma, were assessed on the basis of prior biological information. Hierarchical clustering analysis and principal component analysis were used to divide the cell lines into subgroups on the basis of similarities between their protein expression profiles. The majority of cell lines were grouped according to their organ of origin or histological background. A machine-learning algorithm selected 32 protein spots that were responsible for the classification. The results indicate that proteomic data generated by 2-D-DIGE can provide a signature of essential cell phenotypes, suggesting that it might be possible to apply this technique to developing tumor markers that could identify the organ of origin of metastatic tumors and contribute to the differential diagnosis of lung cancer.

Cell Line, Tumor↗

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92&#x2009;million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study&#xa0;(GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that&#xa0;guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans↗

Learning yeast gene functions from heterogeneous sources of data using hybrid weighted Bayesian networks.

We developed a machine learning system for determining gene functions from heterogeneous sources of data sets using a Weighted Naive Bayesian Network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or ORFs (Open Reading Frames) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore many functional links would be missing when only one or two source of data is used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

Identification and analysis of key genes related to efferocytosis in colorectal cancer.

UNLABELLED: The impact of efferocytosis-related genes (ERGs) on the diagnosis of colorectal cancer (CRC) remains unclear. In this study, efferocytosis-associated biomarkers for the diagnosis of CRC were identified by integrating data from transcriptome sequencing and public databases. Finally, the expression of biomarkers was validated by real-time quantitative polymerase chain reaction (RT-qPCR). Our study may provide a reference for CRC diagnosis. BACKGROUND: It has been shown that some efferocytosis related genes (ERGs) are associated with the development of cancer. However, it is still uncertain how ERGs may influence the diagnosis of colorectal cancer (CRC). METHODS: In our study, the CRC cohorts were gained from transcriptome sequencing and the gene expression omnibus (GEO) database (GSE71187). Efferocytosis related biomarkers with diagnostic utility for CRC were identified through combining differentially expressed analysis, machine learning algorithms, and receiver operating characteristic (ROC) analysis. Then, infiltration abundance of immune cells between CRC and control was evaluated. The regulatory networks (including mRNA-miRNA-lncRNA and miRNA/transcription factors (TF)-mRNA networks) were created. Finally, the expression of biomarkers was validated via real-time quantitative polymerase chain reaction (RT-qPCR). RESULTS: There were 3 biomarkers (ELMO3, P2RY12, and PDK4) related diagnosis for CRC patients gained. ELMO3 was highly expressed in CRC group, while P2RY12 and PDK4 was lowly expressed. Besides, the infiltrating abundance of 3 immune cells between CRC and control groups was significantly differential, namely activated CD4 memory T cells, macrophages M0, and resting mast cells. We then constructed a mRNA-miRNA-lncRNA network containing 3 mRNAs, 33 miRNAs, and 22 lncRNAs, and a miRNA/TF-mRNA network including 3 mRNAs, 33 miRNAs, and 7 TFs. Additionally, RT-qPCR results revealed that the expression trends of all biomarkers were consistent with the transcriptome sequencing data and GSE71187. CONCLUSION: Taken together, this study provides three efferocytosis related biomarkers (ELMO3, P2RY12, and PDK4) for diagnosis of CRC, providing a scientific reference for further studies of CRC.

Humans↗

Methylation profiling in CNS tumor diagnostics: a single-centre real-world experience from Central Europe.

Genome-wide DNA methylation profiling has transformed neuro-oncology by providing an objective, machine learning-based taxonomy that mitigates interobserver variability and refines the histo-molecular criteria of the current WHO classification. We evaluate the real-world diagnostic performance and clinical utility of this modality in a prospective, consecutively accrued three-year cohort of 291 central nervous system (CNS) tumors across a mixed adult-pediatric population. Successful profiling was completed in 95.9% of cases. Using the Epignostix classifier, a high-confidence diagnostic match (calibrated score [CS]&#x2009;&#x2265;&#x2009;0.84) was achieved in 70.3% of analyzable samples, while 26.5% returned lower-confidence scores (&#x2265;&#x2009;0.3 to <&#x2009;0.84) and only 3.2% remained completely unclassifiable (CS&#x2009;<&#x2009;0.3). When integrated into a comprehensive diagnostic framework, methylation profiling provided clinically useful results in 81.1% of cases, establishing diagnoses in 70 cases submitted for molecular subclassification and resolving diagnostic uncertainty or prompting major revisions in 149 histologically challenging tumors. Within truly ambiguous lesions, integration of methylome data dictated tumor grade modifications in 38.8% of cases (upgrading in 29.4% and downgrading in 9.4%), shifting patient risk stratification. Crucially, over half (52.7%) of the lower-confidence cases yielded meaningful clinical integration when supported by histomorphology and ancillary genetic or immunohistochemical markers, demonstrating that rigid score cutoffs should not dictate assay failure. Discrepant or misleading classifications occurred in 1.9%. Updating bioinformatic pipelines from version 11b4 to 12.8 rescued multiple ambiguous entries, increasing overall clinical utility to 84.1%. These findings demonstrate that integrating computational epigenomics with classical neuropathology enhances diagnostic precision, while highlighting the ongoing need for careful clinical-pathological correlation.

Central nervous system tumors↗

Multi-sensor integration for on-line tool wear estimation through radial basis function networks and fuzzy neural network.

On-line tool wear estimation plays a very critical role in industry automation for higher productivity and product quality. In addition, appropriate and timely decision for tool change is significantly required in the machining systems. Thus, this paper is dedicated to develop an estimation system through integration of two promising technologies, artificial neural networks (ANN) and fuzzy logic. An on-line estimation system consisting of five components: (1) data collection; (2) feature extraction; (3) pattern recognition; (4) multi-sensor integration; and (5) tool/work distance compensation for tool flank wear, is proposed herein. For each sensor, a radial basis function (RBF) network is employed to recognize the extracted features. Thereafter, the decisions from multiple sensors are integrated through a proposed fuzzy neural network (FNN) model. Such a model is self-organizing and self-adjusting, and is able to learn from the experience. Physical experiments for the metal cutting process are implemented to evaluate the proposed system. The results show that the proposed system can significantly increase the accuracy of the product profile.

Journal Article↗

Sex-specific associations of the plasma-proteome with incident coronary artery disease.

AIMS: The etiology of coronary artery Disease (CAD) appears different for men and women, yet insights into underlying sex-specific biological mechanisms are limited. We integrated genomic and proteomic analyses to investigate sex-specific associations of the plasma-proteome with CAD. METHODS AND RESULTS: In 40,829 UK Biobank participants (free-of-CAD, baseline-365 days thereafter; 55% women; mean age 56.9&#x2009;&#xb1;&#x2009;8.1 years), we examined associations between 2,922 plasma proteins and incident CAD over a median follow-up of 13.7 years (IQR 13.1-14.4) using multivariable-adjusted Cox proportional hazards models. Sex-specific analyses identified 440 female exclusive and 32 male exclusive proteins associated with incident CAD (FDR-corrected p&#x2009;<&#x2009;0.05), revealing distinct pathway enrichments, including innate immune response in women and angiogenesis in men. Causality was assessed through combined and sex-stratified two-sample Mendelian randomization (MR) using inverse-variance-weighted analyses with genome wide association summary statistics from 422,108 men (61,969 cases) and 521,695 women (27,128 cases) (UK Biobank, FinnGen freeze 9). Integration of direct sex-protein interaction analyses with sex-combined MR identified 59 proteins with evidence for sex-specific causal effects. Four proteins demonstrated concordant directionality in sex-stratified MR analyses (n&#x2009;=&#x2009;943,803) and multivariable regression models, namely CDKN2D, MYH9, and SKAP2 (women), and CTSH (men). To assess translational relevance, prioritized targets were further evaluated in secondary major adverse cardiovascular events among carotid endarterectomy patients (MACE; Athero-Express) and acute myocardial infarction (AMI; MISSION!) using plasma proteomics and ELISA. After further top-target identification in the context of MACE and AMI, clinical drug candidates were identified through a machine learning framework, including CTSH (men), and TNFRSF4 (both sexes). CONCLUSIONS: We identified sex-specific associations of proteins and biological pathways with incident CAD. Whereas the majority of proteins had consistent associations in both men and women, our findings suggest a degree of sex-specific pathogenesis with evidence for potential causality, opening new alleys for tailored prevention strategies and clinical cardiovascular risk management.

Journal Article↗

Strategies in engineering sustainable biochemical synthesis through microbial systems.

Growing environmental concerns and the urgency to address climate change have increased demand for the development of sustainable alternatives to fossil-derived fuels and chemicals. Microbial systems, possessing inherent biosynthetic capabilities, present a promising approach for achieving this goal. This review discusses the coupling of systems and synthetic biology to enable the elucidation and manipulation of microbial phenotypes for the production of chemicals that can substitute for petroleum-derived counterparts and contribute to advancing green biotechnology. The integration of artificial intelligence with metabolic engineering to facilitate precise and data-driven design of biosynthetic pathways is also discussed, along with the identification of current limitations and proposition of strategies for optimizing biosystems, thereby propelling the field of chemical biology towards sustainable chemical production.

Metabolic Engineering↗

Crosstalk between cysteine and lysine modifications: Integrating redox and metabolic regulation.

Protein post-translational modifications (PTMs) on amino acid residues enable dynamic cellular responses to changes in metabolic and redox state. Cysteine and lysine are among the most extensively modified amino acid residues, with both undergoing a diversity of acylation and oxidative modifications. Indeed, proximal (<10&#x202f;&#xc5;) cysteine and lysine residues may form integration nodes for crosstalk between metabolism and redox homeostasis pathways. This review highlights the interaction of proximal Cys-Lys residues, including influence on residue pKa by local electrostatics, cysteine-to-lysine transfer of PTM moieties, and covalent crosslinking. We discuss candidate Cys-Lys regulatory pairs in proteins involved in redox regulation, proteostasis, metabolic adaptation and inflammation. We further utilize computational modeling to identify proximity between cysteine and lysine residues in proteins known to be regulated by acylation and oxidative PTMs, and to demonstrate changes in these distances and local electrostatic potential due to lysine acetylation. Finally, we review how mass spectrometry-based proteomics and machine-learning PTM predictive tools can enable the identification, validation, and interpretation of proximal Cys-Lys interactions that regulate cellular responses to oxidative challenge and metabolic flux.

Cysteine↗

Molecular mechanisms of aquaporin biogenesis by the endoplasmic reticulum Sec61 translocon.

The past decade has witnessed remarkable advances in our understanding of aquaporin (AQP) structure and function. Much, however, remains to be learned regarding how these unique and vitally important molecules are generated in living cells. A major obstacle in this respect is that AQP biogenesis takes place in a highly specialized and relatively inaccessible environment formed by the ribosome, the Sec61 translocon and the ER membrane. This review will contrast the folding pathways of two AQP family members, AQP1 and AQP4, and attempt to explain how six TM helices can be oriented across and integrated into the ER membrane in the context of current (and somewhat conflicting) translocon models. These studies indicate that AQP biogenesis is intimately linked to translocon function and that the ribosome and translocon form a highly dynamic molecular machine that both interprets and is controlled by specific information encoded within the nascent AQP polypeptide. AQP biogenesis thus has wide ranging implications for mechanisms of translocon function and general membrane protein folding pathways.

Animals↗

On the design of robotic hands for brain-machine interface.

Brain-machine interface (BMI) is the latest solution to a lack of control for paralyzed or prosthetic limbs. In this paper the authors focus on the design of anatomical robotic hands that use BMI as a critical intervention in restorative neurosurgery and they justify the requirement for lower-level neuromusculoskeletal details (relating to biomechanics, muscles, peripheral nerves, and some aspects of the spinal cord) in both mechanical and control systems. A person uses his or her hands for intimate contact and dexterous interactions with objects that require the user to control not only the finger endpoint locations but also the forces and the stiffness of the fingers. To recreate all of these human properties in a robotic hand, the most direct and perhaps the optimal approach is to duplicate the anatomical musculoskeletal structure. When a prosthetic hand is anatomically correct, the input to the device can come from the same neural signals that used to arrive at the muscles in the original hand. The more similar the mechanical structure of a prosthetic hand is to a human hand, the less learning time is required for the user to recreate dexterous behavior. In addition, removing some of the nonlinearity from the relationship between the cortical signals and the finger movements into the peripheral controls and hardware vastly simplifies the needed BMI algorithms. (Nonlinearity refers to a system of equations in which effects are not proportional to their causes. Such a system could be difficult or impossible to model.) Finally, if a prosthetic hand can be built so that it is anatomically correct, subcomponents could be integrated back into remaining portions of the user's hand at any transitional locations. In the near future, anatomically correct prosthetic hands could be used in restorative neurosurgery to satisfy the user's needs for both aesthetics and ease of control while also providing the highest possible degree of dexterity.

Brain↗

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines↗

Oncogenic EME1 promotes tumor progression and immune modulation in human cancers with therapeutic targeting potential.

BACKGROUND: EME1, a critical DNA repair endonuclease, has emerged as a potential oncogene implicated in genome instability and cancer progression. However, its pan-cancer roles, prognostic significance, immune interactions, and therapeutic targeting remain underexplored. METHODS: We conducted a comprehensive pan-cancer analysis integrating multi-omics data from public databases, including TIMER2.0, GEPIA2, TISIDB, and cBioPortal, to evaluate EME1 expression, genetic alterations, and their association with clinical outcomes, immune infiltration, and molecular pathways. Virtual screening of 3180 FDA-approved drugs and molecular dynamics (MD) simulations were employed to identify and validate potential EME1 inhibitors. RESULTS: EME1 was significantly overexpressed in various human cancers and positively associated with advanced tumor grade and stage. High EME1 expression and mutations were linked to poor overall and disease-free survival. Immunogenomic profiling revealed strong positive correlations between EME1 and myeloid-derived suppressor cells (MDSCs), alongside a negative association with endothelial cell function, suggesting immunosuppressive roles. Machine learning models based on EME1-associated genes demonstrated high predictive accuracy for liver hepatocellular carcinoma (AUC&#x2009;>&#x2009;0.90). Virtual screening identified eight promising drug candidates, including Everolimus and Dioscin, with strong binding affinities. MD simulations confirmed the stability of these interactions, particularly for Dioscin. CONCLUSION: This study reveals the multifaceted oncogenic roles of EME1 in tumor progression, immune evasion, and prognosis. It proposes EME1 as a promising biomarker and therapeutic target across multiple cancer types. The identified drug candidates warrant further in vitro and in vivo validation for potential repurposing in EME1-targeted cancer therapy.

EME1↗

Predicting risk of ischemic stroke: A transformer model using genomic data.

BACKGROUND AND OBJECTIVE: Ischemic stroke is a leading cause of mortality and long-term disability worldwide. Genetic factors contribute to IS susceptibility, yet conventional polygenic risk score approaches are primarily based on additive effects and may not fully capture non-linear relationships or positional context and interactions among genetic variants. This study aimed to develop and evaluate a transformer-based genomic model incorporating position-wise genotype embedding for IS risk prediction. METHODS: We conducted a genome-wide association study using the UK Biobank dataset to identify IS-associated loci. Gene prioritisation was subsequently performed using tissue-specific expression quantitative trait locus-based Mendelian randomisation and colocalization analyses in whole blood and brain cortex. We then developed a transformer-based model that encoded genotype and SNP-position information using a position-wise embedding layer. Model performance was evaluated across three UK Biobank control definitions and externally assessed in the independent All of Us cohort. Performance metrics included the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score. RESULTS: Across the three UK Biobank control definitions, the proposed method achieved the numerically highest discrimination among the evaluated models, with AUROCs of 0.8109, 0.7843, and 0.7468 using MRF-negative, combined, and MRF-positive controls, respectively. In the external All of Us cohort, the proposed method achieved an AUROC of 0.7251 and retained the highest AUROC among the evaluated models. In a separate incident-stroke survival analysis, medium- and high-score groups had hazard ratios of 1.13 and 1.21, respectively, relative to the low-score group. A total of 18 IS-associated loci were identified. Among the tissue-specific MR results, EDEM2 in the brain cortex remained significant after Bonferroni correction, while DCHS2 showed a nominal association. CONCLUSIONS: The proposed transformer-based framework provides a genomic modelling approach that achieved the highest discrimination among the evaluated models in this study and retained comparative performance in an independent external cohort. In further applications, integrating this genomic framework with conventional clinical, lifestyle, and environmental risk factors may support more comprehensive and personalised IS risk assessment. Prospective, population-representative, and multi-ancestry validation will be important to establish its potential role in future prevention-oriented risk management.

Genomics and bioinformatics↗

The relations between neuroscience and human behavioral science.

Neuroscience seeks to understand how the human brain, perhaps the most complex electrochemical machine in the universe, works, in terms of molecules, membranes, cells and cell assemblies, development, plasticity, learning, memory, cognition, and behavior. The human behavioral sciences, in particular psychiatry and clinical psychology, deal with disorders of human behavior and mentation. The gap between neuroscience and the human behavioral sciences is still large. However, some major advances in neuroscience over the last two decades have diminished the span. This article reviews the major advances of neuroscience in six areas with relevance to the behavioral sciences: (a) evolution of the nervous system; (b) visualizing activity in the human brain; (c) plasticity of the cerebral cortex; (d) receptors, ion channels, and second/third messengers; (e) molecular genetic approaches; and (f) understanding integrative systems with networks and circadian clocks as examples.

Animals↗

Prediction-based fingerprints of protein-protein interactions.

The recognition of protein interaction sites is an important intermediate step toward identification of functionally relevant residues and understanding protein function, facilitating experimental efforts in that regard. Toward that goal, the authors propose a novel representation for the recognition of protein-protein interaction sites that integrates enhanced relative solvent accessibility (RSA) predictions with high resolution structural data. An observation that RSA predictions are biased toward the level of surface exposure consistent with protein complexes led the authors to investigate the difference between the predicted and actual (i.e., observed in an unbound structure) RSA of an amino acid residue as a fingerprint of interaction sites. The authors demonstrate that RSA prediction-based fingerprints of protein interactions significantly improve the discrimination between interacting and noninteracting sites, compared with evolutionary conservation, physicochemical characteristics, structure-derived and other features considered before. On the basis of these observations, the authors developed a new method for the prediction of protein-protein interaction sites, using machine learning approaches to combine the most informative features into the final predictor. For training and validation, the authors used several large sets of protein complexes and derived from them nonredundant representative chains, with interaction sites mapped from multiple complexes. Alternative machine learning techniques are used, including Support Vector Machines and Neural Networks, so as to evaluate the relative effects of the choice of a representation and a specific learning algorithm. The effects of induced fit and uncertainty of the negative (noninteracting) class assignment are also evaluated. Several representative methods from the literature are reimplemented to enable direct comparison of the results. Using rigorous validation protocols, the authors estimated that the new method yields the overall classification accuracy of about 74% and Matthews correlation coefficients of 0.42, as opposed to up to 70% classification accuracy and up to 0.3 Matthews correlation coefficient for methods that do not utilize RSA prediction-based fingerprints. The new method is available at http://sppider.cchmc.org.

Artificial Intelligence↗