PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5 × whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n = 333) and a temporally independent validation cohort (n = 348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans↗

Foundations of Artificial Intelligence in Hepatology: What a Clinician Needs to Know.

This review focuses on foundational knowledge about artificial intelligence (AI) in hepatology, exploring how AI, including machine learning and deep learning, leverages large-scale clinical data to transform the diagnosis, risk assessment, prognostication, and management of liver diseases. Online resources are described to offer fundamental AI knowledge and essential technical skills and to facilitate clinician participation across the entire AI lifecycle, ensuring they contribute not only as end users but also in development and deployment. Unlike traditional statistical approaches that prioritize interpretable parameters and clinical insight, AI focuses on maximizing predictive accuracy by identifying complex, often non-linear patterns using high-dimensional data, albeit often at the cost of model interpretability. AI is demonstrating clinical utility in liver histopathology and radiological imaging, significantly improving detection accuracy for cirrhosis, clinically significant portal hypertension, and hepatocellular carcinoma. Beyond diagnostics, AI-driven prediction models are emerging to provide personalized risk stratification for the development of liver-related complications and treatment guidance, based on complex data including longitudinal laboratory results, comorbidities, and co-medication use to monitor disease progression and therapy response. The field is rapidly expanding into novel areas such as analyzing patient-reported outcomes, genomic data, and real-time liver function monitoring, offering deeper mechanistic insights alongside clinical tools. Despite the potential to revolutionize hepatology practice and research, successful integration into routine care faces challenges. These include seamless workflow integration with existing electronic health records, establishing clear liability frameworks, and guaranteeing protection of patient privacy. Addressing these hurdles requires collaborative efforts from clinicians, researchers, and regulators to develop best practices and governance. Understanding the transformative capabilities, current applications, emerging frontiers, and essential implementation considerations is crucial for clinicians navigating the evolving AI landscape and responsibly utilizing its power for improved patient outcomes.

PROBAST+AI↗

Information geometry of U-Boost and Bregman divergence.

We aim at an extension of AdaBoost to U-Boost, in the paradigm to build a stronger classification machine from a set of weak learning machines. A geometric understanding of the Bregman divergence defined by a generic convex function U leads to the U-Boost method in the framework of information geometry extended to the space of the finite measures over a label set. We propose two versions of U-Boost learning algorithms by taking account of whether the domain is restricted to the space of probability functions. In the sequential step, we observe that the two adjacent and the initial classifiers are associated with a right triangle in the scale via the Bregman divergence, called the Pythagorean relation. This leads to a mild convergence property of the U-Boost algorithm as seen in the expectation-maximization algorithm. Statistical discussions for consistency and robustness elucidate the properties of the U-Boost methods based on a stochastic assumption for training data.

Algorithms↗

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis↗

Neural networks, nativism, and the plausibility of constructivism.

Recent interest in PDP (parallel distributed processing) models is due in part to the widely held belief that they challenge many of the assumptions of classical cognitive science. In the domain of language acquisition, for example, there has been much interest in the claim that PDP models might undermine nativism. Related arguments based on PDP learning have also been given against Fodor's anti-constructivist position--a position that has contributed to the widespread dismissal of constructivism. A limitation of many of the claims regarding PDP learning, however, is that the principles underlying this learning have not been rigorously characterized. In this paper, I examine PDP models from within the framework of Valiant's PAC (probably approximately correct) model of learning, now the dominant model in machine learning, and which applies naturally to neural network learning. From this perspective, I evaluate the implications of PDP models for nativism and Fodor's influential anti-constructivist position. In particular, I demonstrate that, contrary to a number of claims, PDP models are nativist in a robust sense. I also demonstrate that PDP models actually serve as a good illustration of Fodor's anti-constructivist position. While these results may at first suggest that neural network models in general are incapable of the sort of concept acquisition that is required to refute Fodor's anti-constructivist position, I suggest that there is an alternative form of neural network learning that demonstrates the plausibility of constructivism. This alternative form of learning is a natural interpretation of the constructivist position in terms of neural network learning, as it employs learning algorithms that incorporate the addition of structure in addition to weight modification schemes. By demonstrating that there is a natural and plausible interpretation of constructivism in terms of neural network learning, the position that nativism is the only plausible model of acquisition can no longer be defended. Indeed, I briefly discuss a number of learning-theoretic reasons indicating that constructivist models so characterized uniquely possess a number of important learning characteristics.

Cognition↗

QSAR study of ethyl 2-[(3-methyl-2,5-dioxo(3-pyrrolinyl))amino]-4-(trifluoromethyl) pyrimidine-5-carboxylate: an inhibitor of AP-1 and NF-kappa B mediated gene expression based on support vector machines.

The support vector machine, as a novel type of learning machine, for the first time, was used to develop a QSAR model of 57 analogues of ethyl 2-[(3-methyl-2,5-dioxo(3-pyrrolinyl))amino]-4-(trifluoromethyl)pyrimidine-5-carboxylate (EPC), an inhibitor of AP-1 and NF-kappa B mediated gene expression, based on calculated quantum chemical parameters. The quantum chemical parameters involved in the model are Kier and Hall index (order3) (KHI3), Information content (order 0) (IC0), YZ Shadow (YZS) and Max partial charge for an N atom (MaxPCN), Min partial charge for an N atom (MinPCN). The mean relative error of the training set, the validation set, and the testing set is 1.35%, 1.52%, and 2.23%, respectively, and the maximum relative error is less than 5.00%.

Carboxylic Acids↗

A new algorithm for the evaluation of shotgun peptide sequencing in proteomics: support vector machine classification of peptide MS/MS spectra and SEQUEST scores.

Shotgun tandem mass spectrometry-based peptide sequencing using programs such as SEQUEST allows high-throughput identification of peptides, which in turn allows the identification of corresponding proteins. We have applied a machine learning algorithm, called the support vector machine, to discriminate between correctly and incorrectly identified peptides using SEQUEST output. Each peptide was characterized by SEQUEST-calculated features such as delta Cn and Xcorr, measurements such as precursor ion current and mass, and additional calculated parameters such as the fraction of matched MS/MS peaks. The trained SVM classifier performed significantly better than previous cutoff-based methods at separating positive from negative peptides. Positive and negative peptides were more readily distinguished in training set data acquired on a QTOF, compared to an ion trap mass spectrometer. The use of 13 features, including four new parameters, significantly improved the separation between positive and negative peptides. Use of the support vector machine and these additional parameters resulted in a more accurate interpretation of peptide MS/MS spectra and is an important step toward automated interpretation of peptide tandem mass spectrometry data in proteomics.

Algorithms↗

Predictive models for breast cancer susceptibility from multiple single nucleotide polymorphisms.

Hereditary predisposition and causative environmental exposures have long been recognized in human malignancies. In most instances, cancer cases occur sporadically, suggesting that environmental influences are critical in determining cancer risk. To test the influence of genetic polymorphisms on breast cancer risk, we have measured 98 single nucleotide polymorphisms (SNPs) distributed over 45 genes of potential relevance to breast cancer etiology in 174 patients and have compared these with matched normal controls. Using machine learning techniques such as support vector machines (SVMs), decision trees, and naïve Bayes, we identified a subset of three SNPs as key discriminators between breast cancer and controls. The SVMs performed maximally among predictive models, achieving 69% predictive power in distinguishing between the two groups, compared with a 50% baseline predictive power obtained from the data after repeated random permutation of class labels (individuals with cancer or controls). However, the simpler naïve Bayes model as well as the decision tree model performed quite similarly to the SVM. The three SNP sites most useful in this model were (a) the +4536T/C site of the aldosterone synthase gene CYP11B2 at amino acid residue 386 Val/Ala (T/C) (rs4541); (b) the +4328C/G site of the aryl hydrocarbon hydroxylase CYP1B1 at amino acid residue 293 Leu/Val (C/G) (rs5292); and (c) the +4449C/T site of the transcription factor BCL6 at amino acid 387 Asp/Asp (rs1056932). No single SNP site on its own could achieve more than 60% in predictive accuracy. We have shown that multiple SNP sites from different genes over distant parts of the genome are better at identifying breast cancer patients than any one SNP alone. As high-throughput technology for SNPs improves and as more SNPs are identified, it is likely that much higher predictive accuracy will be achieved and a useful clinical tool developed.

Algorithms↗

Stratifying lung adenocarcinoma: a novel prognostic model based on mitochondrial outer membrane permeabilization activity.

UNLABELLED: Mitochondrial outer membrane permeabilization (MOMP) is a core apoptotic regulatory event that dictates mitochondrial integrity, where full activation drives cell death and sublethal dysregulation contributes to tumor genomic instability. We used the Cancer Genome Atlas lung adenocarcinoma cohort (TCGA-LUAD) as the training cohort and the Gene Expression Omnibus dataset GSE42127 as the validation cohort to identify prognostic genes related to MOMP activity in lung adenocarcinoma (LUAD) and to evaluate their potential biological significance. By intersecting MOMP-related genes with differentially expressed genes, combined with survival analysis, Mendelian randomization analysis, and 101 machine-learning algorithm combinations, seven prognostic genes, namely BIRC5, PSMD11, TNFRSF13C, YWHAZ, YWHAG, CYCS, and LTB, were identified. Next, an optimal prognostic model was constructed based on the gradient boosting machine (GBM) algorithm. Based on the risk score, LUAD patients were stratified into high- and low-risk groups, and patients in the high-risk group exhibited poorer overall survival in both the training and validation cohorts. Furthermore, a nomogram integrating the risk score and clinicopathological factors was developed and showed favorable predictive performance for 1-, 3-, and 5-year survival. Meanwhile, functional and immune analyses revealed that the high-risk group was enriched in DNA replication-related pathways and demonstrated a higher tumor mutation burden (TMB). Correlation analysis indicated that TNFRSF13C was positively correlated with activated B cells, whereas BIRC5 was negatively correlated with eosinophils, suggesting that MOMP-related genes might be involved in remodeling the immune microenvironment of LUAD. Drug sensitivity analysis showed differences in predicted half-maximal inhibitory concentration (IC50) values between the risk groups, suggesting the potential value of this model in assisting therapeutic stratification. Single-cell RNA sequencing (scRNA-seq) further identified T lymphocytes as a key cell type, with numerous prognostic genes exhibiting differential expression in T cells or dynamic changes during differentiation. We suggest that the MOMP-related signature established in this study may provide a reference for prognostic stratification in LUAD and offers candidate prognostic genes for subsequent experimental and clinical validation. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05058-6.

Lung adenocarcinoma↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, &#x3008;MAE&#x3009; = 0.11 eV, and &#x3008;RMSE&#x3009; = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by &#x223c;10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping↗

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform↗

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Support vector machines for predicting rRNA-, RNA-, and DNA-binding proteins from amino acid sequence.

Classification of gene function remains one of the most important and demanding tasks in the post-genome era. Most of the current predictive computer methods rely on comparing features that are essentially linear to the protein sequence. However, features of a protein nonlinear to the sequence may also be predictive to its function. Machine learning methods, for instance the Support Vector Machines (SVMs), are particularly suitable for exploiting such features. In this work we introduce SVM and the pseudo-amino acid composition, a collection of nonlinear features extractable from protein sequence, to the field of protein function prediction. We have developed prototype SVMs for binary classification of rRNA-, RNA-, and DNA-binding proteins. Using a protein's amino acid composition and limited range correlation of hydrophobicity and solvent accessible surface area as input, each of the SVMs predicts whether the protein belongs to one of the three classes. In self-consistency and cross-validation tests, which measures the success of learning and prediction, respectively, the rRNA-binding SVM has consistently achieved >95% accuracy. The RNA- and DNA-binding SVMs demonstrate more diverse accuracy, ranging from approximately 76% to approximately 97%. Analysis of the test results suggests the directions of improving the SVMs.

Computational Biology↗

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics↗

Support vector machines for prediction of protein subcellular location by incorporating quasi-sequence-order effect.

Support Vector Machine (SVM), which is one class of learning machines, was applied to predict the subcellular location of proteins by incorporating the quasi-sequence-order effect (Chou [2000] Biochem. Biophys. Res. Commun. 278:477-483). In this study, the proteins are classified into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracellular, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane, and (12) vacuole, which account for most organelles and subcellular compartments in an animal or plant cell. Examinations for self-consistency and jackknife testing of the SVMs method were conducted for three sets consisting of 1,911, 2,044, and 2,191 proteins. The correct rates for self-consistency and the jackknife test values achieved with these protein sets were 94 and 83% for 1,911 proteins, 92 and 78% for 2,044 proteins, and 89 and 75% for 2,191 proteins, respectively. Furthermore, tests for correct prediction rates were undertaken with three independent testing datasets containing 2,148 proteins, 2,417 proteins, and 2,494 proteins producing values of 84, 77, and 74%, respectively.

Proteins↗

Support vector machines for prediction of protein subcellular location.

Support Vector Machine (SVM), which is one kind of learning machines, was applied to predict the subcellular location of proteins from their amino acid composition. In this research, the proteins are classified into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracall, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane, and (12) vacuole, which have covered almost all the organelles and subcellular compartments in an animal or plant cell. The examination for the self-consistency and the jackknife test of the SVMs method was tested for the three sets: 2022 proteins, 2161 proteins, and 2319 proteins. As a result, the correct rate of self-consistency and jackknife test reaches 91 and 82% for 2022 proteins, 89 and 75% for 2161 proteins, and 85 and 73% for 2319 proteins, respectively. Furthermore, the predicting rate was tested by the three independent testing datasets containing 2240 proteins, 2513 proteins, and 2591 proteins. The correct prediction rates reach 82, 75, and 73% for 2240 proteins, 2513 proteins, and 2591 proteins, respectively.

Algorithms↗

Prediction of protein structural classes by support vector machines.

In this paper, we apply a new machine learning method which is called support vector machine to approach the prediction of protein structural class. The support vector machine method is performed based on the database derived from SCOP which is based upon domains of known structure and the evolutionary relationships and the principles that govern their 3D structure. As a result, high rates of both self-consistency and jackknife test are obtained. This indicates that the structural class of a protein inconsiderably correlated with its amino and composition, and the support vector machine can be referred as a powerful computational tool for predicting the structural classes of proteins.

Artificial Intelligence↗