PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans↗

Classifying vertical facial deformity using supervised and unsupervised learning.

OBJECTIVES: To evaluate the potential for machine learning techniques to identify objective criteria for classifying vertical facial deformity. METHODS: 19 parameters were determined from 131 lateral skull radiographs. Classifications were induced from raw data with simple visualisation, C5.0 and Kohonen feature maps; and using a Point Distribution Model (PDM) of shape templates comprising points taken from digitised radiographs. RESULTS: The induced decision trees enable a direct comparison of clinicians' idiosyncrasies in classification. Unsupervised algorithms induce models that are potentially more objective, but their blackbox nature makes them unsuitable for clinical application. The PDM methodology gives dramatic visualisations of two modes separating horizontal and vertical facial growth. Kohonen feature maps favour one clinician and PDM the other. Clinical response suggests that while Clinician 1 places greater weight on 5 of 6 parameters, Clinician 2 relies on more parameters that capture facial shape. CONCLUSIONS: While machine learning and statistical analyses classify subjects for vertical facial height, they have limited application in their present form. The supervised learning algorithm C5.0 is effective for generating rules for individual clinicians but its inherent bias invalidates its use for objective classification of facial form for research purposes. On the other hand, promising results from unsupervised strategies (especially the PDM) suggest a potential use for objective classification and further identification and analysis of ambiguous cases. At present, such methodologies may be unsuitable for clinical application because of the invisibility of their underlying processes. Further study is required with additional patient data and a wider group of clinicians.

Algorithms↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 688 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

DNA methylation↗

Multiclass cancer classification using gene expression profiling and probabilistic neural networks.

Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 or more. In addition, microarray data exhibit a high degree of noise. Most of the discussed methods do not adequately address the problem of dimensionality and noise. Furthermore, although machine learning and data mining methods are based on statistics, most such techniques do not address the biologist's requirement for sound mathematical confidence measures. Finally, most machine learning and data mining classification methods fail to incorporate misclassification costs, i.e. they are indifferent to the costs associated with false positive and false negative classifications. In this paper, we present a probabilistic neural network (PNN) model that addresses all these issues. The PNN model provides sound statistical confidences for its decisions, and it is able to model asymmetrical misclassification costs. Furthermore, we demonstrate the performance of the PNN for multiclass gene expression data sets. Here, we compare the performance of the PNN with two machine learning methods, a decision tree and a neural network. To assess and evaluate the performance of the classifiers, we use a lift-based scoring system that allows a fair comparison of different models. The PNN clearly outperformed the other models. The results demonstrate the successful application of the PNN model for multiclass cancer classification.

Artificial Intelligence↗

Disease candidate genes prediction using positive labeled and unlabeled instances.

Identifying disease genes and understanding their performance is critical in producing drugs for genetic diseases. Nowadays, laboratory approaches are not only used for disease gene identification but also using computational approaches like machine learning are becoming considerable for this purpose. In machine learning methods, researchers can only use two data types (disease genes and unknown genes) to predict disease candidate genes. Notably, there is no source for the negative data set. The proposed method is a two-step process: The first step is the extraction of reliable negative genes from a set of unlabeled genes by one-class learning and a filter based on distance indicators from known disease genes; this step is performed separately for each disease. The second step is the learning of a binary model using causing genes of each disease as a positive learning set and the reliable negative genes extracted from that disease. Each gene in the unlabeled gene's production and ranking step is assigned a normalized score using two filters and a learned model. Consequently, disease genes are predicted and ranked. The proposed method evaluation of various six diseases and Cancer class indicates better results than other studies.

Humans↗

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine↗

Structure-based discovery of inhibitors of Mac1 domain of nonstructural protein-3 of SARS-CoV-2 by machine learning-augmented screening of chemical space.

Significant efforts have been recently dedicated to the discovery of small molecule inhibitors against the Macrodomain 1 (Mac1) of nonstructural protein 3 (NSP3) as potential antivirals for SARS-CoV-2. Thus, Mac1 has also been selected as the target for the Critical Assessment of Hit-finding Experiments (CACHE) challenge #3. As contestants in that challenge, we developed a computational strategy that ranked on the top among all 23 participants in the competition and resulted in the discovery of a novel chemical series of non-charged Mac1 inhibitors. Those have been identified through the combination of machine learning-accelerated virtual screening of Enamine REAL Diversity Subset of approximately 25 million compounds and consequent hit expansion into the entire Enamine REAL Space library. In particular, the initially identified hit compound CACHE3-HI_1706_56 (KD = 20 &#x3bc;M) was explored by probing 17 close analogues from a library of 44 billion molecules from the Enamine REAL. All those analogues effectively displaced the Mac1-binding ADP-ribose peptide, and 12 were confirmed to engage with Mac1 by the Surface Plasmon Resonance experiments, revealing a new chemical series of compounds for hit-to-lead optimization. The structure of the CACHE3-HI_1706_56-Mac1 complex was further determined at high resolution with crystallography, confirming initial computational predictions. Our results illustrate the effectiveness of ML-accelerated docking to rapidly identify novel chemical series and provide a strong foundation for the development of SARS-CoV-2 NSP3 Mac1 inhibitors.

CACHE challenge↗

Machine learning-assisted Mn-N-C nanozyme colorimetric sensor array for trace-level detection of biogenic amines in meat.

Accurate detection of biogenic amines (BAs) in meat remains challenging due to their high structural similarity and co-occurrence. Herein, an Mn-N-C nanozyme was synthesized via a metal-organic framework confined pyrolysis strategy, possessing excellent oxidase (OXD)- and peroxidase (POD)-like activities. The dual enzyme-like activity showed Km values of 0.1584&#xa0;mM (OXD) and 0.1498&#xa0;mM (POD), respectively, in detection system. Leveraging these properties, a colorimetric sensor array was constructed, enabling the detection of four representative BAs within a concentration range of 2-10&#xa0;ppm with 100% classification accuracy. In addition, a concentration independent recognition model based on an artificial neural network was developed to address signal nonlinearity interference in meat. The integrated system achieved accurate trace-level identification of BAs in perishable fish, pork, and chicken, demonstrating its applicability for early-stage BAs monitoring and quality deterioration warning during storage and transportation.

Biogenic Amines↗

Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics.

BACKGROUND: Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction. METHODS: Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32&#x2009;330) and internal validation (n=13&#x2009;857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel. RESULTS: Across cohorts, the median age was 58&#x2009;years and &#x223c;45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information. CONCLUSIONS: Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.

Humans↗

Data-driven consideration of genetic disorders for global genomic newborn screening programs.

PURPOSE: Over 30 international studies are exploring newborn sequencing (NBSeq) to expand the range of genetic disorders included in newborn screening. Substantial variability in gene selection across programs exists, highlighting the need for a systematic approach to prioritize genes. METHODS: We assembled a data set comprising 25 characteristics about each of the 4390 genes included in 27 NBSeq programs. We used regression analysis to identify several predictors of inclusion and developed a machine learning model to rank genes for public health consideration. RESULTS: Among 27 NBSeq programs, the number of genes analyzed ranged from 134 to 4299, with only 74 (1.7%) genes included by over 80% of programs. The most significant associations with gene inclusion across programs were presence on the US Recommended Uniform Screening Panel (inclusion increase of 74.7%, CI: 71.0%-78.4%), robust evidence on the natural history (29.5%, CI: 24.6%-34.4%), and treatment efficacy (17.0%, CI: 12.3%-21.7%) of the associated genetic disease. A boosted trees machine learning model using 13 predictors achieved high accuracy in predicting gene inclusion across programs (area under the curve = 0.915, R2 = 84%). CONCLUSION: The machine learning model developed here provides a ranked list of genes that can adapt to emerging evidence and regional needs, enabling more consistent and informed gene selection in NBSeq initiatives.

Humans↗

Comparative experiments on learning information extractors for proteins and their interactions.

OBJECTIVE: Automatically extracting information from biomedical text holds the promise of easily consolidating large amounts of biological knowledge in computer-accessible form. This strategy is particularly attractive for extracting data relevant to genes of the human genome from the 11 million abstracts in Medline. However, extraction efforts have been frustrated by the lack of conventions for describing human genes and proteins. We have developed and evaluated a variety of learned information extraction systems for identifying human protein names in Medline abstracts and subsequently extracting information on interactions between the proteins. METHODS AND MATERIAL: We used a variety of machine learning methods to automatically develop information extraction systems for extracting information on gene/protein name, function and interactions from Medline abstracts. We present cross-validated results on identifying human proteins and their interactions by training and testing on a set of approximately 1000 manually-annotated Medline abstracts that discuss human genes/proteins. RESULTS: We demonstrate that machine learning approaches using support vector machines and maximum entropy are able to identify human proteins with higher accuracy than several previous approaches. We also demonstrate that various rule induction methods are able to identify protein interactions with higher precision than manually-developed rules. CONCLUSION: Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus.

Algorithms↗

Mining protein function from text using term-based support vector machines.

BACKGROUND: Text mining has spurred huge interest in the domain of biology. The goal of the BioCreAtIvE exercise was to evaluate the performance of current text mining systems. We participated in Task 2, which addressed assigning Gene Ontology terms to human proteins and selecting relevant evidence from full-text documents. We approached it as a modified form of the document classification task. We used a supervised machine-learning approach (based on support vector machines) to assign protein function and select passages that support the assignments. As classification features, we used a protein's co-occurring terms that were automatically extracted from documents. RESULTS: The results evaluated by curators were modest, and quite variable for different problems: in many cases we have relatively good assignment of GO terms to proteins, but the selected supporting text was typically non-relevant (precision spanning from 3% to 50%). The method appears to work best when a substantial set of relevant documents is obtained, while it works poorly on single documents and/or short passages. The initial results suggest that our approach can also mine annotations from text even when an explicit statement relating a protein to a GO term is absent. CONCLUSION: A machine learning approach to mining protein function predictions from text can yield good performance only if sufficient training data is available, and significant amount of supporting data is used for prediction. The most promising results are for combined document retrieval and GO term assignment, which calls for the integration of methods developed in BioCreAtIvE Task 1 and Task 2.

Computational Biology↗

Automatic MeSH term assignment and quality assessment.

For computational purposes documents or other objects are most often represented by a collection of individual attributes that may be strings or numbers. Such attributes are often called features and success in solving a given problem can depend critically on the nature of the features selected to represent documents. Feature selection has received considerable attention in the machine learning literature. In the area of document retrieval we refer to feature selection as indexing. Indexing has not traditionally been evaluated by the same methods used in machine learning feature selection. Here we show how indexing quality may be evaluated in a machine learning setting and apply this methodology to results of the Indexing Initiative at the National Library of Medicine.

Abstracting and Indexing↗

Identification and validation of the important role of KIF11 in the development and progression of endometrial cancer.

BACKGROUND: Human kinesin family member 11 (KIF11) plays a vital role in regulating the cell cycle and is implicated in the tumorigenesis and progression of various cancers, but its role in endometrial cancer (EC) is still unclear. Our current research explored the prognostic value, biological function and targeting strategy of KIF11 in EC through approaches including bioinformatics, machine learning and experimental studies. METHODS: The GSE17025 dataset from the GEO database was analyzed via the limma package to identify differentially expressed genes (DEGs) in EC. Functional enrichment analysis of the DEGs was conducted using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. DEGs were further screened for hub genes through protein-protein interaction (PPI) network analysis and machine learning. The role of the hub gene KIF11 in EC was analyzed using clinical data from the TCGA database. The expression of KIF11 in EC was subsequently validated in clinical samples. In vitro experiments were utilized to evaluate the effects of KIF11 on biological functions such as proliferation, migration, apoptosis, and the cell cycle in endometrial cancer cells. RESULTS: A total of 877 DEGs, which are widely involved in important biological processes such as cell division, tubulin binding, and the cell cycle, were identified. Through PPI network analysis and machine learning, KIF11 was selected as the hub gene for subsequent analysis and experimental validation. An analysis of TCGA data revealed that KIF11 is highly expressed in EC and is associated with tumor grade, stage, and a low survival rate. The overexpression of KIF11 in tumor tissues was further confirmed in EC patient samples. KIF11 knockdown had inhibitory effects on cell proliferation, migration and invasion. Flow cytometry analysis revealed that KIF11 knockdown induced G2/M phase arrest and promoted apoptosis in EC cells. CONCLUSION: Our study demonstrated that KIF11 was upregulated in EC and was strongly associated with a poor prognosis. Notably, we found that reduced KIF11 expression inhibited EC cell proliferation, migration and invasion. KIF11 knockdown caused more EC cells to arrest in the G2/M phase and undergo apoptosis. The findings of our study emphasized that KIF11 may be a promising prognostic biomarker and therapeutic target for EC patients.

Humans↗

Classifying 'drug-likeness' with kernel-based learning methods.

In this article we report about a successful application of modern machine learning technology, namely Support Vector Machines, to the problem of assessing the 'drug-likeness' of a chemical from a given set of descriptors of the substance. We were able to drastically improve the recent result by Byvatov et al. (2003) on this task and achieved an error rate of about 7% on unseen compounds using Support Vector Machines. We see a very high potential of such machine learning techniques for a variety of computational chemistry problems that occur in the drug discovery and drug design process.

Artificial Intelligence↗

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus↗

Systematic learning of gene functional classes from DNA array expression data by using multilayer perceptrons.

Recent advances in microarray technology have opened new ways for functional annotation of previously uncharacterised genes on a genomic scale. This has been demonstrated by unsupervised clustering of co-expressed genes and, more importantly, by supervised learning algorithms. Using prior knowledge, these algorithms can assign functional annotations based on more complex expression signatures found in existing functional classes. Previously, support vector machines (SVMs) and other machine-learning methods have been applied to a limited number of functional classes for this purpose. Here we present, for the first time, the comprehensive application of supervised neural networks (SNNs) for functional annotation. Our study is novel in that we report systematic results for ~100 classes in the Munich Information Center for Protein Sequences (MIPS) functional catalog. We found that only ~10% of these are learnable (based on the rate of false negatives). A closer analysis reveals that false positives (and negatives) in a machine-learning context are not necessarily "false" in a biological sense. We show that the high degree of interconnections among functional classes confounds the signatures that ought to be learned for a unique class. We term this the "Borges effect" and introduce two new numerical indices for its quantification. Our analysis indicates that classification systems with a lower Borges effect are better suitable for machine learning. Furthermore, we introduce a learning procedure for combining false positives with the original class. We show that in a few iterations this process converges to a gene set that is learnable with considerably low rates of false positives and negatives and contains genes that are biologically related to the original class, allowing for a coarse reconstruction of the interactions between associated biological pathways. We exemplify this methodology using the well-studied tricarboxylic acid cycle.

Algorithms↗

Data mining in bioinformatics using Weka.

UNLABELLED: The Weka machine learning workbench provides a general-purpose environment for automatic classification, regression, clustering and feature selection-common data mining problems in bioinformatics research. It contains an extensive collection of machine learning algorithms and data pre-processing methods complemented by graphical user interfaces for data exploration and the experimental comparison of different machine learning techniques on the same problem. Weka can process data given in the form of a single relational table. Its main objectives are to (a) assist users in extracting useful information from data and (b) enable them to easily identify a suitable algorithm for generating an accurate predictive model from it. AVAILABILITY: http://www.cs.waikato.ac.nz/ml/weka.

Algorithms↗