PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Symbiogenesis in learning classifier systems.

Symbiosis is the phenomenon in which organisms of different species live together in close association, resulting in a raised level of fitness for one or more of the organisms. Symbiogenesis is the name given to the process by which symbiotic partners combine and unify, that is, become genetically linked, giving rise to new morphologies and physiologies evolutionarily more advanced than their constituents. The importance of this process in the evolution of complexity is now well established. Learning classifier systems are a machine learning technique that uses both evolutionary computing techniques and reinforcement learning to develop a population of cooperative rules to solve a given task. In this article we examine the use of symbiogenesis within the classifier system rule base to improve their performance. Results show that incorporating simple rule linkage does not give any benefits. The concept of (temporal) encapsulation is then added to the symbiotic rules and shown to improve performance in ambiguous/non-Markov environments.

Algorithms↗

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer↗

Automatic knowledge base refinement: learning from examples and deep knowledge in rheumatology.

MESICAR is a second generation expert system which contains very general descriptions of rheumatological disorders in the primary medical care field. With the help of a detailed hierarchical description of the human anatomy the system is able to support diagnostic decisions. The paper describes how machine learning techniques are used to automatically construct more specific disease descriptions for common, frequently occurring cases. The system MESICAR-LEARN implements a learning method which integrates analytical and empirical learning techniques. Cases diagnosed by MESICAR form the training examples, and MESICAR's knowledge base is used as domain theory. The learned concepts are integrated into a hierarchy of disease descriptions. They support efficient and fast reasoning on common cases in addition to the general diagnostic support afforded by MESICAR's deep knowledge.

Algorithms↗

Selection of patient samples and genes for outcome prediction.

Gene expression profiles with clinical outcome data enable monitoring of disease progression and prediction of patient survival at the molecular level. We present a new computational method for outcome prediction. Our idea is to use an informative subset of original training samples. This subset consists of only short-term survivors who died within a short period and long-term survivors who were still alive after a long follow-up time. These extreme training samples yield a clear platform to identify genes whose expression is related to survival. To find relevant genes, we combine two feature selection methods -- entropy measure and Wilcoxon rank sum test -- so that a set of sharp discriminating features are identified. The selected training samples and genes are then integrated by a support vector machine to build a prediction model, by which each validation sample is assigned a survival/relapse risk score for drawing Kaplan-Meier survival curves. We apply this method to two data sets: diffuse large-B-cell lymphoma (DLBCL) and primary lung adenocarcinoma. In both cases, patients in high and low risk groups stratified by our risk scores are clearly distinguishable. We also compare our risk scores to some clinical factors, such as International Prognostic Index score for DLBCL analysis and tumor stage information for lung adenocarcinoma. Our results indicate that gene expression profiles combined with carefully chosen learning algorithms can predict patient survival for certain diseases.

Biomarkers, Tumor↗

ABNER: an open source tool for automatically tagging genes, proteins and other entity names in text.

ABNER (A Biomedical Named Entity Recognizer) is an open source software tool for molecular biology text mining. At its core is a machine learning system using conditional random fields with a variety of orthographic and contextual features. The latest version is 1.5, which has an intuitive graphical interface and includes two modules for tagging entities (e.g. protein and cell line) trained on standard corpora, for which performance is roughly state of the art. It also includes a Java application programming interface allowing users to incorporate ABNER into their own systems and train models on new corpora.

Algorithms↗

Rule induction and instance-based learning applied in medical diagnosis.

Machine learning methods have been applied in a variety of medical domains in order to improve medical decision making. Improved medical diagnosis and prognosis can be achieved through automatic analysis of patient data stored in medical records, i.e., by learning from past experience. Given patient records with corresponding diagnoses, machine learning methods are able to classify new cases either through constructing explicit rules that generalize the training cases (e.g., rule induction) or by storing (some of) the training cases for reference (instance-based learning). This paper presents the methodologies of rule induction and instance-based learning and their application to medical diagnosis, in particular, the problem of early diagnosis of rheumatic diseases. It also discusses the possibility to use existing expert knowledge to support the learning process and the utility of such knowledge.

Algorithms↗

Combination of a naive Bayes classifier with consensus scoring improves enrichment of high-throughput docking results.

We have previously shown that a machine learning technique can improve the enrichment of high-throughput docking (HTD) results. In the previous cases studied, however, the application of a naive Bayes classifier failed to improve enrichment for instances where HTD alone was unable to generate an acceptable enrichment. We present here a protocol to rescue poor docking results a priori using a combination of rank-by-median consensus scoring and naive Bayesian categorization.

Algorithms↗

On the design of robotic hands for brain-machine interface.

Brain-machine interface (BMI) is the latest solution to a lack of control for paralyzed or prosthetic limbs. In this paper the authors focus on the design of anatomical robotic hands that use BMI as a critical intervention in restorative neurosurgery and they justify the requirement for lower-level neuromusculoskeletal details (relating to biomechanics, muscles, peripheral nerves, and some aspects of the spinal cord) in both mechanical and control systems. A person uses his or her hands for intimate contact and dexterous interactions with objects that require the user to control not only the finger endpoint locations but also the forces and the stiffness of the fingers. To recreate all of these human properties in a robotic hand, the most direct and perhaps the optimal approach is to duplicate the anatomical musculoskeletal structure. When a prosthetic hand is anatomically correct, the input to the device can come from the same neural signals that used to arrive at the muscles in the original hand. The more similar the mechanical structure of a prosthetic hand is to a human hand, the less learning time is required for the user to recreate dexterous behavior. In addition, removing some of the nonlinearity from the relationship between the cortical signals and the finger movements into the peripheral controls and hardware vastly simplifies the needed BMI algorithms. (Nonlinearity refers to a system of equations in which effects are not proportional to their causes. Such a system could be difficult or impossible to model.) Finally, if a prosthetic hand can be built so that it is anatomically correct, subcomponents could be integrated back into remaining portions of the user's hand at any transitional locations. In the near future, anatomically correct prosthetic hands could be used in restorative neurosurgery to satisfy the user's needs for both aesthetics and ease of control while also providing the highest possible degree of dexterity.

Brain↗

Exploiting temporal information in functional magnetic resonance imaging brain data.

Functional Magnetic Resonance Imaging(fMRI) has enabled scientists to look into the active human brain, leading to a flood of new data, thus encouraging the development of new data analysis methods. In this paper, we contribute a comprehensive framework for spatial and temporal exploration of fMRI data, and apply it to a challenging case study: separating drug addicted subjects from healthy non-drug-using controls. To our knowledge, this is the first time that learning on fMRI data is performed explicitly on temporal information for classification in such applications. Experimental results demonstrate that, by selecting discriminative features, group classification can be successfully performed on our case study although training data are exceptionally high dimensional, sparse and noisy fMRI sequences. The classification performance can be significantly improved by incorporating temporal information into machine learning. Both statistical and neuroscientific validation of the method's generalization ability are provided. We demonstrate that incorporation of computer science principles into functional neuroimaging clinical studies, facilitates deduction about the behavioral probes from the brain activation data, thus providing a valid tool that incorporates objective brain imaging data into clinical classification of psychopathologies and identification of genetic vulnerabilities.

Algorithms↗

Support vector machines for predicting the specificity of GalNAc-transferase.

Support Vector Machines (SVMs) which is one kind of learning machines, was applied to predict the specificity of GalNAc-transferase. The examination for the self-consistency and the jackknife test of the SVMs method were tested for the training dataset (305 oligopeptides), the correct rate of self-consistency and jackknife test reaches 100% and 84.9%, respectively. Furthermore, the prediction of the independent testing dataset (30 oligopeptides) was tested, the rate reaches 76.67%.

Algorithms↗

Data mining as a tool for research and knowledge development in nursing.

The ability to collect and store data has grown at a dramatic rate in all disciplines over the past two decades. Healthcare has been no exception. The shift toward evidence-based practice and outcomes research presents significant opportunities and challenges to extract meaningful information from massive amounts of clinical data to transform it into the best available knowledge to guide nursing practice. Data mining, a step in the process of Knowledge Discovery in Databases, is a method of unearthing information from large data sets. Built upon statistical analysis, artificial intelligence, and machine learning technologies, data mining can analyze massive amounts of data and provide useful and interesting information about patterns and relationships that exist within the data that might otherwise be missed. As domain experts, nurse researchers are in ideal positions to use this proven technology to transform the information that is available in existing data repositories into useful and understandable knowledge to guide nursing practice and for active interdisciplinary collaboration and research.

Algorithms↗

Neural-network-based adaptive UPFC for improving transient stability performance of power system.

This paper uses the recently proposed H(infinity)-learning method, for updating the parameter of the radial basis function neural network (RBFNN) used as a control scheme for the unified power flow controller (UPFC) to improve the transient stability performance of a multimachine power system. The RBFNN uses a single neuron architecture whose input is proportional to the difference in error and the updating of its parameters is carried via a proportional value of the error. Also, the coefficients of the difference of error, error, and auxiliary signal used for improving damping performance are depicted by a genetic algorithm. The performance of the newly designed controller is evaluated in a four-machine power system subjected to different types of disturbances. The newly designed single-neuron RBFNN-based UPFC exhibits better damping performance compared to the conventional PID as well as the extended Kalman filter (EKF) updating-based RBFNN scheme, making the unstable cases stable. Its simple architecture reduces the computational burden, thereby making it attractive for real-time implementation. Also, all the machines are being equipped with the conventional power system stabilizer (PSS) to study the coordinated effect of UPFC and PSS in the system.

Algorithms↗

Simple decision rules for classifying human cancers from gene expression profiles.

MOTIVATION: Various studies have shown that cancer tissue samples can be successfully detected and classified by their gene expression patterns using machine learning approaches. One of the challenges in applying these techniques for classifying gene expression data is to extract accurate, readily interpretable rules providing biological insight as to how classification is performed. Current methods generate classifiers that are accurate but difficult to interpret. This is the trade-off between credibility and comprehensibility of the classifiers. Here, we introduce a new classifier in order to address these problems. It is referred to as k-TSP (k-Top Scoring Pairs) and is based on the concept of 'relative expression reversals'. This method generates simple and accurate decision rules that only involve a small number of gene-to-gene expression comparisons, thereby facilitating follow-up studies. RESULTS: In this study, we have compared our approach to other machine learning techniques for class prediction in 19 binary and multi-class gene expression datasets involving human cancers. The k-TSP classifier performs as efficiently as Prediction Analysis of Microarray and support vector machine, and outperforms other learning methods (decision trees, k-nearest neighbour and naïve Bayes). Our approach is easy to interpret as the classifier involves only a small number of informative genes. For these reasons, we consider the k-TSP method to be a useful tool for cancer classification from microarray gene expression data. AVAILABILITY: The software and datasets are available at http://www.ccbm.jhu.edu CONTACT: actan@jhu.edu.

Algorithms↗

Support vector machines for predicting protein structural class.

BACKGROUND: We apply a new machine learning method, the so-called Support Vector Machine method, to predict the protein structural class. Support Vector Machine method is performed based on the database derived from SCOP, in which protein domains are classified based on known structures and the evolutionary relationships and the principles that govern their 3-D structure. RESULTS: High rates of both self-consistency and jackknife tests are obtained. The good results indicate that the structural class of a protein is considerably correlated with its amino acid composition. CONCLUSIONS: It is expected that the Support Vector Machine method and the elegant component-coupled method, also named as the covariant discrimination algorithm, if complemented with each other, can provide a powerful computational tool for predicting the structural classes of proteins.

Algorithms↗

Integrated analysis of plasma metabolomics and proteomics reveals the biological characteristics of damp-heat and stasis-toxin syndrome in colorectal cancer.

OBJECTIVE: To investigate the biological attributes of core syndromes in colorectal cancer, namely, the damp-heat and stasis-toxin syndrome (SRYD). METHODS: Between October 2021 and October 2022, a cohort comprising 40 patients with colorectal cancer (CRC) diagnosed with damp-heat and stasis-toxin syndrome (SRYD group), 40 patients with CRC without this syndrome (non-SRYD group), and 40 healthy controls (Normal group) was recruited at Jiangsu Province Hospital of Chinese Medicine. Untargeted metabolomics analysis was conducted on plasma samples from all 120 participants, while differential protein analysis using four-dimensional data-independent acquisition proteomics was performed on 20 randomly selected samples per group. A combined analysis of proteomics and metabolomics data followed, and the identified potential diagnostic biomarkers were subsequently used to train and validate multiple machine learning models. RESULTS: Proteomic analysis revealed 130 differential proteins in the colorectal cancer with damp-heat and stasis-toxin syndrome (CRC-SRYD) group, enriched in pathways including complement and coagulation cascades, as well as nuclear factor kappa-B (NF-κB) signaling. Metabolomic analysis identified 584 differential metabolites within the same group, showing enrichment in pathways such as primary bile acid biosynthesis, central carbon metabolism in cancer, and glucagon signaling. Integrated pathway analysis indicated heightened activity of the NF-κB signaling pathway in the CRC-SRYD group. A biomarker panel, comprising 6 proteins and 9 metabolites selected through the ReliefF algorithm, was used to construct a diagnostic model with random forest, achieving an accuracy of 93.33%, sensitivity of 80.00%, and specificity of 100%. CONCLUSION: This study systematically elucidates plasma metabolomic and proteomic alterations in patients with CRC, establishing a robust diagnostic model for CRC syndrome (CRC-SRYD). Further investigation is warranted to clarify the underlying molecular mechanisms and biological foundations.

Humans↗

A Bayesian approach to joint feature selection and classifier design.

This paper adopts a Bayesian approach to simultaneously learn both an optimal nonlinear classifier and a subset of predictor variables (or features) that are most relevant to the classification task. The approach uses heavy-tailed priors to promote sparsity in the utilization of both basis functions and features; these priors act as regularizers for the likelihood function that rewards good classification on the training data. We derive an expectation-maximization (EM) algorithm to efficiently compute a maximum a posteriori (MAP) point estimate of the various parameters. The algorithm is an extension of recent state-of-the-art sparse Bayesian classifiers, which in turn can be seen as Bayesian counterparts of support vector machines. Experimental comparisons using kernel classifiers demonstrate both parsimonious feature selection and excellent classification accuracy on a range of synthetic and benchmark data sets.

Algorithms↗

Dimension reduction-based penalized logistic regression for cancer classification using microarray data.

The use of penalized logistic regression for cancer classification using microarray expression data is presented. Two dimension reduction methods are respectively combined with the penalized logistic regression so that both the classification accuracy and computational speed are enhanced. Two other machine-learning methods, support vector machines and least-squares regression, have been chosen for comparison. It is shown that our methods have achieved at least equal or better results. They also have the advantage that the output probability can be explicitly given and the regression coefficients are easier to interpret. Several other aspects, such as the selection of penalty parameters and components, pertinent to the application of our methods for cancer classification are also discussed.

Algorithms↗

An evaluation of contrast enhancement techniques for mammographic breast masses.

The main aim of this paper is to propose a novel set of metrics that measure the quality of the image enhancement of mammographic images in a computer-aided detection framework aimed at automatically finding masses using machine learning techniques. Our methodology includes a novel mechanism for the combination of the metrics proposed into a single quantitative measure. We have evaluated our methodology on 200 images from the publicly available digital database for screening mammograms. We show that the quantitative measures help us select the best suited image enhancement on a per mammogram basis, which improves the quality of subsequent image segmentation much better than using the same enhancement method for all mammograms.

Algorithms↗