PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

From latent disseminated cells to overt metastasis: genetic analysis of systemic breast cancer progression.

According to the present view, metastasis marks the end in a sequence of genomic changes underlying the progression of an epithelial cell to a lethal cancer. Here, we aimed to find out at what stage of tumor development transformed cells leave the primary tumor and whether a defined genotype corresponds to metastatic disease. To this end, we isolated single disseminated cancer cells from bone marrow of breast cancer patients and performed single-cell comparative genomic hybridization. We analyzed disseminated tumor cells from patients after curative resection of the primary tumor (stage M0), as presumptive progenitors of manifest metastasis, and from patients with manifest metastasis (stage M1). Their genomic data were compared with those from microdissected areas of matched primary tumors. Disseminated cells from M0-stage patients displayed significantly fewer chromosomal aberrations than primary tumors or cells from M1-stage patients (P < 0.008 and P < 0.0001, respectively), and their aberrations appeared to be randomly generated. In contrast, primary tumors and M1 cells harbored different and characteristic chromosomal imbalances. Moreover, applying machine-learning methods for the classification of the genotypes, we could correctly identify the presence or absence of metastatic disease in a patient on the basis of a single-cell genome. We suggest that in breast cancer, tumor cells may disseminate in a far less progressed genomic state than previously thought, and that they acquire genomic aberrations typical of metastatic cells thereafter. Thus, our data challenge the widely held view that the precursors of metastasis are derived from the most advanced clone within the primary tumor.

Algorithms↗

Analysis of alcoholism data using support vector machines.

A supervised learning method, support vector machine, was used to analyze the microsatellite marker dataset of the Collaborative Study on the Genetics of Alcoholism Problem 1 for the Genetic Analysis Workshop 14. Twelve binary-valued phenotype variables were chosen for analyses using the markers from all autosomal chromosomes. Using various polynomial kernel functions of the support vector machine and randomly divided genome regions, we were able to observe the association of some marker sets with the chosen phenotypes and thus reduce the size of the dataset. The successful classifications established with the chosen support vector machine kernel function had high levels of correctness for each prediction, e.g., 96% in the fourfold cross-validations. However, owing to the limited sample data, we were not able to test the predictions of the classifiers in the new sample data.

Alcoholism↗

Selection and combination of machine learning classifiers for prediction of linear B-cell epitopes on proteins.

Recently, new machine learning classifiers for the prediction of linear B-cell epitopes were presented. Here we show the application of Receiver Operator Characteristics (ROC) convex hulls to select optimal classifiers as well as possibilities to improve the post test probability (PTP) to meet real world requirements such as high throughput epitope screening of whole proteomes. The major finding is that ROC convex hulls present an easy to use way to rank classifiers based on their prediction conservativity as well as to select candidates for ensemble classifiers when validating against the antigenicity profile of 10 HIV-1 proteins. We also show that linear models are at least equally efficient to model the available data when compared to multi-layer feed-forward neural networks.

Algorithms↗

Generalized discriminant analysis using a kernel approach.

We present a new method that we call generalized discriminant analysis (GDA) to deal with nonlinear discriminant analysis using kernel function operator. The underlying theory is close to the support vector machines (SVM) insofar as the GDA method provides a mapping of the input vectors into high-dimensional feature space. In the transformed space, linear properties make it easy to extend and generalize the classical linear discriminant analysis (LDA) to nonlinear discriminant analysis. The formulation is expressed as an eigenvalue problem resolution. Using a different kernel, one can cover a wide class of nonlinearities. For both simulated data and alternate kernels, we give classification results, as well as the shape of the decision function. The results are confirmed using real data to perform seed classification.

Algorithms↗

Different classification techniques considering brain computer interface applications.

In this work the application of different machine learning techniques for classification of mental tasks from electroencephalograph (EEG) signals is investigated. The main application for this research is the improvement of brain computer interface (BCI) systems. For this purpose, Bayesian graphical network, neural network, Bayesian quadratic, Fisher linear and hidden Markov model classifiers are applied to two known EEG datasets in the BCI field. The Bayesian network classifier is used for the first time in this work for classification of EEG signals. The Bayesian network appeared to have a significant accuracy and more consistent classification compared to the other four methods. In addition to classical correct classification accuracy criteria, the mutual information is also used to compare the classification results with other BCI groups.

Algorithms↗

Modular DAG-RNN architectures for assembling coarse protein structures.

We develop and test machine learning methods for the prediction of coarse 3D protein structures, where a protein is represented by a set of rigid rods associated with its secondary structure elements (alpha-helices and beta-strands). First, we employ cascades of recursive neural networks derived from graphical models to predict the relative placements of segments. These are represented as discretized distance and angle maps, and the discretization levels are statistically inferred from a large and curated dataset. Coarse 3D folds of proteins are then assembled starting from topological information predicted in the first stage. Reconstruction is carried out by minimizing a cost function taking the form of a purely geometrical potential. We show that the proposed architecture outperforms simpler alternatives and can accurately predict binary and multiclass coarse maps. The reconstruction procedure proves to be fast and often leads to topologically correct coarse structures that could be exploited as a starting point for various protein modeling strategies. The fully integrated rod-shaped protein builder (predictor of contact maps + reconstruction algorithm) can be accessed at http://distill.ucd.ie/.

Algorithms↗

Prediction of contact maps by GIOHMMs and recurrent neural networks using lateral propagation from all four cardinal corners.

MOTIVATION: Accurate prediction of protein contact maps is an important step in computational structural proteomics. Because contact maps provide a translation and rotation invariant topological representation of a protein, they can be used as a fundamental intermediary step in protein structure prediction. RESULTS: We develop a new set of flexible machine learning architectures for the prediction of contact maps, as well as other information processing and pattern recognition tasks. The architectures can be viewed as recurrent neural network implemantations of a class of Bayesian networks we call generalized input-output HMMs (GIOHMMs). For the specific case of contact maps, contextual information is propagated laterally through four hidden planes, one for each cardinal corner. We show that these architectures can be trained from examples and yield contact map predictors that outperform previously reported methods. While several extensions and improvements are in progress, the current version can accurately predict 60.5% of contacts at a distance cutoff of 8 A and 45% of distant contacts at 10 A, for proteins of length up to 300.

Algorithms↗

A comparative study of machine-learning methods to predict the effects of single nucleotide polymorphisms on protein function.

MOTIVATION: The large volume of single nucleotide polymorphism data now available motivates the development of methods for distinguishing neutral changes from those which have real biological effects. Here, two different machine-learning methods, decision trees and support vector machines (SVMs), are applied for the first time to this problem. In common with most other methods, only non-synonymous changes in protein coding regions of the genome are considered. RESULTS: In detailed cross-validation analysis, both learning methods are shown to compete well with existing methods, and to out-perform them in some key tests. SVMs show better generalization performance, but decision trees have the advantage of generating interpretable rules with robust estimates of prediction confidence. It is shown that the inclusion of protein structure information produces more accurate methods, in agreement with other recent studies, and the effect of using predicted rather than actual structure is evaluated. AVAILABILITY: Software is available on request from the authors.

Algorithms↗

The application of artificial intelligence in healthcare practice: A mapping review of systematic reviews.

Artificial intelligence (AI) is rapidly transforming healthcare practice, with growing evidence supporting its use in diagnosis, prognosis, treatment planning, and operational decision-making. The proliferation of systematic reviews in recent years underscores the need for an updated synthesis of the literature to inform research, policy, and practice. We searched PubMed, Web of Science, Scopus, IEEE Xplore, and CINAHL for systematic reviews and meta-analyses published between 2019 and February 2026. Eligible reviews focused on AI applications in healthcare practice, were peer-reviewed, and written in English. A total of 368 reviews met the inclusion criteria. Publication volume increased steadily, peaking in 2025. AI research was concentrated in high-density domains, such as radiology, oncology, and critical care. Across reviews, diagnostic imaging, electronic health record (EHR) data, and biomarkers/laboratory results accounted for 68% of training data sources, though newer data types, such as wearable device and sensor data, emerged from 2022 onward. Diagnosis, prognosis, and treatment comprised over 80% of AI applications, with novel uses emerging in recent years, such as AI-assisted clinical documentation (e.g., ambient documentation tools) and patient education. Ethical concerns were reported in 78.5% of reviews, with privacy, model accuracy, data and algorithmic bias, and explainability as recurrent themes. The proportion of reviews reporting ethical concerns increased from 2021 to 2025. AI applications in healthcare are expanding in scope, diversifying in data sources, and evolving toward novel clinical and operational uses. The human-centered AI or augmented intelligence paradigm, integrating computational precision with clinical expertise, holds significant promise but will require parallel advances in governance, regulatory frameworks, and ethical oversight to ensure safe adoption.

Artificial Intelligence↗

A machine learning approach to the analysis of time-frequency maps, and its application to neural dynamics.

The statistical analysis of experimentally recorded brain activity patterns may require comparisons between large sets of complex signals in order to find meaningful similarities and differences between signals with large variability. High-level representations such as time-frequency maps convey a wealth of useful information, but they involve a large number of parameters that make statistical investigations of many signals difficult at present. In this paper, we describe a method that performs drastic reduction in the complexity of time-frequency representations through a modelling of the maps by elementary functions. The method is validated on artificial signals and subsequently applied to electrophysiological brain signals (local field potential) recorded from the olfactory bulb of rats while they are trained to recognize odours. From hundreds of experimental recordings, reproducible time-frequency events are detected, and relevant features are extracted, which allow further information processing, such as automatic classification.

Algorithms↗

Decision tree induction in the diagnosis of otoneurological diseases.

Expert systems have been applied in medicine as diagnostic aids and education tools. The construction of a knowledge base for an expert system may be a difficult task; to automate this task several machine learning methods have been developed. These methods can be also used in the refinement of knowledge bases for removing inconsistencies and redundancies, and for simplifying decision rules. In this study, decision tree induction was employed to acquire diagnostic knowledge for otoneurological diseases and to extract relevant parameters from the database of an otoneurological expert system ONE. The records of patients with benign positional vertigo, Meniere's disease, sudden deafness, traumatic vertigo, vestibular neuritis and vestibular schwannoma were retrieved from the database of ONE, and for each disease, decision trees were constructed. The study shows that decision tree induction is a useful technique for acquiring diagnostic knowledge for otoneurological diseases and for extracting relevant parameters from a large set of parameters.

Algorithms↗

A machine learning strategy to identify candidate binding sites in human protein-coding sequence.

BACKGROUND: The splicing of RNA transcripts is thought to be partly promoted and regulated by sequences embedded within exons. Known sequences include binding sites for SR proteins, which are thought to mediate interactions between splicing factors bound to the 5' and 3' splice sites. It would be useful to identify further candidate sequences, however identifying them computationally is hard since exon sequences are also constrained by their functional role in coding for proteins. RESULTS: This strategy identified a collection of motifs including several previously reported splice enhancer elements. Although only trained on coding exons, the model discriminates both coding and non-coding exons from intragenic sequence. CONCLUSION: We have trained a computational model able to detect signals in coding exons which seem to be orthogonal to the sequences' primary function of coding for proteins. We believe that many of the motifs detected here represent binding sites for both previously unrecognized proteins which influence RNA splicing as well as other regulatory elements.

Algorithms↗

An Equivalence Between Sparse Approximation and Support Vector Machines.

This article shows a relationship between two different approximation techniques: the support vector machines (SVM), proposed by V. Vapnik (1995) and a sparse approximation scheme that resembles the basis pursuit denoising algorithm (Chen, 1995; Chen, Donoho, and Saunders, 1995). SVM is a technique that can be derived from the structural risk minimization principle (Vapnik, 1982) and can be used to estimate the parameters of several different approximation schemes, including radial basis functions, algebraic and trigonometric polynomials, B-splines, and some forms of multilayer perceptrons. Basis pursuit denoising is a sparse approximation technique in which a function is reconstructed by using a small number of basis functions chosen from a large set (the dictionary). We show that if the data are noiseless, the modified version of basis pursuit denoising proposed in this article is equivalent to SVM in the following sense: if applied to the same data set, the two techniques give the same solution, which is obtained by solving the same quadratic programming problem. In the appendix, we present a derivation of the SVM technique in one framework of regularization theory, rather than statistical learning theory, establishing a connection between SVM, sparse approximation, and regularization theory.

Journal Article↗

Classifying "kinase inhibitor-likeness" by using machine-learning methods.

By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued--support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GA/kNN), and recursive partitioning (RP)--proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.

Algorithms↗

A bottom-up method for simplifying support vector solutions.

The high generalization ability of support vector machines (SVMs) has been shown in many practical applications, however, they are considerably slower in test phase than other learning approaches due to the possibly big number of support vectors comprised in their solution. In this letter, we describe a method to reduce such number of support vectors. The reduction process iteratively selects two nearest support vectors belonging to the same class and replaces them by a newly constructed one. Through the analysis of relation between vectors in input and feature spaces, we present the construction of the new vectors that requires to find the unique maximum point of a one-variable function on (0,1), not to minimize a function of many variables with local minima in previous reduced set methods. Experimental results on real life dataset show that the proposed method is effective in reducing number of support vectors and preserving machine's generalization performance.

Algorithms↗

Learning from imbalanced data in surveillance of nosocomial infection.

OBJECTIVE: An important problem that arises in hospitals is the monitoring and detection of nosocomial or hospital acquired infections (NIs). This paper describes a retrospective analysis of a prevalence survey of NIs done in the Geneva University Hospital. Our goal is to identify patients with one or more NIs on the basis of clinical and other data collected during the survey. METHODS AND MATERIAL: Standard surveillance strategies are time-consuming and cannot be applied hospital-wide; alternative methods are required. In NI detection viewed as a classification task, the main difficulty resides in the significant imbalance between positive or infected (11%) and negative (89%) cases. To remedy class imbalance, we explore two distinct avenues: (1) a new re-sampling approach in which both over-sampling of rare positives and under-sampling of the noninfected majority rely on synthetic cases (prototypes) generated via class-specific sub-clustering, and (2) a support vector algorithm in which asymmetrical margins are tuned to improve recognition of rare positive cases. RESULTS AND CONCLUSION: Experiments have shown both approaches to be effective for the NI detection problem. Our novel re-sampling strategies perform remarkably better than classical random re-sampling. However, they are outperformed by asymmetrical soft margin support vector machines which attained a sensitivity rate of 92%, significantly better than the highest sensitivity (87%) obtained via prototype-based re-sampling.

Algorithms↗

Modelling of classification rules on metabolic patterns including machine learning and expert knowledge.

Machine learning has a great potential to mine potential markers from high-dimensional metabolic data without any a priori knowledge. Exemplarily, we investigated metabolic patterns of three severe metabolic disorders, PAHD, MCADD, and 3-MCCD, on which we constructed classification models for disease screening and diagnosis using a decision tree paradigm and logistic regression analysis (LRA). For the LRA model-building process we assessed the relevance of established diagnostic flags, which have been developed from the biochemical knowledge of newborn metabolism, and compared the models' error rates with those of the decision tree classifier. Both approaches yielded comparable classification accuracy in terms of sensitivity (>95.2%), while the LRA models built on flags showed significantly enhanced specificity. The number of false positive cases did not exceed 0.001%.

Algorithms↗

Application of machine learning to improve the results of high-throughput docking against the HIV-1 protease.

We have previously reported that the application of a Laplacian-modified naive Bayesian (NB) classifier may be used to improve the ranking of known inhibitors from a random database of compounds after High-Throughput Docking (HTD). The method relies upon the frequency of substructural features among the active and inactive compounds from 2D fingerprint information of the compounds. Here we present an investigation of the role of extended connectivity fingerprints in training the NB classifier against HTD studies on the HIV-1 protease using three docking programs: Glide, FlexX, and GOLD. The results show that the performance of the NB classifier is due to the presence of a large number of features common to the set of known active compounds rather than a single structural or substructural scaffold. We demonstrate that the Laplacian-modified naive Bayesian classifier trained with data from high-throughput docking is superior at identifying active compounds from a target database in comparison to conventional two-dimensional substructure search methods alone.

Algorithms↗