PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

A scalable machine-learning approach to recognize chemical names within large text databases.

MOTIVATION: The use or study of chemical compounds permeates almost every scientific field and in each of them, the amount of textual information is growing rapidly. There is a need to accurately identify chemical names within text for a number of informatics efforts such as database curation, report summarization, tagging of named entities and keywords, or the development/curation of reference databases. RESULTS: A first-order Markov Model (MM) was evaluated for its ability to distinguish chemical names from words, yielding approximately 93% recall in recognizing chemical terms and approximately 99% precision in rejecting non-chemical terms on smaller test sets. However, because total false-positive events increase with the number of words analyzed, the scalability of name recognition was measured by processing 13.1 million MEDLINE records. The method yielded precision ranges from 54.7% to 100%, depending upon the cutoff score used, averaging 82.7% for approximately 1.05 million putative chemical terms extracted. Extracted chemical terms were analyzed to estimate the number of spelling variants per term, which correlated with the total number of times the chemical name appeared in MEDLINE. This variability in term construction was found to affect both information retrieval and term mapping when using PubMed and Ovid.

Computational Biology↗

Automating the assignment of diagnosis codes to patient encounters using example-based and machine learning techniques.

OBJECTIVE: Human classification of diagnoses is a labor intensive process that consumes significant resources. Most medical practices use specially trained medical coders to categorize diagnoses for billing and research purposes. METHODS: We have developed an automated coding system designed to assign codes to clinical diagnoses. The system uses the notion of certainty to recommend subsequent processing. Codes with the highest certainty are generated by matching the diagnostic text to frequent examples in a database of 22 million manually coded entries. These code assignments are not subject to subsequent manual review. Codes at a lower certainty level are assigned by matching to previously infrequently coded examples. The least certain codes are generated by a naïve Bayes classifier. The latter two types of codes are subsequently manually reviewed. MEASUREMENTS: Standard information retrieval accuracy measurements of precision, recall and f-measure were used. Micro- and macro-averaged results were computed. RESULTS At least 48% of all EMR problem list entries at the Mayo Clinic can be automatically classified with macro-averaged 98.0% precision, 98.3% recall and an f-score of 98.2%. An additional 34% of the entries are classified with macro-averaged 90.1% precision, 95.6% recall and 93.1% f-score. The remaining 18% of the entries are classified with macro-averaged 58.5%. CONCLUSION: Over two thirds of all diagnoses are coded automatically with high accuracy. The system has been successfully implemented at the Mayo Clinic, which resulted in a reduction of staff engaged in manual coding from thirty-four coders to seven verifiers.

Abstracting and Indexing↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

Machine learning for the quality of life in inflammatory bowel disease.

Presence of a chronic disease influences patients' lives and reinforces demands to accept and then cope with the illness. In the case of inflammatory bowel disease, quality of life greatly differs through phases of remissions and relapses. Could the quality of life questionnaire tell the difference? In this study we are disclosing possibilities of assessing patients' perspectives by analysing analogue scale statements regarding concerns and worries related to ulcerative colitis. Some two hundred Swedish patients, 3/4 in remission and 1/4 in relapse, filled out a booklet containing 36 statements. To characterise the disease activity, we have used multivariate discrimination. To structure and describe in details paths distinguishing the remission from relapse, we have used an artificial intelligence procedure. Applications of the CART (Classification And Regression Trees) algorithm resulted in a set of classifiers which are, based on the similar subsets of significant variables, i.e. statements. Best reached classification accuracy did not exceed 80% in any case. Other classifiers namely, K-nearest-neighbour (KNN), Learning Vector Quantization (LVQ) and Back Propagation Neural Network (BPNN) confirmed that outcome. An expectation that the disease activity should clearly speak throughout the questionnaire held for a certain number of the observations such as pain and suffering, loss of bowel control, dying early, feeling alone, ability to have children, being treated as different and concerns regarding the medication. To highlight the difference of incorrect 20%, K-means clustering was performed. The results settled a basis for a hypothesis that the studied quality of life instrument captures more than the disease activity.

Algorithms↗

The effect of sample size and disease prevalence on supervised machine learning of narrative data.

This paper examines the independent effects of outcome prevalence and training sample sizes on inductive learning performance. We trained 3 inductive learning algorithms (MC4, IB, and Naïve-Bayes) on 60 simulated datasets of parsed radiology text reports labeled with 6 disease states. Data sets were constructed to define positive outcome states at 4 prevalence rates (1, 5, 10, 25, and 50%) in training set sizes of 200 and 2,000 cases. We found that the effect of outcome prevalence is significant when outcome classes drop below 10% of cases. The effect appeared independent of sample size, induction algorithm used, or class label. Work is needed to identify methods of improving classifier performance when output classes are rare.

Algorithms↗

On selecting features from splice junctions: an analysis using information theoretic and machine learning approaches.

The computational recognition of precise splice junctions is a challenge faced in the analysis of newly sequenced genomes. This is challenging due to the fact that the distribution of sequence patterns in these regions is not always distinct. Our objective is to understand the sequence signatures at the splice junctions, not simply to create an artificial recognition system. We use a combination of a neural network based calliper randomization approach and an information theoretic based feature selection approach for this purpose. This has been done in an effort to understand regions that harbor information content and to extract features relevant for the prediction of splice junctions. The analysis using the neural network based calliper randomization approach revealed regions important in the internal representation of the network model. The calliper approach captured both correlated as well as independently important features. The feature selection approach captures features that are independently informative. The two different methods can capture features with different properties. Comparative analysis of the results using both the methods help to infer about the kind of information present in the region.

Alternative Splicing↗

A machine learning approach to predicting peptide fragmentation spectra.

Accurate peptide identification from tandem mass spectrometry experiments is the cornerstone of proteomics. Although various approaches for matching database sequences with experimental spectra have been developed to date (e.g. Sequest, Mascot) the sensitivity and specificity of peptide identification have not yet reached their full potential. This is in part due to the tradeoffs between robustness and accuracy of the existing methods with respect to the non-uniform nature of peptide fragmentation and bond cleavages induced by different mass spectrometers. Accordingly, it is expected that new approaches to de novo predicting peptide fragmentation spectra will enable more accurate peptide identification. To address this problem, here we used a data-driven approach to learn peptide fragmentation rules in mass spectrometry, in the form of posterior probabilities, for various fragment-ion types of doubly and triply charged precursor ions. We show that the accuracy of our neural-network based methodology is useful for subsequent peptide database searches and that the most useful rules of fragmentation significantly differ across ion and precursor types.

Amino Acids↗

The use of machine learning program LERS-LB 2.5 in knowledge acquisition for expert system development in nursing.

LERS-LB (Learning from Examples using Rough Sets Lower Boundaries) is a computer program based on rough set theory for knowledge acquisition, which extracts patterns from real-world data in generating production rules for expert system development. From LERS-LB evaluation of an SPSS-X data file containing data for recovery room patients, it was concluded that both statistical data files and existing databases can be converted to decision-table format needed by LERS-LB, but it is less desirable to work with statistical files than a well-developed database. It was also concluded that choosing a well-developed database and checking it thoroughly for accuracy and completeness should be done before running LERS-LB, or other learning programs, to avoid problems with data errors. Using rough set theory and a technique called 'dropping conditions' LERS-LB offers, at least in theory, a possible method for identifying which data items are critical to nursing practice. Further research and continued LERS-LB program enhancements still may help with identifying critical data items versus redundant data for nursing practice. LERS-LB, and other learning programs, offer techniques which will help reduce the knowledge acquisition bottleneck in nursing expert system development. It is doubtful, however, that learning programs will eliminate the need for involving domain experts in evaluating rules and expert systems for clinical decision support.

Artificial Intelligence↗

Machine learning in quantitative histopathology.

The role of expert systems functioning as process controllers in learning image understanding systems is discussed. Numeric learning systems already have found a number of applications in cytologic and histopathologic diagnosis. Depending on the required capabilities, systems of increasing complexity are needed. Expert systems to guide scene segmentation in histopathologic imagery require model-based reasoning. Diagnostic image interpretation with learning capability demands a full model of the human expert's competence, including a considerable variety of knowledge representation schemes and inference strategies, coordinated by a meta-process controller.

Artificial Intelligence↗

Two-sample comparison based on prediction error, with applications to candidate gene association studies.

To take advantage of the increasingly available high-density SNP maps across the genome, various tests that compare multilocus genotypes or estimated haplotypes between cases and controls have been developed for candidate gene association studies. Here we view this two-sample testing problem from the perspective of supervised machine learning and propose a new association test. The approach adopts the flexible and easy-to-understand classification tree model as the learning machine, and uses the estimated prediction error of the resulting prediction rule as the test statistic. This procedure not only provides an association test but also generates a prediction rule that can be useful in understanding the mechanisms underlying complex disease. Under the set-up of a haplotype-based transmission/disequilibrium test (TDT) type of analysis, we find through simulation studies that the proposed procedure has the correct type I error rates and is robust to population stratification. The power of the proposed procedure is sensitive to the chosen prediction error estimator. Among commonly used prediction error estimators, the .632+ estimator results in a test that has the best overall performance. We also find that the test using the .632+ estimator is more powerful than the standard single-point TDT analysis, the Pearson's goodness-of-fit test based on estimated haplotype frequencies, and two haplotype-based global tests implemented in the genetic analysis package FBAT. To illustrate the application of the proposed method in population-based association studies, we use the procedure to study the association between non-Hodgkin lymphoma and the IL10 gene.

Adult↗

Active learning with support vector machine applied to gene expression data for cancer classification.

There is growing interest in the application of machine learning techniques in bioinformatics. The supervised machine learning approach has been widely applied to bioinformatics and gained a lot of success in this research area. With this learning approach researchers first develop a large training set, which is a time-consuming and costly process. Moreover, the proportion of the positive examples and negative examples in the training set may not represent the real-world data distribution, which causes concept drift. Active learning avoids these problems. Unlike most conventional learning methods where the training set used to derive the model remains static, the classifier can actively choose the training data and the size of training set increases. We introduced an algorithm for performing active learning with support vector machine and applied the algorithm to gene expression profiles of colon cancer, lung cancer, and prostate cancer samples. We compared the classification performance of active learning with that of passive learning. The results showed that employing the active learning method can achieve high accuracy and significantly reduce the need for labeled training instances. For lung cancer classification, to achieve 96% of the total positives, only 31 labeled examples were needed in active learning whereas in passive learning 174 labeled examples were required. That meant over 82% reduction was realized by active learning. In active learning the areas under the receiver operating characteristic (ROC) curves were over 0.81, while in passive learning the areas under the ROC curves were below 0.50.

Artificial Intelligence↗

Self-organizing learning array.

A new machine learning concept--self-organizing learning array (SOLAR)--is presented. It is a sparsely connected, information theory-based learning machine, with a multilayer structure. It has reconfigurable processing units (neurons) and an evolvable system structure, which makes it an adaptive classification system for a variety of machine learning problems. Its multilayer structure can handle complex problems. Based on the entropy estimation, information theory-based learning is performed locally at each neuron. Neural parameters and connections that correspond to minimum entropy are adaptively set for each neuron. By choosing connections for each neuron, the system sets up its wiring and completes its self-organization. SOLAR classifies input data based on the weighted statistical information from all the neurons. The system classification ability has been simulated and experiments were conducted using test-bench data. Results show a very good performance compared to other classification methods. An important advantage of this structure is its scalability to a large system and ease of hardware implementation on regular arrays of cells.

Learning↗

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans↗

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans↗

Machine learning-assisted plasma PEA proteomics enables differential diagnosis of melancholic depression and bipolar disorder.

Differentiating bipolar disorder (BD) from major depressive disorder (MDD) remains a critical unmet need in psychiatry due to overlapping clinical presentations and the absence of reliable biological markers. In this study, we assessed the capacity of multivariate machine learning models to accurately differentiate BD from MDD with melancholic features using plasma proteomic profiles obtained via Proximity Extension Assay (PEA) technology. A total of 67 participants were included (23 BD, 20 MDD, and 24 HC), and plasma protein expression was assessed using the Olink Target 96 Neurology panel. Differential proteomic analysis revealed distinct disorder-specific expression patterns, identifying 21 differentially expressed proteins in BD versus MDD, 18 in BD versus healthy controls, and 7 in MDD versus healthy controls. Using a stepwise feature reduction strategy, machine learning models were trained on three feature sets comprising all proteins, the top 20 most informative proteins, and the top 5 most beneficial proteins, and evaluated across BD-MDD, BD-HC, and MDD-HC classification tasks using five algorithms. For BD-MDD discrimination, the Random Forest model achieved the highest performance when trained on the top 5 protein set (LXN, HAGH, MATN3, PLXNB1, and CTSC), yielding an AUC of 0.905, with similarly strong performance observed using the top 20 protein set. Feature importance analysis highlighted proteins involved in neurodevelopmental processes, immune regulation, and extracellular matrix organization. Overall, these findings demonstrate that integrating plasma proteomics with machine learning enables robust differentiation between BD and MDD with melancholic features, supporting the development of scalable and biologically informed diagnostic tools for precision psychiatry.

Bipolar disorder↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗