PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Relevant information for decision support systems: application in cardiology.

In the paper we show information theory tools for extracting relevant information for decision support systems from medical databases. Each proposed algorithm for selecting a set of relevant features has a specific score function defined by means of information-theoretical characteristics. Then algorithms are classified according to the primary criterion, that can lead to influence-preferring algorithms or weight-preferring algorithms. Other type of classification can be based on the way of selecting of features as forward, backward or combined algorithms. The software package called CORE (COnstitution and REduction) that supports the process of selection of features relevant for a decision making problem is described. Application on data about 1417 middle age men collected in the twenty years lasting interventional study of cardiovascular risk factors in middle aged men and for decision support in primary care are shown. However, the methodology presented is applicable for any decision making problem where extracting relevant information from data is required.

Adult↗

Discrete serum protein signatures discriminate between human retrovirus-associated hematologic and neurologic disease.

The human T-cell leukemia virus type I (HTLV-I) is the causative agent for adult T-cell leukemia (ATL) and HTLV-I-associated myelopathy/tropical spastic paraparesis (HAM/TSP). Approximately 5% of infected individuals will develop either disease and currently there are no diagnostic tools for early detection or accurate assessment of disease state. We have employed high-throughput expression profiling of serum proteins using mass spectrometry to identify protein expression patterns that can discern between disease states of HTLV-I-infected individuals. Our study group consisted of 42 ATL, 50 HAM/TSP, and 38 normal controls. Spectral peaks corresponding to peptide ions were generated from MS-TOF data. We applied Classification and Regression Tree analysis to build a decision algorithm, which achieved 77% correct classification rate across the three groups. A second cohort of 10 ATL, 10 HAM and 10 control samples was used to validate this result. Linear discriminate analysis was performed to verify and visualize class separation. Affinity and sizing chromatography coupled with tandem mass spectrometry was used to identify three peaks specifically overexpressed in ATL: an 11.7 kDa fragment of alpha trypsin inhibitor, and two contiguous fragments (19.9 and 11.9 kDa) of haproglobin-2. To the best of our knowledge, this is the first application of protein profiling to distinguish between two disease states resulting from a single infectious agent.

Adolescent↗

Classification of gasoline data obtained by gas chromatography using a piecewise alignment algorithm combined with feature selection and principal component analysis.

A fast and objective chemometric classification method is developed and applied to the analysis of gas chromatography (GC) data from five commercial gasoline samples. The gasoline samples serve as model mixtures, whereas the focus is on the development and demonstration of the classification method. The method is based on objective retention time alignment (referred to as piecewise alignment) coupled with analysis of variance (ANOVA) feature selection prior to classification by principal component analysis (PCA) using optimal parameters. The degree-of-class-separation is used as a metric to objectively optimize the alignment and feature selection parameters using a suitable training set thereby reducing user subjectivity, as well as to indicate the success of the PCA clustering and classification. The degree-of-class-separation is calculated using Euclidean distances between the PCA scores of a subset of the replicate runs from two of the five fuel types, i.e., the training set. The unaligned training set that was directly submitted to PCA had a low degree-of-class-separation (0.4), and the PCA scores plot for the raw training set combined with the raw test set failed to correctly cluster the five sample types. After submitting the training set to piecewise alignment, the degree-of-class-separation increased (1.2), but when the same alignment parameters were applied to the training set combined with the test set, the scores plot clustering still did not yield five distinct groups. Applying feature selection to the unaligned training set increased the degree-of-class-separation (4.8), but chemical variations were still obscured by retention time variation and when the same feature selection conditions were used for the training set combined with the test set, only one of the five fuels was clustered correctly. However, piecewise alignment coupled with feature selection yielded a reasonably optimal degree-of-class-separation for the training set (9.2), and when the same alignment and ANOVA parameters were applied to the training set combined with the test set, the PCA scores plot correctly classified the gasoline fingerprints into five distinct clusters.

Algorithms↗

Cardiac arrhythmia classification using autoregressive modeling.

BACKGROUND: Computer-assisted arrhythmia recognition is critical for the management of cardiac disorders. Various techniques have been utilized to classify arrhythmias. Generally, these techniques classify two or three arrhythmias or have significantly large processing times. A simpler autoregressive modeling (AR) technique is proposed to classify normal sinus rhythm (NSR) and various cardiac arrhythmias including atrial premature contraction (APC), premature ventricular contraction (PVC), superventricular tachycardia (SVT), ventricular tachycardia (VT) and ventricular fibrillation (VF). METHODS: AR Modeling was performed on ECG data from normal sinus rhythm as well as various arrhythmias. The AR coefficients were computed using Burg's algorithm. The AR coefficients were classified using a generalized linear model (GLM) based algorithm in various stages. RESULTS: AR modeling results showed that an order of four was sufficient for modeling the ECG signals. The accuracy of detecting NSR, APC, PVC, SVT, VT and VF were 93.2% to 100% using the GLM based classification algorithm. CONCLUSION: The results show that AR modeling is useful for the classification of cardiac arrhythmias, with reasonably high accuracies. Further validation of the proposed technique will yield acceptable results for clinical implementation.

Arrhythmias, Cardiac↗

Simple models for estimating dementia severity using machine learning.

Estimating dementia severity using the Clinical Dementia Rating (CDR) Scale is a two-stage process that currently is costly and impractical in community settings, and at best has an interrater reliability of 80%. Because staging of dementia severity is economically and clinically important, we used Machine Learning (ML) algorithms with an Electronic Medical Record (EMR) to identify simpler models for estimating total CDR scores. Compared to a gold standard, which required 34 attributes to derive total CDR scores, ML algorithms identified models with as few as seven attributes. The classification accuracy varied with the algorithm used with naïve Bayes giving the highest. (76%) The mildly demented severity class was the only one with significantly reduced accuracy (59%). If one groups the severity classes into normal, very mild-to-mildly demented, and moderate-to-severely demented, then classification accuracies are clinically acceptable (85%). These simple models can be used in community settings where it is currently not possible to estimate dementia severity due to time and cost constraints.

Algorithms↗

Dynamic topology representing networks.

In the present paper, we propose a new algorithm, namely the Dynamic Topology Representing Networks (DTRN) for learning both topology and clustering information from input data. In contrast to other models with adaptive architecture of this kind, the DTRN algorithm adaptively grows the number of output nodes by applying a vigilance test. The clustering procedure is based on a winner-take-quota learning strategy in conjunction with an annealing process in order to minimize the associated mean square error. A competitive Hebbian rule is applied to learn the global topology information concurrently with the clustering process. The topology information learned is also utilized for dynamically deleting the nodes and for the annealing process. Properties of the DTRN algorithm will be discussed. Extensive simulations will be provided to characterize the effectiveness of the new algorithm in topology preserving, learning speed, and classification tasks as compared to other algorithms of the same nature.

Algorithms↗

Cervical precancer detection using a multivariate statistical algorithm based on laser-induced fluorescence spectra at multiple excitation wavelengths.

A portable fluorimeter was developed and utilized to acquire fluorescence spectra from 381 cervical sites in 95 patients at 337, 380 and 460 nm excitation immediately prior to colposcopy. A multivariate statistical algorithm was used to extract clinically useful information from tissue spectra acquired in vivo. Two full-parameter algorithms were developed using tissue fluorescence emission spectra at all three excitation wavelengths (161 excitation-emission wavelength pairs) for cervical precancer (squamous intraepithelial lesion [SIL]) detection: a screening algorithm that discriminates between SIL and non-SIL with a sensitivity of 82 +/- 1.4% and specificity of 68 +/- 0.0%, and a diagnostic algorithm that differentiates high-grade SIL from non-high-grade SIL with a sensitivity and specificity of 79 +/- 2% and 78 +/- 6%, respectively. Multivariate statistical analysis was also employed to reduce the number of fluorescence excitation-emission wavelength pairs needed to redevelop algorithms that demonstrate a minimum decrease in classification accuracy. Two reduced-parameter algorithms that employ fluorescence intensities at only 15 excitation-emission wavelength pairs were developed: the screening algorithm differentiates SIL from non-SIL with a sensitivity of 84 +/- 1.5% and specificity of 65 +/- 2% and the diagnostic algorithm discriminates high-grade SIL from non-high-grade SIL with a sensitivity and specificity of 78 +/- 0.7% and 74 +/- 2%, respectively. Both the full-parameter and reduced-parameter screening algorithms discriminate between SIL and non-SIL with a similar specificity (+/-5%) and a substantially improved sensitivity relative to Pap smear screening. A comparison of the full-parameter and reduced-parameter diagnostic algorithms to colposcopy in expert hands indicates that all three have a very similar sensitivity and specificity for differentiating high-grade SIL from non-high-grade SIL.

Algorithms↗

Diagnosis of Burkitt lymphoma in due time: a practical approach.

The quick diagnosis of Burkitt lymphoma (BL) and its clear-cut differentiation from diffuse large B-cell lymphoma (DLBCL) is of great clinical importance because treatment strategies for these two disease entities differ markedly. As these two lymphomas are difficult to distinguish using the current World Health Organization classification, we studied 39 cases of highly proliferative peripheral blastic B-cell lymphoma (HPBCL) to establish a practical differential-diagnostic algorithm. Characteristics set for BL were a typical morphology, a mature B-cell phenotype of CD10+, Bcl-6+ and Bcl-2- tumour cells, a proliferation rate of >95%, and the presence of C-MYC rearrangements in the absence of t(14;18)(q32;q21). Altogether, these characteristics were found in only five of 39 cases, whereas the majority of tumours revealed mosaic features. We then followed a pragmatic stepwise approach for a classification algorithm that included the assessment of C-MYC status to stratify HPBCL into four predefined diagnostic categories (DC), namely DC I (5/39, 12.8%): 'classical BL', DC II (11/39, 28.2%): 'atypical BL', DC III (9/39, 23.1%): 'C-MYC+ DLBCL' and DC IV (14/39, 35.9%): 'C-MYC- HPBCL'. This proposal may serve as a robust and objective operational basis for therapeutic decisions for HPBCL within 1 week and is applicable to be evaluated for its prognostic relevance in clinical trials with uniformly treated patients.

Adolescent↗

Reliable classification of two-class cancer data using evolutionary algorithms.

In the area of bioinformatics, the identification of gene subsets responsible for classifying available disease samples to two or more of its variants is an important task. Such problems have been solved in the past by means of unsupervised learning methods (hierarchical clustering, self-organizing maps, k-mean clustering, etc.) and supervised learning methods (weighted voting approach, k-nearest neighbor method, support vector machine method, etc.). Such problems can also be posed as optimization problems of minimizing gene subset size to achieve reliable and accurate classification. The main difficulties in solving the resulting optimization problem are the availability of only a few samples compared to the number of genes in the samples and the exorbitantly large search space of solutions. Although there exist a few applications of evolutionary algorithms (EAs) for this task, here we treat the problem as a multiobjective optimization problem of minimizing the gene subset size and minimizing the number of misclassified samples. Moreover, for a more reliable classification, we consider multiple training sets in evaluating a classifier. Contrary to the past studies, the use of a multiobjective EA (NSGA-II) has enabled us to discover a smaller gene subset size (such as four or five) to correctly classify 100% or near 100% samples for three cancer samples (Leukemia, Lymphoma, and Colon). We have also extended the NSGA-II to obtain multiple non-dominated solutions discovering as much as 352 different three-gene combinations providing a 100% correct classification to the Leukemia data. In order to have further confidence in the identification task, we have also introduced a prediction strength threshold for determining a sample's belonging to one class or the other. All simulation results show consistent gene subset identifications on three disease samples and exhibit the flexibilities and efficacies in using a multiobjective EA for the gene subset identification task.

Algorithms↗

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans↗

Machine learning approaches for prediction of linear B-cell epitopes on proteins.

Identification and characterization of antigenic determinants on proteins has received considerable attention utilizing both, experimental as well as computational methods. For computational routines mostly structural as well as physicochemical parameters have been utilized for predicting the antigenic propensity of protein sites. However, the performance of computational routines has been low when compared to experimental alternatives. Here we describe the construction of machine learning based classifiers to enhance the prediction quality for identifying linear B-cell epitopes on proteins. Our approach combines several parameters previously associated with antigenicity, and includes novel parameters based on frequencies of amino acids and amino acid neighborhood propensities. We utilized machine learning algorithms for deriving antigenicity classification functions assigning antigenic propensities to each amino acid of a given protein sequence. We compared the prediction quality of the novel classifiers with respect to established routines for epitope scoring, and tested prediction accuracy on experimental data available for HIV proteins. The major finding is that machine learning classifiers clearly outperform the reference classification systems on the HIV epitope validation set.

Algorithms↗

Prediction of gene expression using histone modification patterns extracted by Particle Swarm Optimization.

MOTIVATION: Histone modifications play an important role in transcription regulation. Although the general importance of some histone modifications for transcription regulation has been previously established, the relevance of others and their interaction is subject to ongoing research. By training Machine Learning models to predict a gene's expression and explaining their decision making process, we can get hints on how histone modifications affect transcription. In previous studies, trained models were either hardly explainable or the models were trained solely on the abundance of histone modifications. Based on other studies, which used histone modification patterns, rather than their abundance, to identify potential regulatory elements, we hypothesize the histone modification pattern in a gene's promoter to be more predictive for gene expression. We used an optimization algorithm to extract predictive histone modification profiles. RESULTS: Our algorithm called PatternChrome achieved an average area under curve (AUC) score of 0.9029 over 56 samples for binary classification, outperforming all previous algorithms for the same task. We explained the models decisions to deduce the effect of specific features, certain histone modifications or promoter positions on transcription regulation. Although the predictive histone modification patterns were extracted for each sample separately, they can be used to predict gene expression in other samples, implying that the created patterns are largely generalizable. Interestingly, the impact of histone modifications on gene regulation appears predominantly indifferent to cellular specificity. Through explanation of the classifier's decisions, we substantiate established literature knowledge while concurrently revealing novel insights into the intricate landscape of transcriptional regulation via histone modification. AVAILABILITY AND IMPLEMENTATION: The code for the PatternChrome algorithm, the scripts for the analyses and the required data can be found at (https://gitlab.gwdg.de/MedBioinf/generegulation/patternchrome).

Humans↗

Prediction of pancreatic cancer by serum biomarkers using surface-enhanced laser desorption/ionization-based decision tree classification.

OBJECTIVE: In order to improve the prognosis of pancreatic cancer patients, it is crucial to explore novel tools for its early diagnosis. Here, we attempted to screen serum biomarkers to distinguish pancreatic cancer from non-cancer individuals. METHODS: 47 serum samples from pancreatic cancer patients, 39 of whom had small surgically resectable cancers, were collected before surgery, and an additional 53 serum samples from age- and sex-matched individuals without cancer were used as controls. The surface-enhanced laser desorption/ionization (SELDI) ProteinChip was applied to analyze serum protein profiling. 54 samples (27 with pancreatic cancer and 27 controls) were analyzed in the training set by a decision tree algorithm to be able to separate pancreatic cancer from controls. A double-blind test was used to determine the sensitivity and specificity of the classification model. RESULTS: A panel of six biomarkers was selected to set up a decision tree as the classification model. The model separated effectively pancreatic cancer from control samples, achieving a sensitivity of 88.9% and a specificity of 74.1%. The double-blind test challenged the model with a sensitivity of 80% and a specificity of 84.6%. CONCLUSION: The SELDI ProteinChip combined with an artificial intelligence classification algorithm shows great potential for the diagnosis of pancreatic cancer.

Adult↗

Building an asynchronous web-based tool for machine learning classification.

Various unsupervised and supervised learning methods including support vector machines, classification trees, linear discriminant analysis and nearest neighbor classifiers have been used to classify high-throughput gene expression data. Simpler and more widely accepted statistical tools have not yet been used for this purpose, hence proper comparisons between classification methods have not been conducted. We developed free software that implements logistic regression with stepwise variable selection as a quick and simple method for initial exploration of important genetic markers in disease classification. To implement the algorithm and allow our collaborators in remote locations to evaluate and compare its results against those of other methods, we developed a user-friendly asynchronous web-based application with a minimal amount of programming using free, downloadable software tools. With this program, we show that classification using logistic regression can perform as well as other more sophisticated algorithms, and it has the advantages of being easy to interpret and reproduce. By making the tool freely and easily available, we hope to promote the comparison of classification methods. In addition, we believe our web application can be used as a model for other bioinformatics laboratories that need to develop web-based analysis tools in a short amount of time and on a limited budget.

Algorithms↗

An expert diagnostic system based on neural networks and image analysis techniques in the field of automated cytogenetics.

In this study, we introduce an expert system for intelligent chromosome recognition and classification based on artificial neural networks (ANN) and features obtained by automated image analysis techniques. A microscope equipped with a CCTV camera, integrated with an IBM-PC compatible computer environment including a frame grabber, is used for image data acquisition. Features of the chromosomes are obtained directly from the digital chromosome images. Two new algorithms for automated object detection and object skeletonizing constitute the basis of the feature extraction phase which constructs the components of the input vector to the ANN part of the system. This first version of our intelligent diagnostic system uses a trained unsupervised neural network structure and an original rule-based classification algorithm to find a karyotyped form of randomly distributed chromosomes over a complete metaphase. We investigate the effects of network parameters on the classification performance and discuss the adaptability and flexibility of the neural system in order to reach a structure giving an output including information about both structural and numerical abnormalities. Moreover, the classification performances of neural and rule-based system are compared for each class of chromosome.

Algorithms↗

A comparison of siRNA efficacy predictors.

Short interfering RNA (siRNA) efficacy prediction algorithms aim to increase the probability of selecting target sites that are applicable for gene silencing by RNA interference. Many algorithms have been published recently, and they base their predictions on such different features as duplex stability, sequence characteristics, mRNA secondary structure, and target site uniqueness. We compare the performance of the algorithms on a collection of publicly available siRNAs. First, we show that our regularized genetic programming algorithm GPboost appears to have a higher and more stable performance than other algorithms on the collected datasets. Second, several algorithms gave close to random classification on unseen data, and only GPboost and three other algorithms have a reasonably high and stable performance on all parts of the dataset. Third, the results indicate that the siRNAs' sequence is sufficient input to siRNA efficacy algorithms, and that other features that have been suggested to be important may be indirectly captured by the sequence.

Algorithms↗

Machine learning for the quality of life in inflammatory bowel disease.

Presence of a chronic disease influences patients' lives and reinforces demands to accept and then cope with the illness. In the case of inflammatory bowel disease, quality of life greatly differs through phases of remissions and relapses. Could the quality of life questionnaire tell the difference? In this study we are disclosing possibilities of assessing patients' perspectives by analysing analogue scale statements regarding concerns and worries related to ulcerative colitis. Some two hundred Swedish patients, 3/4 in remission and 1/4 in relapse, filled out a booklet containing 36 statements. To characterise the disease activity, we have used multivariate discrimination. To structure and describe in details paths distinguishing the remission from relapse, we have used an artificial intelligence procedure. Applications of the CART (Classification And Regression Trees) algorithm resulted in a set of classifiers which are, based on the similar subsets of significant variables, i.e. statements. Best reached classification accuracy did not exceed 80% in any case. Other classifiers namely, K-nearest-neighbour (KNN), Learning Vector Quantization (LVQ) and Back Propagation Neural Network (BPNN) confirmed that outcome. An expectation that the disease activity should clearly speak throughout the questionnaire held for a certain number of the observations such as pain and suffering, loss of bowel control, dying early, feeling alone, ability to have children, being treated as different and concerns regarding the medication. To highlight the difference of incorrect 20%, K-means clustering was performed. The results settled a basis for a hypothesis that the studied quality of life instrument captures more than the disease activity.

Algorithms↗

Pulmonary CT image classification with evolutionary programming.

RATIONALE AND OBJECTIVES: It is often difficult to classify information in medical images from derived features. The purpose of this research was to investigate the use of evolutionary programming as a tool for selecting important features and generating algorithms to classify computed tomographic (CT) images of the lung. MATERIALS AND METHODS: Training and test sets consisting of 11 features derived from multiple lung CT images were generated, along with an indicator of the target area from which features originated. The images included five parameters based on histogram analysis, 11 parameters based on run length and co-occurrence matrix measures, and the fractal dimension. Two classification experiments were performed. In the first, the classification task was to distinguish between the subtle but known differences between anterior and posterior portions of transverse lung CT sections. The second classification task was to distinguish normal lung CT images from emphysematous images. The performance of the evolutionary programming approach was compared with that of three statistical classifiers that used the same training and test sets. RESULTS: Evolutionary programming produced solutions that compared favorably with those of the statistical classifiers. In separating the anterior from the posterior lung sections, the evolutionary programming results were better than two of the three statistical approaches. The evolutionary programming approach correctly identified all the normal and abnormal lung images and accomplished this by using less features than the best statistical method. CONCLUSION: The results of this study demonstrate the utility of evolutionary programming as a tool for developing classification algorithms.

Algorithms↗