PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Classification of a large anticancer data set by adaptive fuzzy partition.

An Adaptive Fuzzy Partition (AFP) algorithm, derived from Fuzzy Logic concepts, was used to classify an anticancer data set, including about 1300 compounds subdivided into eight mechanisms of action. AFP classification builds relationships between molecular descriptors and bio-activities by dynamically dividing the descriptor hyperspace into a set of fuzzy subspaces. These subspaces are described by simple linguistic rules, from which scores ranging between 0 and 1 can be derived. The latter values define, for each compound, the degrees of membership of the different mechanisms analyzed. A particular attention was devoted to develop structure-activity relations that have a real utility. Then, well-defined and widely accepted protocols were used to validate the models by defining their robustness and prediction ability. More particularly, after selecting the most relevant descriptors with help of a genetic algorithm, a training set of 640 compounds was isolated by a rational procedure based on Self-Organizing Maps. The related AFP model was then validated with help of a validation set and, above all, of cross-validation and Y-randomization procedures. Good validation scores of about 80% were obtained, underlining the robustness of the model. Moreover, the prediction ability was evaluated with 374 test compounds that had not been used to establish the model and 77% of them were predicted correctly.

Algorithms↗

Automated lung nodule classification following automated nodule detection on CT: a serial approach.

We have evaluated the performance of an automated classifier applied to the task of differentiating malignant and benign lung nodules in low-dose helical computed tomography (CT) scans acquired as part of a lung cancer screening program. The nodules classified in this manner were initially identified by our automated lung nodule detection method, so that the output of automated lung nodule detection was used as input to automated lung nodule classification. This study begins to narrow the distinction between the "detection task" and the "classification task." Automated lung nodule detection is based on two- and three-dimensional analyses of the CT image data. Gray-level-thresholding techniques are used to identify initial lung nodule candidates, for which morphological and gray-level features are computed. A rule-based approach is applied to reduce the number of nodule candidates that correspond to non-nodules, and the features of remaining candidates are merged through linear discriminant analysis to obtain final detection results. Automated lung nodule classification merges the features of the lung nodule candidates identified by the detection algorithm that correspond to actual nodules through another linear discriminant classifier to distinguish between malignant and benign nodules. The automated classification method was applied to the computerized detection results obtained from a database of 393 low-dose thoracic CT scans containing 470 confirmed lung nodules (69 malignant and 401 benign nodules). Receiver operating characteristic (ROC) analysis was used to evaluate the ability of the classifier to differentiate between nodule candidates that correspond to malignant nodules and nodule candidates that correspond to benign lesions. The area under the ROC curve for this classification task attained a value of 0.79 during a leave-one-out evaluation.

Adult↗

Computer-assisted analysis of medulloblastoma. A cytologic study.

OBJECTIVE: To explore data from a set of cases of medulloblastoma to see whether quantitative image analysis might suggest evidence for the existence of lower and higher grade lesions. STUDY DESIGN: Fourteen consecutive cases of medulloblastoma were obtained. Smears were stained with toluidine blue. For each case, 50 nuclei were measured and a number of densitometric features extracted. RESULTS: The existence of two subgroups of cases, identified as lower and higher grade groups, was suggested by a plot of the total optical density versus nuclear area. Two nuclear texture features--the number of pixels with the same optical density value occurring consecutively in the nucleus and the proportion of pixels in the high optical density range--divided the cases into the same subgroups. The use of a clustering algorithm established two clusters that corresponded to that subgrouping except for one case. Discriminant analysis gave an identical classification, with the misplaced case having a borderline discriminant function score. An unsupervised learning algorithm based on an adaptive distance metric formed two clusters and assigned the borderline case to the low grade subgroup. The grouping obtained by quantitative analysis was only partly related to the grade of nuclear atypia subjectively evaluated. CONCLUSION: In our series of medulloblastomas, quantitative analysis provided a means of detecting differences in the nuclear size and texture that allowed the classification of cases into two subgroups.

Adult↗

Continuous personnel scheduling algorithms: a literature review.

Hospitals frequently use personnel scheduling options as recruiting and retention instruments. The successful application of these personnel scheduling tools, whether developed in-house or purchased from vendors, requires appreciation of the interrelationships of three basic manpower decisions--staffing, personnel scheduling, and allocation. This article introduces these basic relationships and their influence on the development of satisfactory personnel schedules. Next, it reviews published personnel scheduling algorithms, applicable to hospital operations, within the context of the three manpower decisions. It is proposed that scheduling algorithms be classified by type of schedule produced (cyclic or noncyclic); and technique used (heuristic, mathematical programming, or self-scheduling). The characteristics of each classification are discussed. Considerations for the development of new personnel scheduling algorithms are also presented.

Algorithms↗

SMO algorithm for least-squares SVM formulations.

This article extends the well-known SMO algorithm of support vector machines (SVMs) to least-squares SVM formulations that include LS-SVM classification, kernel ridge regression, and a particular form of regularized kernel Fisher discriminant. The algorithm is shown to be asymptotically convergent. It is also extremely easy to implement. Computational experiments show that the algorithm is fast and scales efficiently (quadratically) as a function of the number of examples.

Algorithms↗

A computer-aided diagnostic system to characterize CT focal liver lesions: design and optimization of a neural network classifier.

In this paper, a computer-aided diagnostic (CAD) system for the classification of hepatic lesions from computed tomography (CT) images is presented. Regions of interest (ROIs) taken from nonenhanced CT images of normal liver, hepatic cysts, hemangiomas, and hepatocellular carcinomas have been used as input to the system. The proposed system consists of two modules: the feature extraction and the classification modules. The feature extraction module calculates the average gray level and 48 texture characteristics, which are derived from the spatial gray-level co-occurrence matrices, obtained from the ROIs. The classifier module consists of three sequentially placed feed-forward neural networks (NNs). The first NN classifies into normal or pathological liver regions. The pathological liver regions are characterized by the second NN as cyst or "other disease." The third NN classifies "other disease" into hemangioma or hepatocellular carcinoma. Three feature selection techniques have been applied to each individual NN: the sequential forward selection, the sequential floating forward selection, and a genetic algorithm for feature selection. The comparative study of the above dimensionality reduction methods shows that genetic algorithms result in lower dimension feature vectors and improved classification performance.

Algorithms↗

Classification of in vivo autofluorescence spectra using support vector machines.

An algorithm based on support vector machines (SVM), the most recent advance in pattern recognition, is presented for use in classifying light-induced autofluorescence collected from cancerous and normal tissues. The in vivo autofluorescence spectra used for development and evaluation of SVM diagnostic algorithms were measured from 85 nasopharyngeal carcinoma (NPC) lesions and 131 normal tissue sites from 59 subjects during routine nasal endoscopy. Leave-one-out cross-validation was used to evaluate the performance of the algorithms. An overall diagnostic accuracy of 96%, a sensitivity of 94%, and a specificity of 97% for discriminating nasopharyngeal carcinomas from normal tissues were achieved using a linear SVM algorithm. A diagnostic accuracy of 98%, a sensitivity of 95%, and a specificity of 99% for detecting NPC were achieved with a nonlinear SVM algorithm. In a comparison with previously developed algorithms using the same dataset and the principal component analysis (PCA) technique, the SVM algorithms produced better diagnostic accuracy in all instances. In addition, we investigated a method combining PCA and SVM techniques for reducing the complexity of the SVM algorithms.

Algorithms↗

Training nu-support vector classifiers: theory and algorithms.

The nu-support vector machine (nu-SVM) for classification proposed by Schölkopf, Smola, Williamson, and Bartlett (2000) has the advantage of using a parameter nu on controlling the number of support vectors. In this article, we investigate the relation between nu-SVM and C-SVM in detail. We show that in general they are two different problems with the same optimal solution set. Hence, we may expect that many numerical aspects of solving them are similar. However, compared to regular C-SVM, the formulation of nu-SVM is more complicated, so up to now there have been no effective methods for solving large-scale nu-SVM. We propose a decomposition method for nu-SVM that is competitive with existing methods for C-SVM. We also discuss the behavior of nu-SVM by some numerical experiments.

Journal Article↗

Numerical and chemical classification of Streptosporangium and some related actinomycetes.

One hundred and seventeen streptosporangia from soil were compared with marker strains of the family Streptosporangiaceae for many phenotypic properties. The data were examined using the Jaccard, pattern and simple matching coefficients with clustering achieved using average, complete and single linkage algorithms. Particular confidence was placed in the product of the pattern, average linkage analysis given the sharp definition of aggregate groups and clusters and a combination of low test error and high cophenetic correlation values. The test strains were assigned to five aggregate groups that were equated with the genera Streptosporangium (group A), Microbispora (group B), Planobispora and Planomonospora (Group C), Kutzneria (neé Streptosporangium viridogriseum (group D), and Microtetraspora (group E). The streptosporangia, both isolates and marker strains, were assigned to 5 major, 7 minor and 18 single membered clusters. Representative streptosporangia examined for chemical markers were characterised by the presence of meso-diaminopimelic acid in whole-organism hydrolysates, complex mixtures of straight- and branched chain fatty acids, di- and tetrahydrogenated menaquinones as predominant isoprenologues, and complex polar lipid patterns containing diphosphatidylglycerol, phosphatidylethanolamine, phosphatidylglycerol, phosphatidylmethylethanolamine, phosphatidylinositol, phosphatidylinositol mannosides and uncharacterised components. The chemical and numerical data support the taxonomic integrity of the validly described species of Streptosporangium and suggest that the genus is markedly underspeciated.

Actinomycetales↗

acmgscaler: an R package and Colab for standardized gene-level variant effect score calibration within the ACMG/AMP framework.

MOTIVATION: A genome-wide variant effect calibration method was recently developed under the guidelines of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP), following ClinGen recommendations for variant classification. While genome-wide approaches offer clinical utility, emerging evidence highlights the need for gene- and context-specific calibration to improve accuracy. Building on previous work, we have developed an algorithm tailored to converting functional scores from both multiplexed assays of variant effects (MAVEs) and computational variant effect predictors (VEPs) into ACMG/AMP evidence strengths. RESULTS: Our method is designed to deliver consistent performance across different genes and score distributions, with all variables adaptively determined from the input data, preventing selective adjustments or overfitting that could inflate evidence strengths beyond empirical support. To facilitate adoption, we introduce acmgscaler, a lightweight R package and a plug-and-play Google Colab notebook for the calibration of custom datasets. This algorithmic framework bridges the gap between MAVEs/VEPs and clinically actionable variant classification. AVAILABILITY AND IMPLEMENTATION: The R package and Colab notebook are available at https://github.com/badonyi/acmgscaler.

Software↗

A one-layer model of laser-induced fluorescence for diagnosis of disease in human tissue: applications to atherosclerosis.

This paper describes a general model of tissue fluorescence which can be used both to: 1) determine chemical and physical properties of the tissue, and 2) design an optimal algorithm for clinical diagnosis of tissue composition. This model is based on a picture of tissue as a single, optically thick layer, in which fluorophores and absorbing species are homogeneously distributed. As a specific example, the model is applied to the laser induced fluorescence (LIF) of normal and atherosclerotic human aorta using 476 nm excitation. Methods for determining the relevant attenuation and fluorescence lineshapes are detailed, and these lineshapes are used to apply the model to data from 148 samples. The model parameters are related to the concentrations of the major arterial chromophores: structural proteins, hemoglobin and ceroid. In addition, the model parameters are used to derive diagnostic algorithms for the presence of atherosclerosis. Utilizing a binary classification scheme, the presence or absence of pathology was determined correctly in 88 percent of cases.

Algorithms↗

Bayesian approach to feature selection and parameter tuning for support vector machine classifiers.

A Bayesian point of view of SVM classifiers allows the definition of a quantity analogous to the evidence in probabilistic models. By maximizing this one can systematically tune hyperparameters and, via automatic relevance determination (ARD), select relevant input features. Evidence gradients are expressed as averages over the associated posterior and can be approximated using Hybrid Monte Carlo (HMC) sampling. We describe how a Nyström approximation of the Gram matrix can be used to speed up sampling times significantly while maintaining almost unchanged classification accuracy. In experiments on classification problems with a significant number of irrelevant features this approach to ARD can give a significant improvement in classification performance over more traditional, non-ARD, SVM systems. The final tuned hyperparameter values provide a useful criterion for pruning irrelevant features, and we define a measure of relevance with which to determine systematically how many features should be removed. This use of ARD for hard feature selection can improve classification accuracy in non-ARD SVMs. In the majority of cases, however, we find that in data sets constructed by human domain experts the performance of non-ARD SVMs is largely insensitive to the presence of some less relevant features. Eliminating such features via ARD then does not improve classification accuracy, but leads to impressive reductions in the number of features required, by up to 75%.

Algorithms↗

Acoustic-phonetic features for the automatic classification of fricatives.

In this article, the acoustic-phonetic characteristics of the American English fricative consonants are investigated from the automatic classification standpoint. The features studied in the literature are evaluated and new features are proposed. To test the value of the extracted features, a statistically guided, knowledge-based, acoustic-phonetic system for the automatic classification of fricatives in speaker-independent continuous speech is proposed. The system uses an auditory-based front-end processing system and incorporates new algorithms for the extraction and manipulation of the acoustic-phonetic features that proved to be rich in their information content. Classification experiments are performed using hard-decision algorithms on fricatives extracted from the TIMIT database continuous speech of 60 speakers (not used in the design/training process) from seven different dialects of American English. An accuracy of 93% is obtained for voicing detection, 91% for place of articulation detection, and 87% for the overall classification of fricatives.

Algorithms↗

Fourier transform infrared (FT-IR) spectroscopy in bacteriology: towards a reference method for bacteria discrimination.

Rapid and reliable discrimination among clinically relevant pathogenic organisms is a crucial task in microbiology. Microorganism resistance to antimicrobial agents increases prevalence of infections. The possibility of Fourier transform infrared (FT-IR) spectroscopy to assess the overall molecular composition of microbial cells in a non-destructive manner is reflected in the specific spectral fingerprints highly typical for different microorganisms. With the objective of using FT-IR spectroscopy for discrimination between diverse microbial species and strains on a routine basis, a wide range of chemometrics techniques need to be applied. Still a major issue in using FT-IR for successful bacteria characterization is the method for spectra pre-processing. We analyzed different spectra pre-processing methods and their impact on the reduction of spectral variability and on the increase of robustness of chemometrics models. Different types of the Enterococcus faecium bacterial strain were classified according to chromosomal DNA restriction patterns produced by pulsed-field gel electrophoresis (PFGE). Samples were collected from human patients. Collected FT-IR spectra were used to verify if the same classification was obtained. In order to further optimize bacteria classification we investigated whether a selected combination of the most discriminative spectral regions could improve results. Two different variable selection methods (genetic algorithms (GAs) and bootstrapping) were investigated and their relative merit for bacteria classification is reported by comparing with results obtained using the entire spectra. Discriminant partial least-squares (Di-PLS) models based on corrected spectra showed improved predictive ability up to 40% when compared to equivalent models using the entire spectral range. The uncertainty in estimating scores was reduced by about 50% when compared to models with all wavelengths. Spectral ranges with relevant chemical information for Enterococcus faecium bacteria discrimination were outlined.

Bacterial Typing Techniques↗

Adaptive brightness transfer functions in echocardiography.

Despite the clear advantages of echocardiography as a diagnostic tool, its images tend to be noisy and unclear. This paper presents an innovative algorithm, called ABTF (adaptive brightness transfer function), designed to optimally adjust the gray-levels used in echocardiography. The algorithm is aimed at aiding in visual tissue classification and texture-based visual tissue tracking in echocardiographic images. The ABTF method is based on fitting the cine-loop's gray-level histogram to a sum of three Gaussian functions, each of which relates to a different region within the image, the left ventricular cavity, the relatively dark regions within the cardiac muscle and the bright regions within the cardiac muscle. The procedure's feasibility has been supported by a test-set, including 23 echocardiographic cine-loops from 10 different patients. The resulting image quality appears to be superior to that of the original images, tending to show better contrast and a higher dynamic range of gray-levels within the cardiac muscle. According to two expert cardiologists, who have blindly ranked the image quality of each cine-loop on a scale from 1 to 10, where 10 corresponds to the highest possible image quality, the mean score of the original cine-loops is 7.1 +/- 1.1, while the mean score of the cine-loops to which ABTF has been applied is 8.0 +/- 1.2.

Algorithms↗

Lifelong menstrual histories are typically erratic and trending: a taxonomy.

OBJECTIVE: Menstrual cycles are composites of complex events; the data describing them are correspondingly rich. We seek to quantitatively represent menstrual histories from menarche to menopause and to evaluate the clinical belief that regular and stable cycle lengths are the most normative histories. DESIGN: Using prospective data from the Tremin Trust, we classified the menstrual histories of 628 women as very stable (type I), stable but with greater variability in cycle lengths (type II), oscillating and erratic with a downward trend in cycle length (type III), oscillating and erratic with no downward trend in cycle length (type IV), or highly erratic and variable (type V). Classification criteria were created by examining basic summary statistics of menstrual cycle lengths. Specifically, we identified key features describing variability of median cycle length, the mean of the interquartile range, the consistency of the interquartile range, the slope of median cycle lengths, and the number of stable 5-year intervals between ages 15 and 45+. RESULTS: We present the first characterization of full menstrual histories. Our taxonomy captures the essential features of menstrual bleeding patterns for a heterogeneous population. Persistently stable histories (types I and II) were seen in only 28% of the women; erratic histories (types III through V) characterized 72%. When examining all participants, significant differences were seen in age at menarche (P < 0.05), age at menopause (P < 0.01), and number of births (P < 0.01) between these stable and erratic groups. CONCLUSIONS: Although clinicians have traditionally thought of "normal" menstrual histories as being regular and stable, the distribution of women in our five categories suggest that variable histories are most common. Clinically, these results may suggest the need for a paradigm shift in what gynecologists view as normal and abnormal menstrual cycle histories.

Algorithms↗

A nonparametric scoring algorithm for identifying informative genes from microarray data.

Microarray data routinely contain gene expression levels of thousands of genes. In the context of medical diagnostics, an important problem is to find the genes that are correlated with given phenotypes. These genes may reveal insights to biological processes and may be used to predict the phenotypes of new samples. In most cases, while the gene expression levels are available for a large number of genes, only a small fraction of these genes may be informative in classification with statistical significance. We introduce a nonparametric scoring algorithm that assigns a score to each gene based on samples with known classes. Based on these scores, we can find a small set of genes which are informative of their class, and subsequent analysis can be carried out with this set. This procedure is robust to outliers and different normalization schemes, and immediately reduces the size of the data with little loss of information. We study the properties of this algorithm and apply it to the data set from cancer patients. We quantify the information in a given set of genes by comparing its distribution of the score statistics to a set of distributions generated by permutations that preserve the correlation structure among the genes.

Algorithms↗

Evaluation of an algorithm for treatment of status epilepticus in adult patients undergoing video/EEG monitoring.

Convulsive or generalized tonic clonic status epilepticus (SE) is a neurological emergency that can lead to transient or permanent brain damage or even death. An algorithm was designed to aid nursing and medical staff members in decision making about the type of SE and pharmacological intervention needed to stop prolonged or repetitive seizures. Fifteen registered nurses at a northern New England medical center's epilepsy unit participated in educational sessions on classification of seizures and status epilepticus prior to use of the algorithm. A pretest-posttest design with an investigator-developed tool was used to measure SE knowledge before and after educational intervention. There was a significant improvement in scores on the posttest of the classification of status epilepticus (Z = -2.93, p = .003). Twenty-nine medical records of patients who had experienced SE between February 1992 and December 1997 were reviewed. Nineteen patients experienced SE before the algorithm was implemented, and 10 patients experienced SE after the algorithm was implemented. A total of 16 patients experienced generalized convulsive SE with 12 episodes occurring before and 4 episodes after algorithm implementation. The mean time taken to stop the episode of SE after pharmacologic treatment began was compared in both groups using a t-test. The mean difference between the groups was 235 minutes (t = 2.57, p = .026). The findings of this project demonstrate that combining a treatment algorithm with education of staff members on its use has benefits in the practice setting of an inpatient comprehensive epilepsy program. Episodes of SE are more accurately classified and successful treatment of the episodes occurs earlier.

Adolescent↗