PubMed HealthSearch

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Multivariate procedures to describe clinical staging of melanoma.

Analyzing multivariate clinical data to identify subclasses of patients being treated for a specific disease may improve patient management and increase understanding of the behavior of disease under clinical conditions. In some cases, patients have been classified on prognostic characteristics using standard risk assessment procedures (e.g., Cox' regression). This requires long term follow-up, differentiates patients only on attributes relevant to survival, and assumes that patients are sampled from a common population. Other approaches involve the use of clustering algorithms to classify patients into categories based on multiple clinical attributes. We illustrate the use of a multivariate statistical procedure to directly characterize patients on multiple clinical characteristics. The procedure is designed to analyze discrete response data with parameters representing individual differences within groups. Its use is illustrated for patients with Stage I melanoma in determining how age is related to treatment response in different patient groups.

Adult

From biopsy to automatic diagnosis.

High resolution two-dimensional gel electrophoresis is a very powerful biochemical tool for analysis of complex protein mixtures. In well defined situations, protein maps, obtained from tissue biopsies or biological fluids by this technique, can be automatically analyzed by computer. Some polypeptide patterns are the fingerprints of diseases. Applying clustering algorithm and learning techniques, the prototype expert system MELANIE recognized patterns and associated the correct diagnosis to the specific pattern.

Diagnosis, Computer-Assisted

[Phylogenetic analysis of partial nucleotide sequences of 18S rRNA for 14 plant species].

The variable 260 base long region from the interior of 18S rRNA of 14 plant species was determined by chain termination method with the use of reverse transcriptase. The hairpin revealed in this region appeared to be conservative in all species compared. Thermodynamic stability of such hairpin is lower than of an alternative structure with different base pairing mode. From sequence data dendrograms were produced by clustering algorithms and by the compatibility method. In addition to the plant sequences these dendrograms included also the homologous regions from yeast and Xenopus 18S rRNAs. The compatibility method seems to be more reliable. Inferences were drawn on relations between gymnosperms and angiosperms, monocots and dicots on the bases of the analysis of this tree.

Base Sequence

Quantification of progressive diabetic macular nonperfusion.

We used the IS-2000 Image Analyzer to estimate the extent of progressive diabetic macular nonperfusion in a patient by means of an automatic clustering algorithm applied to digitized fluorescein angiograms of the patient's macula taken over time. This method may provide an objective and reproducible quantification of progressive macular nonperfusion.

Adult

Quantification of gray/white matter in neonates and adults.

Quantitation of gray/white matter is important in evaluation of cerebral blood flow, atrophy, and development of the brain. First-order statistical analysis of neonatal computed tomographic (CT) images revealed that there was only a 6 Hounsfield unit (H) difference between gray and white matter compared with the observed 3 H for the standard deviation over the field of a skull water phantom. Scene segmentation methods based on first-order statistics proved unsuccessful in separating gray and white matter. A new regional clustering algorithm based on local textural properties was developed for separation of these structures.

Adult

Automatic classification of EEG segments and extraction of representative ones by dynamic clusters method.

Use of the dynamic clusters method for automatic extraction of compressed information about recorded EEG signal is presented. The computer first divides the record into quasi-stationary segments by means of adaptive segmentation. Second, the extracted segments are classified by a method of dynamic clusters into homogeneous classes. One part of the used clustering algorithm permits to specify and draw the most typical class members, which may represent the whole studied EEG signal and may be used as input for the further phase of the automatic EEG analysis, i.e. for the classification of the whole EEG records. The above procedure was applied to a 75 sec long EEG record of anaesthetized cat intoxicated by CO.

Computers

Numerical taxonomy of staphylococci.

Over two hundred staphylococci from human and animal sources and representatives of established species of Staphylococcus, Micrococcus and Planococcus were compared in a numerical phenetic survey using 115 unit characters. Data were analyzed using the Jaccard coefficient and the unweighted pair group method with averages algorithm. Cluster composition was not markedly affected by test error, estimated as 3.49%. The staphylococci were assigned to eighteen clusters containing four or more strains and to three single member clusters. Most of the clusters were distinct and homogeneous though two were divided into subclusters. Some of the clusters and subclusters were equated with the established taxa S. aureus, S. capitis, S. cohnii, S. epidermidis, S. haemolyticus, S. hominis, S. hyicus, S. saprophyticus, S. sciuri subspecies lentus, S. sciuri subspecies sciuri, S. simulans, S. warneri and S. xylosus, the remaining ones may represent the nuclei of additional centres of variation. The numerical data also cast doubts upon the reliability of some of the tests recommended for the identification of coagulase-negative staphylococci.

Computers

[Sectorization of the central 10 degrees visual field in open-angle glaucoma].

In an attempt to determine an optimal sector pattern of the central 10 degrees visual field in glaucoma, we applied the VARCLUS procedure, a new clustering algorithm provided by SAS Institute, to 379 glaucoma visual fields of the central 10-2 program of the Humphrey visual field analyzer. The subjects were 211 normal-tension glaucoma (NTG) and 168 primary open angle glaucoma (POAG) eyes with early to moderately advanced stages of visual field defects. The 68 2-degree grid test points in the central 10 degrees visual field were divided into 10 sectors. The sector pattern was compatible with the projection of nerve fiber layers and no sectors extended over the horizontal meridian. Comparison of the mean of total deviation in the sector of 76 eyes of 76 POAG patients with the maximum intraocular pressure: (IOP) > or = 25 mmHg (high-tension group) and 85 eyes of 85 NTG patients with the maximum IOP < or = 18 mmHg (low-tension group) revealed that four sectors nasal superior to the fixation were significantly more damaged in the low-tension group. We suggest that the sector pattern obtained here is useful in studying the visual field of glaucoma and that the damaging processes in the optic nervehead are not the same in the high-tension and low-tension groups.

Aged

A detection algorithm for multiform premature ventricular contractions.

This paper reports an algorithm developed to identify and quantify multiform PVCs. The algorithm clusters PVCs of similar morphology using a combination of time-domain and frequency-domain analysis. Initially, PVCs are grouped together on the basis of four time-domain-based morphological feature measurements. However, these time-domain-based clusters many times are nonunique because commonly encountered signal changes can cause substantial variations in the feature measurements of clinically similar beats. These redundant clusters are consolidated using two frequency-domain parameters: The First Spectral Moment (FSM) (center of gravity) of the amplitude spectrum, and the 5-Hz phase angle.

Cardiac Complexes, Premature

Automatic analysis of protein conformational changes by multiple linkage clustering.

An automatic algorithm is presented for analyzing protein conformational changes such as those occurring upon substrate binding or in different crystal forms of the same protein. Using, as sole information, the atomic coordinates of a pair of protein structures, the procedure first generates structure alignments, which optimize the root-mean-square deviation of the backbone atoms. To this end, equivalent secondary structures and/or loops from both proteins are combined by a multiple linkage hierarchic clustering algorithm, which generates several intertwined clustering trees. Automatic analysis of these clustering trees is used to dissect the mechanism of the conformational change. It allows the identification of the static core, representing the collection of secondary structures which undergo no structural changes, as well as other entities which move like rigid bodies. It also permits the description of the movement of secondary structures or loops relative to this core or entities. USing this information, it can be inferred whether a particular conformational change involves shear or hinge motion, or components of both. The algorithm is applied to the analysis of the conformational changes of citrate synthase, lactate dehydrogenase, lactoferrin and beta-glucosyltransferase, representing typical examples of shear- and hinge-type mechanisms, and a varied range in movement size. The results are shown to be in excellent agreement with previous analyses, and to provide additional information which gives a more complete and objective picture of the conformational change. Using our automatic algorithm, we find that any conformational change may be viewed as having components of both shear- and hinge-type motion. Determining which of these is most appropriate requires the combination of the information provided by our procedure with detailed knowledge of the protein tertiary structures.

Algorithms

Receiver-operated characteristic curve analysis of two algorithms assessing human growth hormone pulsatile secretion (PULSAR, CLUSTER): comparison of peak detection efficacy.

Computer-based peak identification algorithms reduce observer bias in the analysis of pulsatile hormone secretion. With increasing peak detection stringency, an algorithm will detect varying proportions of true-positive and false-positive peaks, determining its receiver-operated characteristics (ROC). To demonstrate that ROC curve analysis can characterize algorithm performance, we analyzed growth hormone (hGH) profiles from 94 children obtained with different hGH assay techniques [radioimmunoassay (RIA), immunoradiometric assay (IRMA)] at different sampling intervals (group A: 1 h/RIA; group B: 1 h/IRMA; group C: 20 min/RIA; group D: 20 min/IRMA), using the PULSAR and CLUSTER algorithms. The area under the ROC curve (AUC) was taken to compare the efficacy of both algorithms over a range of peak recognition stringency thresholds kept constant between algorithms, using hGH noise series for threshold calibration and the results of multiple visual inspection as reference standards. AUC by PULSAR ranged from 0.926 (group C) to 0.961 (group A), indicating good algorithm performance. AUC by CLUSTER ranged from 0.869 (group B) to 0.916 (group D) in the 20-min series, decreasing to 0.756 (group C) and 0.868 (group A) in the 1-hour series. At lower sampling intensity, significant discordant sensitivity existed between algorithms for RIA (p < 0.001) and IRMA (p < 0.0026). When adjusted to a high, assay-specific, comparable stringency, and employed on 20-min sampling hGH data, both the CLUSTER and PULSAR algorithm operated at a similarly high peak detection efficacy. The PULSAR algorithm appears to be more robust when hGH series with lower sampling intensities are analyzed.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

Performance of multi-layer feedforward neural networks to predict liver transplantation outcome.

A novel multisolutional clustering and quantization (MCQ) algorithm has been developed that provides a flexible way to preprocess data. It was tested whether it would impact the neural network's performance favorably and whether the employment of the proposed algorithm would enable neural networks to handle missing data. This was assessed by comparing the performance of neural networks using a well-documented data set to predict outcome following liver transplantation. This new approach to data preprocessing leads to a statistically significant improvement in network performance when compared to simple linear scaling. The obtained results also showed that coding missing data as zeroes in combination with the MCQ algorithm, leads to a significant improvement in neural network performance on a data set containing missing values in 59.4% of cases when compared to replacement of missing values with either series means or medians.

Algorithms

Multiple sequence alignment with hierarchical clustering.

An algorithm is presented for the multiple alignment of sequences, either proteins or nucleic acids, that is both accurate and easy to use on microcomputers. The approach is based on the conventional dynamic-programming method of pairwise alignment. Initially, a hierarchical clustering of the sequences is performed using the matrix of the pairwise alignment scores. The closest sequences are aligned creating groups of aligned sequences. Then close groups are aligned until all sequences are aligned in one group. The pairwise alignments included in the multiple alignment form a new matrix that is used to produce a hierarchical clustering. If it is different from the first one, iteration of the process can be performed. The method is illustrated by an example: a global alignment of 39 sequences of cytochrome c.

Algorithms

Clustering cDNA sequences.

A set of programs has been written to quantify the similarities between large numbers of cDNA sequences. This information is used to cluster similar sequences together. The main program can cluster thousands of cDNA sequences per day using a novel, computationally inexpensive algorithm. The clustering information is kept in a small index file so that disk storage requirements are negligible. Using this index file, subsidiary programs create various views and statistical summaries of the entire cDNA sequence collection.

Algorithms

Algorithm for the detection of fine clustered calcifications on film mammograms.

An algorithmic process for the detection and marking of clustered calcifications in digitized film-screen mammograms has been applied to mammograms from 50 clinical cases sampled at two digitization levels, in both the craniocaudal and mediolateral views. In all but one case the detector accurately located suggestive clusters found by radiologists in normal screening. In five cases additional clusters were also found by the detector. The detector has a negligible false-positive rate for the detection of clustered calcifications, although it is sensitive to clusters of emulsion defects displayed as artifactual calcification densities in the original film. The detector is flexible in structure and is easily adapted to various calcification/cluster criteria. The detector shows considerable promise when applied to clinical examples but will require refinement before formal testing.

Algorithms

Atlas-level single-cell integration and clustering-free differential expression analysis with GEDI 2.0.

MOTIVATION: GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS: We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98&#xd7; mean single-threaded speedups, and up to 11.5&#xd7; acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.

Single-Cell Analysis

Applying watershed algorithms to the segmentation of clustered nuclei.

Cluster division is a critical issue in fluorescence microscopy-based analytical cytology when preparation protocols do not provide appropriate separation of objects. Overlooking clustered nuclei and analyzing only isolated nuclei may dramatically increase analysis time or affect the statistical validation of the results. Automatic segmentation of clustered nuclei requires the implementation of specific image segmentation tools. Most algorithms are inspired by one of the two following strategies: 1) cluster division by the detection of internuclei gradients; or 2) division by definition of domains of influence (geometrical approach). Both strategies lead to completely different implementations, and usually algorithms based on a single view strategy fail to correctly segment most clustered nuclei, or perform well just for a specific type of sample. An algorithm based on morphological watersheds has been implemented and tested on the segmentation of microscopic nuclei clusters. This algorithm provides a tool that can be used for the implementation of both gradient- and domain-based algorithms, and, more importantly, for the implementation of mixed (gradient- and shape-based) algorithms. Using this algorithm, almost 90% of the test clusters were correctly segmented in peripheral blood and bone marrow preparations. The algorithm was valid for both types of samples, using the appropriate markers and transformations.

Algorithms

A scan statistic with a variable window.

Given N points or events occurring according to some probability distribution in the unit interval (0, 1), the simple scan statistic is defined to be the maximum number of points in any sub-interval of length d. In many areas, as in epidemiology, it is used to test the null hypothesis that the events are random, against the alternative that they cluster within some window of fixed width d. Since d must be chosen without snooping at the data, the test restricts the alternative to clusters of a fixed size. In this paper, we propose a scan statistic with a variable window, whose size does not need to be chosen a priori. This test is the generalized likelihood ratio test for a uniform null distribution against an alternative of non-random clustering which allows for clusters of variable width. A simple algorithm for the implementation of the method is given and applied to birth defects data previously analysed by a simple scan statistic.

Algorithms