PubMed HealthSearch

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Particle track structure and its correlation with radiobiological endpoint.

One of the possible ways to classify track structures is application of the conventional partition techniques of analysis of multidimensional data to the track structure. Using these cluster algorithms this paper attempts to find characteristics of radiation reflecting the spatial distribution of ionizations in the primary particle track. Absolute frequency distributions of clusters giving the mean number of clusters produced by radiation per unit of deposited energy have been computed for radiation of different qualities. The results were compared with the published experimental data of cell inactivation. For particular biological objects the critical properties of radiation correlating with the cell inactivation can be found and it seems that the occurrence of a cluster of at least four ionizations formed in a domain of approximately 2-3 nm correlates with the induction of double strand break.

Ions

Cluster analysis applied to symptom ratings of psychiatric patients: an evaluation of its predictive ability.

Rating on 39 symptoms were examined for patients admitted to the Neuropsychiatric Institute of the University of Michigan Medical Center. A detailed evaluation was made of the clusters derived by a hierarchical clustering algorithm, using complete linkage and a simple matching coefficient on the binary variables of presence or absence of symptoms. The four groups of patients suggested by the cluster analysis can be characterized as follows: (1) generalized multiplicity of symptoms; (2) capacity to cope except for orientation apart from generally held norms; (3) activity level and thought processes speeded up, intensified, and unselected; (4) inwardly punitive, slowed down and distressed. It is shown that these groups received significantly different treatment and that the effect of treatment was significantly different, while no such differences were noted for groups defined in terms of diagnoses. By means of linear discriminant functions, rules are suggested for assigning other psychiatric patients to one of these four groups.

Antipsychotic Agents

Analysis of methotrexate treatment effect in a longitudinal observational study: utility of cluster analysis.

We studied 235 patients with rheumatoid arthritis (RA) beginning therapy with methotrexate utilizing a k-means clustering algorithm. Four groups were identified: mild RA (Group 3), very severe RA (Group 4), and 2 groups intermediate in severity (Groups 1 and 2). Group 2, the largest of the clusters (n = 89), appeared to have greater tolerability of RA as measured by severity and psychological variables, and took the drug almost twice as long as other groups, although improvement was not greater nor side effects fewer. All groups improved over a mean of 1.9 years, and the degree of improvement was not related to the initial severity classification. Improvement occurred almost equally in all clusters, and the relative ranking of the groups was maintained at study closure.

Arthritis, Rheumatoid

5S rRNA sequences of representatives of the genera Chlorobium, Prosthecochloris, Thermomicrobium, Cytophaga, Flavobacterium, Flexibacter and Saprospira and a discussion of the evolution of eubacteria in general.

5S rRNA sequences were determined for the green sulphur bacteria Chlorobium limicola, Chlorobium phaeobacteroides and Prosthecochloris aestuarii, for Thermomicrobium roseum, which is a relative of the green non-sulphur bacteria, and for Cytophaga aquatilis, Cytophaga heparina, Cytophaga johnsonae, Flavobacterium breve, Flexibacter sp. and Saprospira grandis, organisms allotted to the phylum 'Bacteroides-Cytophaga-Flavobacterium' and relatives as determined by 16S rRNA analyses. By using a clustering algorithm a dendrogram was constructed from these sequences and from all other known eubacterial 5S RNA sequences. The dendrogram showed differences, as well as similarities, with respect to results obtained by 16S RNA analyses. The 5S RNA sequences of green sulphur bacteria were closely related to one another, and to a cluster containing 5S RNA sequences from Bacteroides and its relatives, including Cytophaga aquatilis. 5S RNA sequences of all other representatives of the 'Bacteroides-Cytophaga-Flavobacterium' phylum as distinguished by 16S RNA analysis failed to group with Bacteroides and related clusters. On the basis of 5S RNA sequences, Thermomicrobium roseum clustered with Chloroflexus aurantiacus, as was expected from 16S RNA analysis.

Base Sequence

Considerations in applying clustering techniques to speaker-independent word recognition.

Recent work at Bell Laboratories has demonstrated the utility of applying sophisticated pattern recognition techniques to obtain a set of speaker-independent word templates for an isolated word recognition system [Levinson et al.,IEEE Trans. Acoust. Speech Signal Process. ASSP-27 (2), 134--141 (1979); Rabiner et al., IEEE Trans. Acoust. Speech Signal Process.(in press)]. In these studies, it was shown that a careful experimenter could guide the clustering algorithms to choose a small set of templates that were representative of a large number of replications for each word in the vocabulary. Subsequent word recognition tests verified that the templates chosen were indeed representative of a fairly large population of talkers. Given the success of this approach, the next important step is to investigate fully automatic techniques for clustering multiple versions of a single word into a set of speaker-independent word templates. Two such techniques are described in this paper. The first method uses distance data (between replications of a word) to segment the population into stable clusters. The word template is obtained as either the cluster minimax, or as an averaged version of all the elements in the cluster. The second method is a variation of the one described by Rabiner [IEEE Trans. Acoust. Speech Signal Process. ASSP-26 (3), 34--42 (1978)] in which averaging techniques are directly combined with the nearest neighbor rule to simultaneously define both the word template (i.e., the cluster center) and the elements in the cluster. Experimental data show the first method to be superior to the second method when three or more clusters per word are used in the recognition task.

Humans

An algorithm for comparing RNA secondary structures and searching for similar substructures.

To access the functional informations carried by RNA molecules at the level of their secondary structure interactions, we propose a comparison method based on a tree edit algorithm which takes into account the tree structure of RNA foldings. Any secondary structure is translated into a tree involving all its elementary substructures; then a shorter condensed tree is built in which any unbranched helix interspersed with bulges and interior loops is taken as a single node. This method includes several parameters: a comparison matrix between structural units, gap penalties, and the scoring between nodes of the condensed trees. Their effects have been analysed using as a model a rapidly divergent domain of the large ribosomal RNA, for which structural variation during evolution is well known. This method allows one to recognize precisely, in large target molecules, definite substructures that present with the query molecules only a limited set of closely related secondary structure features; it is still efficient if intervening features, which can correspond to insertion/deletion of entire stem regions, separate such structural elements. When coupled with a hierarchical clustering algorithm, this method is suitable for classifying RNA molecules according to their secondary structure homologies.

Algorithms

A new approach to analysis of synchronized sympathetic nerve activity.

Renal sympathetic nerve activity (RSNA) recorded from the multifiber preparation is a continuously fluctuating variable in terms of period and amplitude, reflecting a coordinated tonic level of output from the vasomotor center. Yet current methods of analysis cannot simultaneously measure both of these parameters. A new accurate technique for assessing changes in global sympathetic activity is required. We made a novel application of a computerized peak detection algorithm (Cluster program) to recordings of synchronized sympathetic nerve discharges. The procedure was applied to this new area to retrieve information about the characteristics of synchronized RSNA. Peaks in synchronized RSNA activity were detected from short-term (20 ms) integrated recordings in which voltage changes had been digitized at 200 Hz and stored on computer. The program scanned the data series for significant increases followed by significant decreases in a small cluster of voltage values. The program permits the input of the cluster sample sizes for the test peaks and pre- and postpeak nadirs and also the minimum height to be defined as a peak. Once each synchronized RSNA peak had been detected, its corresponding amplitude, width, and peak-to-peak interval were calculated. The program successfully characterized RSNA in a group of eight cats and yielded results comparable to other analysis techniques. The peak-to-peak interval period showed two modes of synchronized discharge, one related to the cardiac cycle and a faster 8- to 14-Hz frequency. The synchronized peak amplitude and width showed unimodal frequency distributions. The relationship between each of the three variables was examined; only the peak height and width were significantly related to each other.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

A typology of parasuicide.

Parasuicide is not a single syndrome. Subtypes at present recognized are based largely on clinically derived stereotypes. When considering a series of patients, the clinician is unable to handle more than a few attributes at a time. This paper describes the application of three very different clustering algorithms to a material of 350 treated parasuicide patients. Mathematically, three types emerge. Clinically, two of these are interpretable and make sense. The types established are: I (n = 107) a group not characterized by any of the variables we examined; this group is a puzzle, mainly because the reasons for the parasuicidal act are not clear. II (n = 132) a depressed, alienated group with high life-endangerment. III (n = III) a group whose act was highly operant: they felt alienated and were angry with others. These groups did not differ significantly on demographic variables. The usefulness of this typology, particularly for management, after-care and prevention, has now to be assessed.

Anger

Visual and auditory association areas of the cat's posterior ectosylvian gyrus: thalamic afferents.

The feline posterior ectosylvian gyrus contains a broad band of association cortex that is bounded anteriorly by tonotopic auditory areas and posteriorly by retinotopic visual areas. To characterize the possible functions of this cortex and to throw light on its pattern of internal divisions, we have carried out an analysis of its thalamic afferents. Deposits of differentiable retrograde tracers were placed at 17 cortical sites in nine cats. The deposit sites spanned the crown of the posterior ectosylvian gyrus and adjacent cortex in the suprasylvian sulcus. We compiled counts of retrogradely labeled neurons in 12 thalamic nuclei delineated by use of Nissl and acetylcholinesterase stains. We then employed a statistical clustering algorithm to identify groups of injections that gave rise to similar patterns of thalamic labeling. The results suggest that the posterior ectosylvian gyrus contains 3 fundamentally different cortical districts that have the form of parallel vertical bands. Very anterior cortex, overlapping previously identified tonotopic auditory areas (AI, P and VP) receives a dense projection from the laminated division of the medial geniculate body (MGl). An intermediate strip, to which we refer as the auditory belt, is innervated by axons from nontonotopic divisions of the medial geniculate body (MGds, MGvl, MGm, and MGd), from the lateral division of the posterior group (Pol), and from the posterior suprageniculate nucleus (SGp). A posterior strip, to which we refer as EPp, receives strong projections from the LM-SG complex (LM-SGa and LMp), and lighter projections from the intralaminar and lateroposterior (LPm and LPl) nuclei. On grounds of thalamic connectivity, EPp is not obviously distinguishable from adjacent retinotopic visual areas (PLLS, DLS, and VLS), and may be regarded as forming, together with these areas, a connectionally homogeneous visual belt.

Animals

Effects of opioid receptor blockade on luteinizing hormone (LH) pulses and interpulse LH concentrations in normal women during the early phase of the menstrual cycle.

To determine the role of endogenous opioid peptides in regulating pulsatile luteinizing hormone (LH) release in the early follicular phase of the menstrual cycle of eumenorrheic women, we evaluated serum LH concentrations in blood collected every 10 min for 12 h in 27 women each studied during two menstrual cycles: (1) without pretreatment and (2) following oral administration of naltrexone, a mu opiate receptor blocking agent, at a dose of 1.0 mg/kg. Pulsatile LH release was assessed by the CLUSTER algorithm. The mean (+/- SE) integrated serum LH concentration (IU/L/min) increased following the administration of naltrexone (4715 +/- 298) in comparison to the control day (3997 +/- 381; p = 0.0008). The mean number of LH pulses (/12 h) detected on the naltrexone day (10.3 +/- 0.3) was higher than on the control day (8.9 +/- 0.4; p = 0.0068). Mean maximal LH peak height (IU/L) was greater on the naltrexone (7.8 +/- 0.5) vs control (6.7 +/- 0.5) days (p = 0.0064) as was the interpulse valley mean serum LH concentration (IU/L; 6.3 +/- 0.4 vs 5.0 +/- 0.4; p = 0.0013). No difference was noted in the mean incremental LH pulse amplitude (IU/L; 1.9 +/- 0.1 vs 2.1 +/- 0.1; p = 0.13), or peak duration (min; 40 +/- 1.8 vs 45.0 +/- 2.4; p = 0.06). Mean LH peak area (IU/L/min) was greater on the control (45.0 +/- 2.4) vs naltrexone (40 +/- 1.8) days (p = 0.0475).(ABSTRACT TRUNCATED AT 250 WORDS)

Administration, Oral

Androgen-dependent somatotroph function in a hypogonadal adolescent male: evidence for control of exogenous androgens on growth hormone release.

A 14(10/12)-year-old white male with primary gonadal failure following testicular irradiation for acute lymphocytic leukemia was evaluated for poor growth. He had received 2400 rad of prophylactic cranial irradiation. The growth velocity had decelerated from 7 to 3.2 cm/yr over 3 years. His bone age was 12(0/12) years (by TW2-RUS), and his peak growth hormone (GH) response to provocative stimuli was 1.4 ng/mL. The 24-hour GH secretion was studied by drawing blood every 20 minutes for 24 hours. The resulting GH profile was analyzed by a computerized pulse detection algorithm, CLUSTER. Timed serum GH samples were also obtained after a 1 microgram/kg IV bolus injection of the GH releasing factor (GRH). The studies showed a flat 24-hour profile and a peak GH response to GRH of 3.9 ng/ml. Testosterone enanthate treatment was started, 100 mg IM every 4 weeks. Ten months after the initiation of therapy the calculated growth rate was 8.6 cm/yr. The 24-hour GH study and GRH responses were repeated at the time, showing a remarkably normal 24-hour GH secretory pattern and a peak GH response to GRH of 14.4 ng/mL. Testosterone therapy was discontinued, and 4 months later similar studies were repeated. A marked decrease in the mean 24-hour GH secretion and mean peak height occurred, but with maintenance of the GH pulse frequency. The GH response to GRH was intermediate, with a peak of 8 ng/mL. There was no further growth during those 4 months despite open epiphyses.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent

How do androgens affect episodic gonadotrophin secretion in postmenopausal women?

In the absence of any significant ovarian oestrogen secretion, as in post-menopausal women, the hypothalamic-pituitary axis may still be influenced by the androgens which continue to be produced. The episodic secretion of luteinizing hormone (LH) and follicle-stimulating hormone (FSH) by postmenopausal women was accordingly assessed following short-term androgen antagonism induced by flutamide, a specific androgen receptor blocker. Blood samples were collected at 10-min intervals for 10 h in nine women before and during flutamide administration (750 mg/day for 6 days) for the determination of gonadotrophin and sex hormone concentrations by radioimmunoassay. On both occasions, 25 micrograms of gonadotrophin-releasing-hormone (GnRH) was injected intravenously 8 h after initiation of the blood collections. Flutamide administration decreased (P less than 0.01 or less) androgen concentrations (testosterone, androstenedione and dehydroepiandrosterone sulphate) in relation to baseline values, but did not alter oestrogen (oestrone and oestradiol) or sex-hormone-binding globulin levels. The LH and FSH pulse characteristics (frequency, amplitude, interpulse interval and transverse mean levels) determined by a cluster algorithm in the gonadotrophin secretory profiles did not differ before and during androgen blockade. By contrast, androgen antagonism increased LH (P less than 0.01) and tended to enhance FSH (P = 0.10) FSH release in response to GnRH stimulation. Hence, short-term androgen receptor blockade with flutamide did not greatly affect episodic gonadotrophin secretion. However, the combined evidence of the enhanced gonadotrophin release observed in response to GnRH stimulation and the unchanged gonadotrophin secretion during androgen antagonism suggests that alterations in the magnitude, but not the frequency, of hypothalamic GnRH release had occurred. Even in the presence of substantial serum androgen concentrations, the gonadotrophin pulse rhythm in hypogonadal women constitutes the maximal-rate GnRH-LH release pattern.

Androgens

Multiomics Integration Identifies a Molecular Subtype of Intrahepatic Cholangiocarcinoma With Enhanced Benefit From Adjuvant Therapy.

Intrahepatic cholangiocarcinoma (iCCA) is a molecularly heterogeneous liver cancer with a poor prognosis. Improved stratification is needed to guide postoperative therapy. In this study, we applied integrative multiomics analysis to classify iCCA and identify biomarkers predictive of adjuvant treatment benefit. Using publicly available datasets (including whole exome sequencing, RNA sequencing, proteomics, and phosphoproteomics from FU-iCCA cohort and a transcriptomic cohort GSE244807), we defined 3 robust molecular subtypes of iCCA. These subtypes exhibited distinct genomic alterations, pathway activation, and immune microenvironments, with significant differences in overall survival (OS). Through protein-protein interaction network analysis and consensus feature selection using 10 clustering algorithms, we prioritized 8 marker genes distinguishing the subtypes. A Cox proportional-hazards model constructed from these markers stratified patients into high- and low-risk groups. High-risk iCCA, characterized by elevated expression of markers such as CLDN18, MUC1, and MUC5AC, had significantly worse OS in the absence of adjuvant therapy. Notably, in an independent validation of 174 patients with iCCA who underwent resection (single-center cohort), high expression of any of these 3 markers were associated with markedly prolonged OS in patients who received adjuvant chemotherapy or chemoembolization, compared with those who did not. In contrast, marker-negative patients showed no clear benefit from adjuvant therapy. In conclusion, our multiomics approach identified a high-risk, mucin-enriched subtype of iCCA. CLDN18, MUC1, and MUC5AC emerge as candidate predictive biomarkers for adjuvant chemotherapy benefit in iCCA, warranting prospective validation to improve personalized postoperative management.

Humans

Empirically derived personality types among male and female college students.

Data on the 16 PF obtained from 130 male and female college students were cluster analyzed to produce an empirical personality typology. Two different clustering algorithms were compared. Seven personality types emerged: 1) Well-Adjusted Conservative, 2) Ego-Involved Neurotic, 3) Norm Independent, 4) Socially-Detached Neurotic, 5) Superego Controlled, 6) Self-Assured Experimenter, and 7) Tough-Minded Controlled. Not only did the types differ significantly in personality, but they also were found to be significantly different on nine different measures of interpersonal orientation. Since the types did give intuitive insight into the nature of personality as particular combinations of personality traits and also were different on variables other than those used for the classification, the scientific utility of the typological approach received support.

Ego

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.

MOTIVATION: Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP technologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. RESULTS: Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis. AVAILABILITY AND IMPLEMENTATION: The source code of stDyer-image and detailed tutorials are available at https://github.com/ericcombiolab/stDyer-image.

Proteomics

Relative changes in LH pulsatility during the menstrual cycle: using data from hypogonadal women as a reference point.

The basic premise of this study is that the GnRH-LH pulsatile activity, particularly its frequency characteristics, constitutes, in the absence of any considerable ovarian feedback, the intrinsic rhythm of the hypothalamic-pituitary unit at its maximal rate. Thus, LH pulse attributes determined in postpubertal hypogonadal subjects may be used as a reference in assessing the degree of influence exerted by endocrine factors that modulate GnRH-LH pulses. Accordingly, serum LH levels were determined in samples obtained at 15-min intervals for 24 h in 20 hypogonadal women: 13 postmenopausal women (PMW) and seven women with premature ovarian failure (POF). Similar measurements were performed in 60 normally cycling women: 25 in the early follicular phase (EFP), 13 in the late follicular phase (LFP), seven at midcycle surge (LH surge) and 15 in the midluteal phase (MLP). Significant pulses were identified by the cluster algorithm utilizing factors appropriate for 24 h data series of a sampling frequency of 15-min intervals. The results show a 24-h mean (+/- SE) LH pulse frequency of 78.2 +/- 2.8 and 85.5 +/- 2.4 min per pulse for young (POF) and older (PMW) hypogonadal women, respectively. During the follicular phase of the cycle, the LH pulse frequency is not significantly different from that of hypogonadal women, but there is a significant (P less than 0.05) increase from early to late follicular phases (95.4 +/- 3.3 vs 78.8 +/- 2.2 min per pulse). However, when the sleep periods are excluded from the 24-h data series because of the associated decrease of LH pulse frequency in EFP women, the resulting pulse frequencies are almost identical for EFP, LFP and PMW. An elevation beyond the basic pulse rhythm determined in PMW or POF is not observed in any phase of the menstrual cycle studied, including the midcycle surge. The decrease in LH pulse frequency during the luteal phase of the cycle (151.8 +/- 8.0 min per pulse, P less than 0.001 vs hypogonadal women) beyond the reference pulse frequency of hypogonadal women is unequivocal. By contrast, the pulse amplitude varies markedly among the groups with the largest found in POF (36.6 +/- 4.5 IU/l). It follows, in descending order, PMW (22.7 +/- 3.1 IU/l), midcycle surge (17.3 +/- 2.8 IU/l), MLP women (7.0 +/- 1.3 IU/l) and the EFP (4.9 +/- 0.3 IU/l) and LFP (4.0 +/- 0.4 IU/l).(ABSTRACT TRUNCATED AT 250 WORDS)

Adult