PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Identification of carboxypeptidase E and gamma-glutamyl hydrolase as biomarkers for pulmonary neuroendocrine tumors by cDNA microarray.

Pulmonary neuroendocrine tumors vary dramatically in their malignant behavior. Their classification, based on histological examination, is often difficult. In search of molecular and prognostic markers for these tumors, we used cDNA microarray analysis of human transcripts against reference RNA from a well-characterized immortalized bronchial epithelial cell line, BEAS-2B. Tumor cells were isolated by laser-capture microdissection from primary tumors of 17 typical carcinoids, small cell lung cancers, and large cell neuroendocrine carcinomas. An unsupervised, hierarchical clustering algorithm resulted in a precise classification of each tumor subtype according to the proposed histological classification. Selection of genes, using supervised analysis, resulted in the identification of 198 statistically significant genes (P <.004) that also accurately discriminated between 3 predefined tumor subtypes. Two-by-two comparisons of these genes identified classifier genes that distinguished each tumor subtype from the others. Changes in expression of selected differentially expressed genes for each tumor subtype were internally validated by real-time reverse-transcription polymerase chain reaction. Expression of 2 potential classifier gene products, carboxypeptidase E (CPE) and gamma-glutamyl hydrolase (GGH), was validated by immunohistochemistry and cross-validated on additional archival samples of pulmonary neuroendocrine tumors. Kaplan-Meier survival analysis revealed that immunostaining for CPE was a statistically significant predictor of good prognosis, whereas GGH expression correlated with poor prognosis. Thus, cDNA microarray analysis led to the identification of 2 novel biomarkers that should facilitate molecular diagnosis and further study of pulmonary neuroendocrine tumors.

Biomarkers, Tumor↗

Correlation between MR imaging-derived nasopharyngeal carcinoma tumor volume and TNM system.

PURPOSE: To measure nasopharyngeal carcinoma tumor volume based on magnetic resonance images using a validated semiautomated measurement methodology and correlate tumor volume with TNM T classification. METHODS AND MATERIALS: The study population consisted of 206 consecutive nasopharyngeal carcinoma patients who had magnetic resonance imaging staging scans. Tumor volume was measured using a semisupervised knowledge-based fuzzy clustering algorithm. Patients were divided into 4 groups according to TNM T classification. The difference in tumor volumes among the various TNM T-classification groups was examined. RESULTS: The mean tumor volume in each T-classification group is as follows: T1, 8.6 mL +/- 5.0 (standard deviation [SD]); T2, 18.1 mL +/- 8.1 (SD); T3, 25.8 mL +/- 14.1 (SD); and T4, 36.2 mL +/- 18.9 (SD). The mean tumor volume increased significantly with advancing T classification (p < 0.0001). Tumor volume in a more advanced T group was significantly larger than that in an adjacent early T group (p < 0.01). CONCLUSION: Validated magnetic resonance imaging-based tumor volume shows positive correlation between tumor volume and advancing T-classification groups. It may be possible to incorporate tumor volume as an additional prognostic factor into the existing TNM system.

Adolescent↗

Tissue microarray study for classification of breast tumors.

Clinical and pathological heterogeneity of breast cancer hinders selection of appropriate treatment for individual cases. Molecular profiling at gene or protein levels may elucidate the biological variance of tumors and provide a new classification system that correlates better with biological, clinical and prognostic parameters. We studied the immunohistochemical profile of a panel of seven important biomarkers using tumor tissue arrays. The tumor samples were then classified with a monothetic (binary variables) clustering algorithm. Two distinct groups of tumors are characterized by the estrogen receptor (ER) status and tumor grade (p = 0.0026). Four biomarkers, c-erbB2, Cox-2, p53 and VEGF, were significantly overexpressed in tumors with the ER-negative (ER-) phenotype. Eight subsets of tumors were further identified according to the expression status of VEGF, c-erbB2 and p53. The malignant potential of the ER-/VEGF+ subgroup was associated with the strong correlations of Cox-2 and c-erbB2 with VEGF. Our results indicate that this molecular classification system, based on the statistical analysis of immunohistochemical profiling, is a useful approach for tumor grouping. Some of these subgroups have a relative genetic homogeneity that may allow further study of specific genetically-controlled metabolic pathways. This approach may hold great promise in rationalizing the application of different therapeutic strategies for different subgroups of breast tumors.

Algorithms↗

Genetic fuzzy modelling and control of bispectral index (BIS) for general intravenous anaesthesia.

Based on an adaptive genetic fuzzy clustering algorithm, a derived fuzzy knowledge model is proposed for quantitatively estimating the systolic arterial pressure (SAP), heart rate (HR), and bispectral index (BIS) using 12 patients and it validates them according to pharmacological reasoning. Also, a genetic proportional integral derivative controller (GPIDC) to adaptive three controller parameters and a genetic fuzzy logic controller (GFLC) to adaptive controller rules using genetic algorithms (GAs) were simulated and compared each other in a patient model using the BIS value as a controlled variable. Each controller was tested using a set of 12 virtual patients undergoing a Gaussian random surgical disturbance repeated with BIS targets set at 40, 50, and 60. Controller performance was assessed using mean absolute error (MAE) of the BIS target, the percentage of time with acceptable BIS control (PTABC), and drug consumption (DC). It was found that the MAE value of the BIS target was significantly lower (P < 0.05) and the values of PTABC and DC of BIS target were significantly higher (P < 0.05) in BIS targets set at 40 than at 50 or 60 in both GPIDC and GFLC. However, when compared with two controllers in terms of the values of MAE, PTABC, and DC each other in BIS targets set at 40, 50, and 60, there were no significant differences (P > 0.05). Furthermore, when the simulation results in these two controllers were compared with routine standard practice of 12 clinical trials (i.e., manual control) in BIS target set at 50, the values of PTABC in both GPIDC and GFLC groups were significantly higher (P < 0.05) than in the manual control group. In contrast, there were no significant differences (P > 0.05) for these three groups in terms of drug consumption. This indicates that either GPIDC or GFLC can control the BIS target set at 50 better than manual control, although the similar drug consumption is used.

Anesthesia, General↗

Multiomics Integration Identifies a Molecular Subtype of Intrahepatic Cholangiocarcinoma With Enhanced Benefit From Adjuvant Therapy.

Intrahepatic cholangiocarcinoma (iCCA) is a molecularly heterogeneous liver cancer with a poor prognosis. Improved stratification is needed to guide postoperative therapy. In this study, we applied integrative multiomics analysis to classify iCCA and identify biomarkers predictive of adjuvant treatment benefit. Using publicly available datasets (including whole exome sequencing, RNA sequencing, proteomics, and phosphoproteomics from FU-iCCA cohort and a transcriptomic cohort GSE244807), we defined 3 robust molecular subtypes of iCCA. These subtypes exhibited distinct genomic alterations, pathway activation, and immune microenvironments, with significant differences in overall survival (OS). Through protein-protein interaction network analysis and consensus feature selection using 10 clustering algorithms, we prioritized 8 marker genes distinguishing the subtypes. A Cox proportional-hazards model constructed from these markers stratified patients into high- and low-risk groups. High-risk iCCA, characterized by elevated expression of markers such as CLDN18, MUC1, and MUC5AC, had significantly worse OS in the absence of adjuvant therapy. Notably, in an independent validation of 174 patients with iCCA who underwent resection (single-center cohort), high expression of any of these 3 markers were associated with markedly prolonged OS in patients who received adjuvant chemotherapy or chemoembolization, compared with those who did not. In contrast, marker-negative patients showed no clear benefit from adjuvant therapy. In conclusion, our multiomics approach identified a high-risk, mucin-enriched subtype of iCCA. CLDN18, MUC1, and MUC5AC emerge as candidate predictive biomarkers for adjuvant chemotherapy benefit in iCCA, warranting prospective validation to improve personalized postoperative management.

Humans↗

Unraveling substantia nigra sequential gene expression in a progressive MPTP-lesioned macaque model of Parkinson's disease.

Taking advantage of a progressive nonhuman primate model mimicking Parkinson's disease (PD) evolution, we monitored transcriptional fluctuations in the substantia nigra using Affymetrix microarrays in control (normal), saline-treated (normal), 6 days-treated (asymptomatic with 20% cell loss), 12 days-treated (asymptomatic with 40% cell loss) and 25 days-treated animals (fully parkinsonian with 85% cell loss). Two statistical methods were used to ascertain the regulation and real-time quantitative PCR was used to confirm their regulation. Surprisingly, the number of deregulated transcripts is limited at all time points and five clusters exhibiting different profiles were defined using a hierarchical clustering algorithm. Such profiles are likely to represent activation/deactivation of mechanisms of different nature. We briefly speculate about (i) the existence of yet unknown compensatory mechanisms is unraveled, (ii) the putative triggering of a developmental program in the mature brain in reaction to progressing degeneration and finally, (iii) the activation of mechanisms leading eventually to death in final stage. These data should help development of new therapeutic approaches either aimed at enhancing existing compensatory mechanisms or at protecting dopamine neurons.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Batch and median neural gas.

Neural Gas (NG) constitutes a very robust clustering algorithm given Euclidean data which does not suffer from the problem of local minima like simple vector quantization, or topological restrictions like the self-organizing map. Based on the cost function of NG, we introduce a batch variant of NG which shows much faster convergence and which can be interpreted as an optimization of the cost function by the Newton method. This formulation has the additional benefit that, based on the notion of the generalized median in analogy to Median SOM, a variant for non-vectorial proximity data can be introduced. We prove convergence of batch and median versions of NG, SOM, and k-means in a unified formulation, and we investigate the behavior of the algorithms in several experiments.

Algorithms↗

Heterogeneous distribution of taste cells in facial and vagal nerve-innervated taste buds.

Input from the three gustatory nerves of vertebrates is used to evaluate the nutritional quality of food. In some species, these cranial nerves are modified to accomplish additional specific functions. For example, the facial nerve innervated taste buds distributed over the body surface of catfish aid food search. Physiological studies indicate that this extra-oral taste pathway is more sensitive to amino acids than either the glossopharyngeal or vagal systems of the oral cavity. The current investigation seeks to determine if differences in taste cell subtypes might contribute to the observed differences in sensitivity. The distributions of five low molecular weight metabolites, L-alanine, L-aspartate, L-glutamate, GABA, taurine and the tripeptide glutathione, were examined in 2118 individual taste cells innervated by either the facial or vagal nerve of the channel catfish, Ictalurus punctatus. The metabolite profiles of these cells were determined immunocytochemically and subjected to a k-means clustering algorithm. Fifteen cell classes with quantitatively different patterns of metabolite co-localization were identified. All but one small class of two cells were found in both facial and vagal nerve-innervated taste buds. Four classes (9% of the total cells) had high, two classes (17%) had intermediate and the remaining nine classes (74%) had low levels of GABA immunoreactivity. While the functional significance of differences in metabolite profile remains to be determined, taste cell classes were not uniformly distributed across vagal and facial nerve innervated taste buds and may provide an anatomical basis for previously reported differences in gustatory sensitivity.

Algorithms↗

Bioinformatics in proteomics: application, terminology, and pitfalls.

Bioinformatics applies data mining, i.e., modern computer-based statistics, to biomedical data. It leverages on machine learning approaches, such as artificial neural networks, decision trees and clustering algorithms, and is ideally suited for handling huge data amounts. In this article, we review the analysis of mass spectrometry data in proteomics, starting with common pre-processing steps and using single decision trees and decision tree ensembles for classification. Special emphasis is put on the pitfall of overfitting, i.e., of generating too complex single decision trees. Finally, we discuss the pros and cons of the two different decision tree usages.

Computational Biology↗

A fuzzy index model for trophic status evaluation of reservoir waters.

An index model for quality evaluation based on the formula of similarity membership functions in the fuzzy c-means (FCM) clustering algorithm is proposed. Summing up the weighted similarity degrees between an observation and designed specific quality-ordered levels develops an alternative overall index. Stretching the values of the controlling parameters in the formula of the similarity membership functions causes diverse patterns of overall index models. Applying this proposed fuzzy index model to the trophic evaluation of reservoir waters is studied to demonstrate the practical application of this index. Every measurement of the variables is standardized by the membership function of quality evaluation on the interval [0,1], referring to the trophic status clarified in the Carlson Trophic State Index. The sensitivity analyses are studied both in the proposed index system and the Carlson Trophic State Index. Besides, a case study of the trophic status evaluation of the Feitsui Reservoir from 1987 to 2003 is presented to demonstrate the feasibility of applying the proposed evaluation model.

Chlorophyll↗

The S. purpuratus genome: a comparative perspective.

The predicted gene models derived from the sea urchin genome were compared to the gene catalogs derived from other completed genomes. The models were categorized by their best match to conserved protein domains. Identification of potential orthologs and assignment of sea urchin gene models to groups of homologous genes was accomplished by BLAST alignment and through the use of a clustering algorithm. For the first time, an overview of the sea urchin genetic toolkit emerges and by extension a more precise view of the features shared among the gene catalogs that characterize the super-clades of animals: metazoans, bilaterians, chordate and non-chordate deuterostomes, ecdysozoan and lophotrochozoan protostomes. About one third of the 40 most prevalent domains in the sea urchin gene models are not as abundant in the other genomes and thus constitute expansions that are specific at least to sea urchins if not to all echinoderms. A number of homologous groups of genes previously restricted to vertebrates have sea urchin representatives thus expanding the deuterostome complement. Obversely, the absence of representatives in the sea urchin confirms a number of chordate specific inventions. The specific complement of genes in the sea urchin genome results largely from minor expansions and contractions of existing families already found in the common metazoan "toolkit" of genes. However, several striking expansions shed light on how the sea urchin lives and develops.

Animals↗

Characteristics of luteinizing hormone secretion in younger versus older premenopausal women.

OBJECTIVES: The objectives of this study were to document specific attributes of pulsatile luteinizing hormone secretion in middle-aged women before discernible alterations in their menstrual cycles and to compare the results to corresponding data obtained in younger women. STUDY DESIGN: After documenting normal cycle length, biphasic basal body temperatures, and normal midluteal progesterone in younger and middle-aged women during an initial cycle, daily blood samples and samples withdrawn at 10-minute intervals for 8 hours during the midfollicular phase were obtained during a subsequent cycle. RESULTS: Assessment of luteinizing hormone pulses with the pulse detection algorithm Cluster demonstrated a prolonged interpulse interval and increased pulse width in the older women. Assessment of luteinizing hormone secretory bursts and half-life with the deconvolution analysis procedure demonstrated a prolonged interburst interval and half-life in the older women. Appraisal of approximate entropy revealed greater orderliness of luteinizing hormone release in the older women. CONCLUSIONS: Middle-aged women exhibit alterations in hypothalamic-pituitary function that may account in part for age-related changes in reproductive potential.

Adult↗

Expression profiling of renal epithelial neoplasms: a method for tumor classification and discovery of diagnostic molecular markers.

The expression patterns of 7075 genes were analyzed in four conventional (clear cell) renal cell carcinomas (RCC), one chromophobe RCC, and two oncocytomas using cDNA microarrays. Expression profiles were compared among tumors using various clustering algorithms, thereby separating the tumors into two categories consistent with corresponding histopathological diagnoses. Specifically, conventional RCCs were distinguished from chromophobe RCC/oncocytomas based on large-scale gene expression patterns. Chromophobe RCC/oncocytomas displayed similar expression profiles, including genes involved with oxidative phosphorylation and genes expressed normally by distal nephron, consistent with the mitochondrion-rich morphology of these tumors and the theory that both lesions are related histogenetically to distal nephron epithelium. Conventional RCCs underexpressed mitochondrial and distal nephron genes, and were further distinguished from chromophobe RCC/oncocytomas by overexpression of vimentin and class II major histocompatibility complex-related molecules. Novel, tumor-specific expression of four genes-vimentin, class II major histocompatibility complex-associated invariant chain (CD74), parvalbumin, and galectin-3-was confirmed in an independent tumor series by immunohistochemistry. Vimentin was a sensitive, specific marker for conventional RCCs, and parvalbumin was detected primarily in chromophobe RCC/oncocytomas. In conclusion, histopathological subtypes of renal epithelial neoplasia were characterized by distinct patterns of gene expression. Expression patterns were useful for identifying novel molecular markers with potential diagnostic utility.

Adult↗

Winner take all experts network for sensor validation.

The validation of sensor measurements has become an integral part of the operation and control of modern industrial equipment. The sensor under harsh environment must be shown to consistently provide the correct measurements. Analysis of the validation hardware or software should trigger an alarm when the sensor signals deviate appreciably from the correct values. Neural network based models can be used to on-line estimate critical sensor values when neighboring sensor measurements are used as inputs. The underlying assumption is that the neighboring sensors share an analytical relationship. The discrepancy between the measured and predicted sensor values may then be used as an indicator for sensor health. The proposed Winner Take All Experts (WTAE) network based on a 'divide and conquer' strategy significantly reduces the computational time required to train the neural network. It employs a growing fuzzy clustering algorithm to divide a complicated problem into a series of simpler sub-problems and assigns an expert to each of them locally. After the sensor approximation, the outputs from the estimator and the real sensor readings are compared both in the time domain and the frequency domain. Three fault indicators are used to provide analytical redundancy to detect the sensor failure. In the decision stage, the intersection of three fuzzy sets accomplishes a decision level fusion, which indicates the confidence level of the sensor health. Two data sets, the Spectra Quest Machinery Fault Simulator data set and the Westland vibration data set, were used in simulations to demonstrate the performance of the proposed WTAE network. The simulation results show the proposed WTAE is competitive with or even superior to the existing approaches.

Journal Article↗

Microarray analysis of bacterial pathogenicity.

The DNA microarray, a surface that contains an ordered arrangement of each identified open reading frame of a sequenced genome, is the engine of functional genomics. Its output, the expression profile, provides a genome wide snap-shot of the transcriptome. Refined by array-specific statistical instruments and data-mined by clustering algorithms and metabolic pathway databases, the expression profile discloses, at the transcriptional level, how the microbe adapts to new conditions of growth--the regulatory networks that govern the adaptive response and the metabolic and biosynthetic pathways that effect the new phenotype. Adaptation to host microenvironments underlies the capacity of infectious agents to persist in and damage host tissues. While monitoring the whole genome transcriptional response of bacterial pathogens within infected tissues has not been achieved, it is likely that the complex, tissue-specific response is but the sum of individual responses of the bacteria to specific physicochemical features that characterize the host milieu. These are amenable to experimentation in vitro and whole-genome expression studies of this kind have defined the transcriptional response to iron starvation, low oxygen, acid pH, quorum-sensing pheromones and reactive oxygen intermediates. These have disclosed new information about even well-studied processes and provide a portrait of the adapting bacterium as a 'system', rather than the product of a few genes or even a few regulons. Amongst the regulated genes that compose this adaptive system are transcription factors. Expression profiling experiments of transcription factor mutants delineate the corresponding regulatory cascade. The genetic basis for pathogenicity can also be studied by using microarray-based comparative genomics to characterize and quantify the extent of genetic variability within natural populations at the gene level of resolution. Also identified are differences between pathogen and commensal that point to possible virulence determinants or disclose evolutionary history. The host vigorously engages the pathogen; expression studies using host genome microarrays and bacterially infected cell cultures show that the initial host reaction is dominated by the innate immune response. However, within the complex expression profile of the host cell are components mediated by pathogen-specific determinants. In the future, the combined use of bacterial and host microarrays to study the same infected tissue will reveal the dialogue between pathogen and host in a gene-by-gene and site- and time-specific manner. Translating this conversation will not be easy and will probably require a combination of powerful bioinformatic tools and traditional experimental approaches--and considerable effort and time.

Animals↗

Exploring trafficking GTPase function by mRNA expression profiling: use of the SymAtlas web-application and the Membrome datasets.

Despite complete sequencing of the human and mouse genomes, functional annotation of novel gene function still remains a major challenge in mammalian biology. Emerging strategies to help elucidate unknown gene function include the analysis of tissue-specific patterns of mRNA expression. A recent study investigated the steady-state mRNA expression profiling of the vast majority of protein-encoding human and mouse genes across a panel of 79 human and 61 mouse nonredundant tissues. The microarray data from this study constitutes the Genomics Institute of Novartis Foundation (GNF) Human and Mouse Gene Atlases and is publicly available for exploration through the SymAtlas web-application (http://symatlas.gnf.org/). We have recently reported the use of these data and hierarchical clustering algorithms to generate a global overview of the distribution of Rabs, SNAREs, and coat machinery components, as well as their respective adaptors, effectors, and regulators. This systems biology approach led us to propose Rab-centric protein activity hubs as a framework for an integrated coding system, the membrome network, which orchestrates the dynamics of specialized membrane architecture of differentiated cells. Here, we describe the use of the SymAtlas web-application and the Membrome datasets to help explore trafficking GTPase function. The human and mouse membrome datasets are available through the Membrome homepage (http://www.membrome.org/) and correspond to subsets of the SymAtlas content restricted to known membrane trafficking components. Considering the fragmentary nature of the current reductionist approaches in elucidating trafficking component functions, the membrome datasets provide a more focused systems biology perspective that not only complements our current understanding of transport in complex tissues but also provides an integrated perspective of Rab activity in controlling membrane architecture.

Animals↗

Genetic heterogeneity of single disseminated tumour cells in minimal residual cancer.

BACKGROUND: Because cancer patients with small tumours often relapse despite local and systemic treatment, we investigated the genetic variation of the precursors of distant metastasis at the stage of minimal residual disease. Disseminated tumour cells can be detected by epithelial markers in mesenchymal tissues and represent targets for adjuvant therapies. METHODS: We screened 525 bone-marrow, blood, and lymph-node samples from 474 patients with breast, prostate, and gastrointestinal cancers for single disseminated cancer cells by immunocytochemistry with epithelial-specific markers. 71 (14%) of the samples contained two or more tumour cells whose genomic organisation we studied by single cell genomic hybridisation. In addition, we tested whether TP53 was mutated. Hierarchical clustering algorithms were used to determine the degree of clonal relatedness of sister cells that were isolated from individual patients. FINDINGS: Irrespective of cancer type, we saw an unexpectedly high genetic divergence in minimal residual cancer, particularly at the level of chromosomal imbalances. Although few disseminated cells harboured TP53 mutations at this stage of disease, we also saw microheterogeneity of the TP53 genotype. The genetic heterogeneity was strikingly reduced with the emergence of clinically evident metastasis. INTERPRETATION: Although the heterogeneity of primary tumours has long been known, we show here that early disseminated cancer cells are genomically very unstable as well. Selection of clonally expanding cells leading to metastasis seems to occur after dissemination has taken place. Therefore, adjuvant therapies are confronted with an extremely large reservoir of variant cells from which resistant tumour cells can be selected.

Bone Marrow↗

A self calibration method using a soft clustering procedure for eye movement recordings.

A nearly automatic method for calibrating eye movement records has been developed. This very robust method is based on a soft clustering algorithm which allows exploration of the whole range of eye movement records for reliable calibration. In contrast to many other methods which carry out the calibration on several discrete points, this method is suitable for continuous determination of the transfer function of the eye movement transducer. Moreover it simultaneously uses the combined properties of vestibulo-ocular reflex, neck-to-eye reflex and smooth pursuit system to reach approximately a unity gain and zero phase lag (in subjects with no severe vestibular disorders or ocumomotor palsy). In addition, this method does not rely heavily on the degree of attention of the subject. The method is particularly suited for the calibration of non linear or noisy transducers like Electro Oculography (EOG). Calibration is performed within a few seconds. So when necessary in clinical applications it is possible to repeat calibrations frequently.

Algorithms↗