PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Automated particle classification based on digital acquisition and analysis of flow cytometric pulse waveforms.

In flow cytometry, the typical use of front-end analog processing limits the pulse waveform features that can be measured to pulse integral, height, and width. Direct digitizing of the waveforms provides a means for the extraction of additional features, for example, pulse skewness and kurtosis, and Fourier properties. In this work, we have first demonstrated that the Fourier properties of the pulse can be employed usefully for discrimination between different types of cells that otherwise cannot be classified by using only time-domain features of the pulse. We then implemented and evaluated automatic procedures for cell classification based on neural networks. We established that neural networks could provide an efficient means of classification of cell types without the need for user interaction. The neural networks were also employed in an innovative manner for analysis of the digital flow cytometric data without feature extraction. The performance of the neural networks was compared with that of a more conventional means of classification, the K-means clustering algorithm. Neural networks can be realized in hardware, and this, in addition to their highly parallel architecture, makes them an important potential part of real-time analysis systems. These results are discussed in terms of the design of a real-time digital data acquisition system for flow cytometry.

Animals↗

Apolipoprotein E and other cerebrospinal fluid proteins differentiate ante mortem variant Creutzfeldt-Jakob disease from ante mortem sporadic Creutzfeldt-Jakob disease.

The ability to perform an ante mortem differential diagnosis of Creutzfeld-Jakob disease (CJD) is aided by several clinical and molecular tests. There is a need for molecular tests which can reliably distinguish ante mortem variant CJD (vCJD) from ante mortem sporadic CJD (spCJD). A proteomics approach employing two-dimensional protein electrophoresis is applied to the study of ante mortem CSF samples obtained in collaboration with the CJD Surveillance Unit and the National Hospital for Neurology and Neurosurgery. The sample set includes two cases of vCJD, three cases of spCJD and three neurologic controls. Preliminary data using a panel of seven molecular markers is able to distinguish vCJD from spCJD using a heuristic clustering algorithm. One of the molecular markers has been identified as apolipoprotein E which appears to be upregulated in the cerebrospiral fluid (CSF) of patients with vCJD as compared to spCJD. Analysis of ante mortem CSF may help to differentiate patients with vCJD from those patients with spCJD.

Adult↗

Molecular scanner experiment with human plasma: improving protein identification by using intensity distributions of matching peptide masses.

The development of high throughput utilities to identify proteins is a major challenge in present research in the field of proteomics. One such utility, the molecular scanner, uses proteins separated by two-dimensional polyacrylamide gel electrophoresis that are digested in the gel and during transfer onto a collecting membrane. After adding a matrix, the membrane is inserted into a matrix-assisted laser desorption/ionization-time of flight mass spectrometer and a peptide mass fingerprint (PMF) is measured for every scanned site. Since the spacing between scanned sites is much smaller than the size of the most abundant protein spots, there is a certain redundancy in the data that was used in an earlier experiment with Escherichia coli [1] to improve mass calibration and PMF identification results. It was observed that the signal intensity of a peptide mass as a function of the position on the membrane showed similar patterns if peptides stemmed from the same protein. Taking account of these similarities a clustering algorithm was used to find lists of experimental masses with similar intensity distributions, which provided clearer identification of the corresponding proteins. Here, these methods are applied to a human plasma scan, where proteins were highly modified and less separated. The presence of very abundant proteins like albumin and immunoglobulins added another difficulty. The calibration of the initial PMFs was not satisfactory and masses had to be recalibrated. After discarding chemical noise, the membrane was partitioned into regions and for each region protein identification was carried out separately. A new scoring method was used, where the PMF score was multiplied by a factor that measures the similarity of matching peptides. This method proved to be more robust than the method developed in [1] if the region where a protein was found had an extended, nonspherical shape and strong overlap with regions of other proteins. Many proteins annotated on the SWISS-2D PAGE human plasma master gel could be clearly identified and many interesting properties were observed.

Algorithms↗

Argentine population genetic structure: large variance in Amerindian contribution.

Argentine population genetic structure was examined using a set of 78 ancestry informative markers (AIMs) to assess the contributions of European, Amerindian, and African ancestry in 94 individuals members of this population. Using the Bayesian clustering algorithm STRUCTURE, the mean European contribution was 78%, the Amerindian contribution was 19.4%, and the African contribution was 2.5%. Similar results were found using weighted least mean square method: European, 80.2%; Amerindian, 18.1%; and African, 1.7%. Consistent with previous studies the current results showed very few individuals (four of 94) with greater than 10% African admixture. Notably, when individual admixture was examined, the Amerindian and European admixture showed a very large variance and individual Amerindian contribution ranged from 1.5 to 84.5% in the 94 individual Argentine subjects. These results indicate that admixture must be considered when clinical epidemiology or case control genetic analyses are studied in this population. Moreover, the current study provides a set of informative SNPs that can be used to ascertain or control for this potentially hidden stratification. In addition, the large variance in admixture proportions in individual Argentine subjects shown by this study suggests that this population is appropriate for future admixture mapping studies.

Asian People↗

Coexpression of the type 1 growth factor receptor family members HER-1, HER-2, and HER-3 has a synergistic negative prognostic effect on breast carcinoma survival.

BACKGROUND: The clinical significance of coexpression of type 1 growth factor receptor (T1GFR) family members remains largely unknown. The objective of the current study was to determine the frequency and the possible prognostic effect of coexpression of HER-1, HER-2, HER-3, and HER-4 by breast carcinoma. METHODS: Tissue microarrays were constructed using clinically annotated formalin-fixed, paraffin-embedded tumor samples from 242 patients with invasive breast carcinomas with a median 15-year follow-up. The levels of TIGFR family members (HER-1-HER-4) were measured by immunohistochemistry. K-means clustering algorithm, as well as univariate (Kaplan-Meier, log-rank test) and multivariate (Cox regression) survival analyses were applied to the data set. RESULTS: Using univariate analysis, expression of HER-1, HER-2, and HER-3, but not HER-4, was significantly associated with decreased patient disease-specific survival (P < 0.05). Kaplan-Meier survival analysis showed that coexpression of >/= 2 of HER-1, HER-2, and HER-3 in any combination was associated with reduced patient disease-specific survival compared with single marker expression or no expression (35% vs. 65% vs. 78% 10-year survival rates, P = 0.001). Using multivariate analysis, expression of >/= 2 of HER-1, HER-2, and HER-3 was independent of lymph node status and tumor size. CONCLUSIONS: In a cohort of patients with breast carcinoma, the authors observed T1GFR family member coexpression (HER-1, HER-2, and HER-3) to have a negative synergistic effect on patient outcome, independent of tumor size or lymph node status. Thus, coexpression of T1GFR family members identified a subset of patients with a poor disease prognosis who may potentially benefit from therapy simultaneously targeting several T1GFR family members.

Biomarkers, Tumor↗

Visual and auditory association areas of the cat's posterior ectosylvian gyrus: thalamic afferents.

The feline posterior ectosylvian gyrus contains a broad band of association cortex that is bounded anteriorly by tonotopic auditory areas and posteriorly by retinotopic visual areas. To characterize the possible functions of this cortex and to throw light on its pattern of internal divisions, we have carried out an analysis of its thalamic afferents. Deposits of differentiable retrograde tracers were placed at 17 cortical sites in nine cats. The deposit sites spanned the crown of the posterior ectosylvian gyrus and adjacent cortex in the suprasylvian sulcus. We compiled counts of retrogradely labeled neurons in 12 thalamic nuclei delineated by use of Nissl and acetylcholinesterase stains. We then employed a statistical clustering algorithm to identify groups of injections that gave rise to similar patterns of thalamic labeling. The results suggest that the posterior ectosylvian gyrus contains 3 fundamentally different cortical districts that have the form of parallel vertical bands. Very anterior cortex, overlapping previously identified tonotopic auditory areas (AI, P and VP) receives a dense projection from the laminated division of the medial geniculate body (MGl). An intermediate strip, to which we refer as the auditory belt, is innervated by axons from nontonotopic divisions of the medial geniculate body (MGds, MGvl, MGm, and MGd), from the lateral division of the posterior group (Pol), and from the posterior suprageniculate nucleus (SGp). A posterior strip, to which we refer as EPp, receives strong projections from the LM-SG complex (LM-SGa and LMp), and lighter projections from the intralaminar and lateroposterior (LPm and LPl) nuclei. On grounds of thalamic connectivity, EPp is not obviously distinguishable from adjacent retinotopic visual areas (PLLS, DLS, and VLS), and may be regarded as forming, together with these areas, a connectionally homogeneous visual belt.

Animals↗

Genetic association mapping under founder heterogeneity via weighted haplotype similarity analysis in candidate genes.

Taking advantage of increasingly available high-density single nucleotide polymorphism (SNP) markers within genes and across genomes, more and more genetic association studies began to use multiple closely linked markers in candidate genes. A practical analytical challenge arising in such studies is the possibility that not all case chromosomes have inherited disease-causing mutations from a common ancestral chromosome (founder heterogeneity). To alleviate the problem, we propose a method that applies a clustering algorithm to haplotype similarity analysis. The method identifies a sequence of nested subsets of case chromosomes by a peeling procedure, where each subset is relatively homogeneous. The average similarity score estimated from each subset in the sequence is compared to that estimated in controls, and a raw (unadjusted for multiple comparisons) P value is obtained. The test for the association between the trait and the candidate gene is based on the minimum raw P value observed in the comparison sequence, with its significance level estimated by a permutation procedure. The method can be applied to both haplotype and genotype data. Simulation studies suggest that our method has the correct type I error rate, and is generally more powerful than existing methods of haplotype similarity analysis.

Alleles↗

Unsupervised measurement of brain tumor volume on MR images.

We examined unsupervised methods of segmentation of MR images of the brain for measuring tumor volume in response to treatment. Two clustering methods were used: fuzzy c-means and a nonfuzzy clustering algorithm. Results were compared with volume segmentations by two supervised methods, k-nearest neighbors and region growing, and all results were compared with manual labelings. Results of individual segmentations are presented as well as comparisons on the application of the different methods with 10 data sets of patients with brain tumors. Unsupervised segmentation is preferred for measuring tumor volumes in response to treatment, as it eliminates operator dependency and may be adequate for delineation of the target volume in radiation therapy. Some obstacles need to be overcome, in particular regarding the detection of anatomically relevant tissue classes. This study shows that these improvements are possible.

Adult↗

Expression profiling of the developing testis in wild-type and Dazl knockout mice.

Genetic understanding of male-factor infertility requires knowledge of gene expression patterns associated with normal germ cell differentiation. The mouse is one of the best models of mammalian fertility due to its well-characterized genetics and the existence of many infertile mutants both naturally occurring and experimentally induced. We used cDNA microarrays firstly to investigate normal gene expression in the wild-type (wt) testis and secondly to gain a better insight into the effect of the disruption of the Dazl gene on spermatogenesis. We constructed a cDNA microarray from a subtracted and normalized adult testis library and focused on six developmental time-points during the initial synchronous wave of spermatogenesis. The results suggest that in the wild-type testis, 89.5% of genes on our chip change expression dramatically during the time-course. To identify patterns in the gene-expression data, a k-means clustering algorithm and principal component analysis were used. In the Dazl knockout testes, the majority of genes remain at baseline levels of expression, because absence of Dazl has a severe effect on cell-types present in the testis. Although in the prepubescent Dazl-null mice the final point reached in germ cell development is the leptotene-zygotene stage, the microarray results suggest that lack of Dazl expression has a detectable effect on the mRNA complement of germ cells as early as day 5 when only type A spermatogonia are present. Mol. Reprod. Dev. 67: 26-54, 2004.

Algorithms↗

Dual contrast TrueFISP imaging for left ventricular segmentation.

Based on varying tissue contrasts at different RF flip angles, a new TrueFISP imaging strategy for cardiac function measurement is presented. A single breath-hold dual RF flip angle cine multi-slice TrueFISP imaging sequence was implemented which provides a significant increase in signal contrast between blood and myocardium. The increase in image contrast combined with different characteristics in RF response facilitates the delineation of cardiovascular borders. Based on this imaging strategy it is demonstrated how a simple 2D histogram clustering algorithm can be used for the fully automatic segmentation of the left ventricular (LV) blood pool. The method is validated with data acquired from 10 asymptomatic subjects, and the results are shown to be comparable to that of manual delineation by experienced observers.

Algorithms↗

Automatic annotation of protein function based on family identification.

Although genomes are being sequenced at an impressive rate, the information generated tells us little about protein function, which is slow to characterize by traditional methods. Automatic protein function annotation based on computational methods has alleviated this imbalance. The most powerful current approach for inferring the function of new proteins is by studying the annotations of their homologues, since their common origin is assumed to be reflected in their structure and function. Unfortunately, as proteins evolve they acquire new functions, so annotation based on homology must be carried out in the context of orthologues or subfamilies. Evolution adds new complications through domain shuffling: homology (or orthology) frequently corresponds to domains rather than complete proteins. Moreover, the function of a protein may be seen as the result of combining the functions of its domains. Additionally, automatic annotation has to deal with problems related to the annotations in the databases: errors (which are likely to be propagated), inconsistencies, or different degrees of function specification. We describe a method that addresses these difficulties for the annotation of protein function. Sequence relationships are detected and measured to obtain a map of the sequence space, which is searched for differentiated groups of proteins (similar to islands on the map), which are expected to have a common function and correspond to groups of orthologues or subfamilies. This mapmaking is done by applying a clustering algorithm based on Normalized cuts in graphs. The domain problem is addressed in a simple way: pairwise local alignments are analyzed to determine the extent to which they cover the entire sequence lengths of the two proteins. This analysis determines both what homologues are preferred for functional inheritance and the level of confidence of the annotation. To alleviate the problems associated with database annotations, the information on all the homologues that are grouped together with the query protein are taken into account to select the most representative functional descriptors. This method has been applied for the annotation of the genome of Buchnera aphidicola (specific host Baizongia pistaciae). Human inspection of the annotations allowed an estimation of accuracy of 94%; the different kinds of error that may appear when using this approach are described. Results can be accessed at http://www.pdg.cnb.uam.es/funcut.html. The programs are available upon request, although installation in other systems may be complicated.

Algorithms↗

A 3D building blocks approach to analyzing and predicting structure of proteins.

A new approach is introduced for analyzing and ultimately predicting protein structures, defined at the level of C alpha coordinates. We analyze hexamers (oligopeptides of six amino acid residues) and show that their structure tends to concentrate in specific clusters rather than vary continuously. Thus, we can use a limited set of standard structural building blocks taken from these clusters as representatives of the repertoire of observed hexamers. We demonstrate that protein structures can be approximated by concatenating such building blocks. We have identified about 100 building blocks by applying clustering algorithms, and have shown that they can "replace" about 76% of all hexamers in well-refined known proteins with an error of less than 1 A, and can be joined together to cover 99% of the residues. After replacing each hexamer by a standard building block with similar conformation, we can approximately reconstruct the actual structure by smoothly joining the overlapping building blocks into a full protein. The reconstructed structures show, in most cases, high resemblance to the original structure, although using a limited number of building blocks and local criteria of concatenating them is not likely to produce a very precise global match. Since these building blocks reflect, in many cases, some sequence dependency, it may be possible to use the results of this study as a basis for a protein structure prediction procedure.

Algorithms↗

Male vocal imitation produces call convergence during pair bonding in budgerigars, Melopsittacus undulatus.

The budgerigar, a small species of parrot, can learn new vocalizations throughout life and is therefore widely used as a model system for studying various aspects of vocal learning. It is not known, however, why parrots imitate sounds. To test the hypothesis that vocal imitation in budgerigars is related to pair bonding, we recorded approximately 100 contact calls from each of nine male and nine female adult budgerigars that were unfamiliar with one another and then placed them into pairs. We sampled their contact call repertoire weekly and conducted twice-weekly behavioural observation sessions. We compared contact calls by sonagram cross-correlation and classified them by means of a hierarchical clustering algorithm. This analysis showed that all pairs developed a shared call within an average of 2.1 weeks. Further analysis revealed that eight of the nine male budgerigars imitated the contact calls of their assigned mates, while none of the females imitated the calls of their males. We conclude that contact call imitation in adult budgerigars probably contributes to pair bond formation and maintenance. Prior studies on budgerigars were limited by the lack of a behavioural paradigm to elicit vocal imitation reliably. Our study remedies this and thereby serves as a foundation for future studies on vocal learning in adult animals. Copyright 2000 The Association for the Study of Animal Behaviour.

Journal Article↗

Gene expression profiling in postmortem Rett Syndrome brain: differential gene expression and patient classification.

The identification of mutations in the transcriptional repressor methyl-CpG-binding protein 2 (MECP2) gene in Rett Syndrome (RTT) suggests that an inappropriate release of transcriptional silencing may give rise to RTT neuropathology. Despite this progress, the molecular basis of RTT neuropathogenesis remains unclear. Using multiple cDNA microarray technologies, subtractive hybridization, and conventional biochemistry, we generated comprehensive gene expression profiles of postmortem brain tissue from RTT patients and matched controls. Many glial transcripts involved in known neuropathological mechanisms were found to have increased expression in RTT brain, while decreases were observed in the expression of multiple neuron-specific mRNAs. Dramatic and consistent decreases in transcripts encoding presynaptic markers indicated a specific deficit in presynaptic development. Employing multiple clustering algorithms, it was possible to accurately segregate RTT from control brain tissue samples based solely on gene expression profile. Although previously achieved in cancers, our results constitute the first report of human disease classification using gene expression profiling in a complex tissue source such as brain.

Adolescent↗

Heterogeneity of Escherichia coli derived from artiodactyla animals analyzed with the use of rep-PCR fingerprinting.

Genetic polymorphism of 83 isolates of E. coli, derived from 4 species of artiodactyla animals living in a relatively close contact on the grounds of a theme park ZOO Safarii Swierkocin (Poland) was determined using the rep-PCR fingerprinting method, which utilizes oligonucleotide primers matching interspersed repetitive DNA sequences in PCR reaction to yield DNA fingerprints of individual bacterial isolates based on repetitive extragenic palindrome (REP) primers. The fingerprint patterns demonstrated the essential polymorphism of distribution of REP sequences in genomes of the examined isolates. The arithmetic averages clustering algorithm (UPGMA) statistical analysis of fingerprints with the use of the Jaccard similarity coefficient differentiated E. coli isolates into three similarity groups containing various numbers of isolates. The groups comprised isolates derived from two, three and four species of the source animals. The isolates derived from each source segregated in the dendrogram in a different way, both within the similarity groups and among them, indicating an individual repertoire of E. coli in the examined species of animals. The similarity relations among E. coli derived from the same source, illustrated in a dendrogram with a number of subclusters of a low mutual similarity (< or = 20%), indicated an essential interstrain differentiation in terms of the distribution of REP sequences. Our results confirmed the hypothesis of the oligoclonal characters of populations obtained from particular sources. The rep-PCR fingerprinting method with REP primers is simple and highly differentiating and can be recommended for use in explorations of large groups of animals and monitoring the variability of strains.

Animals↗

Effects of opioid receptor blockade on luteinizing hormone (LH) pulses and interpulse LH concentrations in normal women during the early phase of the menstrual cycle.

To determine the role of endogenous opioid peptides in regulating pulsatile luteinizing hormone (LH) release in the early follicular phase of the menstrual cycle of eumenorrheic women, we evaluated serum LH concentrations in blood collected every 10 min for 12 h in 27 women each studied during two menstrual cycles: (1) without pretreatment and (2) following oral administration of naltrexone, a mu opiate receptor blocking agent, at a dose of 1.0 mg/kg. Pulsatile LH release was assessed by the CLUSTER algorithm. The mean (+/- SE) integrated serum LH concentration (IU/L/min) increased following the administration of naltrexone (4715 +/- 298) in comparison to the control day (3997 +/- 381; p = 0.0008). The mean number of LH pulses (/12 h) detected on the naltrexone day (10.3 +/- 0.3) was higher than on the control day (8.9 +/- 0.4; p = 0.0068). Mean maximal LH peak height (IU/L) was greater on the naltrexone (7.8 +/- 0.5) vs control (6.7 +/- 0.5) days (p = 0.0064) as was the interpulse valley mean serum LH concentration (IU/L; 6.3 +/- 0.4 vs 5.0 +/- 0.4; p = 0.0013). No difference was noted in the mean incremental LH pulse amplitude (IU/L; 1.9 +/- 0.1 vs 2.1 +/- 0.1; p = 0.13), or peak duration (min; 40 +/- 1.8 vs 45.0 +/- 2.4; p = 0.06). Mean LH peak area (IU/L/min) was greater on the control (45.0 +/- 2.4) vs naltrexone (40 +/- 1.8) days (p = 0.0475).(ABSTRACT TRUNCATED AT 250 WORDS)

Administration, Oral↗

Effect of glycemic control on growth hormone and IGFBP-1 secretion in patients with type I diabetes mellitus.

Growth hormone (GH) secretion disorders have been reported in poorly controlled type I diabetes mellitus patients. Our work was aimed to evaluate GH secretion in 9 type I young diabetes mellitus patients as well as the low molecular weight IGF-binding protein secretion (IGFBP-1) in 5 of them. The patients did not show any signs of malnutrition or neurovascular complications, neither were they on any medication except for insulin. The study protocol included blood samples collection during a 24-h period for measurement of glucose, glycated hemoglobin, GH IGF-I and IGFBP-1 levels under two situations: on poor glycemic control and after 2-3 months on better control through systematic diet, low in carbohydrates and increase in insulin dosage. GH secretion data were analyzed by Cluster algorithm for pulsatility parameters; for rhythm assessment Cosinor method was used. The first study (poor control) reported significant increase of GH maximal and incremental amplitude and duration pulse values, when compared to the second study (better control). Mean 24-h secretion values as well mean GH for interpulse intervals (valleys) decreased, although not statistically significant. The fraction of pulsatile GH/24 h GH did not change significantly with better glycemic control. No changes in pulse frequency were observed. Mean IGF-I concentrations were significantly higher when patients were on better glycemic control. An ultradian variation for GH secretion was noticed in the first study (poor control) and a circadian variation in the second one (better control). IGFBP-1 analysis showed significant decrease of the mean 24-h values under better glycemic control. Linear regression analysis demonstrated a correlation between IGFBP-1 levels and fasting glucose levels. A circadian variation was present in IGFBP-1 secretion, irrespective of glycemic control. Therefore, we concluded that for type I diabetic patients: 1. GH secretion is increased on poor control, through maximal, incremental amplitude and pulse duration values; 2. IGFBP-1 values were significantly reduced and IGF-1 levels significantly higher after better glycemic control; 4. GH ultradian secretion is reported on poor control, and circadian on the better one, 5. IGFBP-1 circadian secretion occurred irrespective of glycemic control.

Adolescent↗

Multiresolution fuzzy clustering of functional MRI data.

Recent developments in the analysis of functional MRI data reveal a shift from hypothesis-driven statistical tests to unsupervised strategies. One of the most promising approaches is the fuzzy clustering algorithm (FCA), whose potential to detect activation patterns has already been demonstrated. But the FCA suffers from three drawbacks: first the computational complexity, second the higher sensitivity to noise and third the dependence on the random initialization. With the multiresolution approach presented here, these weak points are significantly improved, as is demonstrated in our tests with simulated and real functional MRI data.

Algorithms↗