PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Global landscape of protein complexes in the yeast Saccharomyces cerevisiae.

Identification of protein-protein interactions often provides insight into protein function, and many cellular processes are performed by stable protein complexes. We used tandem affinity purification to process 4,562 different tagged proteins of the yeast Saccharomyces cerevisiae. Each preparation was analysed by both matrix-assisted laser desorption/ionization-time of flight mass spectrometry and liquid chromatography tandem mass spectrometry to increase coverage and accuracy. Machine learning was used to integrate the mass spectrometry scores and assign probabilities to the protein-protein interactions. Among 4,087 different proteins identified with high confidence by mass spectrometry from 2,357 successful purifications, our core data set (median precision of 0.69) comprises 7,123 protein-protein interactions involving 2,708 proteins. A Markov clustering algorithm organized these interactions into 547 protein complexes averaging 4.9 subunits per complex, about half of them absent from the MIPS database, as well as 429 additional interactions between pairs of complexes. The data (all of which are available online) will help future studies on individual proteins as well as functional genomics and systems biology.

Biological Evolution↗

MLL translocations specify a distinct gene expression profile that distinguishes a unique leukemia.

Acute lymphoblastic leukemias carrying a chromosomal translocation involving the mixed-lineage leukemia gene (MLL, ALL1, HRX) have a particularly poor prognosis. Here we show that they have a characteristic, highly distinct gene expression profile that is consistent with an early hematopoietic progenitor expressing select multilineage markers and individual HOX genes. Clustering algorithms reveal that lymphoblastic leukemias with MLL translocations can clearly be separated from conventional acute lymphoblastic and acute myelogenous leukemias. We propose that they constitute a distinct disease, denoted here as MLL, and show that the differences in gene expression are robust enough to classify leukemias correctly as MLL, acute lymphoblastic leukemia or acute myelogenous leukemia. Establishing that MLL is a unique entity is critical, as it mandates the examination of selectively expressed genes for urgently needed molecular targets.

Acute Disease↗

Distinctive gene expression profiles associated with Hepatitis B virus x protein.

Hepatitis B virus (HBV) is a major risk factor for the development of hepatocellular carcinoma (HCC). HBV encodes the potentially oncogenic HBx protein, which mainly functions as a transcriptional co-activator involving in multiple gene deregulations. However, mechanisms underlying HBx-mediated oncogenicity remain unclear. To determine the role(s) of HBx in the early genesis of HCC, we utilized the NCI Oncochip microarray that contains 2208 human cDNA clones to examine the gene expression profiles in either freshly isolated normal primary adult human hepatocytes (Hhep) or an HCC cell line (SK-Hep-1) ecotopically expressing HBx via an adenoviral system. The gene expression profiles also were determined in liver samples from HBV-infected chronic active hepatitis patients when compared with normal liver samples. The microarray results were validated through Northern blot analysis of the expression of selected genes. Using reciprocally labeling hybridizations, scatterplot analysis of gene expression ratios in human primary hepatocytes expressing HBx demonstrates that microarrays are highly reproducible. The comparison of gene expression profiles between HBx-expressing primary hepatocytes and HBV-infected liver samples shows a consistent alteration of many cellular genes including a subset of oncogenes (such as c-myc and c-myb) and tumor suppressor genes (such as APC, p53, WAF1 and WT1). Furthermore, clustering algorithm analysis showed distinctive gene expression profiles in Hhep and SK-Hep-1 cells. Our findings are consistent with the hypothesis that the deregulation of cellular genes by oncogenic HBx may be an early event that favors hepatocyte proliferation during liver carcinogenesis.

Adult↗

Tetanus antigen modulates the gene expression profile of aluminum phosphate adjuvant in spleen lymphocytes in vivo.

Adjuvants play an important role in stimulation of the immune response to antigens. Very little is known about the molecular mechanisms of this stimulation. Here we address this issue by studying gene expression profiles from spleen lymphocytes after in vivo immunization of mice with a clinically relevant vaccine, tetanus toxoid formulated with aluminum phosphate as adjuvant (TT(ADJ)), or the adjuvant alone (ADJ). The Th1/Th2 response to TT(ADJ) was obtained from a combination of up- and downstream markers to conventional cytokines, which were in good agreement with cytokine protein levels. A clustering algorithm revealed that ADJ elicited expression of 47 genes active in cytotoxic lymphocytes, inflammation, oncogenesis, stress, toxicity and cell cycle regulation. In TT(ADJ) these adjuvant-elicited genes were expressed at lower levels and a compensatory onset of protective and inhibitory genes was observed. We conclude that the antigen, to a larger extent than previously recognized, modulates the molecular mechanism of the aluminum phosphate adjuvant and that the identified genes may serve as predictive biomarkers in the development of new adjuvants and vaccines.

Adjuvants, Immunologic↗

Evaluation of forest nutrition based on large-scale foliar surveys: are nutrition profiles the way of the future?

This paper introduces the use of nutrition profiles as a first step in the development of a concept that is suitable for evaluating forest nutrition on the basis of large-scale foliar surveys. Nutrition profiles of a tree or stand were defined as the nutrient status, which accounts for all element concentrations, contents and interactions between two or more elements. Therefore a nutrition profile overcomes the shortcomings associated with the commonly used concepts for evaluating forest nutrition. Nutrition profiles can be calculated by means of a neural network, i.e. a self-organizing map, and an agglomerative clustering algorithm with pruning. As an example, nutrition profiles were calculated to describe the temporal variation in the mineral composition of Scots pine and Norway spruce needles in Finland between 1987 and 2000. The temporal trends in the frequency distribution of the nutrition profiles of Scots pine indicated that, between 1987 and 2000, the N, S, P, K, Ca, Mg and Al decreased, whereas the needle mass (NM) increased or remained unchanged. As there were no temporal trends in the frequency distribution of the nutrition profiles of Norway spruce, the mineral composition of the needles of Norway spruce needles subsequently did not change. Interpretation of the (lack of) temporal trends was outside the scope of this example. However, nutrition profiles prove to be a new and better concept for the evaluation of the mineral composition of large-scale surveys only when a biological interpretation of the nutrition profiles can be provided.

Algorithms↗

Hierarchical genetic structure of the introduced wasp Vespula germanica in Australia.

The wasp Vespula germanica is a highly successful invasive pest. This study examined the population genetic structure of V. germanica in its introduced range in Australia. We sampled 1320 workers and 376 males from 141 nests obtained from three widely separated geographical areas on the Australian mainland and one on the island of Tasmania. The genotypes of all wasps were assayed at three polymorphic DNA microsatellite markers. Our analyses uncovered significant allelic differentiation among all four V. germanica populations. Pairwise estimates of genetic divergence between populations agreed with the results of a model-based clustering algorithm which indicated that the Tasmanian population was particularly distinct from the other populations. Within-population analyses revealed that genetic similarity declined with spatial distance, indicating that wasps from nests separated by more than approximately 25 km belonged to separate mating pools. We suggest that the observed genetic patterns resulted from frequent bottlenecks experienced by the V. germanica populations during their colonization of Australia.

Animals↗

[Identification of preeclampsia by cDNA-gene expression profiling in human placentas and serum -- a pilot study].

OBJECTIVE: Preeclampsia is associated with significant maternal and fetal morbidity and mortality. The etiology remains unclear. For the accurate diagnosis and the prevention of preeclampsia it seems to be important to find a diagnostic tool that identifies risk patients before symptoms occur. With a new approach, the cDNA-Array analysis, human placentas and blood from preeclamptic and healthy pregnant women were examined for differentially expressed genes to find typical genes expression profiles. MATERIAL AND METHODS: In this pilot study, cDNA array analysis with a 19 200 gene array of placenta and blood samples from three preeclamptic patients have been performed to classify this samples based on expression patterns. RESULTS: Comparing normal placenta and blood from healthy delivered women (n = 4), a subset of 200 genes repeatedly found to be differentially expressed in preeclampsia. The placenta and blood samples from preeclampsia were accurately grouped by their individual gene expression patterns. CONCLUSIONS: These results suggest that the use of cDNA array is a tool to identify gene expression patterns in preeclampsia. With this set of differentially expressed genes in conjunction with sample clustering algorithms the identification of preeclampsia in placenta or blood samples is possible.

Adult↗

Spectral analysis of EEG in the late course of primary generalized myoclonic-astatic epilepsy. II. Cluster analysis of the power spectra.

Cluster analysis is applied to power spectra of the EEG of 38 patients with a primary generalized myoclonic astatic epilepsy (Gundel et al. 1981). The tendency of the data to form clusters within this group is indicated by a random experiment which has been performed with the data. The clustering algorithm divided the material in seven distinct groups which may be combined to three main types of power spectra. These types are power spectra with a 10 cps peak, a 4-7 cps peak and power spectra without remarkable rhythmization. The comparison of EEG types and clinical data shows a correlation of 4-7 cps rhythms with the occurrence of seizures. 4-7 cps rhythms are interpreted as a symptom of "centrencephalic" convulsibility.

Adolescent↗

Measuring the similarity between trajectories using clustering techniques.

A clustering method has been developed to group signals that display similar dynamic behavior. The procedure involves using the method of time delay embedding to construct a trajectory in state space from a time series. Certain features that characterize the geometry of the trajectory have been defined. These features were subjected to a series of statistical tests to determine their usefulness in a hierarchical clustering analysis. The latter is aimed at finding groups of similar trajectories. The trajectory-based clustering algorithm has been applied to simulated data, which included both stochastic data generated by a linear AR model, and nonlinear data generated by a Duffing oscillator. The results show that the algorithm works reliably in both cases.

Journal Article↗

Bioaccumulation of elements in bryophytes from Serra da Estrela, Portugal, and Veluwezoom, the Netherlands.

BACKGROUND, AIMS AND SCOPE: Pollution by heavy metals over large areas and long periods of time may cause chronic damage to living organisms and must be carefully controlled. One way to determine the extent of environmental contamination is by measuring the levels of contaminants in plants. The use of mosses as biomonitors is a convenient method to determine levels of (atmospheric) deposition, as terrestrial mosses obtain most of their supply of mineral elements from precipitation and dry deposition of airborne particles. Mosses have therefore received increasing attention as a suitable tool for monitoring regional patterns of elemental deposition from the atmosphere in large-scale studies in various countries, in areas close to industrial installations as well as in areas not expected to be contaminated. Although this technique is widely known, ecological studies of this type have rarely been done in Portugal. The aim of this paper is to evaluate and compare the spatial distribution of heavy metals in Hypnum cupressiforme, Pleurozium schreberi, Dicranum scoparium and Polytrichum piliferum collected from the Serra da Estrela natural park in Portugal and in the Veluwezoom natural park in the Netherlands. The selected species are the most widely used bryophytes for biomonitoring in the boreal region. The popularity of these species for this purpose is due to their wide ecological amplitude and distribution. METHODS: At 54 sampling sites in both nature parks, samples of Hypnum cupressiforme, Pleurozium schreberi, Dicranum scoparium and Polytrichum piliferum were collected. Plant digests were analysed for Al, Ba, Ca, Cr, Fe, K, Mg, Mn, Ni, Sr, V, Zn, Pb, Cu, Cd, N and P. Differentiations between sampling sites in terms of concentrations of elements in mosses were evaluated by ANOVA and the least significant difference was calculated. The normality of the analysed features was checked with the chi square test. After standardization, the matrix of 54 samples and 10 heavy metals was subjected to numerical classification to detect groups of samples with similar patterns of metal concentrations. The clustering algorithm was prepared with Ward's method, and the City Block Manhattan method was used for the similarity measure. Metals and samples were also subjected to ordination to reveal possible gradients of heavy metal levels, using PCA. Correlations were calculated between concentrations of metals and factors 1 and 2, allowing the dependence between the concentration of metals and factors (factor loading) to be estimated. RESULTS AND DISCUSSION: All species examined in both areas contained elevated levels of Mn and Pb. For each particular species, concentrations of N, P and Pb were significantly higher at Serra da Estrela, while concentrations of Cu were significantly higher at the Veluwezoom. Mosses from Portugal and the Netherlands differed significantly mainly in the concentrations of Al, Ba, Cr, Fe, Mn, Ni, Pb and V. This differentiation did not exceed that within the mosses from Portugal. CONCLUSIONS: Mosses from Portugal and the Netherlands differ significantly mainly in the concentrations of Al, Ba, Cr, Fe, Mn, Ni, Pb and V. This differentiation does not exceed the differentiation within the mosses from Portugal. RECOMMENDATION AND OUTLOOK: Further research is required into the origin and deposition of the polluting elements in other environmental compartments.

Bryophyta↗

Genetics and imaging to assess oocyte and preimplantation embryo health.

Two major criteria are currently used in human assisted reproductive technologies (ART) to evaluate oocyte and preimplantation embryo health: (1) rate of preimplantation embryonic development; and (2) overall morphology. A major gene that regulates the rate of preimplantation development is the preimplantation embryo development (Ped) gene, discovered in our laboratory. In mice, presence of the Ped gene product, Qa-2 protein, results in a fast rate of preimplantation embryonic development, compared with a slow rate of preimplantation embryonic development for embryos that are lacking Qa-2 protein. Moreover, mice that express Qa-2 protein have an overall reproductive advantage that extends beyond the preimplantation period, including higher survival to birth, higher birthweight, and higher survival to weaning. Data are presented that suggest that Qa-2 increases the rate of development of early embryos by acting as a cell-signalling molecule and that phosphatidylinositol-32 kinase is involved in the cell-signalling pathway. The most likely human homologue of Qa-2 has recently been identified as human leukocyte antigen (HLA)-G. Data are presented which show that HLA-G, like Qa-2, is located in lipid rafts, implying that HLA-G also acts as a signalling molecule. In order to better evaluate the second criterion used in ART (i.e. overall morphology), a unique and innovative imaging microscope has been constructed, the Keck 3-D fusion microscope (Keck 3DFM). The Keck 3DFM combines five different microscopic modes into a single platform, allowing multi-modal imaging of the specimen. One of the modes, the quadrature tomographic microscope (QTM), creates digital images of non-stained transparent cells by measuring changes in the index of refraction. Quadrature tomographic microscope images of oocytes and preimplantation mouse embryos are presented for the first time. The digital information from the QTM images should allow the number of cells in a preimplantation embryo to be counted non-invasively. The Keck 3DFM is also being used to assess mitochondrial distribution in mouse oocytes and embryos by using the k-means clustering algorithm. Both the number of cells in preimplantation embryos and mitochondrial distribution are related to oocyte and embryo health. New imaging data obtained from the Keck 3DFM, combined with genetic and biochemical approaches, have the promise of being able to distinguish healthy from unhealthy oocytes and embryos in a non-invasive manner. The goal is to apply the information from our mouse model system to the clinic in order to identify one and only one healthy embryo for transfer back to the mother undergoing an ART procedure. This approach has the potential to increase the success rate of ART and to decrease the high, and undesirable, multiple birth rate presently associated with ART.

Animals↗

DNA microarray analysis of the uninoculated eye following anterior chamber inoculation of HSV-1.

PURPOSE: To use DNA microarray to analyze the expression patterns of genes in the uninoculated eye following uniocular anterior chamber inoculation of HSV-1. METHODS: On Day 9 following inoculation of 2 x 10( 4) PFU of HSV-1 (KOS strain) or an equivalent volume of tissue culture medium into one anterior chamber of BALB/c mice, the uninoculated eyes were enucleated, pooled, and total RNA was isolated. cDNA was synthesized from the total RNA. The gene expression patterns were inferred based on the hybridization intensities of the probes on the cDNA array. The hybridization signals were globally normalized and filtered. The data were analyzed using hierarchical and gene tree clustering algorithms. Additional uninoculated eyes collected on Day 9 p.i. were stained for F4/80 and CD19. RESULTS: Compared with the uninoculated eye of control mice, 3800 genes were upregulated at least twofold in the contralateral eye of HSV-1-infected mice. Among the 10 most upregulated genes, T cell-specific protein, MHC II antigen A, and MHC II k region locus 2 were upregulated 179-, 164-, and 162-fold, respectively. Ten T-cell receptor-related genes, 61 cytokine and chemokine genes, and 16 MHC genes were upregulated. Furthermore, 11 immunoglobulin and B cell genes and 11 macrophage-related genes were also upregulated. F4/80+ and CD19+ cells were observed on Day 9 p.i. CONCLUSIONS: The DNA microarray results support the idea that T cells and immunomodulatory factors (cytokines, chemokines) are likely to be involved in HSV-1 retinitis. These results also suggest that B cells and/or macrophages play a role in the pathogenesis of HSV-1 retinitis.

Animals↗

Empirically derived personality types among male and female college students.

Data on the 16 PF obtained from 130 male and female college students were cluster analyzed to produce an empirical personality typology. Two different clustering algorithms were compared. Seven personality types emerged: 1) Well-Adjusted Conservative, 2) Ego-Involved Neurotic, 3) Norm Independent, 4) Socially-Detached Neurotic, 5) Superego Controlled, 6) Self-Assured Experimenter, and 7) Tough-Minded Controlled. Not only did the types differ significantly in personality, but they also were found to be significantly different on nine different measures of interpersonal orientation. Since the types did give intuitive insight into the nature of personality as particular combinations of personality traits and also were different on variables other than those used for the classification, the scientific utility of the typological approach received support.

Ego↗

Audiometric configurations of hearing impaired children in Hong Kong: implications for amplification.

PURPOSE: Children with hearing loss who require special school placement may have a wide range of audiometric configurations. Since such children will vary in auditory status their amplification requirements may also be diverse. This study examined the audiological records of 231 children attending four schools for hearing impaired children in Hong Kong to gain an understanding of common audiometric patterns found in the school children and their auditory rehabilitation needs. METHOD: Data on the children's aetiology of hearing loss, hearing status, tympanometric findings and the electroacoustic characteristics of their hearing aids were obtained. For 424 children's ears considered having essentially sensorineural hearing loss, k-means cluster analysis methods were used to categorize audiometric configuration groups. RESULTS: Cluster analysis that indicated that five distinct audiometric configurations could be found among the school children. Different clusters contained children who had differing amplification needs. The study analysed a number of parameters to check fitting outcomes, including average prescribed gain, frequency-specific measured versus prescribed gain, prescribed frequency response, measured versus prescribed frequency response and the predicted aided thresholds for the children. CONCLUSION: The amplification needs associated with these five configurations, including recommended prescription gain, maximum power output and possible signal processing strategies, were considered. The clustering algorithm approach proved useful as a means of grouping distinctive audiometric profiles.

Analysis of Variance↗

An image-analysis system based on support vector machines for automatic grade diagnosis of brain-tumour astrocytomas in clinical routine.

An image-analysis system based on the concept of Support Vector Machines (SVM) was developed to assist in grade diagnosis of brain tumour astrocytomas in clinical routine. One hundred and forty biopsies of astrocytomas were characterized according to the WHO system as grade II, III and IV. Images from biopsies were digitized, and cell nuclei regions were automatically detected by encoding texture variations in a set of wavelet, autocorrelation and parzen estimated descriptors and using an unsupervised SVM clustering methodology. Based on morphological and textural nuclear features, a decision-tree classification scheme distinguished between different grades of tumours employing an SVM classifier. The system was validated for clinical material collected from two different hospitals. On average, the SVM clustering algorithm correctly identified and accurately delineated 95% of all nuclei. Low-grade tumours were distinguished from high-grade tumours with an accuracy of 90.2% and grade III from grade IV with an accuracy of 88.3% The system was tested in a new clinical data set, and the classification rates were 87.5 and 83.8%, respectively. Segmentation and classification results are very encouraging, considering that the method was developed based on every-day clinical standards. The proposed methodology might be used in parallel with conventional grading to support the regular diagnostic procedure and reduce subjectivity in astrocytomas grading.

Astrocytoma↗

Transcriptome profiling to identify genes involved in peroxisome assembly and function.

Yeast cells were induced to proliferate peroxisomes, and microarray transcriptional profiling was used to identify PEX genes encoding peroxins involved in peroxisome assembly and genes involved in peroxisome function. Clustering algorithms identified 224 genes with expression profiles similar to those of genes encoding peroxisomal proteins and genes involved in peroxisome biogenesis. Several previously uncharacterized genes were identified, two of which, YPL112c and YOR084w, encode proteins of the peroxisomal membrane and matrix, respectively. Ypl112p, renamed Pex25p, is a novel peroxin required for the regulation of peroxisome size and maintenance. These studies demonstrate the utility of comparative gene profiling as an alternative to functional assays to identify genes with roles in peroxisome biogenesis.

Gene Expression Profiling↗

Population structure, admixture, and aging-related phenotypes in African American adults: the Cardiovascular Health Study.

U.S. populations are genetically admixed, but surprisingly little empirical data exists documenting the impact of such heterogeneity on type I and type II error in genetic-association studies of unrelated individuals. By applying several complementary analytical techniques, we characterize genetic background heterogeneity among 810 self-identified African American subjects sampled as part of a multisite cohort study of cardiovascular disease in older adults. On the basis of the typing of 24 ancestry-informative biallelic single-nucleotide-polymorphism markers, there was evidence of substantial population substructure and admixture. We used an allele-sharing-based clustering algorithm to infer evidence for four genetically distinct subpopulations. Using multivariable regression models, we demonstrate the complex interplay of genetic and socioeconomic factors on quantitative phenotypes related to cardiovascular disease and aging. Blood glucose level correlated with individual African ancestry, whereas body mass index was associated more strongly with genetic similarity. Blood pressure, HDL cholesterol level, C-reactive protein level, and carotid wall thickness were not associated with genetic background. Blood pressure and HDL cholesterol level varied by geographic site, whereas C-reactive protein level differed by occupation. Both ancestry and genetic similarity predicted the number and quality of years lived during follow-up, but socioeconomic factors largely accounted for these associations. When the 24 genetic markers were tested individually, there were an excess number of marker-trait associations, most of which were attenuated by adjustment for genetic ancestry. We conclude that the genetic demography underlying older individuals who self identify as African American is complex, and that controlling for both genetic admixture and socioeconomic characteristics will be required in assessing genetic associations with chronic-disease-related traits in African Americans. Complementary methods that identify discrete subgroups on the basis of genetic similarity may help to further characterize the complex biodemographic structure of human populations.

Black or African American↗

Identifying and quantifying sources of variation in microarray data using high-density cDNA membrane arrays.

Microarray experiments involve many steps, including spotting cDNA, extracting RNA, labeling targets, hybridizing, scanning, and analyzing images. Each step introduces variability, confounding our ability to obtain accurate estimates of the biological differences between samples. We ran repeated experiments using high-density cDNA microarray membranes (Research Genetics Human GeneFilters Microarrays Version I) and 33P-labeled targets. Total RNA was extracted from a Burkitt lymphoma cell line (GA-10). We estimated the components of variation coming from: (1) image analysis, (2) exposure time to PhosphorImager screens, (3) differences in membranes, (4) reuse of membranes, and (5) differences in targets prepared from two independent RNA extractions. Variation was assessed qualitatively using a clustering algorithm and quantitatively using a version of ANOVA adapted to multivariate microarray data. The largest contribution to variation came from reusing membranes, which contributed 38% of the total variation. Differences in membranes and in exposure time each contributed about 10%. Differences in target preparations contributed less than 5%. The effect of image quantification was negligible. Much of the effect from reusing membranes was attributable to increasing levels of background radiation and can be reduced by using membranes at most four times. The effects of exposure time, which were partly attributable to variation in the scanning process, can be minimized by using the same exposure time for all experiments.

Algorithms↗