PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Gene expression data analysis of human lymphoma using support vector machines and output coding ensembles.

The large amount of data generated by DNA microarrays was originally analysed using unsupervised methods, such as clustering or self-organizing maps. Recently supervised methods such as decision trees, dot-product support vector machines (SVM) and multi-layer perceptrons (MLP) have been applied in order to classify normal and tumoural tissues. We propose methods based on non-linear SVM with polynomial and Gaussian kernels, and output coding (OC) ensembles of learning machines to separate normal from malignant tissues, to classify different types of lymphoma and to analyse the role of sets of coordinately expressed genes in carcinogenic processes of lymphoid tissues. Using gene expression data from "Lymphochip", a specialised DNA microarray developed at Stanford University School of Medicine, we show that SVM can correctly separate normal from tumoural tissues, and OC ensembles can be successfully used to classify different types of lymphoma. Moreover, we identify a group of coordinately expressed genes related to the separation of two distinct subgroups inside diffuse large B-cell lymphoma (DLBCL), validating a previous Alizadeh's hypothesis about the existence of two distinct diseases inside DLBCL.

Artificial Intelligence↗

The pattern of malignant tumours: tumour registry data analysis, AFIP, Rawalpindi, Pakistan (1992-2001).

OBJECTIVE: To provide information regarding frequency of malignant tumours through data retrieved from pathology based tumour registry of AFIP, Rawalpindi, Pakistan. METHODS: All malignant tumours recorded with the AFIP tumour registry over a period of 10 years (1992-2001) were analysed in terms of age group, gender and type of tumour with relation to site. A comparison with the previously published material from same setting, national and international studies were also done. RESULTS: The total malignant tumours in the 10 years period were 21,168. Out of these, 12584 (59.5%) were seen in male patients while 8584 (40.5%) were in females. Total malignant tumours in pediatric age group were 927 (4.4%). The common malignant tumours in males in order of decreasing frequency were, those of prostate, skin, lymph node, leukaemia, urinary bladder, colorectum, bone, lung, stomach and liver. In females, breast carcinoma was on top followed by skin, leukaemia, ovary, coloretum, lymph node, bone, liver, cervix and gall bladder. In females, contrary to the Western studies and India, ovarian tumours were more frequent than cervical cancers. Comparison of this analysis with our previous analysis, national and international studies showed some interesting features. CONCLUSION: It was found that in males, tumours of the prostate were the most frequent as compared to the previous study, which showed lymphomas and leukemias to be the most common. On the other hand in females, tumours of the breast remained to be consistently most frequent.

Female↗

Multilocus sequence typing: Data analysis in clinical microbiology and public health.

Numerous computer-based statistical packages have been developed in recent years and it has become easier to analyze nucleotide sequence data and gather subsequent information that would not normally be available. Multilocus sequence typing (MLST) is used for characterizing isolates of bacterial and fungal species and uses nucleotide sequences of internal fragments of housekeeping genes. This method is finding a place in clinical microbiology and public health by providing data for epidemiological surveillance and development of vaccine policy. It adds greatly to our knowledge of the genetic variation that can occur within a species and has therefore been used for studies of population biology. Analysis requires the detailed interpretation of nucleotide sequence data obtained from housekeeping and nonhousekeeping genes. This is due to the amount of data generated from nucleotide sequencing and the information generated from an array of analytical tools improves our understanding of bacterial pathogens. This can benefit public health interventions and the development of enhanced therapies and vaccines. This review concentrates on the analytical tools used in MLST and their use in the clinical microbiology and public health fields.

Bacteria↗

Quantification of metabolites from single-voxel in vivo 1H NMR data of normal human brain by means of time-domain data analysis.

We present here a combination of time-domain signal analysis procedures for quantification of human brain in vivo 1H NMR spectroscopy (MRS) data. The method is based on a separate removal of a residual water resonance followed by a frequency-selective time-domain line-shape fitting analysis of metabolite signals. Calculation of absolute metabolite concentrations was based on the internal water concentration as a reference. The estimated average metabolite concentrations acquired from six regions of normal human brain with a single-voxel spin-echo technique for the N-acetylaspartate, creatine, and choline-containing compounds were 11.4 +/- 1.0, 6.5 +/- 0.5, and 1.7 +/- 0.2 mumol kg-1 wet weight, respectively. The time-domain analyses of in vivo 1H MRS data from different brain regions with their specific characteristics demonstrate a case in which the use of frequency-domain methods pose serious difficulties.

Aspartic Acid↗

Sorting points into neighborhoods (SPIN): data analysis and visualization by ordering distance matrices.

SUMMARY: We introduce a novel unsupervised approach for the organization and visualization of multidimensional data. At the heart of the method is a presentation of the full pairwise distance matrix of the data points, viewed in pseudocolor. The ordering of points is iteratively permuted in search of a linear ordering, which can be used to study embedded shapes. Several examples indicate how the shapes of certain structures in the data (elongated, circular and compact) manifest themselves visually in our permuted distance matrix. It is important to identify the elongated objects since they are often associated with a set of hidden variables, underlying continuous variation in the data. The problem of determining an optimal linear ordering is shown to be NP-Complete, and therefore an iterative search algorithm with O(n3) step-complexity is suggested. By using sorting points into neighborhoods, i.e. SPIN to analyze colon cancer expression data we were able to address the serious problem of sample heterogeneity, which hinders identification of metastasis related genes in our data. Our methodology brings to light the continuous variation of heterogeneity--starting with homogeneous tumor samples and gradually increasing the amount of another tissue. Ordering the samples according to their degree of contamination by unrelated tissue allows the separation of genes associated with irrelevant contamination from those related to cancer progression. AVAILABILITY: Software package will be available for academic users upon request.

Algorithms↗

An air quality data analysis system for interrelating effects, standards, and needed source reductions: Part 13--Applying the EPA Proposed Guidelines for Carcinogen Risk Assessment to a set of asbestos lung cancer mortality data.

The Clean Air Act Amendments of 1990 (CAAA-90) list 189 hazardous air pollutants (HAPs) for which "safe" ambient concentrations are to be determined. The primary purpose of this paper is to develop two mathematical models, lognormal and logarithmic, that effectively express excess lung cancer mortality as a function of asbestos concentration for an example set of data and also to suggest using these two models for additional HAPs. The secondary purpose of this paper is to calculate a "safe" asbestos concentration by first assuming a default linear extrapolation (to one excess death per million people, as specified for carcinogenic HAPs). The resulting "safe" concentration is an impossible-to-achieve 1/1000 of present background asbestos concentrations. A letter to the editor and a response in this Journal issue use additional asbestos data that suggest that the "safe" concentration should be about 730 times higher than first calculated here and that a default nonlinear extrapolation should be used instead, with the "safe" concentration proportional to the desired mortality level raised to the 0.39 power. These results suggest that the most important problem in setting a "safe" concentration for each carcinogenic HAP is to determine the correct nonlinear extrapolation to use for each HAP.

Asbestos↗

Safety of long-term therapy with ciprofloxacin: data analysis of controlled clinical trials and review.

We reviewed the literature and the manufacturer's U.S. clinical data pool for safety data on long-term administration of ciprofloxacin (Bayer, West Haven, CT). Only controlled clinical trials including patients treated for >30 days were selected. We identified 636 patients by literature search and 413 patients in the Bayer U.S. database who fulfilled our search criteria; the average treatment duration for these patients was 130 and 80 days, respectively. Main indications for long-term therapy were osteomyelitis, skin and soft-tissue infection, prophylaxis for urinary tract infection, mycobacterial infections, and inflammatory bowel disease. Adverse events, premature discontinuation of therapy, and deaths occurred at a similar frequency in both treatment arms. Most adverse events occurred early during therapy with little increase in frequency over time. As with short-term therapy, gastrointestinal events were more frequent than central nervous system or skin reactions, but pseudomembranous colitis was not observed. No previously unknown adverse events were noted. We conclude that ciprofloxacin is tolerated as well as other antibiotics when extended courses of therapy are required.

Arthritis, Reactive↗

Sparse drug concentration data analysis using a population approach: a valuable tool in clinical pharmacology.

1. Drug concentration or pharmacological effect data collected from patients during therapy or as part of a Phase III or post-marketing study are generally sparse (i.e. one or a few observations per patient) in nature. 2. The population approach to analysing sparse drug concentration data provides a valuable tool for obtaining information about the pharmacokinetics of drugs in special patient groups (neonates, aged or critically ill), the importance of drug interactions in the clinic (using routinely collected blood concentration data) and for conducting a 'pharmacokinetic screen' in patients during early phase efficacy trials. 3. Using a population approach to analyse drug concentration-time data collected from patients during therapy or during an early phase investigation can complement information obtained from traditional pharmacokinetic/dynamic investigations to help gain a further insight into the factors that influence dosing guidelines.

Clinical Trials as Topic↗

GeneMerge--post-genomic analysis, data mining, and hypothesis testing.

SUMMARY: GeneMerge is a web-based and standalone program written in PERL that returns a range of functional and genomic data for a given set of study genes and provides statistical rank scores for over-representation of particular functions or categories in the data set. Functional or categorical data of all kinds can be analyzed with GeneMerge, facilitating regulatory and metabolic pathway analysis, tests of population genetic hypotheses, cross-experiment comparisons, and tests of chromosomal clustering, among others. GeneMerge can perform analyses on a wide variety of genomic data quickly and easily and facilitates both data mining and hypothesis testing. AVAILABILITY: GeneMerge is available free of charge for academic use over the web and for download from: http://www.oeb.harvard.edu/hartl/lab/publications/GeneMerge.html.

Algorithms↗

Laser light scattering immunoassay: an improved data analysis by CONTIN method.

Laser light scattering immunoassay (LIA) is a diagnostic method for the detection of antibody by monitoring the agglutination of antigen carrier particles mediated by antibody, using dynamic light scattering (DLS) as probe. We have used this method for the detection of antibody to P. falciparum that cause malaria. The data were analysed using CONTIN method and the superiority of the distribution analysis over the conventional interpretation of the data in terms of mean diffusion coefficient or hydrodynamic radius is discussed in detail.

Agglutination↗

Improved R-factors for diffraction data analysis in macromolecular crystallography.

The quantity Rsym (also called Rmerge) is almost universally used for describing X-ray diffraction data quality. Here, we prove that Rsym is seriously flawed, because it has an implicit dependence on the redundance of the data. A corrected R-factor, Rmeas, is introduced as the equivalent robust indicator of data consistency. In addition, we introduce Rmrgd an R-factor that reflects the gain in accuracy upon averaging of equivalent reflections, as a useful indicator of the quality of reduced data. These new data quality indicators better reveal the benefits of highly redundant data and should stimulate improvements in data quality through increased merging of data from multiple crystals.

Crystallography, X-Ray↗

Complex sampling: implications for data analysis.

Investigators in dental public health often use strategies other than simple random sampling to identify potential subjects; however, their statistical analyses do not always take into account the complex sampling mechanism. Often it is not clear whether a given strategy requires adjustment for stratification and/or cluster sampling of observations. We propose that the need for such adjustment depends on the primary study objective. As a general rule, we recommend that if the study goal is to estimate the magnitude of either a population value of interest (e.g., prevalence), or an established exposure-outcome association, adjustment of variances to reflect complex sampling is essential because obtaining appropriate variance estimates is a priority. However, if the study goal is to establish the presence of an association, especially in a preliminary investigation of novel conditions or understudied populations, obtaining appropriate variance estimates may not be of primary importance; hence, adjustment of variances for complex sampling is not always required, but often is recommended. This paper describes several types of complex sampling designs, methods of adjusting for complex sampling strategies, examples illustrating the effect of adjustment, and alternative approaches for analysis of complex samples.

Analysis of Variance↗

[Heterozygote thalassemias screening. Contribution of data analysis methods (author's transl)].

In order to reduce the cost of a systematic heterozygote thalassemias screening, and using a discriminant analysis, the authors propose a pre-selection performed on the erythrocytometric data. In the present conditions of applicability, the methodology we used appears to be very efficient in the pre-screening process. An electrophoresis with the determination of hemoglobin A2 is then carried out on the pre-selected individuals as a confirmation of the diagnosis.

Adolescent↗