PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Lessons learned from the data analysis of the second harvest (1998-2001) of the Society of Thoracic Surgeons (STS) Congenital Heart Surgery Database.

OBJECTIVE: The analysis of the second harvest of the STS Congenital Heart Surgery Database produced meaningful outcome data and several critical lessons relevant to congenital heart surgery outcomes analysis worldwide. METHODS: This data harvest represents the first STS multi-institutional experience with software utilizing the nomenclature and database requirements adopted by the STS and EACTS (April 2000 Annals of Thoracic Surgery). Members of the STS Congenital Heart Committee analyzed the STS data. RESULTS: This STS harvest includes data from 16 centers (12787 cases, 2881 neonates, 4124 infants). In 2002, the EACTS reported similar outcome data utilizing the same database definitions (41 centers, 12736 cases, 2245 neonates, 4195 infants). Lessons from the analysis include: (1) Death must be clearly defined. (2) The Primary Procedure in a given operation must be documented. (3) Inclusionary and exclusionary criteria for all diagnoses and procedures must be agreed upon. (4) Missing data values remain an issue for the database. (5) Generic terms in the nomenclature lists, that is terms ending in Not Otherwise Specified (NOS), are redundant and decrease the clarity of data analysis. (6) Methodology needs to be developed and implemented to assure and verify data completeness and data accuracy. 'Operative Mortality' and 'Mortality Assigned to this Operation' were defined by the STS and EACTS; these definitions were not utilized uniformly. 'Thirty Day Mortality' was problematic because some centers did not track mortality after hospital discharge. Only 'Mortality Prior to Discharge' was consistently reported. Designation of Primary Procedure for a given operation determines its location for analysis. Until Complexity Scores lead to automated methodology for choosing the Primary Procedure, the surgeon must designate the Primary Procedure. Inclusionary and exclusionary criteria for all diagnoses and procedures have been developed in an effort to define acceptable concomitant diagnoses and procedures for each analysis. Improvements in data completeness can be achieved using a variety of techniques including developing more functional techniques of data entry at individual institutions and software improvements. Future versions of the STS Congenital Database will request that the coding of diagnoses and procedures avoid the terms ending in NOS. CONCLUSIONS: Lessons from this data harvest should improve congenital heart surgery outcome analysis.

Databases, Factual↗

[Investigation of fuzzy-clustering in octane number prediction model based on detailed hydrocarbon analysis data].

A method to establish octane number prediction model based on detailed hydrocarbon analysis (DHA) data is presented. The techniques of fuzzy-clustering and the Euclidian distance are employed to select the samples needed in pattern establishment. One hundred and fifty gasoline samples and an amount of 140 characteristic components in the DHA chromatogram of each sample are used for the fuzzy-clustering research. It is found that the 3 - 10 samples, which have the nearest Euclidian distance ( < 1.5) to the prediction sample in the same cluster, are enough to build the octane number prediction model. The experimental results proved that the model obtained according to the above method has more predictable accuracy, wider application range and higher data resource utility compared with the current prediction method.

Cluster Analysis↗

The effect of aging on functional decline among older Japanese living in a community: a 5-year longitudinal data analysis.

BACKGROUND AND AIMS: Using longitudinal data analyses, we examined the effects of aging on functional decline, based on activities of daily living (ADL) and instrumental activities of daily living (IADL) during a 5-year follow-up among older people living in a community in Japan. METHODS: The baseline survey in July 1988 involved all elderly residents aged 60 or older in Saku City, Nagano, Japan (N=13418). All survivors of this cohort were asked to participate in follow-up surveys conducted in 1989, 1990, 1991, 1992 and 1993. Five items of ADL and five of IADL were measured on each survey. A generalized estimating equations (GEE) analysis was used to examine the effects of aging on the increase of the proportion of subjects with functional dependence. RESULTS: These results indicated that the proportion of subjects who were dependent in ADL increased during the 5-year period by 2.2 times (p<0.001) and the proportion of those who were dependent in either ADL or IADL increased during the same period by 1.8 times (p<0.001). Gender did not appear to be significantly associated with functional decline. CONCLUSIONS: The GEE analysis in this study identified the statistically significant effect of aging on the increase of the proportion of subjects with functional dependence based on ADL and IADL.

Activities of Daily Living↗

["Pyromania" and arson. A psychiatric and criminologic data analysis].

We analyzed psychiatric and criminological data from 103 arsonists. The following criticisms of the definition of pyromania according to DSM-III-R and IDC-10 seem appropriate. First, the categoric exclusion of aggressive motives does not seem very promising, since approximately one fourth of arsonists whose firesetting is based on motives quoted in DSM-III-R may also have an aggressive motive. Second, ICD-10 gives being drunk and alcoholism as a criterion for the exclusion of pyromania. This seems untenable, since the behavior classed as pyromania is largely a product of alcohol misuse. Repeated firesetting, resulting from being fascinated by fire etc., may be less a disturbance of impulse control but rather the manifestation of a psychoinfantilism, which, supported by alcohol abuse, extends into older age. The mean age of such arsonists is slightly above 20 years. The tendency for relapses after imprisonment seems to be low; this tendency probably decreases spontaneously in older age. The mean age of arsonists with aggressive motives is a little below 30 years, those setting fire with suicidal motives have a mean age of 35, deluded arsonists have a mean age of 40 years. Concrete sexual motives are relatively rare. Approximately 50% of arsonists have a purely aggressive motive. Retaliation is a rare cause, however, since most of them do not even know the victims. One third of these persons set the fire in their own homes. Most arsonists show a personality disorder, with insecurity and narcissism predominating. Data on firesetting are to be treated with caution, since two thirds of all cases are newer resolved; one fourth of cases concern minors, and in Central Europe arsonists with rational motives are hardly ever referred to psychiatrists.

Adult↗

Oesophageal cancer treatment in North East Thames region, 1981: medical audit using Hospital Activity Analysis data.

Figures from the Hospital Activity Analysis in the North East Thames region in 1981 were used to perform a medical audit on oesophageal cancer treatment. Four hundred and forty four patients were admitted with this diagnosis; 80 had been intubated without a thoracotomy or laparotomy, and 73 had had surgery (two thirds radical and one third palliative) with an overall operative mortality of 33%. Fifty five patients had had radiotherapy and 179 patients had no recorded operation or investigation. One hundred and seventy seven different consultants had looked after all these inpatients, most being general surgeons. Only five consultants had looked after 10 or more patients each year. From a calculated estimate of a total 286 patients in the region, 28% had palliative intubation and 25% had surgery; 20% of all the patients had radiotherapy either as a radical or palliative treatment, the remainder having no recorded therapeutic procedure. One hundred and eighty seven patients (66% of the calculated total) died in hospital. Investigation and treatment do not seem to be limited by lack of money, but money is being wasted by admitting patients for terminal care into acute hospital beds. It would be more humane for these patients to die at home or in a hospice if they wished.

Aged↗

Interpreter of maladies: redescription mining applied to biomedical data analysis.

Comprehensive, systematic and integrated data-centric statistical approaches to disease modeling can provide powerful frameworks for understanding disease etiology. Here, one such computational framework based on redescription mining in both its incarnations, static and dynamic, is discussed. The static framework provides bioinformatic tools applicable to multifaceted datasets, containing genetic, transcriptomic, proteomic, and clinical data for diseased patients and normal subjects. The dynamic redescription framework provides systems biology tools to model complex sets of regulatory, metabolic and signaling pathways in the initiation and progression of a disease. As an example, the case of chronic fatigue syndrome (CFS) is considered, which has so far remained intractable and unpredictable in its etiology and nosology. The redescription mining approaches can be applied to the Centers for Disease Control and Prevention's Wichita (KS, USA) dataset, integrating transcriptomic, epidemiological and clinical data, and can also be used to study how pathways in the hypothalamic-pituitary-adrenal axis affect CFS patients.

Algorithms↗

Permutation methods for the structured exploratory data analysis (SEDA) of total cholesterol measured in five Israeli populations.

Three structured exploratory data analysis-functionals are applied to plasma total cholesterol concentrations measured for 2,480 young men and women aged 17-18 years and living in Jerusalem, and for their parents. These triad families are divided into five groups according to whether both parents were born in Asia, North Africa, Europe-America, or Israel or whether they were of mixed "origins." The significances of the functionals were determined by a spectrum of permutation techniques that selectively shuffled the trait values across families in order to systematically alter certain family structure relationships while keeping other familial relationships intact. These analyses suggest that generational differences and various distributional effects influence patterns of spouse and parent-offspring interactions within these families and that the nature and forms of these effects and interactions may differ according to the origin of the parents. Results are discussed in relationship to historical and cultural differences among groups.

Adolescent↗

Customized dual data entry for computerized data analysis.

A major responsibility of any Quality Assurance Unit (QUA) is to ensure data integrity. Errors made during data entry can lead to many problems in the study review process and decrease the quality, accuracy, and overall efficiency of data management. One technique that can reduce the number of data entry errors in computer data sets is the use of a dual entry data system. Currently available software allows creation of customized data entry screens that either closely resemble or duplicate the data collection forms used during studies. Two data entry operators enter data into two independent data sets. The use of an on-screen display that resembles the data collection form reduces the potential for keypunch errors. The two data sets can then be electronically compared. The comparison reports differences between the two data sets. When differences exist, the correct values can be determined by reference to the original data sheets and the two data files can then be corrected. Theoretically, the only key punch errors that will exist after making these corrections are when the two independent entry operators make the same exact data entry error. Typically, the time required for two people to enter data is minimal compared to the time required to manually identify and correct data entry discrepancies. With error-free data entry, we have found that electronic data quality, accuracy, and audit efficiency are improved at every subsequent step of data management, analysis, quality assurance auditing, and report generation.

Information Systems↗

Multiscale and Bayesian approaches to data analysis in genomics high-throughput screening.

Tremendous amounts of data are produced by high-throughput screening methods currently employed in drug discovery and product development. A typical cDNA microarray or oligonucleotide-based gene chip experiment easily generates over 10,000 data points for each array or chip. The challenge of inferring meaningful information is formidable given the size and number of these datasets. This paper reviews the current status of statistical tools available for gene expression analysis, with emphasis on Bayesian approaches and multiscale wavelet filtering. Fundamental concepts of Bayesian and multiscale modeling are discussed from the perspective of their potential to address important issues related to the analysis of gene expression data, such as the fact that genomic data often have non-Gaussian distributions and feature localization and multiple scales in both frequency and measurement dimension. Recent publications in these areas are reviewed. Wavelet filtering and the advantages of multiscale methods are demonstrated by application to publicly available gene expression data from the National Cancer Institute (NCI). Multiscale methods, including multiscale principal component analysis (MSPCA), are applied to extract gene subsets and to visualize data in multidimensions for comparisons. Similarity in cell lines and gene selection are effectively visualized and quantitatively compared.

Animals↗

Quantified neurophysiology with mapping: statistical inference, exploratory and confirmatory data analysis.

Topographic mapping of brain electrical activity has become a commonly used method in the clinical as well as research laboratory. To enhance analytic power and accuracy, mapping applications often involve statistical paradigms for the detection of abnormality or difference. Because mapping studies involve many measurements and variables, the appearance of a large data dimensionality may be created. If abnormality is sought by statistical mapping procedures and if the many variables are uncorrelated, certain positive findings could be attributable to chance. To protect against this undesirable possibility we advocate the replication of initial findings on independent data sets. Statistical difference attributable to chance will not replicate, whereas real difference will reproduce. Clinical studies must, therefore, provide for repeat measurements and research studies must involve analysis of second populations. Furthermore, Principal Components Analysis can be employed to demonstrate that variables derived from mapping studies are highly intercorrelated and data dimensionality substantially less than the total number of variables initially created. This reduces the likelihood of capitalization on chance. The need to constrain alpha levels is not necessary when dimensionality is low and/or a second data set is available. When only one data set is available in research applications, techniques such as the Bonferroni correction, the "leave-one-out" method, and Descriptive Data Analysis (DDA) are available. These techniques are discussed, clinical and research examples are given, and differences between Exploratory (EDA) and Confirmatory Data Analysis (EDA) are reviewed.

Brain↗

The use of a personal computer for trend data analysis with the Ohmeda 3700 pulse oximeter.

The Ohmeda 3700 pulse oximeter provides trend data storage of arterial oxygen saturation (SaO2) and pulse rate measurements for a maximum of 8 hours. This feature allows the oximeter to be used as a stand-alone unit for overnight studies of saturation during sleep. Subsequent transfer and processing of the stored SaO2 data requires additional software. We present a program for data processing that uses the Lotus 1-2-3 program on IBM and compatible microcomputers and that performs a data distribution and statistical analysis on these trend data and presents them in graphic form. Processing SaO2 trend data within the Lotus 1-2-3 worksheet format allows the user to easily add to or modify the program presented here, depending on individual needs.

Adult↗

Exploratory data analysis and the use of the hazard function for interpreting survival data: an investigator's primer.

This report discusses how one can use the hazard function to gain important insights on the patterns of failure in clinical studies when the principal endpoint is a time metric. These new insights may help gain increased understanding into the pathogenesis of a chronic disease and how it is affected by treatment intervention. The qualitative behavior of the hazard function can reveal whether mortality is increasing, decreasing, or is constant over time. Simple graphic plots are all that is necessary to show characteristic failure patterns. These informal procedures are in the spirit of carrying out exploratory analyses on the data. This report discusses the organization of clinical data using a "branch and leaf" plot, outlines the calculation of the hazard function and life table, and uses examples from lung cancer and uveal melanoma to illustrate calculations and ways of interpreting hazard functions.

Adenocarcinoma↗

Multiclass Decision Forest--a novel pattern recognition method for multiclass classification in microarray data analysis.

The wealth of knowledge imbedded in gene expression data from DNA microarrays portends rapid advances in both research and clinic. Turning the prodigious and noisy data into knowledge is a challenge to the field of bioinformatics, and development of classifiers using supervised learning techniques is the primary methodological approach for clinical application using gene expression data. In this paper, we present a novel classification method, multiclass Decision Forest (DF), that is the direct extension of the two-class DF previously developed in our lab. Central to DF is the synergistic combining of multiple heterogenic but comparable decision trees to reach a more accurate and robust classification model. The computationally inexpensive multiclass DF algorithm integrates gene selection and model development, and thus eliminates the bias of gene preselection in crossvalidation. Importantly, the method provides several statistical means for assessment of prediction accuracy, prediction confidence, and diagnostic capability. We demonstrate the method by application to gene expression data for 83 small round blue-cell tumors (SRBCTs) samples belonging to one of four different classes. Based on 500 runs of 10-fold crossvalidation, tumor prediction accuracy was approximately 97%, sensitivity was approximately 95%, diagnostic sensitivity was approximately 91%, and diagnostic accuracy was approximately 99.5%. Among 25 genes selected to distinguish tumor class, 12 have functional information in the literature implicating their involvement in cancer. The four types of SRBCTs samples are also distinguishable in a clustering analysis based on the expression profiles of these 25 genes. The results demonstrated that the multiclass DF is an effective classification method for analysis of gene expression data for the purpose of molecular diagnostics.

Carcinoma, Small Cell↗

An international data analysis on the level of maternal and child health in relation to socioeconomic factors.

International data on health and socioeconomic factors were analyzed to understand the trends and the determinants of maternal and infant mortality in the late years. Multivariate analyses were carried out to summarize the structure of the data. Multiple regression analyses were also carried out with these two mortality rates as dependent variables. The range of independent variables included health resource availability, immunization, GNP, illiteracy rates, distribution in working area, the indicators of living standards such as percentage of telephone lines and television sets per capita and the percentages of working children, population with access to safe water and sanitation, people living in urban areas, among others. In the preliminary analysis the indicators of living standards appeared highly correlated to maternal and infant mortality. Working area (industrial or agricultural) showed also an important correlation. In factor analysis indirect variables (economic and living condition) were summarized into two factors. Two regression analyses were executed. In the first the variables were used directly, while factors obtained by the factor analysis were used in the second. The second analysis confirmed the previous analysis: fertility rate, immunization and urbanization appeared as determinants of maternal mortality. Birth rate, percentage of females working in agriculture and total illiteracy appeared as determinants of infant mortality. The factors extracted in the factor analysis made a significant contribution to the second regression analysis. We concluded: 1) The factors extracted by factor analyses from indirect variables had high explanatory ability on infant mortality rates, 2) The presence of immunization together with birth rate and fertility rate in the regression models pointed out the importance of investing in birth rate reduction and disease prevention methods.

Child↗

Mayday--a microarray data analysis workbench.

UNLABELLED: Mayday is a workbench for visualization, analysis and storage of microarray data. It features a graphical user interface and supports the development and integration of existing and new analysis methods. Besides the infrastructural core functionality, Mayday offers a variety of plug-ins, such as various interactive viewers, a connection to the R statistical environment, a connection to SQL-based databases and different data mining methods, including WEKA-library based methods for classification and various clustering methods. In addition, so-called meta information objects are provided for annotation of the microarray data allowing integration of data from different sources, which is a feature that, for instance, is employed in the enhanced heatmap visualization. SUPPLEMENTARY INFORMATION: The software and more detailed information including screenshots and a user guide as well as test data can be found on the Mayday home page http://www.zbit.uni-tuebingen.de/pas/mayday. The core is published under the GPL (GNU Public License) and the associated plug-ins under the LGPL (Lesser GNU Public License).

Computer Graphics↗

Data analysis of the Second International Workshop on Small Cell Lung Cancer Antigens.

Methods of data collection for the 2nd Small Cell Lung Cancer Workshop are described, and data reliability is reviewed. The method of cluster analysis of the workshop antibodies is described and discussed. Of the 27,111 results submitted 20,705 were judged to be reliable for analysis and 13,802 of these came from immunohistology experiments. Data derived from immunocytochemistry experiments were somewhat less reproducible than flow cytometry, immunohistology and ELISA experiments. The cluster analysis was developed from methods employed in the leucocyte antigens workshops. Several checks on the methods of cluster analysis and the transformation of data did not substantially alter the final groupings. The workshop confirms that, although there are some methodological difficulties, the cluster analysis can successfully be applied to data derived largely from immunohistology, and thus has applicability to other tumour types.

Antibodies, Monoclonal↗

Longitudinal data analysis for discrete and continuous outcomes.

Longitudinal data sets are comprised of repeated observations of an outcome and a set of covariates for each of many subjects. One objective of statistical analysis is to describe the marginal expectation of the outcome variable as a function of the covariates while accounting for the correlation among the repeated observations for a given subject. This paper proposes a unifying approach to such analysis for a variety of discrete and continuous outcomes. A class of generalized estimating equations (GEEs) for the regression parameters is proposed. The equations are extensions of those used in quasi-likelihood (Wedderburn, 1974, Biometrika 61, 439-447) methods. The GEEs have solutions which are consistent and asymptotically Gaussian even when the time dependence is misspecified as we often expect. A consistent variance estimate is presented. We illustrate the use of the GEE approach with longitudinal data from a study of the effect of mothers' stress on children's morbidity.

Child↗

Exploratory data analysis groupware for qualitative and quantitative electrophoretic gel analysis over the Internet-WebGel.

Many scientists use quantitative measurements to compare the presence and amount, of various proteins and nucleotides among series of one- and two-dimensional (1-D and 2-D) electrophoretic gels. These gels are often scanned into digital image files. Gel spots are then quantified using stand-alone analysis software. However, as more research collaborations take place over the Internet, it has become useful to share intermediate quantitative data between researchers. This allows research group members to investigate their data and share their work in progress. We developed a World Wide Web group-accessible software system, WebGel, for interactively exploring qualitative and quantitative differences between electrophoretic gels. Such Internet databases are useful for publishing quantitative data and allow other researchers to explore the data with respect to their own research. Because intermediate results of one user may be shared with their collaborators using WebGel, this form of active data-sharing constitutes a groupware method for enhancing collaborative research. Quantitative and image gel data from a stand-alone gel image processing system are copied to a database accessible on the WebGel Web server. These data are then available for analysis by the WebGel database program residing on that server. Visualization is critical for better understanding of the data. WebGel helps organize labeled gel images into montages of corresponding spots as seen in these different gels. Various views of multiple gel images, including sets of spots, normalization spots, labeled spots, segmented gels, etc. may also be displayed. These displays are active and may be used for performing database operations directly on individual protein spots by simply clicking on them. Corresponding regions between sets of gels may be visually analyzed using Flicker-comparison (Electrophoresis 1997, 18, 122-140) as one of the WebGel methods for qualitative analysis. Quantitative exploratory data analysis can be performed by comparing protein concentration values between corresponding spots for multiple samples run in separate gels. These data are then used to generate reports on statistical differences between sets of gels (e.g., between different disease states such as benign or metastatic cancers, etc.). Using combined visual and quantitative methods, WebGel can help bridge the analysis of dissimilar gels which are difficult to analyze with stand-alone systems and can serve as a collaborative Internet tool in a groupware setting.

Electrophoresis↗