PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Guide for Exploring Pleiotropic Associations in Genome-Wide Association Studies Using Summary Statistics.

Genome-wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross-phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic-based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.

Genome-Wide Association Study

Direct estimation and inference of higher-level correlations from lower-level measurements with applications in gene-pathway and proteomics studies.

This paper tackles the challenge of estimating correlations between higher-level biological variables (e.g. proteins and gene pathways) when only lower-level measurements are directly observed (e.g. peptides and individual genes). Existing methods typically aggregate lower-level data into higher-level variables and then estimate correlations based on the aggregated data. However, different data aggregation methods can yield varying correlation estimates as they target different higher-level quantities. Our solution is a latent factor model that directly estimates these higher-level correlations from lower-level data without the need for data aggregation. We further introduce a shrinkage estimator to ensure the positive definiteness and improve the accuracy of the estimated correlation matrix. Furthermore, we establish the asymptotic normality of our estimator, enabling efficient computation of P-values for the identification of significant correlations. The effectiveness of our approach is demonstrated through comprehensive simulations and the analysis of proteomics and gene expression datasets. We develop the R package highcor for implementing our method.

Proteomics

Optimal Control of Directional False Discovery Rates in Large-Scale Testing.

The high-throughput biomedical technology enables measurement of thousands of gene expression levels contemporaneously. A major task in analyzing these gene expression data is to identify both over-expressed and under-expressed genes. The popular two-group models select the non-null genes without further classifying them as overexpression or underexpression. Consequently, two-group decision rules are unable to constrain the numbers of falsely discovered over-expressed or under-expressed genes respectively. We propose a general three-group model that allows dependence between the test statistics and develop a decision rule that separately controls the two types of false discoveries. We show that the optimal decision rule in our three-group model has a special monotonic structure. By making use of this monotonic structure, we can linearize the two-directional false discovery rate constraints. We prove that our decision rule optimizes the expected number of true discoveries while controlling the proportions of falsely discovered over-expressed and under-expressed genes at desired levels simultaneously. The data-driven versions of the proposed procedures are suggested, and their consistency is established. Comparisons with state-of-the-art approaches and applications to genomic studies show that our procedures work well.

Humans

Analysis of beat-to-beat cardiovascular hemodynamic variables obtained from long-term biotelemetry.

Manual methods of large volume data storage, retrieval, and analysis are difficult, time consuming, and present numerous opportunities for calculation errors. We have designed and implemented a comprehensive computer-based system for performing these functions. Development of this system was necessary since left ventricular (LV) blood pressure and two regional LV wall thickness measurements were obtained during long-term extracorporeal biotelemetry of miniswine for 24-h periods. During a single recording period over 100,000 individual cardiac cycles were recorded on analog tape and later analysed for determination of global myocardial oxygen demand and regional myocardial function. In addition, custom designed software was developed to determine the extent and duration of myocardial dysfunction. Batch file commands enabled the customized software to operate without prompting by the user thus optimizing the time usage of the computer, and the computer based data acquisition and analysis system. Although this system was designed specifically for analysing cardiovascular hemodynamic variables, it is flexible and can be applied to other experimental applications.

Analog-Digital Conversion

Family-Wise Error Rate Control in Clinical Trials With Overlapping Populations.

We consider clinical trials with multiple, overlapping patient populations that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect several populations. For type I error control, often the family-wise error rate (FWER) is controlled, which is the probability to reject at least one true null hypothesis. If the joint distribution of the test statistics is known, the FWER level can be exhausted by determining critical values or adjusted-levels. The adjustment is typically done under the common ANOVA assumptions. However, the performed tests are then only valid under the rather strong assumption of homogeneous null effects, that is, when the null hypothesis applies to all subpopulations and their intersections. We show that under cancelling null effects, when heterogeneous effects cancel out in some or all subpopulations, this procedure does not provide FWER control. We also suggest different alternatives and compare them in terms of FWER control and their power.

Humans

[Additional coding for legally required ICD-9 as the basis for high quality patient record searches].

We are convinced, that the updating of the ID DIACOS diagnosis catalogue fits better to the today's medical language, than the former one. By using a consequent additional code we could eliminate several lacks of the ICD-9. Valid and reproducible diagnosis statistics are now possible. Subjective coding mistakes are largely avoidable. Inquiries of certain diagnosis can be answered by using the one of the two codes, that describes the diagnosis in the best form. Regarding the allotments of the best form. Regarding the allotments of the catalogue of orthopedic and traumatic diagnosis including the possibility of double-coding we got a special better version of ID DIACOS than before.

Data Collection

[Quality of data provided by VESKA medical statistics: the case of the fractured proximal femur].

Within the framework of a retrospective study of the incidence of hip fractures in the canton of Vaud (Switzerland), all cases of hip fracture occurring among the resident population in 1986 and treated in the hospitals of the canton were identified from among five different information sources. Relevant data were then extracted from the medical records. At least two sources of information were used to identify cases in each hospital, among them the statistics of the Swiss Hospital Association (VESKA). These statistics were available for 9 of the 18 hospitals in the canton that participated in the study. The number of cases identified from the VESKA statistics was compared to the total number of cases for each hospital. For the 9 hospitals the number of cases in the VESKA statistics was 407, whereas, after having excluded diagnoses that were actually "status after fracture" and double entries, the total for these hospitals was 392, that is 4% less than the VESKA statistics indicate. It is concluded that the VESKA statistics provide a good approximation of the actual number of cases treated in these hospitals, with a tendency to overestimate this number. In order to use these statistics for calculating incidence figures, however, it is imperative that a greater proportion of all hospitals (50% presently in the canton, 35% nationwide) participate in these statistics.

Data Interpretation, Statistical

Importance of trends in the interpretation of an overall odds ratio in the meta-analysis of clinical trials.

This paper contains a proposition related to the publication of meta-analyses of clinical trials. We consider the situation where the results of a number of trials are summarized by a common or typical odds ratio. We show that stating such an odds ratio as the summary of evidence from a number of trials can be misleading if certain systematic differences between trials exist. In such cases the author should state not just one odds ratio but also its dependence on the relevant characteristics of the trials. In particular, we propose that those reporting a meta-analysis state in advance a (limited) number of variables to be considered for potential interaction with the exposure (risk factor or treatment) of interest. The list might include centre size and the odds in the placebo or control group if such an effect is a priori clinically plausible. The trials should be ordered according to each of these variables and a trend test for the odds ratio should be computed. Apart from a 'genuine' effect, an appreciable interaction could also be indicative of the (multiplicative) odds ratio being an inappropriate measure for the particular meta-analysis. Without any consideration as to the possibility of interaction, the meta-analysis should be considered incomplete. If such an interaction exists, the odds ratio should be stated as a function of the interacting variable, either as a formula or (preferably) in a table stating the odds ratio for a number of different values of the interacting variable, and not as a single summary statistic.

Clinical Trials as Topic