PubMed Health⌕ Search

PubMed · 16401270

Generalization of the Mantel-Haenszel estimating function for sparse clustered binary data.

Abstract

We extend the Mantel-Haenszel estimating function to estimate both the intra-cluster pairwise correlation and the main effects for sparse clustered binary data. We propose both a composite likelihood approach and an estimating function approach for the analysis of such data. The proposed estimators are consistent and asymptotically normally distributed. Simulation results demonstrate that the two approaches are comparable in terms of bias and efficiency; however, the estimating equation approach is computationally simpler. Analysis of the Georgia High Blood Pressure survey is used for illustration.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Molin Wang, John M Williamson. 2005. Generalization of the Mantel-Haenszel estimating function for sparse clustered binary data.. https://doi.org/10.1111/j.1541-0420.2005.00362.x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Analysis of cluster randomized cross-over trial data: a comparison of methods.

In a cluster randomized cross-over trial, all participating clusters receive both intervention and control treatments consecutively, in separate time periods. Patients recruited by each cluster within the same time period receive the same intervention, and randomization determines order of treatment within a cluster. Such a design has been used on a number of occasions. For analysis of the trial data, the approach of analysing cluster-level summary measures is appealing on the grounds of simplicity, while hierarchical modelling allows for the correlation of patients within periods within clusters and offers flexibility in the model assumptions. We consider several cluster-level approaches and hierarchical models and make comparison in terms of empirical precision, coverage, and practical considerations. The motivation for a cluster randomized trial to employ cross-over of trial arms is particularly strong when the number of clusters available is small, so we examine performance of the methods under small, medium and large (6, 18, 30) numbers of clusters. One hierarchical model and two cluster-level methods were found to perform consistently well across the designs considered. These three methods are efficient, provide appropriate standard errors and coverage, and continue to perform well when incorporating adjustment for an individual-level covariate. We conclude that choice between hierarchical models and cluster-level methods should be influenced by the extent of complexity in the planned analysis.

Cluster Analysis↗

Improved hypothesis testing for coefficients in generalized estimating equations with small samples of clusters.

The sandwich standard error estimator is commonly used for making inferences about parameter estimates found as solutions to generalized estimating equations (GEE) for clustered data. The sandwich tends to underestimate the variability in the parameter estimates when the number of clusters is small, and reference distributions commonly used for hypothesis testing poorly approximate the distribution of Wald test statistics. Consequently, tests have greater than nominal type I error rates. We propose tests that use bias-reduced linearization, BRL, to adjust the sandwich estimator and Satterthwaite or saddlepoint approximations for the reference distribution of resulting Wald t-tests. We conducted a large simulation study of tests using a variety of estimators (traditional sandwich, BRL, Mancl and DeRouen's BC estimator, and a modification of an estimator proposed by Kott) and approximations to reference distributions under diverse settings that varied the distribution of the explanatory variables, the values of coefficients, and the degree of intra-cluster correlation (ICC). Our new method generally worked well, providing accurate estimates of the variability of fitted coefficients and tests with near-nominal type I error rates when the ICC is small. Our method works less well when the ICC is large, but it continues to out-perform the traditional sandwich and other alternatives.

Cluster Analysis↗