PubMed Health⌕ Search

Biomedical subjects

Glen A Satten

Publications and source records attributed to Glen A Satten.

At least 19 recordsLinked to original sources

Improved association analyses of disease subtypes in case-parent triads.

The sampling of case-parent triads is an appealing strategy for conducting association analyses of complex diseases. In certain situations, one may have interest in using the triads to identify genetic variants that are associated with a specific subtype of disease, perhaps related to a characteristic cluster of symptoms. A straightforward strategy for conducting such a subtype analysis would be to analyze only those triads with the subtype of interest. While such a strategy is valid, we show that triads without the subtype of interest can provide additional genetic information that increases power to detect association with the subtype of interest. We incorporate this additional information using a likelihood-based framework that permits flexible modeling and estimation of allelic effects on disease subtypes and also allows for missing parental data. Using simulated data under a variety of genetic models, we show that our proposed association test consistently outperforms association tests that only analyze triads with the subtype of interest. We also apply our method to a triad study of attention-deficit hyperactivity disorder and identify a genetic variant in the dopamine transporter gene that is associated with a subtype characterized by extreme levels of both inattentive and hyperactive-impulsive symptoms.

Attention Deficit and Disruptive Behavior Disorder↗

Robust testing of haplotype/disease association.

Haplotypes, the combination of closely linked alleles that fall on the same chromosome, show great promise for studying the genetic components of complex diseases. However, when only multilocus genotype data are available, statistical approaches need to be employed to resolve haplotype phase ambiguity. Recently, we have proposed an approach to testing and estimating haplotype/disease association that is invariant to any existing genetic structure in the population. Here we evaluate this approach by applying it to the Genetic Analysis Workshop 14 simulated data.

Bias↗

Genetic association analysis using data from triads and unrelated subjects.

The selection of an appropriate control sample for use in association mapping requires serious deliberation. Unrelated controls are generally easy to collect, but the resulting analyses are susceptible to spurious association arising from population stratification. Parental controls are popular, since triads comprising a case and two parents can be used in analyses that are robust to this stratification. However, parental controls are often expensive and difficult to collect. In some situations, studies may have both parental and unrelated controls available for analysis. For example, a candidate-gene study may analyze triads but may have an additional sample of unrelated controls for examination of background linkage disequilibrium in genomic regions. Also, studies may collect a sample of triads to confirm results initially found using a traditional case-control study. Initial association studies also may collect each type of control, to provide insurance against the weaknesses of the other type. In these situations, resulting samples will consist of some triads, some unrelated controls, and, possibly, some unrelated cases. Rather than analyze the triads and unrelated subjects separately, we present a likelihood-based approach for combining their information in a single combined association analysis. Our approach allows for joint analysis of data from both triad and case-control study designs. Simulations indicate that our proposed approach is more powerful than association tests that are based on each separate sample. Our approach also allows for flexible modeling and estimation of allele effects, as well as for missing parental data. We illustrate the usefulness of our approach using SNP data from a candidate-gene study of psoriasis.

Algorithms↗

Standardization and denoising algorithms for mass spectra to classify whole-organism bacterial specimens.

MOTIVATION: Application of mass spectrometry in proteomics is a breakthrough in high-throughput analyses. Early applications have focused on protein expression profiles to differentiate among various types of tissue samples (e.g. normal versus tumor). Here our goal is to use mass spectra to differentiate bacterial species using whole-organism samples. The raw spectra are similar to spectra of tissue samples, raising some of the same statistical issues (e.g. non-uniform baselines and higher noise associated with higher baseline), but are substantially noisier. As a result, new preprocessing procedures are required before these spectra can be used for statistical classification. RESULTS: In this study, we introduce novel preprocessing steps that can be used with any mass spectra. These comprise a standardization step and a denoising step. The noise level for each spectrum is determined using only data from that spectrum. Only spectral features that exceed a threshold defined by the noise level are subsequently used for classification. Using this approach, we trained the Random Forest program to classify 240 mass spectra into four bacterial types. The method resulted in zero prediction errors in the training samples and in two test datasets having 240 and 300 spectra, respectively.

Algorithms↗

An empirical bayes adjustment to increase the sensitivity of detecting differentially expressed genes in microarray experiments.

MOTIVATION: Detection of differentially expressed genes is one of the major goals of microarray experiments. Pairwise comparison for each gene is not appropriate without controlling the overall (experimentwise) type 1 error rate. Dudoit et al. have advocated use of permutation-based step-down P-value adjustments to correct the observed significance levels for the individual (i.e. for each gene) two sample t-tests. RESULTS: In this paper, we consider an ANOVA formulation of the gene expression levels corresponding to multiple tissue types. We provide resampling-based step-down adjustments to correct the observed significance levels for the individual ANOVA t-tests for each gene and for each pair of tissue type comparisons. More importantly, we introduce a novel empirical Bayes adjustment to the t-test statistics that can be incorporated into the step-down procedure. Using simulated data, we show that the empirical Bayes adjustment improved the sensitivity of detecting differentially expressed genes up to 16%, while maintaining a high level of specificity. This adjustment also reduces the false non-discovery rate to some degree at the cost of a modest increase in the false discovery rate. We illustrate our approach using a human colon cancer dataset consisting of oligonucleotide arrays of normal, adenoma and carcinoma cells. The number of genes with differential expression level declared statistically significant was about 50 when comparing normal to adenoma cells and about five when comparing adenoma to carcinoma cells. This list includes genes previously known to be associated with colon cancer as well as some novel genes. AVAILABILITY: R code for the empirical Bayes adjustment and step-down P-value calculation via resampling are available from the supplementary web-site. SUPPLEMENTARY INFORMATION: http://www.mathstat.gsu.edu/~matsnd/EB/supp.htm

Algorithms↗

Comparison of prospective and retrospective methods for haplotype inference in case-control studies.

We compare bias and power of three methods for haplotype inference on disease risk using unphased genotype data from a case-control study. We examine the prospective score test of Schaid et al., a novel modification of the prospective estimating equations of Zhao et al. and the retrospective likelihood of Epstein and Satten. We find that all three approaches are roughly comparable when the haplotype effect on disease odds follows a multiplicative model. However, for dominant and recessive models of haplotype effect, the retrospective-likelihood method has increased efficiency with respect to the prospective methods. As all three methods assume haplotype frequencies are in Hardy-Weinberg Equilibrium (HWE), we compare the robustness of each procedure to departures from HWE. We find the prospective methods are robust to departure from HWE, while the retrospective-likelihood method is biased for dominant and recessive models of haplotype effect. To remedy this limitation of the retrospective-likelihood method, we propose a modification that allows for a non-negative fixation index (common to all haplotype pairs) and show it dramatically reduces the bias of the retrospective likelihood when HWE is violated.

Algorithms↗

How special is a 'special' interval: modeling departure from length-biased sampling in renewal processes.

Length-biased sampling occurs in renewal processes when the probability that an interval is selected is proportional to the length of the interval. This can occur when intervals are selected because they contain an event that is independent of the renewal process and occurs with constant hazard. For example, if the times between donations for repeat blood donors are independent and identically distributed, and if the donor seroconverts to HIV (develops antibodies that indicate infection with human immunodeficiency virus), then the interval between the last HIV seronegative and first HIV seropositive test is expected to be longer than that donor's previous time intervals between donations. We develop hypothesis tests to determine if the relationship between the typical and length-biased intervals is as expected, or if there is departure from length-biased sampling. We further develop a regression method to determine if there are covariates that explain the departure from length-biased sampling. Our approach is motivated by the question of whether there is evidence that repeat blood donors who develop antibodies to HIV or other viral infections change their donation pattern in some way because of seroconversion.

Blood Donors↗

Bootstrap calibration of TRANSMIT for informative missingness of parental genotype data.

Informative missingness of parental genotype data occurs when the genotype of a parent influences the probability of the parent's genotype data being observed. Informative missingness can occur in a number of plausible ways and can affect both the validity and power of procedures that assume the data are missing at random (MAR). We propose a bootstrap calibration of MAR procedures to account for informative missingness and apply our methodology to refine the approach implemented in the TRANSMIT program. We illustrate this approach by applying it to data on hypertensive probands and their parents who participated in the Framingham Heart Study.

Adult Children↗

Inference on haplotype effects in case-control studies using unphased genotype data.

A variety of statistical methods exist for detecting haplotype-disease association through use of genetic data from a case-control study. Since such data often consist of unphased genotypes (resulting in haplotype ambiguity), such statistical methods typically apply the expectation-maximization (EM) algorithm for inference. However, the majority of these methods fail to perform inference on the effect of particular haplotypes or haplotype features on disease risk. Since such inference is valuable, we develop a retrospective likelihood for estimating and testing the effects of specific features of single-nucleotide polymorphism (SNP)-based haplotypes on disease risk using unphased genotype data from a case-control study. Our proposed method has a flexible structure that allows, among other choices, modeling of multiplicative, dominant, and recessive effects of specific haplotype features on disease risk. In addition, our method relaxes the requirement of Hardy-Weinberg equilibrium of haplotype frequencies in case subjects, which is typically required of EM-based haplotype methods. Also, our method easily accommodates missing SNP information. Finally, our method allows for asymptotic, permutation-based, or bootstrap inference. We apply our method to case-control SNP genotype data from the Finland-United States Investigation of Non-Insulin-Dependent Diabetes Mellitus (FUSION) Genetics study and identify two haplotypes that appear to be significantly associated with type 2 diabetes. Using the FUSION data, we assess the accuracy of asymptotic P values by comparing them with P values obtained from a permutation procedure. We also assess the accuracy of asymptotic confidence intervals for relative-risk parameters for haplotype effects, by a simulation study based on the FUSION data.

Algorithms↗

Performance characteristics of a new less sensitive HIV-1 enzyme immunoassay for use in estimating HIV seroincidence.

Less sensitive (LS) HIV-1 enzyme immunoassays (EIAs) have significantly improved the quantity and quality of HIV surveillance data. The first LS-HIV-1 EIA, the Abbott 3A11-LS, provided reliable incidence data, but the assay required specialized equipment, and the lack of available reagents made testing difficult. This study evaluated the use of an alternate assay, a modified version of the Vironostika HIV-1 EIA (Vironostika-LS), to be used for LS testing. The Vironostika-LS has similar performance characteristics to the Abbott 3A11-LS with additional advantages. This 96-well formatted assay is commonly found in public health laboratories for routine HIV-1 testing and can be used with both serum and dried blood spot specimens. The estimated mean time from seroconversion (defined using a standardized optical density cutoff of 1.0) with the Vironostika-LS was 170 days (95% CI, 145-200 days). When the Vironostika-LS was applied to a matched serum set previously tested with the Abbott 3A11-LS, the Vironostika-LS accurately identified 97% of specimens with recent or long-standing HIV infection. The paper also reports Vironostika-LS quality control guidelines and the results from 3 rounds of proficiency testing.

AIDS Serodiagnosis↗

Informative missingness in genetic association studies: case-parent designs.

We consider the effect of informative missingness on association tests that use parental genotypes as controls and that allow for missing parental data. Parental data can be informatively missing when the probability of a parent being available for study is related to that parent's genotype; when this occurs, the distribution of genotypes among observed parents is not representative of the distribution of genotypes among the missing parents. Many previously proposed procedures that allow for missing parental data assume that these distributions are the same. We propose association tests that behave well when parental data are informatively missing, under the assumption that, for a given trio of paternal, maternal, and affected offspring genotypes, the genotypes of the parents and the sex of the missing parents, but not the genotype of the affected offspring, can affect parental missingness. (This same assumption is required for validity of an analysis that ignores incomplete parent-offspring trios.) We use simulations to compare our approach with previously proposed procedures, and we show that if even small amounts of informative missingness are not taken into account, they can have large, deleterious effects on the performance of tests.

Computer Simulation↗

Random error and undercounting in birth defects surveillance data: implications for inference.

BACKGROUND: There has been an ongoing debate among birth defects investigators about whether or not to publish estimates of rates of birth defects with confidence intervals to allow for comparisons of rates across regions and time. A major impediment in resolving this debate has been the lack of a framework for quantifying uncertainties in the data that can be applied uniformly to birth defects surveillance programs. This report presents an overview of random error and ascertainment bias in birth defects surveillance data, and of the implications of these errors for estimation and comparisons of birth defects rates. METHODS: We consider when confidence intervals can be used as part of a strategy to make inference on rates, as well as ratios of or differences between two rates. Worth noting is that confidence intervals only address random error in the data. In the presence of undercounting of cases, estimation of rates and confidence intervals requires knowledge or an estimate of the extent of underascertainment. Rate estimates and confidence intervals that ignore such bias can be misleading. However, if it is reasonable to assume that the ascertainment bias is constant over time (or across regions), then it is possible to make valid comparisons of rates over time (or across regions) using ratio or difference estimators, even when lack of knowledge of the extent of undercounting makes estimating the absolute rate and its confidence interval problematic. Finally, sensitivity analyses can use confidence limits to determine the difference in ascertainment bias necessary to explain an apparent difference in rates. CONCLUSION: Because birth defects surveillance systems have evolved in the absence of agreed upon standards to guide the process, it is difficult to determine the extent to which the variability in rates of birth defects across programs or over time is real or due to differences in surveillance methods. Efforts to develop standards for birth defects surveillance may help to minimize the variability in prevalence of birth defects due to differences in case ascertainment methods and allow for evaluations of real temporal and spatial variations in environmental effects. In the meantime, if comparisons of rates need to be made to address public health concerns, it would be prudent to conduct only such comparisons between regions or across time when the degree of case ascertainment can be assumed to be relatively constant across regions and time.

Birth Certificates↗

Marginal analyses of clustered data when cluster size is informative.

We propose a new approach to fitting marginal models to clustered data when cluster size is informative. This approach uses a generalized estimating equation (GEE) that is weighted inversely with the cluster size. We show that our approach is asymptotically equivalent to within-cluster resampling (Hoffman, Sen, and Weinberg, 2001, Biometrika 73, 13-22), a computationally intensive approach in which replicate data sets containing a randomly selected observation from each cluster are analyzed, and the resulting estimates averaged. Using simulated data and an example involving dental health, we show the superior performance of our approach compared to unweighted GEE, the equivalence of our approach with WCR for large sample sizes, and the superior performance of our approach compared with WCR when sample sizes are small.

Cluster Analysis↗

HIV seroincidence among patients at clinics for sexually transmitted diseases in nine cities in the United States.

Although the numbers of newly reported diagnoses of AIDS decreased in the 1990s, it is not clear whether they reflect a decreasing number of new HIV infections. Direct measurement of HIV incidence through follow-up cohort studies is difficult and costly. We estimated HIV incidence and trends in incidence among men who have sex with men (MSM) and heterosexual men and women at clinics for sexually transmitted diseases (STDs) by using a recently developed serologic testing algorithm that requires only a single blood specimen. Cross-sectional anonymous serosurveys were conducted at 13 STD clinics in nine cities in the United States from 1991 through 1997. Before anonymous HIV testing, demographic and clinical information was abstracted. Of 129,774 specimens tested, 362 (0.28%) were from persons estimated to be recently infected. Incidence among MSM was 7.1% (95% confidence interval (CI): 4.8-10.3), 14 times higher than that among heterosexuals, which was 0.5% (CI: 0.4- 0.7). Incidence among MSM and heterosexuals remained unchanged during the time studied. Decreasing rates of new AIDS diagnoses in the 1990s do not reflect stable rates of new HIV infections among MSM and heterosexual patients attending these clinics.

Adult↗

Subtype-specific transmission probabilities for human immunodeficiency virus type 1 among injecting drug users in Bangkok, Thailand.

The Bangkok (Thailand) Metropolitan Administration cohort of injecting drug users (IDUs) consisted of 1,209 IDUs initially seronegative for human immunodeficiency virus (HIV) who were followed from 1995 to 1998 at 15 Administration drug treatment clinics. At enrollment and approximately every 4 months thereafter, participants were assessed for HIV seropositivity. As of December 1998, there were 133 HIV type 1 seroconversions and approximately 2,300 person-years of follow-up. Of the 133 observed seroconversions, specimens from 126 persons were available for subtyping (27 subtype B, 99 subtype E). In this analysis, the authors assessed differences in subtype-specific transmission while controlling for important risk factors. The methodology used accounts for left truncation, interval censoring, and competing risks as well as for time-varying covariates such as each IDU's history of reported frequency of injection and of incarceration. Using plausible epidemiologic assumptions and controlling for behavioral risks, the authors found that a significantly higher transmission probability was associated with subtype E compared with subtype B in this population. Since many epidemiologic, virologic, and host factors can influence HIV transmission, it was difficult to conclude whether these differences in transmission probabilities were due to biologic properties associated with subtype.

Adolescent↗

Marginal estimation for multi-stage models: waiting time distributions and competing risks analyses.

We provide non-parametric estimates of the marginal cumulative distribution of stage occupation times (waiting times) and non-parametric estimates of marginal cumulative incidence function (proportion of persons who leave stage j for stage j' within time t of entering stage j) using right-censored data from a multi-stage model. We allow for stage and path dependent censoring where the censoring hazard for an individual may depend on his or her natural covariate history such as the collection of stages visited before the current stage and their occupation times. Additional external time dependent covariates that may induce dependent censoring can also be incorporated into our estimates, if available. Our approach requires modelling the censoring hazard so that an estimate of the integrated censoring hazard can be used in constructing the estimates of the waiting times distributions. For this purpose, we propose the use of an additive hazard model which results in very flexible (robust) estimates. Examples based on data from burn patients and simulated data with tracking are also provided to demonstrate the performance of our estimators.

Burns↗

HIV seroconverting donors delay their return: screening test implications.

BACKGROUND: The yield of HIV p24 antigen testing implemented in March 1996 has been lower than projected. One possible explanation is that HIV seroconverting donors delay their return because of the recent practice of risk behaviors and/or signs and symptoms associated with primary infection. STUDY DESIGN AND METHODS: From a database of 6.8-million allogeneic donations collected at five U.S. blood centers from 1991 to 1997, 49 HIV, 21 HCV, 32 HTLV, and 44 HBsAg seroconverters with at least three donations were identified. A statistical method was developed to investigate whether the time between a donor's last negative donation and their positive donation was significantly longer than expected based on their previous return history. RESULTS: HIV seroconverters returned on average 42 percent later than expected (p < 0.01). Although not significant, HCV seroconverters donated on average 43 percent earlier than expected. HTLV and HBsAg seroconverters did not appear to change their donation pattern around the time of seroconversion. Sixty-three percent of the HIV seroconverters later acknowledged practicing a high-risk behavior. CONCLUSIONS: HIV seroconverters delay their return around the time of seroconversion and are thus less likely to be recently infected. Unique among HIV seroconverters, this observation provides a possible explanation for the lower than expected yield of HIV p24 antigen testing and suggests that NAT may have a similar low yield.

Blood Donors↗

Estimation of integrated transition hazards and stage occupation probabilities for non-Markov systems under dependent censoring.

We propose nonparametric estimators of the stage occupation probabilities and transition hazards for a multistage system that is not necessarily Markovian, using data that are subject to dependent right censoring. We assume that the hazard of being censored at a given instant depends on a possibly time-dependent covariate process as opposed to assuming a fixed censoring hazard (independent censoring). The estimator of the integrated transition hazard matrix has a Nelson-Aalen form where each of the counting processes counting the number of transitions between states and the risk sets for leaving each stage have an IPCW (inverse probability of censoring weighted) form. We estimate these weights using Aalen's linear hazard model. Finally, the stage occupation probabilities are obtained from the estimated integrated transition hazard matrix via product integration. Consistency of these estimators under the general paradigm of non-Markov models is established and asymptotic variance formulas are provided. Simulation results show satisfactory performance of these estimators. An analysis of data on graft-versus-host disease for bone marrow transplant patients is used as an illustration.

Biometry↗