PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “covariance matrix estimation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Average-intensity reconstruction and Wiener reconstruction of bioelectric current distribution based on its estimated covariance matrix.

This paper proposes two methods for reconstructing current distributions from biomagnetic measurements. Both of these methods are based on estimating the source-current covariance matrix from the measured-data covariance matrix. One method is the reconstruction of average current intensity distributions. This method first estimates the source-current covariance matrix and, using its diagonal terms, it reconstructs current intensity distributions averaged over a certain time. Although the method does not reconstruct the orientation of each current element at each time instant, it can retrieve information regarding the current time-averaged intensity at each voxel location using extremely low SNR data. The second method is Wiener reconstruction using the estimated source-current covariance matrix. Unlike the first method, this Wiener reconstruction can provide a current distribution with its orientation at each time instant. Computer simulation shows that the Wiener method is less affected by the choice of the regularization parameter, resulting in a method that is more effective than the conventional minimum-norm method when the SNR of the measurement is low.

Computer Simulation↗

A shrinkage approach to large-scale covariance matrix estimation and implications for functional genomics.

Inferring large-scale covariance matrices from sparse genomic data is an ubiquitous problem in bioinformatics. Clearly, the widely used standard covariance and correlation estimators are ill-suited for this purpose. As statistically efficient and computationally fast alternative we propose a novel shrinkage covariance estimator that exploits the Ledoit-Wolf (2003) lemma for analytic calculation of the optimal shrinkage intensity. Subsequently, we apply this improved covariance estimator (which has guaranteed minimum mean squared error, is well-conditioned, and is always positive definite even for small sample sizes) to the problem of inferring large-scale gene association networks. We show that it performs very favorably compared to competing approaches both in simulations as well as in application to real expression data.

Journal Article↗

On the singularity of the covariance matrix for estimates of multinomial proportions.

It is well known that the covariance matrix for the multinomial distribution is singular and, therefore, does not have a unique inverse. If, however, any row and corresponding column are removed, the reduced matrix is nonsingular and the unique inverse has a closed form. We elucidate some of the properties of the multinomial covariance matrix and its reduced forms. We state and prove a theorem that gives insight into the singularity and its removal. Based on these results, we establish that the covariance matrix for the multinomial distribution is positive semidefinite and that the reduced matrix is positive definite. In addition, we show that the determinant of the reduced matrix is invariant to the particular row and column that are removed. Goodness-of-fit statistics, including Pearson's chi-square, and justification of the degrees of freedom follow from the multivariate central limit theorem once the singularity is removed.

Analysis of Variance↗

Accurate critical constants for the one-sided approximate likelihood ratio test of a normal mean vector when the covariance matrix is estimated.

Tang, Gnecco, and Geller (1989, Biometrika 76, 577-583) proposed an approximate likelihood ratio (ALR) test of the null hypothesis that a normal mean vector equals a null vector against the alternative that all of its components are nonnegative with at least one strictly positive. This test is useful for comparing a treatment group with a control group on multiple endpoints, and the data from the two groups are assumed to follow multivariate normal distributions with different mean vectors and a common covariance matrix (the homoscedastic case). Tang et al. derived the test statistic and its null distribution assuming a known covariance matrix. In practice, when the covariance matrix is estimated, the critical constants tabulated by Tang et al. result in a highly liberal test. To deal with this problem, we derive an accurate small-sample approximation to the null distribution of the ALR test statistic by using the moment matching method. The proposed approximation is then extended to the heteroscedastic case. The accuracy of both the approximations is verified by simulations. A real data example is given to illustrate the use of the approximations.

Biometry↗

On the calculation of a robust S-estimator of a covariance matrix.

An S-estimator of multivariate location and scale minimizes the determinant of the covariance matrix, subject to a constraint on the magnitudes of the corresponding Mahalanobis distances. The relationship between S-estimators and w-estimators of multivariate location and scale can be used to calculate robust estimates of covariance matrices. Elemental subsets of observations are generated to derive initial estimates of means and covariances, and the w-estimator equations are then iterated until convergence to obtain the S-estimates. An example shows that converging to a (local) minimum from the initial estimates from the elemental subsets is an effective way of determining the overall minimum. None of the estimates gained from the elemental samples is close to the final solution.

Biometry↗

Induced smoothing for rank regression with censored survival times.

Adaptions of weighted rank regression to the accelerated failure time model for censored survival data have been successful in yielding asymptotically normal estimates and flexible weighting schemes to increase statistical efficiencies. However, for only one simple weighting scheme, Gehan or Wilcoxon weights, are estimating equations guaranteed to be monotone in parameter components, and even in this case are step functions, requiring the equivalent of linear programming for computation. The lack of smoothness makes standard error or covariance matrix estimation even more difficult. An induced smoothing technique overcame these difficulties in various problems involving monotone but pure jump estimating equations, including conventional rank regression. The present paper applies induced smoothing to the Gehan-Wilcoxon weighted rank regression for the accelerated failure time model, for the more difficult case of survival time data subject to censoring, where the inapplicability of permutation arguments necessitates a new method of estimating null variance of estimating functions. Smooth monotone parameter estimation and rapid, reliable standard error or covariance matrix estimation is obtained.

Data Interpretation, Statistical↗

Optimal temperature input design for estimation of the square root model parameters: parameter accuracy and model validity restrictions.

As part of the model building process, parameter estimation is of great importance in view of accurate prediction making. Confidence limits on the predicted model output are largely determined by the parameter estimation accuracy that is reflected by its parameter estimation covariance matrix. In view of the accurate estimation of the Square Root model parameters, Bernaerts et al. have successfully applied the techniques of optimal experiment design for parameter estimation [Int. J. Food Microbiol. 54 (1-2) (2000) 27]. Simulation-based results have proved that dynamic (i.e., time-varying) temperature conditions characterised by a large abrupt temperature increase yield highly informative cell density data enabling precise estimation of the Square Root model parameters. In this study, it is shown by bioreactor experiments with detailed and precise sampling that extreme temperature shifts disturb the exponential growth of Escherichia coli K12. A too large shift results in an intermediate lag phase. Because common growth models lack the ability to model this intermediate lag phase, temperature conditions should be designed such that exponential growth persist even though the temperature may be changing. The current publication presents (i) the design of an optimal temperature input guaranteeing model validity yet yielding accurate Square Root model parameters, and (ii) the experimental implementation of the optimal input in a computer-controlled bioreactor. Starting values for the experiment design are generated by a traditional two-step procedure based on static experiments. Opposed to the single step temperature profile, the novel temperature input comprises a sequence of smaller temperature increments. The structural development of the temperature input is extensively explained. High quality data of E. coli K12 under optimally varying temperature conditions realised in a computer-controlled bioreactor yield accurate estimates for the Square Root model parameters. The latter is illustrated by means of the individual confidence intervals and the joint confidence region.

Bioreactors↗

Adjusting for time trends when estimating the relationship between dietary intake obtained from a food frequency questionnaire and true average intake.

In measuring food intake, three common methods are used: 24-hour recalls, food frequency questionnaires and food records. Food records or 24-hour recalls are often thought to be the most reliable, but they are difficult and expensive to obtain. The question of interest to us is to use the food records or 24-hour recalls to examine possible systematic biases in questionnaires as a measure of usual food intake. In Freedman, et al. (1991), this problem is addressed through a linear errors in variables analysis. Their model assumes that all measurements on a given individual have the same mean and variance. However, such assumptions may be violated in at least two circumstances, as in for example the Women's Health Trial Vanguard Study and in the Finnish Smokers' Study. First, some studies occur over a period of years, and diets may change over the course of the study. Second, measurements might be taken at different times of the year, and it is known that diets differ on the basis of seasonal factors. In this paper, we will suggest new models incorporating mean and variance offsets, i.e., changes in the population mean and variance for observations taken at different time points. The parameters in the model are estimated by simple methods, and the theory of unbiased estimating equations (M-estimates) is used to derive asymptotic covariance matrix estimates. The methods are illustrated with data from the Women's Health Trial Vanguard Study.

Analysis of Variance↗

Some applications of the analysis of multivariate normal data with missing observations.

Methods are proposed for the analysis of data from a multivariate normal distribution when observations are missing completely at random on some of the variates. Park (1) proved the equivalence of the solutions given by maximum likelihood and generalized estimating equations when data are complete and an unstructured covariance is assumed. He suggested that generalized estimating equations may be used if sample sizes are large relative to the amount of missing data and the estimated covariance matrix is positive definite. We give several examples indicating that the estimating equations give results similar to those of maximum likelihood when smoothing of the covariance matrix to eliminate nonpositive definiteness is not encountered as, for example, under an assumption of exchangeable correlation. Generalized linear models are formulated that are appropriate for a wide class of experimental plans.

Blood Coagulation↗

Ability of geometric morphometric methods to estimate a known covariance matrix.

Landmark-based morphometric methods must estimate the amounts of translation, rotation, and scaling (or, nuisance) parameters to remove nonshape variation from a set of digitized figures. Errors in estimates of these nuisance variables will be reflected in the covariance structure of the coordinates, such as the residuals from a superimposition, or any linear combination of the coordinates, such as the partial warp and standard uniform scores. A simulation experiment was used to compare the ability of the generalized resistant fit (GRF) and a relative warp analysis (RWA) to estimate known covariance matrices with various correlations and variance structures. Random covariance matrices were perturbed so as to vary the magnitude of the average correlation among coordinates, the number of landmarks with excessive variance, and the magnitude of the excessive variance. The covariance structure was applied to random figures with between 6 and 20 landmarks. The results show the expected performance of GRF and RWA across a broad spectrum of conditions. The performance of both GRF and RWA depended most strongly on the number of landmarks. RWA performance decreased slightly when one or a few landmarks had excessive variance. GRF performance peaked when approximately 25% of the landmarks had excessive variance. In general, both RWA and GRF performed better at estimating the direction of the first principal axis of the covariance matrix than the structure of the entire covariance matrix. RWA tended to outperform GRF when > approximately 75% of the coordinates had excessive variance. When < 75% of the coordinates had excessive variance, the relative performance of RWA and GRF depended on the magnitude of the excessive variance; when the landmarks with excessive variance had standard deviations (sigma) > or = 4 sigma minimum, GRF regularly outperformed RWA.

Analysis of Variance↗

A note on variance estimation in random effects meta-regression.

For random effects meta-regression inference, variance estimation for the parameter estimates is discussed. Because estimated weights are used for meta-regression analysis in practice, the assumed or estimated covariance matrix used in meta-regression is not strictly correct, due to possible errors in estimating the weights. Therefore, this note investigates the use of a robust variance estimation approach for obtaining variances of the parameter estimates in random effects meta-regression inference. This method treats the assumed covariance matrix of the effect measure variables as a working covariance matrix. Using an example of meta-analysis data from clinical trials of a vaccine, the robust variance estimation approach is illustrated in comparison with two other methods of variance estimation. A simulation study is presented, comparing the three methods of variance estimation in terms of bias and coverage probability. We find that, despite the seeming suitability of the robust estimator for random effects meta-regression, the improved variance estimator of Knapp and Hartung (2003) yields the best performance among the three estimators, and thus may provide the best protection against errors in the estimated weights.

Analysis of Variance↗

Comparison of preprocessing procedures for oligo-nucleotide micro-arrays by parametric bootstrap simulation of spike-in experiments.

OBJECTIVE: Due to scarcity of calibration data for microarray experiments, simulation methods are employed to assess preprocessing procedures. Here we analyze several procedures' robustness against increasing numbers of differentially expressed genes and varying proportions of up-regulation. METHODS: Raw probe data from oligo-nucleotide microarrays are assumed to be approximately multivariate normally distributed on the log scale. Chips can be simulated from a multivariate normal distribution with mean and variance-covariance matrix estimated from a real raw data set. A chip effect induces strong positive correlations. In reverse, sampling from a normal distribution with strong correlation variance-covariance matrix generates data exhibiting a chip effect. No explicit model of chip-effect is needed. Differences can be artificially spiked-in according to a given distribution of effect sizes. Thirty preprocessing procedures combining background correction, normalization, perfect match correction and summarization methods available from the BioConductor project were compared. RESULTS: In the symmetrical setting "50% differentially expressed genes, 50% of which up-regulated" background correction reduces bias, but inflates low intensity probe variance as well as the mean squared error of the estimates. Any normalization reduces variance and increases sensitivity with no clear winner. Asymmetry between up and down regulation causes bias in the effect-size estimate of non-differentially expressed genes. This markedly inflates the false positive discovery rates. Variance stabilizing normalization (VSN) behaved best. CONCLUSION: A simple parametric bootstrap was used to simulate oligo-nucleotide micro-array raw data. Current normalization methods inflate the false positive rate when many genes show an effect in the same direction.

Computer Simulation↗

MC-Fit: using Monte-Carlo methods to get accurate confidence limits on enzyme parameters.

A program is described for estimating enzymatic parameters from experimental data using Apple Macintosh computers. MC-Fit uses iterative least-square fitting and Monte-Carlo sampling to get accurate estimates of the confidence limits. This approach is more robust than the conventional covariance matrix estimation, especially in cases where experimental data is partially lacking or when the standard error on individual measurements is large. This happens quite often when analysing the properties of variant enzymes obtained by mutagenesis, as these can have severely impaired activities and reduced affinities for their substrates.

Confidence Intervals↗

A maximum likelihood method for region-of-interest evaluation in emission tomography.

A maximum likelihood (ML) estimation method, called the ML-ROI algorithm, is presented for the calculation of region-of-interest (ROI) values from emission tomography scans. The EM algorithm is used to directly estimate ROI values from tomographic projection data given the location, size, and shape of all ROIs. The algorithm requires for the specification of a detailed model of the physical factors contributing to projection measurements including resolution and attenuation. The ML-ROI algorithm also provides an estimate of the variability of the ROI estimator (covariance matrix). The algorithm was tested with simulation and phantom data and compared with ROI estimation strategies using filtered backprojection (FBP) images. The ML-ROI estimates were unbiased, i.e., the partial volume effect was eliminated. Except for regions smaller than the detector resolution, the variability of the ML estimates was comparable to or less than the biased FBP estimators. Computation time for the ML-ROI algorithm was between 5 and 10 s/iteration. An evaluation of the sensitivity of the algorithm to misdefinition of the location and size of the ROIs was also performed.

Humans↗

Ultrawideband microwave breast cancer detection: a detection-theoretic approach using the generalized likelihood ratio test.

Microwave imaging has been suggested as a promising modality for early-stage breast cancer detection. In this paper, we propose a statistical microwave imaging technique wherein a set of generalized likelihood ratio tests (GLRT) is applied to microwave backscatter data to determine the presence and location of strong scatterers such as malignant tumors in the breast. The GLRT is formulated assuming that the backscatter data is Gaussian distributed with known covariance matrix. We describe the method for estimating this covariance matrix offline and formulating a GLRT for several heterogeneous two-dimensional (2-D) numerical breast phantoms, several three-dimensional (3-D) experimental breast phantoms, and a 3-D numerical breast phantom with a realistic half-ellipsoid shape. Using the GLRT with the estimated covariance matrix and a threshold chosen to constrain the false discovery rate (FDR) of the image, we show the capability to detect and localize small (<0.6 cm) tumors in our numerical and experimental breast phantoms even when the dielectric contrast of the malignant-to-normal tissue is below 2:1.

Algorithms↗

Localization of brain electrical activity via linearly constrained minimum variance spatial filtering.

A spatial filtering method for localizing sources of brain electrical activity from surface recordings is described and analyzed. The spatial filters are implemented as a weighted sum of the data recorded at different sites. The weights are chosen to minimize the filter output power subject to a linear constraint. The linear constraint forces the filter to pass brain electrical activity from a specified location, while the power minimization attenuates activity originating at other locations. The estimated output power as a function of location is normalized by the estimated noise power as a function of location to obtain a neural activity index map. Locations of source activity correspond to maxima in the neural activity index map. The method does not require any prior assumptions about the number of active sources of their geometry because it exploits the spatial covariance of the source electrical activity. This paper presents a development and analysis of the method and explores its sensitivity to deviations between actual and assumed data models. The effect on the algorithm of covariance matrix estimation, correlation between sources, and choice of reference is discussed. Simulated and measured data is used to illustrate the efficacy of the approach.

Algorithms↗

Variation among early North American crania.

The limited morphometric work on early American crania to date has treated them as a single, temporally defined group. This paper addresses the question of whether there is significant variability among ancient American crania. A sample of 11 crania (Spirit Cave, Wizards Beach, Browns Valley, Pelican Rapids, Prospect, Wet Gravel male, Wet Gravel female, Medicine Crow, Turin, Lime Creek, and Swanson Lake) dating from the early to mid Holocene was available. Some have recent accelerator mass spectrometry (AMS) dates, while others are dated geologically or archaeologically. All are in excess of 4500 BP, and most are 7000 BP or older. Measurements follow the definitions of Howells [(1973) Cranial variation in man, Cambridge: Harvard University). Some crania are incomplete, but 22 measurements were common to all fossils. Cranial variation was examined by calculating the Mahalanobis distance between each pair of fossils, using a pooled within sample covariance matrix estimated from the data of Howells. The distance relationships among crania suggest the presence of at least three distinct groups: 1) a middle Archaic Plains group (Turin and Medicine Crow), 2) a Paleo/Early Archaic Great Lakes/Plains group (Browns Valley, Pelican Rapids, Lime Creek), and 3) a spatially and temporally heterogeneous group that includes the Great Basin/Pacific Coast (Spirit Cave, Wizards Beach, Prospect) and Nebraska (Wet Gravel specimens and Swanson Lake). These crania were also compared to Howells' worldwide recent sample, which was expanded by including six additional American Indian samples. None of the fossils, except for the Wet Gravel male, shows any particular affinity to recent Native Americans; their greatest similarities are with Europe, Polynesia, or East Asia. Several crania would be atypical in any recent population for which we have data. Browns Valley, Pelican Rapids, and Lime Creek are the most distinctive. They provide evidence for the presence of an early population that bears no similarity to the morphometric pattern of recent American Indians or even to crania of comparable date in other regions of the continent. The heterogeneity among early American crania makes it inadvisable to pool them for purposes of morphometric analysis. Whether this heterogeneity results from different early migrations or one highly differentiated population cannot be established from our data. Our results are inconsistent with hypotheses of an ancestor-descendent relationship between early and late Holocene American populations. They suggest that the pattern of cranial variation is of recent origin, at least in the Plains region.

Adult↗