PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Upper bounds on maximum likelihood for phylogenetic trees.

We introduce a mechanism for analytically deriving upper bounds on the maximum likelihood for genetic sequence data on sets of phylogenies. A simple 'partition' bound is introduced for general models. Tighter bounds are developed for the simplest model of evolution, the two state symmetric model of nucleotide substitution under the molecular clock. This follows earlier theoretical work which has been restricted to this model by analytic complexity. A weakness of current numerical computation is that reported 'maximum likelihood' results cannot be guaranteed, both for a specified tree (because of the possibility of multiple maxima) or over the full tree space (as the computation is intractable for large sets of trees). The bounds we develop here can be used to conclusively eliminate large proportions of tree space in the search for the maximum likelihood tree. This is vital in the development of a branch and bound search strategy for identifying the maximum likelihood tree. We report the results from a simulation study of approximately 10(6) data sets generated on clock-like trees of five leaves. In each trial a likelihood value of one specific instance of a parameterised tree is compared to the bound determined for each of the 105 possible rooted binary trees. The proportion of trees that are eliminated from the search for the maximum likelihood tree ranged from 92% to almost 98%, indicating a computational speed-up factor of between 12 and 44.

Algorithms↗

An efficient Monte Carlo approach to assessing statistical significance in genomic studies.

MOTIVATION: Multiple hypothesis testing is a common problem in genome research, particularly in microarray experiments and genomewide association studies. Failure to account for the effects of multiple comparisons would result in an abundance of false positive results. The Bonferroni correction and Holm's step-down procedure are overly conservative, whereas the permutation test is time-consuming and is restricted to simple problems. RESULTS: We developed an efficient Monte Carlo approach to approximating the joint distribution of the test statistics along the genome. We then used the Monte Carlo distribution to evaluate the commonly used criteria for error control, such as familywise error rates and positive false discovery rates. This approach is applicable to any data structures and test statistics. Applications to simulated and real data demonstrate that the proposed approach provides accurate error control, and can be substantially more powerful than the Bonferroni and Holm methods, especially when the test statistics are highly correlated.

Algorithms↗

GALGO: an R package for multivariate variable selection using genetic algorithms.

SUMMARY: The development of statistical models linking the molecular state of a cell to its physiology is one of the most important tasks in the analysis of Functional Genomics data. Because of the large number of variables measured a comprehensive evaluation of variable subsets cannot be performed with available computational resources. It follows that an efficient variable selection strategy is required. However, although software packages for performing univariate variable selection are available, a comprehensive software environment to develop and evaluate multivariate statistical models using a multivariate variable selection strategy is still needed. In order to address this issue, we developed GALGO, an R package based on a genetic algorithm variable selection strategy, primarily designed to develop statistical models from large-scale datasets.

Algorithms↗

Discontinuity indices: a tool for epidemiological studies on breastfeeding.

BACKGROUND: The Discontinuity Index (DI), which measures the percentage of infants who were exclusively breastfed (EBF) at the beginning of a given age interval and had abandoned this mode of feeding at its end, and the relative weight of this discontinuation, was introduced and employed in the National Survey on Breast Feeding and Infant Feeding Practices carried out in Cuba in 1990. The aim of this article is to illustrate, through a specific example, the quality of DI as a simple procedure for assessing breastfeeding trends. METHODS: The prevalence of EBF in the 14 provinces of Cuba at discharge from the maternity services and at 30, 60, 120, and 180 days of age, was obtained using data from a national sample of 6661 infants (4820 urban and 1791 rural) which were processed by means of a logistic regression model. Cumulative DI were calculated for the intervals 0-30, 0-60, 0-120 and 0-180 days, and partial DI for the terms 30-60, 60-120 and 120-180 days, for each province and for the whole country. RESULTS: Cumulative DI show the progress of cessation of breastfeeding and are strongly influenced by previous intervals. The Eastern provinces showed the lowest figures at most of the terms. Discontinuation during the first month of life was particularly high in two Western provinces. Partial DI are more specific and allow discrimination of the intervals at which EBF discontinuation is more frequent. The highest values were observed between 4 and 6 months. CONCLUSIONS: Discontinuity indices are useful complements to prevalence rates in epidemiological studies of breastfeeding. The separate analysis of discontinuation in different periods can be highly useful when comparing trends and in the study of the impact of breastfeeding promotion programmes focused on different age intervals.

Age Factors↗

Contrasts and correlations in theory assessment.

OBJECTIVE: To describe a systematic quantitative approach to assessing the predictions made by competing theories using contrasts and correlational indices of effect sizes. METHODS: We illustrate the use of the contrast F and t to compare and combine predictions when the raw data are continuous scores, and z contrasts when working with frequencies in 2 x k tables of counts. RESULTS: The traditional effect size correlation indicates the magnitude of the effect on individual scores of participants' assignment to particular conditions. The contrast correlation obtained from the contrast F or t is, in some cases, the easiest way of estimating the effect size correlation in designs using more than two groups. The alerting correlation is another way of appraising the predictive power of a contrast and can be used to compute the contrast F from published results when all we have are condition means and the omnibus F from an overall analysis of variance. Omnibus Fs, those with more than 1 df in the numerator, are rarely useful in data analytic work since they address unfocused questions, yielding only vague answers. CONCLUSIONS: Asking focused questions using contrasts increases the clarity of our questions and the clarity and statistical power of our answers.

Adolescent↗

The effect of different needle recording electrodes on somatosensory-evoked potentials and intertrial waveform variation.

This investigation examined the cortical somatosensory-evoked potentials (SEP) waveforms obtained from four sets of commercially available subdermal needle electrodes in 19 normal subjects. The composite materials of the four electrodes were stainless steel and a platinum/iridium alloy. Tibial nerve SEP peak latencies for P37 and N45 as well as P37/N45 amplitudes were recorded from each electrode pair in a random fashion. Using nonparametric analysis, no significant differences of waveform parameters were found between electrode pairs (P greater than 0.01). Correlation evaluation demonstrated values in excess of 0.92. Additionally, intertrial waveform analysis for each of the electrode pairs was performed. Again, nonparametric evaluation demonstrated no statistically significant waveform differences. Correlation coefficients were also highly correlative. Variable temperature response to prolonged tibial nerve stimulation was recorded that did not significantly effect the latencies or amplitudes of the cortical SEP responses. We conclude that within temperature ranges typically encountered in clinical practice, there is no statistically significant waveform differences recorded with commonly available subdermal needle electrodes. Additionally, although intertrial waveform variation may exist during SEP recordings, these differences do not reach statistically significant levels.

Adult↗

The application of multivariate statistical techniques improves single-wavelength anomalous diffraction phasing.

Recently, there has been a resurgence in phasing using the single-wavelength anomalous diffraction (SAD) experiment; data from a single wavelength in combination with techniques such as density modification have been used to solve macromolecular structures, even with a very small anomalous signal. Here, a formulation for SAD phasing and refinement employing multivariate statistical techniques is presented. The equation developed accounts explicitly for the correlations among the observed and calculated Friedel mates in a SAD experiment. The correlated SAD equation has been implemented and test cases performed on real diffraction data have revealed better results compared with currently used programs in terms of correlation with the final map and obtaining more reliable phase probability statistics.

Crystallography, X-Ray↗

Solution structure of a zinc substituted eukaryotic rubredoxin from the cryptomonad alga Guillardia theta.

The rubredoxin from the cryptomonad Guillardia theta is one of the first examples of a rubredoxin encoded in a eukaryotic organism. The structure of a soluble zinc-substituted 70-residue G. theta rubredoxin lacking the membrane anchor and the thylakoid targeting sequence was determined by multidimensional heteronuclear NMR, representing the first three-dimensional (3D) structure of a eukaryotic rubredoxin. For the structure calculation a strategy was applied in which information about hydrogen bonds was directly inferred from a long-range HNCO experiment, and the dynamics of the protein was deduced from heteronuclear nuclear Overhauser effect data and exchange rates of the amide protons. The structure is well defined, exhibiting average root-mean-square deviations of 0.21 A for the backbone heavy atoms and 0.67 A for all heavy atoms of residues 7-56, and an increased flexibility toward the termini. The structure of this core fold is almost identical to that of prokaryotic rubredoxins. There are, however, significant differences with respect to the charge distribution at the protein surface, suggesting that G. theta rubredoxin exerts a different physiological function compared to the structurally characterized prokaryotic rubredoxins. The amino-terminal residues containing the putative signal peptidase recognition/cleavage site show an increased flexibility compared to the core fold, but still adopt a defined 3D orientation, which is mainly stabilized by nonlocal interactions to residues of the carboxy-terminal region. This orientation might reflect the structural elements and charge pattern necessary for correct signal peptidase recognition of the G. theta rubredoxin precursor.

Amino Acid Sequence↗

Molecular characterization and sequence of phosphatidylinositol-specific phospholipase C of Bacillus thuringiensis.

The gene encoding monophosphatidylinositol inositol phosphohydrolase (PI-specific phospholipase C, PI-PLC) of Bacillus thuringiensis was cloned in Staphylococcus carnosus TM300. The complete coding region comprises 987 base pairs corresponding to a precursor protein of 329 amino acids (molecular weight, 38,095). The NH2-terminal sequence of the purified enzyme from Escherichia coli indicated that the mature PI-PLC consists of 299 amino acid residues with a molecular weight of 34,586. Polyacrylamide gel electrophoresis revealed the same molecular weight for the purified enzyme isolated from the DNA-donor strain of B. thuringiensis and from the E. coli clone. By computer analysis, the secondary structure was predicted. The enzyme from the E. coli recombinant shows no activity on other phospholipids and sphingomyelin. The cleaving specificity of PI-PLC was examined by thin layer chromatography.

Amino Acid Sequence↗

Statistical analysis of data derived from clinical variables of plaque and gingivitis.

Selection of suitable subjects and statistical analysis of data derived from clinical trials presents a number of problems. In this trial, clinical data were analysed separately for pooled whole mouth data and for data including positive scores only for the variables of dental plaque and gingivitis. It was demonstrated that comparable data sets and statistical analyses were obtained using both data sets. Furthermore, it was shown that in order to achieve the best possible results in a clinical trial, the variable of gingivitis should be used in preference to plaque scores, that only individuals of high and/or moderate susceptibility to inflammation should be selected for inclusion in the statistical analysis, and that sites which have positive signs of disease only, should be included in the statistical analysis.

Adult↗

Detection of periodontal probing change by analysis of distribution mean and skew.

A problem associated with probing measurement is that important site-specific change may be obscured by measurement variability. Furthermore, if many sites are monitored there is an increasing likelihood that any particular "detected change" might be the result of this measurement error. Since error and multiplicity effects would tend to increase mouth-wise false positive rate, demanding decision thresholds are often set. However, imposition of difficult criteria, increases site-wise false negative rate and therefore reduces utility in the clinical setting. Evaluation of mean change might sometimes be a useful alternative, but also problematic as a few dramatic site-level losses can be inconsequential among many stable or improving sites. A strategy is proposed for evaluating probing changes in a single patient, which is based on the statistical evaluation of two attributes of the change distribution. (1) Disease progression is concluded when mean probing loss increases over time, and (2) a clinically relevant asymmetry is concluded when the distribution tail corresponding to loss is skewed. Computerized simulation was used to determine alpha-error for tests of mean and skew, for three different distributions of probing change. Actual alpha-error was shown to be near nominal levels. Power was estimated as a function of the number and magnitude of sites with probing loss and as a function of whether there was change in both mean and skew or in skew alone. Under most conditions studied, simultaneous tests for skew and mean provided enhanced power relative to a test for loss alone and would appear to offer the clinician an additional statistical context for appraising disease status.

Bias↗

EUD-based margin selection in the presence of set-up uncertainties.

To assess the impact of geometric uncertainties on treatment plan design, we have performed a numerical simulation in which both systematic and random errors were included. A clinical target volume (CTV) with an abutting organ at risk (OAR), both of 50 mm diameter, in a cubic phantom was modeled. A four-field conformal treatment plan was designed in which one pair of parallel-opposed beams traversed the OAR and CTV while the other pair intersected the CTV only. Field size, prescribed (isocenter) dose and systematic set-up uncertainty were varied in two orthogonal directions to examine their impact on the outcome as predicted by the dose volume histogram (DVH) and the phenomenological form of equivalent uniform dose (EUD). Of the systematic uncertainty levels considered (0, 2, 4, and 6 mm standard deviations of a Gaussian distribution), 10 mm margin (CTV-PTV) was adequate to maintain the integrity of the dose distribution within the CTV. However, reducing the margin (and hence field size) without reducing set-up errors required an increase in the isocenter dose to compensate for the loss in EUD. It was found that, in the direction containing both the CTV and OAR, with random and systematic uncertainties of 2 and 4 mm respectively, increasing the isocenter dose by about 3.5 Gy on a 6 mm-margin plan resulted in the statistically equivalent EUD value to that with a 10 mm-margin for the CTV, while the OAR EUD is dropped by 1 Gy. In general, though, the directional sensitivity to geometric uncertainties, and hence the required margin size in different directions, was dependent on beam geometries and the relative positions of the structures under consideration relative to the beam directions. Based on the validity of the EUD concept, our general conclusion is that modest dose escalation may result in plans that better achieve clinical objectives. Also, a simple single number plan quality index such as EUD5%, discussed in the paper, facilitates meaningful statistical comparisons between competing treatment strategies.

Computer Simulation↗

Inhibition of the p53 tumor suppressor gene results in growth of human aortic vascular smooth muscle cells. Potential role of p53 in regulation of vascular smooth muscle cell growth.

Loss of activity of the p53 tumor suppressor gene product has been postulated in the pathogenesis of human restenosis. Although the antioncogenes p53 and retinoblastoma (Rb) susceptibility gene have been reported to play a pivotal role in cell cycle progression in various cells, the role of p53 and Rb in the growth of human vascular smooth muscle cells (VSMC) has not yet been clarified. We used antisense strategy against p53 and Rb genes by the viral envelope-liposomal method. Transfection of antisense p53 oligodeoxynucleotides (ODN) alone resulted in an increase in DNA synthesis compared with control (P<0.01). Similarly, transfection of antisense Rb ODN alone resulted in a higher DNA synthesis rate than control (P<0.01). Moreover, increase in VSMC number was only induced by transfection of antisense p53 ODN alone or cotransfection of p53/Rb ODN (P<0.01), whereas a single transfection of antisense Rb ODN had little effect on cell number. Therefore, we hypothesized that this discrepancy is due to the induction of apoptosis mediated by p53. Interestingly, apoptotic cells were markedly increased in VSMC transfected with antisense Rb ODN alone, accompanied by the induction of p53 protein. The number of apoptotic cells was attenuated by cotransfection of antisense p53 ODN (P<0.01). We finally examined the molecular mechanisms of apoptosis induced by the absence of Rb. In VSMC transfected with antisense Rb ODN, bax, a promoter of apoptosis, was significantly increased in VSMC transfected with antisense Rb ODN (P<0.01), whereas bcl-2 and Fas did not play a pivotal role in the induction of apoptosis. Overall, these data first demonstrated that the antioncogenes p53 and Rb negatively regulated the cell cycle in VSMC, suggesting that the modulation of their activity may mediate VSMC growth such as that in restenosis and atherosclerosis. The presence of p53 plays a pivotal role in the regulation of apoptosis in human VSMC growth, probably through the bax pathway. These results provide evidence that p53 is a functional link between cell growth and apoptosis in VSMC.

Analysis of Variance↗

Statistical issues in assessing powered toothbrushes.

A symposium on powered toothbrushes was held at the 2000 IADR General Session. The author was asked to address statistical issues in conducting clinical studies. The objective was to cover important statistical topics that should be considered in scientific investigations, including designing the study, analyzing the results, and performing a critical scientific review. The American Dental Association (ADA) established guidelines during the 1990s. The author discusses these guidelines and addresses statistical issues they do not cover.

American Dental Association↗

A note on generalized Genome Scan Meta-Analysis statistics.

BACKGROUND: Wise et al. introduced a rank-based statistical technique for meta-analysis of genome scans, the Genome Scan Meta-Analysis (GSMA) method. Levinson et al. recently described two generalizations of the GSMA statistic: (i) a weighted version of the GSMA statistic, so that different studies could be ascribed different weights for analysis; and (ii) an order statistic approach, reflecting the fact that a GSMA statistic can be computed for each chromosomal region or bin width across the various genome scan studies. RESULTS: We provide an Edgeworth approximation to the null distribution of the weighted GSMA statistic, and, we examine the limiting distribution of the GSMA statistics under the order statistic formulation, and quantify the relevance of the pairwise correlations of the GSMA statistics across different bins on this limiting distribution. We also remark on aggregate criteria and multiple testing for determining significance of GSMA results. CONCLUSION: Theoretical considerations detailed herein can lead to clarification and simplification of testing criteria for generalizations of the GSMA statistic.

Chromosome Mapping↗

Meta-analytic methods for pooling rates when follow-up duration varies: a case study.

BACKGROUND: Meta-analysis can be used to pool rate measures across studies, but challenges arise when follow-up duration varies. Our objective was to compare different statistical approaches for pooling count data of varying follow-up times in terms of estimates of effect, precision, and clinical interpretability. METHODS: We examined data from a published Cochrane Review of asthma self-management education in children. We selected two rate measures with the largest number of contributing studies: school absences and emergency room (ER) visits. We estimated fixed- and random-effects standardized weighted mean differences (SMD), stratified incidence rate differences (IRD), and stratified incidence rate ratios (IRR). We also fit Poisson regression models, which allowed for further adjustment for clustering by study. RESULTS: For both outcomes, all methods gave qualitatively similar estimates of effect in favor of the intervention. For school absences, SMD showed modest results in favor of the intervention (SMD -0.14, 95% CI -0.23 to -0.04). IRD implied that the intervention reduced school absences by 1.8 days per year (IRD -0.15 days/child-month, 95% CI -0.19 to -0.11), while IRR suggested a 14% reduction in absences (IRR 0.86, 95% CI 0.83 to 0.90). For ER visits, SMD showed a modest benefit in favor of the intervention (SMD -0.27, 95% CI: -0.45 to -0.09). IRD implied that the intervention reduced ER visits by 1 visit every 2 years (IRD -0.04 visits/child-month, 95% CI: -0.05 to -0.03), while IRR suggested a 34% reduction in ER visits (IRR 0.66, 95% CI 0.59 to 0.74). In Poisson models, adjustment for clustering lowered the precision of the estimates relative to stratified IRR results. For ER visits but not school absences, failure to incorporate study indicators resulted in a different estimate of effect (unadjusted IRR 0.77, 95% CI 0.59 to 0.99). CONCLUSIONS: Choice of method among the ones presented had little effect on inference but affected the clinical interpretability of the findings. Incidence rate methods gave more clinically interpretable results than SMD. Poisson regression allowed for further adjustment for heterogeneity across studies. These data suggest that analysts who want to improve the clinical interpretability of their findings should consider incidence rate methods.

Absenteeism↗

Computational intelligence for laboratory information systems.

Non-linear models, such as given by neural networks and fuzzy logic, have established a good reputation for medical data analysis as computational and logical counterparts to statistical methods. Whereas multilayer perceptrons perform well with large data sets, a combination of neural learning together with fuzzy logical network interpretations provides a network reduction well suited for smaller data sets. The aim of this paper is to present an approach to neural fuzzy systems data analysis and knowledge acquisition in laboratory information systems. We also describe a software system, DiagaiD, which provides an analysis and development workbench involving laboratory data.

Clinical Laboratory Information Systems↗

The controversy of significance testing: misconceptions and alternatives.

The current debate about the merits of null hypothesis significance testing, even though provocative, is not particularly novel. The significance testing approach has had defenders and opponents for decades, especially within the social sciences, where reliance on the use of significance testing has historically been heavy. The primary concerns have been (1) the misuse of significance testing, (2) the misinterpretation of P values, and (3) the lack of accompanying statistics, such as effect sizes and confidence intervals, that would provide a broader picture into the researcher's data analysis and interpretation. This article presents the current thinking, both in favor and against, on significance testing, the virtually unanimous support for reporting effect sizes alongside P values, and the overall implications for practice and application.

Bias↗