PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Tests for differential gene expression using weights in oligonucleotide microarray experiments.

BACKGROUND: Microarray data analysts commonly filter out genes based on a number of ad hoc criteria prior to any high-level statistical analysis. Such ad hoc approaches could lead to conflicting conclusions with no clear guidance as to which method is most likely to be reproducible. Furthermore, the number of tests performed with concomitant inflation in type I error also plagues the statistical analysis of microarray data, since the number of tested quantities in a study significantly affects the family-wise error rate. It would, therefore, be very useful to develop and adopt strategies that allow quantification of the quality of each probeset, to filter out or give little credence to low-quality or unexpressed probesets, and to incorporate these strategies into gene selection within a multiple testing framework. RESULTS: We have proposed a unified scheme for filtering and gene selection. For Affymetrix gene expression microarrays, we developed new methods for measuring the reliability of a particular probeset in a single array, and we used these to develop measures for a set of arrays. These measures are then used as weights in standard t-statistic calculations, and are incorporated into the multiple testing procedures. We demonstrated the advantages of our methods using simulated data, publicly available spiked-in data as well as data comparing normal muscle to muscle from patients with Duchenne muscular dystrophy (DMD), in which a set of truly differentially expressed genes is known. CONCLUSION: Our quality measures provide convenient ways to search for individual genes of high quality. The quality weighting strategies we proposed for testing differential gene expression have demonstrable improvement on the traditional filtering methods, the standard t-statistic and a regularized t-statistic in Affymetrix data analysis.

Data Interpretation, Statistical↗

Structural diversity of protein segments follows a power-law distribution.

The local structures of protein segments were classified and their distribution was analyzed to explore the structural diversity of proteins. Representative proteins were divided into short segments using a sliding L-residue window. Each set of local structures consisting of consecutive 1-31 amino acids was classified using a single-pass clustering method. The results demonstrate that the local structures of proteins are very unevenly distributed in the protein universe. The distribution of local structures of relatively long segments shows a power-law behavior that is formulated well by Zipf's law, implying that a protein structure possesses recursive and fractal characteristics. The degree of effective conformational freedom per residue as well as the structure entropy per residue decreases gradually with an increasing value of L and then converges to constant values. This suggests that the number of protein conformations resides within the range between 1.2L and 1.5L and that 10- to 20-residue segments are already proteinlike in terms of their structural diversity.

Amino Acid Motifs↗

Mapping of dopamine D3 receptor binding site by pharmacological characterization of mutants expressed in CHO cells with the Semliki Forest virus system.

Nine mutants and the wild-type human dopamine D3 receptor were expressed at high levels in BHK and CHO cells using the Semliki Forest virus system and were analysed for receptor binding with several structurally different dopamine D3 ligands. The mutation His349Leu showed a significant decrease in pKi values for raclopride, dopamine and GR218231, but an increase in affinity for GR99841. Thr369Val had an increase in pKi for both GR99841 and 7-OH-DPAT. The receptor modelling based on sequence alignment with bacteriorhodopsin indicated that Thr369 and His349 are located on the inside of the ligand binding pocket and the effect of the mutagenesis was therefore expected. The change in binding affinity for Thr369Val could be due to the location in the transmembrane domain VII close to the aspartate residue in domain III, the postulated counter ion for dopamine.

Amino Acid Sequence↗

Use of reported immunisation performance data in monitoring UIP.

Health information system is not well developed in most parts of the country. We propose to use locally available data on immunization performance so as to make them more informative and useful for implementing and supervisory personnel. Reclassification of data according to estimated number of eligibles, monthly performance and age of children immunized can given impressions about coverage, sustainability of services and their quality. Similarly, area and institution-wise data can be used to identify places needing more attention. Active use of available data will help in improvement of vaccine coverage and control of target diseases.

Data Interpretation, Statistical↗

The taste of monosodium glutamate: membrane receptors in taste buds.

Receptor proteins for photoreception have been studied for several decades. More recently, putative receptors for olfaction have been isolated and characterized. In contrast, no receptors for taste have been identified yet by molecular cloning. This report describes experiments aimed at identifying a receptor responsible for the taste of monosodium glutamate (MSG). Using reverse transcriptase (RT)-PCR, we found that several ionotropic glutamate receptors are present in rat lingual tissues. However, these receptors also could be detected in lingual tissue devoid of taste buds. On the other hand, RT-PCR and RNase protection assays indicated that a G-protein-coupled metabotropic glutamate receptor, mGluR4, also is expressed in lingual tissues and is limited only to taste buds. In situ hybridization demonstrated that mGluR4 is detectable in 40-70% of vallate and foliate taste buds but not in surrounding nonsensory epithelium, confirming the localization of this metabotropic receptor to gustatory cells. Expression of mGluR4 in taste buds is higher in preweaning rats compared with adult rats. This may correspond to the known higher sensitivity to the taste of MSG in juvenile rodents. Finally, behavioral studies have indicated that MSG and L-2-amino-4-phosphonobutyrate (L-AP4), a ligand for mGluR4, elicit similar tastes in rats. We conclude that mGluR4 may be a chemosensory receptor responsible, in part, for the taste of MSG.

Amino Acid Sequence↗

[Interpretation of results of clinical trials in benign prostatic hyperplasia].

The random clinical trial (RCT) is the most suitable study to evaluate the treatment effectiveness in the benign prostatic hyperplasia (BPH). Although most of the urologists will not collaborate in a RCT development, they will treat BPH patients, so it is very important to know if a CRT in BPH is well designed and their conclusions are correct. The aim of this article is to give the basic elements of analysis that urologists need in order to evaluate the quality and the level of evidence of a RCT in BPH. This article emphasizes the three main elements of a RCT: to check if the study has been correctly performed (internal validity), to evaluate if the treatment achieves an important clinical improvement (relevance of the results) and the applicability of the results in our patients (external validity). The article shows that to analyse these elements common sense and clinical judgment are needed rather than statistical knowledge.

Data Interpretation, Statistical↗

An assessment of neural network and statistical approaches for prediction of E. coli promoter sites.

We have constructed a perceptron type neural network for E. coli promoter prediction and improved its ability to generalize with a new technique for selecting the sequence features shown during training. We have also reconstructed five previous prediction methods and compared the effectiveness of those methods and our neural network. Surprisingly, the simple statistical method of Mulligan et al. performed the best amongst the previous methods. Our neural network was comparable to Mulligan's method when false positives were kept low and better than Mulligan's method when false negatives were kept low. We also showed the correlation between the prediction rates of neural networks achieved by previous researchers and the information content of their data sets.

Base Sequence↗

Progress on the CSE diagnostic study. Application of McNemar's test revisited.

The authors describe an extension of McNemar's test that can be used to compare diagnostic performance when multiple statements are obtained from computer analysis or visual interpretation of the ECG. If two or more diagnostic statements are made, by definition only one can be correct for the cases in the CSE pilot database, which were selected for single, clinically well documented, abnormalities. If one statement stood out from others as being made with the highest degree of certainty, that was accepted as the single interpretation, right or wrong. However, when two or more statements were made with the same degree of certainty and only one statement was correct, then in the previous application of McNemar's test that interpretation was given credit for being correct. In the extension of the method presented in this article, such an interpretation is given partial credit for the one correct statement and partial discredit for any incorrect statement, thereby reporting the results more properly in the sensitivity and specificity statistics for the different diagnostic categories.

Algorithms↗

Alternative definitions of comparable case groups and estimates of lead time and benefit time in randomized cancer screening trials.

Randomized screening trials provide the optimal means of assessing the benefit of screening for cancer and other chronic diseases. Unlike therapy trials, however, where strict eligibility criteria assure the comparability of cases of disease in the arms of the trial, the cancer cases identified during follow-up are a subset of all randomized participants. Furthermore, those cases detected by screening tend to arise from length biased sampling which also can bias estimates of the screening benefit and of average lead time. To reduce or eliminate this bias, we propose several methods for defining comparable groups of cases from the trial arms. We examine, via simulation, these methods with respect to their effects on (i). point and interval estimates of average lead time and average benefit time and (ii). the logrank test statistic for a mortality effect of screening. The most successful new method for defining comparable case groups uses an estimate of the mean sojourn time (mean preclinical duration), and results in nearly unbiased estimates of average lead time and average benefit time as well as an unbiased logrank test statistic.

Adult↗

Localized measures for nonstationary time-series of physiological data.

We will discuss localized measures related to the concepts of dimension, Lyapunove exponents (entropy), and recurrence diagrams. We stress the relevance of localized events and coincidences in physiological time series that often are lost when statistical averaging methods are applied. We suggest event-based statistics as an alternative to spectral or averaged-based statistics. The use of wavelets bases for characterizing localized structures is discussed as a potential alternative to Fourier-based analysis. Finally we mention how local domains in state space could be applied as triggers for external stimuli and thereby improve the statistics of ERP recordings.

Animals↗

SWORDS: a statistical tool for analysing large DNA sequences.

In this article, we present some simple yet effective statistical techniques for analysing and comparing large DNA sequences. These techniques are based on frequency distributions of DNA words in a large sequence, and have been packaged into a software called SWORDS. Using sequences available in public domain databases housed in the Internet, we demonstrate how SWORDS can be conveniently used by molecular biologists and geneticists to unmask biologically important features hidden in large sequences and assess their statistical significance.

Animals↗

Hierarchical time-oriented approaches to missing data inference.

In practice clinical data are nearly always incomplete. When confronted with such data, a physician or investigator must make inferences about missing information. Possible strategies for inference include (1) interpolation, (2) extrapolation, (3) repeating the nearest value, (4) repeating the previous value, (5) patient-specific mean values, (6) patient-specific linear regression over time, (7) disease-specific mean values, (8) normal values, and (9) linear regression of correlated co-recorded variables. This study analyzes these strategies in a time-oriented data bank of patients with systemic lupus erythematosus, demonstrating that more accurate inferences of missing data are obtained when (1) strategies are tailored to the characteristics of the individual variable, (2) time-oriented strategies (e.g., interpolation) rather than non-time-oriented strategies (e.g., disease mean) are incorporated, (3) a ranked set of strategies is incorporated in a hierarchical stepwise fashion, and (4) the degree to which missing data are "nonrandomly" missing is assessed to allow estimation of bias. Interpolation is the best single technique with these data while linear regression of correlated co-recorded variables is a relatively weak technique. Inferences made by these hierarchical time-oriented approaches show significantly smaller mean differences from the actual values than do results from typical statistical package strategies.

Data Interpretation, Statistical↗

A demonstration of interval-censored survival analysis.

Interval-censoring occurs in survival analysis when the time until an event of interest is not known precisely (and instead, only is known to fall into a particular interval). Such censoring commonly is produced when periodic assessments (usually clinical or laboratory examinations) are used to assess if the event has occurred. My objectives were to raise awareness about interval-censoring including its existence, the potential ramifications of ignoring its existence, the different types of interval-censored data, and the analytical methods to analyze such data (including availability in standard statistical software). Asynchronous interval-censored survival analysis was demonstrated by parametric evaluation of risk factors for the time to first detected shedding of Salmonella muenster (identified by repeated periodic fecal cultures) for a herd of dairy cows. These results were compared with those from survival analyses which ignored or approximated the interval-censoring. Ignoring or approximating the asynchronous interval-censoring in the survival analysis generally resulted in the risk factors' regression coefficients having the same signs and a decrease (often >50%) in their absolute size. All the standard errors from the three methods of approximating the interval-censoring were <40% of their interval-censored counterparts. The conclusions drawn from the asynchronous interval-censored analysis versus those from the approximations varied dramatically. (The general conclusion from the approximations was that none of the risk factors for this example warranted further consideration.) That ignoring or approximating the left- and interval-censored nature of the dependent variable resulted in biased results was consistent with the literature. In the currently available asynchronous interval-censored models, the inclusion of time-dependent covariates that vary continuously is awkward. Statistical models for the semi-parametric estimation of asynchronous interval-censored survival analysis are not generally available in standard statistical software.

Animals↗

Number-between g-type statistical quality control charts for monitoring adverse events.

Alternate Shewhart-type statistical control charts, called "g" and "h" charts, are developed and evaluated for monitoring the number of cases between hospital-acquired infections and other adverse events, such as heart surgery complications, catheter-related infections, surgical site infections, contaminated needle sticks, and other iatrically induced outcomes. These new charts, based on inverse sampling from geometric and negative binomial distributions, are simple to use and can exhibit significantly greater detection power over conventional binomial-based approaches, particularly for infrequent events and low "defect" rates. A companion article illustrates several interesting properties of these charts and design modifications that significantly can improve their statistical properties, operating characteristics, and sensitivity.

Cross Infection↗

Overcoming feelings of powerlessness in "aging" researchers: a primer on statistical power in analysis of variance designs.

A general rationale and specific procedures for examining the statistical power characteristics of psychology-of-aging empirical studies are provided. First, 4 basic ingredients of statistical hypothesis testing are reviewed. Then, 2 measures of effect size are introduced (standardized mean differences and the proportion of variation accounted for by the effect of interest), and methods are given for estimating these measures from already-completed studies. Power and sample size formulas, examples, and discussion are provided for common comparison-of-means designs, including independent samples I-factor and factorial analysis of variance (ANOVA) design, analysis of covariance designs, repeated measures (correlated samples) ANOVA designs, and split-plot (combined between- and within-subjects) ANOVA designs. Because of past conceptual differences, special attention is given to the power associated with statistical interactions, and cautions about applying the various procedures are indicated. Illustrative power estimations also are applied to a published study from the literature. It is argued that psychology-of-aging researchers will be both better informed consumers of what they read and more "empowered" with respect to what they research by understanding the important roles played by power and sample size in statistical hypothesis testing.

Aged↗