PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Statistical analysis and interpretation in a study of prelingually deaf children implanted before five years of age.

This article describes the major statistical analyses used in a large-scale study of prelingually deaf children implanted before 5 yr of age. Two major challenges posed by the data were the need to reduce the number of outcome measures to a reasonable number while retaining the important information underlying the multiple measures, and the need to determine the unique contributions of multiple sets of predictor variables. The approach taken-principal components analysis of the outcome measures followed by hierarchical multiple regression-provided a compromise between statistical complexity and ease of interpretation.

Child↗

Item response theory: applications of modern test theory in medical education.

CONTEXT: Item response theory (IRT) measurement models are discussed in the context of their potential usefulness in various medical education settings such as assessment of achievement and evaluation of clinical performance. PURPOSE: The purpose of this article is to compare and contrast IRT measurement with the more familiar classical measurement theory (CMT) and to explore the benefits of IRT applications in typical medical education settings. SUMMARY: CMT, the more common measurement model used in medical education, is straightforward and intuitive. Its limitation is that it is sample-dependent, in that all statistics are confounded with the particular sample of examinees who completed the assessment. Examinee scores from IRT are independent of the particular sample of test questions or assessment stimuli. Also, item characteristics, such as item difficulty, are independent of the particular sample of examinees. The IRT characteristic of invariance permits easy equating of examination scores, which places scores on a constant measurement scale and permits the legitimate comparison of student ability change over time. Three common IRT models and their statistical assumptions are discussed. IRT applications in computer-adaptive testing and as a method useful for adjusting rater error in clinical performance assessments are overviewed. CONCLUSIONS: IRT measurement is a powerful tool used to solve a major problem of CMT, that is, the confounding of examinee ability with item characteristics. IRT measurement addresses important issues in medical education, such as eliminating rater error from performance assessments.

Clinical Competence↗

[Additional coding for legally required ICD-9 as the basis for high quality patient record searches].

We are convinced, that the updating of the ID DIACOS diagnosis catalogue fits better to the today's medical language, than the former one. By using a consequent additional code we could eliminate several lacks of the ICD-9. Valid and reproducible diagnosis statistics are now possible. Subjective coding mistakes are largely avoidable. Inquiries of certain diagnosis can be answered by using the one of the two codes, that describes the diagnosis in the best form. Regarding the allotments of the best form. Regarding the allotments of the catalogue of orthopedic and traumatic diagnosis including the possibility of double-coding we got a special better version of ID DIACOS than before.

Data Collection↗

Prediction of protein signal sequences and their cleavage sites by statistical rulers.

Functioning as an "address tag" or "zip code" that guides nascent proteins (newly synthesized proteins in the cytosol) to wherever they are needed, signal peptides (also called targeting signals or signal sequences) have become a crucial tool in finding new drugs or reprogramming cells for gene therapy. To effectively and timely use such a tool, however, the first important thing is to develop an automated method for quickly and accurately identifying the signal peptide for a given nascent protein. With the avalanche of new protein sequences generated in the post-genomic era, the challenge has become even more urgent and critical. In this paper, five statistical rulers were derived via performing a mutual information analysis. By combining these statistical rulers, a new prediction algorithm was established and high success prediction rates were observed. The new algorithm may play a complementary role to the existing algorithms in this area. It is anticipated that the mutual information approach introduced here may be very useful for studying many other sequence-coupling problems in molecular biology as well.

Algorithms↗

Biostatistical approaches to reducing the number of animals used in biomedical research.

For ethical reasons, the least number of animals possible should be used in biomedical research, though not so few as to fail to detect biologically important effects or to necessitate the repetition of experiments. We describe biostatistical approaches that can contribute to either reducing the number of animals in single experiments or to increasing the quality of studies so that fewer subsequent studies (and thus animals) will be needed. The described approaches regard different phases of experimentation, specifically: planning the experimental design and calculating the sample size, controlling variability, choosing the response variable, postulating the statistical hypothesis to be tested, choosing the procedure for analysing data, and interpreting and suitably presenting the results.

Animal Experimentation↗

Funnels, pathways, and the energy landscape of protein folding: a synthesis.

The understanding, and even the description of protein folding is impeded by the complexity of the process. Much of this complexity can be described and understood by taking a statistical approach to the energetics of protein conformation, that is, to the energy landscape. The statistical energy landscape approach explains when and why unique behaviors, such as specific folding pathways, occur in some proteins and more generally explains the distinction between folding processes common to all sequences and those peculiar to individual sequences. This approach also gives new, quantitative insights into the interpretation of experiments and simulations of protein folding thermodynamics and kinetics. Specifically, the picture provides simple explanations for folding as a two-state first-order phase transition, for the origin of metastable collapsed unfolded states and for the curved Arrhenius plots observed in both laboratory experiments and discrete lattice simulations. The relation of these quantitative ideas to folding pathways, to uniexponential vs. multiexponential behavior in protein folding experiments and to the effect of mutations on folding is also discussed. The success of energy landscape ideas in protein structure prediction is also described. The use of the energy landscape approach for analyzing data is illustrated with a quantitative analysis of some recent simulations, and a qualitative analysis of experiments on the folding of three proteins. The work unifies several previously proposed ideas concerning the mechanism protein folding and delimits the regions of validity of these ideas under different thermodynamic conditions.

Amino Acid Sequence↗

The association between size of test chamber and patch test reaction: a statistical reanalysis.

A recent study by Brasch and co-workers reported on the association between size of test chamber and patch test reaction. The investigators interpreted their data on 495 patients as having conclusively shown that standard preparations of fragrance mix, wool wax alcohols, Kathon CG and formaldehyde led to more positive test reactions when large Finn Chambers were used for patch testing. We have scrutinized the statistical aspects of this study and conclude that the authors should have adopted a statistical approach suitable to analyse dependent samples. After explaining the correct methodological way of dealing with quadratic contingency tables formed by 2 dependent samples, we reanalyze the data accordingly and compare the results to those of the original paper. Based on this reanalysis, the conclusions are more complex: the reaction pattern for the fragrance mix and wool wax alcohols is significantly different between small and large test chambers; however, this discrepancy arises primarily from changing weak positive reactions with small chambers to strong positive reactions with large chambers. For formaldehyde, no relationship between chamber size and patch test reaction was found in the data, while for Kathon CG, statistical evidence is borderline that more positive test reactions are yielded by large test chambers than by small ones.

Adult↗

Mastering research critique and statistical interpretation. Guidelines and golden rules.

Mastery of statistical analysis research critique is an important skill for professional nurses. A Guideline for Statistical Analysis and Golden Rules for Statistical Analysis Adequacy are presented and applied to classroom use. Students who learn how to critique research and statistics usage effectively are satisfied consumers and report using knowledge in other clinical courses. Effective strategies to teach statistical analysis critique are discussed.

Data Interpretation, Statistical↗

goCluster integrates statistical analysis and functional interpretation of microarray expression data.

MOTIVATION: Several tools that facilitate the interpretation of transcriptional profiles using gene annotation data are available but most of them combine a particular statistical analysis strategy with functional information. goCluster extends this concept by providing a modular framework that facilitates integration of statistical and functional microarray data analysis with data interpretation. RESULTS: goCluster enables scientists to employ annotation information, clustering algorithms and visualization tools in their array data analysis and interpretation strategy. The package provides four clustering algorithms and GeneOntology terms as prototype annotation data. The functional analysis is based on the hypergeometric distribution whereby the Bonferroni correction or the false discovery rate can be used to correct for multiple testing. The approach implemented in goCluster was successfully applied to interpret the results of complex mammalian and yeast expression data obtained with high density oligonucleotide microarrays (GeneChips). AVAILABILITY: goCluster is available via the BioConductor portal at www.bioconductor.org. The software package, detailed documentation, user- and developer guides as well as other background information are also accessible via a web portal at http://www.bioz.unibas.ch/gocluster CONTACT: michael.primig@unibas.ch

Algorithms↗

[Statistically validated evaluation of clinical trials].

Data of clinical trials of medicinal products must be evaluated in statistically valid models. The statistical validity criteria are defined. Statistically invalid models will result in biased parameter and confidence interval estimations, erroneous statistical inferences and clinical interpretations. Finally, wrong decisions will call forth deleterious consequences in the judgement of the therapeutic effect and the frequency and severity of the adverse reactions of the tested new medicinal, and generic products. Statistically validated analyses will promote the international harmonization of the scientific evaluation of medicinal products according to the idea of the evidence-based-medicine. The study presents examples of clinical trials evaluated with a software checking statistical validity assumptions while performing evaluation of data.

Clinical Trials as Topic↗

Statistical analysis of the 5' untranslated region of human mRNA using "Oligo-Capped" cDNA libraries.

We constructed 34 types of human "full-length enriched" and "5'-end enriched" cDNA libraries based on the "Oligo-Capping" method. We randomly picked and sequenced 10,000 clones from these libraries. BLAST analysis showed that about 50% of the cDNAs were identical to known genes. Among them, we selected 954 species of cDNA that should represent the entire sequence from the mRNA start sites. Compared with previously reported sequences, they were on average 45 bp longer in the 5'-end. Using these cDNA data, we statistically analyzed the sequence features of the 5'UTR. The average length of the 5'UTR was 125 bp, and there was little correlation with the corresponding mRNA length (correlation coefficient = 0.26). Of the 954 species of 5'UTR, 459 contained no in-frame terminator codon, which is against the common belief. Two hundred seventy-eight species contained at least one ATG codon upstream of the initiator ATG codon. We identified 569 upstream ATGs, in total, 63% of which adequately satisfied Kozak's criteria. These findings are contrary to the typical translation initiation model, which states that translation is initiated from the "first" ATG codon.

5' Untranslated Regions↗

A study of the heterogeneity of bacterial fluctuation-test data and the effects of auxotrophic-growth enhancement.

Microtitre fluctuation tests using Salmonella typhimurium TA100 and TA98 and Escherichia coli WP2uvrApKM101 were performed to determine the heterogeneity of spontaneous mutagenicity data and how this affects interpretation of results. Assays were performed in the absence and presence of additional amino acids to simulate the effects of the auxotrophic growth enhancement characteristic of tests of complex biological mixtures. The results indicate that the heterogeneity of the data did not depart significantly from that expected from binomial theory and that two criteria must be satisfied in order to define a positive result: (i) a net increase which significantly exceeds the background mutation and which takes account of its heterogeneity; and (ii) a statistically significant dose response.

Data Interpretation, Statistical↗

Conceptive delays of twin-prone mothers: a demographic epidemiologic approach.

We studied the time interval to the first birth and to the twin birth using statistical and mathematical models in two groups of mothers, those with twins and those with singletons, from the same population. We made use of a pair-matched case-control design. We treated the maternal birth cohort and parity as confounders and thus as controlled. We also investigated the sex of twin pairs as an interactive variable, employing such methods as survival curve testing and using geometric, gamma, and exponential distributions where appropriate. The expectations derived from the mathematical models yield numerical estimates of fertility components. The results suggest that unlike-sex twin-prone mothers have higher fecundity than controls when they conceive singletons. Further, fecundity appears high and unimpaired before the birth of twins. Mothers of like-sex twins experience somewhat shorter and but more variable birth intervals than corresponding controls before the birth of twins, suggesting within-group heterogeneity. Specifically, the birth of like-sex twins is preceded by low fecundity and a short period of postpartum amenorrhea. Biologically, like-sex (presumably monozygotic) twin-prone mothers have a hormonal defect related eventually to menopausal status that interferes with ovulation and perhaps with lactation. As for unlike-sex twin-bearing mothers, they probably experience a displacement of their maximum fertility potential toward early reproductive life and an extension of their menstrual life. From a methodologic standpoint, the study of the fertility of twin-prone mothers cannot proceed without estimates of the fertility components of birth intervals, as these intervals do not lend themselves to straightforward analytical interpretations by statistical analyses.

Adult↗

Searching for statistically significant regulatory modules.

MOTIVATION: The regulatory machinery controlling gene expression is complex, frequently requiring multiple, simultaneous DNA-protein interactions. The rate at which a gene is transcribed may depend upon the presence or absence of a collection of transcription factors bound to the DNA near the gene. Locating transcription factor binding sites in genomic DNA is difficult because the individual sites are small and tend to occur frequently by chance. True binding sites may be identified by their tendency to occur in clusters, sometimes known as regulatory modules. RESULTS: We describe an algorithm for detecting occurrences of regulatory modules in genomic DNA. The algorithm, called mcast, takes as input a DNA database and a collection of binding site motifs that are known to operate in concert. mcast uses a motif-based hidden Markov model with several novel features. The model incorporates motif-specific p-values, thereby allowing scores from motifs of different widths and specificities to be compared directly. The p-value scoring also allows mcast to only accept motif occurrences with significance below a user-specified threshold, while still assigning better scores to motif occurrences with lower p-values. mcast can search long DNA sequences, modeling length distributions between motifs within a regulatory module, but ignoring length distributions between modules. The algorithm produces a list of predicted regulatory modules, ranked by E-value. We validate the algorithm using simulated data as well as real data sets from fruitfly and human. AVAILABILITY: http://meme.sdsc.edu/MCAST/paper

Algorithms↗

The utility of prior information and stratification for parameter estimation with two screening tests but no gold standard.

When a gold standard screening or diagnostic test is not routinely available, it is common to apply two different imperfect tests to subjects from a study population. There is a considerable literature on estimating relevant parameters from the resultant data. In the situation that test sensitivities and specificities are unknown, several inferential strategies have been proposed. One suggestion is to use rough knowledge about the unknown test characteristics as prior information in a Bayesian analysis. Another suggestion is to obtain the statistical advantage of an identified model by splitting the population into two strata with differing disease prevalences. There is some division of opinion in the epidemiological literature on the relative merits of these two approaches. This article aims to shed light on the issue, by applying some recently developed theory on the performance of Bayesian inference in non-identified statistical models.

Bayes Theorem↗

Which measures of skewness and kurtosis are best?

Indices of distributional shape based on linear combinations of order statistics have recently been described by Hosking. Their usefulness as tools for practical data analysis is examined. They are found to have several advantages over the conventional indices of skewness and kurtosis (square root of b1 and b2) and no serious drawbacks. It is proposed, therefore, that they should replace square root of b1 and b2 in routine data analysis. To implement this suggestion, action by the developers of standard statistical software is needed.

Bias↗

A program for measuring timing and accuracy in a complex motor reaction task.

A program was developed to explore a subject's accuracy and response timing in a reaction-time paradigm, which involves a multiple choice along with a complex response. The program can be run on an IBM-PC or compatible computers. The subject's task is: (1) to react to a stimulus (task) pattern by releasing a go-key (release time); the pattern is displayed on a keyboard as a combination of up to five light-emitting diodes associated with five answer keys; (2) then to press correctly the key combination indicated (response completion time). The keys are located at positions corresponding to finger tips. The response is complete at the moment all keys of the required combination are held down simultaneously, the order in response completion is not important. Responses are classified according to possible errors: whenever a key is pressed that does not belong to the required combination or the full key combination is not held until a (preset) waiting time period, the response is classified as incorrect. Various task patterns representing the required response combination of any subset of the five keys can be designed. Both the left and the right hand can be tested separately. In this way, possible distortions in performance caused, for instance, by apraxia can be analyzed. Basic statistical characteristics of the release times and completion times for correct responses are computed from the data and stored. The program is written in MODULA-2, the output files are in the ASCII format, which enables further processing of the data by standard statistical packages.

Apraxias↗