PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

[Choice and consequence: measurement level determines the statistical tool-box].

Various types of data are used in health and behavioural sciences. Some of the variables, such as blood variables, can be measured by standardised, calibrated methods while others are based on subjective judgements from an expert or from the patients. There is a strong link between the measurement process and the choice of statistical toolbox. The statistically important measurement levels are; categorical, ordinal, quantitative discrete and quantitative continuous data. Special attention should be paid to non-negative quantitative data, and to ordinal data which are very common in medicine. In this paper the properties of the measurement levels are given and their consequences on the choice of statistical methods for description and analysis.

Data Interpretation, Statistical↗

Systems for data analysis.

A system for data analysis is the end product of study planning, form design, data entry, data verification, and statistical analysis. This article reviews these steps and considers the fundamental choices in software for data entry and analysis. The appendix includes a listing of general and specialized software for data management and statistical analysis.

Data Interpretation, Statistical↗

Application of statistical methods in the papers published in the East African Medical Journal (EAMJ) since 1923.

The subject Statistics was hardly known in the research world at the time the EAMJ was being started in 1923. It was at this time, scientists working in their field of specialization were busy developing statistical methods with the aim of solving problems affecting them in their research work. Using a random sample of some of the papers published by the EAMJ, this paper evaluates the usage of statistical methods since 1923. Usage was very low before 1965, but started picking-up with a rising trend since then but still not impressive. Scientists should strive to apply statistical methods correctly in their research work. EAMJ should lay more emphasis on statistical refereeing as a policy to raise the quality of published papers in the Journal. If statisticians were not available, which often may be the case, scientists should be encouraged to show their work to other research colleagues working in similar areas, before sending the paper to the EAMJ. By adopting this approach, it is possible that the quality of papers will go up and the usage of statistical methods will increase.

Bias↗

[Examples of pitfalls in statistical analysis--4: What is a type II error?].

When assessing the results of statistical analysis, we have to consider two types of errors, i.e., type I errors, which present the risk of erroneously judging a true hypothesis to be false, and type II errors, which present the risk of erroneously judging a true hypothesis to be true. If we use only type I errors to detect the significances in statistical analysis, we can not avoid making the erroneous judgement that a true hypothesis is false. When the results of analysis are not significant, we have to calculate type II errors and discuss the probabilities of having missed the truth. In this article, I present an example to illustrate two types of errors which can occur in assessing the result of statistical analysis.

Anesthesia, Inhalation↗

Statistical measures of the structure of genomic sequences: entropy, complexity, and position information.

Identifying regions of DNA with extreme statistical characteristics is an important aspect of the structural analysis of complete genomes. Linguistic methods, mainly based on estimating word frequency, can be used for this as they allow for the delineation of regions of low complexity. Low complexity may be due to biased nucleotide composition, by tandem- or dispersed repeats, by palindrome-hairpin structures, as well as by a combination of all these features. We developed software tools in which various numerical measures of text complexity are implemented, including combinatorial and linguistic ones. We also added Hurst exponent estimate to the software to measure dependencies in DNA sequences. By applying these tools to various functional genomic regions, we demonstrate that the complexity of introns and regulatory regions is lower than that of coding regions, whilst Hurst exponent is larger. Further analysis of promoter sequences revealed that the lower complexity of these regions is associated with long-range correlations caused by transcription factor binding sites.

Algorithms↗

The interaction of statistical significance, biology of dose-response and test design in the assessment of genotoxicity data.

The most useful function of statistical significance in genetic toxicology is to separate weakly positive from negative data. To achieve this objective, the experimental system, design and analysis must be considered as a whole and not as separate entities. Assumptions underlying statistical methods should be checked for applicability to the biology of the test. The number and arrangement of samples and dose levels must be organized to suit the selected statistical method. The number of observations in the test and in each sample will depend on the smallest increase of biological importance which should be detected as significant (at a selected value for alpha error) and the maximum acceptable chance of failing to detect this increase as significant (beta error). Generally, a few dose levels and a high degree of replication are desirable with a repeat test if practical. In the analyses of data it is particularly important to adapt the statistics if there are multiple comparisons. Even after correct determination of significance, interpretation can be difficult because of possible experimental artefacts and chance errors plus the impossibility of proving a negative result.

Data Interpretation, Statistical↗

A statistical perspective on gene expression data analysis.

Rapid advances in biotechnology have resulted in an increasing interest in the use of oligonucleotide and spotted cDNA gene expression microarrays for medical research. These arrays are being widely used to understand the underlying genetic structure of various diseases, with the ultimate goal to provide better diagnosis, prevention and cure. This technology allows for measurement of expression levels from several thousands of genes simultaneously, thus resulting in an enormous amount of data. The role of the statistician is critical to the successful design of gene expression studies, and the analysis and interpretation of the resulting voluminous data. This paper discusses hypotheses common to gene expression studies, and describes some of the statistical methods suitable for addressing these hypotheses. S-plus and SAS codes to perform the statistical methods are provided. Gene expression data from an unpublished oncologic study is used to illustrate these methods.

Cluster Analysis↗

Statistical detection of chromosomal homology using shared-gene density alone.

MOTIVATION: Over evolutionary time, various processes including point mutations and insertions, deletions and inversions of variable sized segments progressively degrade the homology of duplicated chromosomal regions making identification of the homologous regions correspondingly difficult. Existing algorithms that attempt to detect homology are based on shared-gene density and colinearity and possibly also strand information. RESULTS: Here, we develop a new algorithm for the statistical detection of chromosomal homology, CloseUp, which uses shared-gene density alone to fully exploit the observation that relaxing colinearity requirements in general is beneficial for homology detection and at the same time optimizes computation time. CloseUp has two components: the identification of candidate homologous regions followed by their statistical evaluation using Monte Carlo methods and data randomization. Using both artificial and real data, we compared CloseUp with two existing programs (ADHoRe and LineUp) for chromosomal homology detection and found that in general CloseUp compares favorably. AVAILABILITY: CloseUp and supplementary information are available at http://www.igb.uci.edu/servers/cgss.html CONTACT: pfbaldi@ics.uci.edu.

Algorithms↗

Non-inferiority trials: the 'at least as good as' criterion with dichotomous data.

The 'at least as good as' criterion, introduced by Laster and Johnson for a continuous response variate, is developed here for applications with dichotomous data. This approach is adaptive in nature, as the margin of non-inferiority is not taken as a fixed difference; it varies as a function of the positive control response. When the non-inferiority margin is referenced as a high fraction of the positive control response, the procedure is seen to be uniformly more efficient than the fixed margin approach, yielding smaller sample sizes when sizing non-inferiority trials under identically specified conditions. Extending this method to proportions is straightforward, but highlights special considerations in the design of non-inferiority trials versus superiority trials, including potential trade-offs in statistical efficiency and interpretability.

Data Interpretation, Statistical↗

Power vectors: an application of Fourier analysis to the description and statistical analysis of refractive error.

The description of sphero-cylinder lenses is approached from the viewpoint of Fourier analysis of the power profile. It is shown that the familiar sine-squared law leads naturally to a Fourier series representation with exactly three Fourier coefficients, representing the natural parameters of a thin lens. The constant term corresponds to the mean spherical equivalent (MSE) power, whereas the amplitude and phase of the harmonic correspond to the power and axis of a Jackson cross-cylinder (JCC) lens, respectively. Expressing the Fourier series in rectangular form leads to the representation of an arbitrary sphero-cylinder lens as the sum of a spherical lens and two cross-cylinders, one at axis 0 degree and the other at axis 45 degrees. The power of these three component lenses may be interpreted as (x,y,z) coordinates of a vector representation of the power profile. Advantages of this power vector representation of a sphero-cylinder lens for numerical and graphical analysis of optometric data are described for problems involving lens combinations, comparison of different lenses, and the statistical distribution of refractive errors.

Computer Graphics↗

Three perspectives on work-related injury surveillance systems.

This paper reviews surveillance approaches for occupational injuries and evaluates three emerging methodologies for the enhancement of work-related injury surveillance: (1) narrative data analysis, (2) data set linkage, and (3) comprehensive company-wide surveillance systems. All three methods are the result of new applications of computer hardware and software that have apparent strengths and limitations. A major strength is the improved description of work exposures and related injuries leading to better understanding of injury etiology. This understanding, however, is limited by the data quality and completeness entered on records at the time of the injury. We recommend (1) more widespread inclusion of narrative text in databases, analyses of which can be a valuable supplement to injury coded data; (2) the increased use of data set linkage studies to combine injury and work-history data; and (3) the development of comprehensive company-wide surveillance systems to expedite the use of epidemiologic data for occupational injury prevention activities. Further development of these methods and others is encouraged, especially in light of technological advancements in data capture, analysis and presentation. Only through such efforts can we best apply epidemiologic principles to preventing injuries in the workplace.

Accidents, Occupational↗

The Pearson product-moment correlation coefficient is better suited for identification of DNA fingerprint profiles than band matching algorithms.

A database of DNA fingerprint profiles from permanently established human and animal cell lines was prepared with a computer program originally designed for numerical taxonomy of bacteria. Identifications of cell line DNA profiles were performed, both by the Pearson product-moment correlation coefficient and by band matching. Under the conditions used the Pearson product-moment correlation coefficient was consistently more reliable.

Algorithms↗

Design of the Fracture Intervention Trial.

The Fracture Intervention Trial (FIT) is a randomized, double-masked, placebo-controlled trial designed to test the hypothesis that alendronate, an amino-bisphosphonate, will reduce the rate of fractures in women aged 55-80 years with low hip bone marrow density (< 0.68 gm/cm2 at the femoral neck). It is being conducted at 11 clinical centers around the United States with a coordinating center at UC San Francisco. The goal was to randomize 6000 women. When recruitment was completed (in May 1993), 6457 women had been randomized, amounting to 108% of goal. The women were assigned to one of two substudies. The first (Vertebral Deformity study) includes 2023 women who have at least one vertebral deformity, and will test the hypothesis that alendronate reduces the rate of new vertebral deformities during 3 years of follow-up. This substudy has a power of 0.90 to detect a 32% reduction in the incidence of new vertebral deformities, assuming a 6.5% annual incidence of new vertebral deformities in the placebo group. The second study (Clinical Fracture study) includes 4434 women without vertebral deformities at baseline and will test the hypothesis that alendronate reduces the rate of clinically recognized fractures of all types over an average of 4.25 years of follow-up. This substudy has a 0.90 power to detect a 25% reduction in the rate of all clinical fractures, assuming 4% annual incidence in the placebo group. To our knowledge, this is the largest prospective, randomized, controlled study undertaken to determine the effectiveness of a treatment in reducing the risk of fractures in postmenopausal women.

Aged↗

The DxDxDG motif for calcium binding: multiple structural contexts and implications for evolution.

Calcium ions regulate many cellular processes and have important structural roles in living organisms. Despite the great variety of calcium-binding proteins (CaBPs), many of them contain the same Ca(2+)-binding helix-loop-helix structure, referred to as the EF-hand. In the canonical EF-hand, the loop contains three calcium-binding aspartic acid residues, which form the DxDxDG sequence motif, and is flanked by two alpha-helices. Recently, other CaBPs containing the same motif, but lacking one or both helices, have been described. Here, structural motif searches were used to analyse the full diversity of structural context in the known set of DxDxDG-containing CaBPs, including those where the structural resemblance of a given DxDxDG motif to that of EF-hands had not been noted. The results obtained indicate that the EF-hand represents but one, among many, structural context for the DxDxDG-like Ca(2+)-binding loops. While the structural similarity of the binuclear calcium-binding sites in anthrax protective antigen and human thrombospondin suggests that they are homologous, evolutionary relationships for mononuclear sites are harder to discern. The possible scenarios for the evolution of DxDxDG motif-containing calcium-binding loops in a variety of non-homologous proteins suggested loop transplant as a mechanism perhaps responsible for much of the diversity in structural contexts of present day DxDxDG-type CaBPs. Additionally, while it can be shown that existence of a DxDxDG sequence is not enough to confer a conformation suitable for calcium binding, local convergent evolution may still have a role. The analysis presented here has consequences for the prediction of calcium binding from sequence alone.

Amino Acid Motifs↗

Analysis of genomic and proteomic data using advanced literature mining.

High-throughput technologies, such as proteomic screening and DNA micro-arrays, produce vast amounts of data requiring comprehensive analytical methods to decipher the biologically relevant results. One approach would be to manually search the biomedical literature; however, this would be an arduous task. We developed an automated literature-mining tool, termed MedGene, which comprehensively summarizes and estimates the relative strengths of all human gene-disease relationships in Medline. Using MedGene, we analyzed a novel micro-array expression dataset comparing breast cancer and normal breast tissue in the context of existing knowledge. We found no correlation between the strength of the literature association and the magnitude of the difference in expression level when considering changes as high as 5-fold; however, a significant correlation was observed (r = 0.41; p = 0.05) among genes showing an expression difference of 10-fold or more. Interestingly, this only held true for estrogen receptor (ER) positive tumors, not ER negative. MedGene identified a set of relatively understudied, yet highly expressed genes in ER negative tumors worthy of further examination.

Abstracting and Indexing↗

True and false gharials: a nuclear gene phylogeny of crocodylia.

The phylogeny of Crocodylia offers an unusual twist on the usual molecules versus morphology story. The true gharial (Gavialis gangeticus) and the false gharial (Tomistoma schlegelii), as their common names imply, have appeared in all cladistic morphological analyses as distantly related species, convergent upon a similar morphology. In contrast, all previous molecular studies have shown them to be sister taxa. We present the first phylogenetic study of Crocodylia using a nuclear gene. We cloned and sequenced the c-myc proto-oncogene from Alligator mississippiensis to facilitate primer design and then sequenced an 1,100-base pair fragment that includes both coding and noncoding regions and informative indels for one species in each extant crocodylian genus and six avian outgroups. Phylogenetic analyses using parsimony, maximum likelihood, and Bayesian inference all strongly agreed on the same tree, which is identical to the tree found in previous molecular analyses: Gavialis and Tomistoma are sister taxa and together are the sister group of Crocodylidae. Kishino-Hasegawa tests rejected the morphological tree in favor of the molecular tree. We excluded long-branch attraction and variation in base composition among taxa as explanations for this topology. To explore the causes of discrepancy between molecular and morphological estimates of crocodylian phylogeny, we examined puzzling features of the morphological data using a priori partitions of the data based on anatomical regions and investigated the effects of different coding schemes for two obvious morphological similarities of the two gharials.

Alligators and Crocodiles↗

TmaDB: a repository for tissue microarray data.

BACKGROUND: Tissue microarray (TMA) technology has been developed to facilitate large, genome-scale molecular pathology studies. This technique provides a high-throughput method for analyzing a large cohort of clinical specimens in a single experiment thereby permitting the parallel analysis of molecular alterations (at the DNA, RNA, or protein level) in thousands of tissue specimens. As a vast quantity of data can be generated in a single TMA experiment a systematic approach is required for the storage and analysis of such data. DESCRIPTION: To analyse TMA output a relational database (known as TmaDB) has been developed to collate all aspects of information relating to TMAs. These data include the TMA construction protocol, experimental protocol and results from the various immunocytological and histochemical staining experiments including the scanned images for each of the TMA cores. Furthermore the database contains pathological information associated with each of the specimens on the TMA slide, the location of the various TMAs and the individual specimen blocks (from which cores were taken) in the laboratory and their current status i.e. if they can be sectioned into further slides or if they are exhausted. TmaDB has been designed to incorporate and extend many of the published common data elements and the XML format for TMA experiments and is therefore compatible with the TMA data exchange specifications developed by the Association for Pathology Informatics community. Finally the design of the database is made flexible such that TMA experiments from several types of cancer can be stored in a single database, which incorporates the national minimum data set required for pathology reports supported by the Royal College of Pathologists (UK). CONCLUSION: TmaDB will provide a comprehensive repository for TMA data such that a large number of results from the numerous immunostaining experiments can be efficiently compared for each of the TMA cores. This will allow a systematic, large-scale comparison of tumour samples to facilitate the identification of gene products of clinical importance such as therapeutic or prognostic markers. In addition this work will contribute to the establishment of a standard for reporting TMA data analogous to MIAME in the description of microarray data.

Data Interpretation, Statistical↗