PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Molecular sequence accuracy: analysing imperfect data.

Molecular sequences are experimentally derived data that can be expected to contain errors as a result of diverse phenomena such as biological variation, molecular cloning artifacts, imperfect sequence determination, and data handling during contig assembly. Errors will affect the reliability of database searches and sequence alignments, but their impact may be minimized by the use of analytical techniques that anticipate that the data will be imperfect.

Amino Acid Sequence

Analysis of beat-to-beat cardiovascular hemodynamic variables obtained from long-term biotelemetry.

Manual methods of large volume data storage, retrieval, and analysis are difficult, time consuming, and present numerous opportunities for calculation errors. We have designed and implemented a comprehensive computer-based system for performing these functions. Development of this system was necessary since left ventricular (LV) blood pressure and two regional LV wall thickness measurements were obtained during long-term extracorporeal biotelemetry of miniswine for 24-h periods. During a single recording period over 100,000 individual cardiac cycles were recorded on analog tape and later analysed for determination of global myocardial oxygen demand and regional myocardial function. In addition, custom designed software was developed to determine the extent and duration of myocardial dysfunction. Batch file commands enabled the customized software to operate without prompting by the user thus optimizing the time usage of the computer, and the computer based data acquisition and analysis system. Although this system was designed specifically for analysing cardiovascular hemodynamic variables, it is flexible and can be applied to other experimental applications.

Analog-Digital Conversion

Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons.

DNA and protein sequence comparisons are performed by a number of computational algorithms. Most of these algorithms search for the alignment of two sequences that optimizes some alignment score. It is an important problem to assess the statistical significance of a given score. In this paper we use newly developed methods for Poisson approximation to derive estimates of the statistical significance of k-word matches on a diagonal of a sequence comparison. We require at least q of the k letters of the words to match where 0 less than q less than or equal to k. The distribution of the number of matches on a diagonal is approximated as well as the distribution of the order statistics of the sizes of clumps of matches on the diagonal. These methods provide an easily computed approximation of the distribution of the longest exact matching word between sequences. The methods are validated using comparisons of vertebrate and E. coli protein sequences. In addition, we compare two HLA class II transplantation antigens by this method and contrast the results with a dynamic programming approach. Several open problems are outlined in the last section.

Algorithms

Enrichment of oligonucleotide sets with transcription control signals. II: Mammalian DNA.

We studied the frequency distribution of oligonucleotides 10 bp long in a sample of 1.6 Mb of mammalian genes, containing 579 sequences from GenBank(R) 55.0, with the aim of detecting transcription control signals. 2216 decamers had a frequency higher than 10 times the mean and were subjected to further statistical analysis. For each of the 2216 decamers (parents), we counted the individual frequencies of the 30 decamers differing from the parent by one base mutation (progeny) and then calculated two variance/mean chi squares for the progeny, with and without the parent. We then studied the distribution of the ratio between the two chi squares. Out of 2216 decamers, 346 had a chi square ratio of 1.9 or larger. In this final set, which corresponds to less than 0.033 per cent of all possible decamers, 18 were found to contain 23 eukaryotic transcription control elements 5-10 bp of length, such as Sp1 and others. Furthermore, when compared to 210 random sets containing 346 decamers, this set contains a highly significant excess of the longer signals.

Algorithms

Statistical issues arising from mouse micronucleus assay experiments of Ashby and co-workers.

Data from a series of mouse micronucleus assays have been reanalysed to illustrate various statistical issues raised by Ashby and co-workers during the development of the assay. Most of the statistical points discussed in these earlier papers can be explained by the stochastic nature of the data. Reanalysis shows that the type of data collected in mammalian micronucleus assays is amenable to analysis by standard biometric methods. It is concluded that statistical analysis has an important role in the exploration and interpretation of data from the micronucleus assay.

Analysis of Variance

Family-Wise Error Rate Control in Clinical Trials With Overlapping Populations.

We consider clinical trials with multiple, overlapping patient populations that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect several populations. For type I error control, often the family-wise error rate (FWER) is controlled, which is the probability to reject at least one true null hypothesis. If the joint distribution of the test statistics is known, the FWER level can be exhausted by determining critical values or adjusted-levels. The adjustment is typically done under the common ANOVA assumptions. However, the performed tests are then only valid under the rather strong assumption of homogeneous null effects, that is, when the null hypothesis applies to all subpopulations and their intersections. We show that under cancelling null effects, when heterogeneous effects cancel out in some or all subpopulations, this procedure does not provide FWER control. We also suggest different alternatives and compare them in terms of FWER control and their power.

Humans

Sequence of a Euplotes crassus macronuclear DNA molecule encoding a protein with homology to a rat form-I phosphoinositide-specific phospholipase C.

A 604-base pair macronuclear DNA molecule from the hypotrichous ciliate Euplotes crassus was cloned and its DNA sequence determined. The DNA sequence contains an open reading frame capable of encoding a protein 141 amino acids in length. The putative protein contains significant sequence similarity to other eukaryotic proteins, including the rat form-I phosphoinositide-specific phospholipase-C.

Amino Acid Sequence

[Data analysis by statistical models].

The basic idea for the realization of effective statistical data analysis is illustrated with an example. The use of statistical models is explained and the feasibility of objective comparison of the models by an information criterion AIC is demonstrated. Further, the possibility of practical use of Bayesian models for complex data analysis is explained. Finally, the necessity of cooperation between the experts of respective fields and statisticians for further development of statistical data analysis is mentioned.

Adult

[Additional coding for legally required ICD-9 as the basis for high quality patient record searches].

We are convinced, that the updating of the ID DIACOS diagnosis catalogue fits better to the today's medical language, than the former one. By using a consequent additional code we could eliminate several lacks of the ICD-9. Valid and reproducible diagnosis statistics are now possible. Subjective coding mistakes are largely avoidable. Inquiries of certain diagnosis can be answered by using the one of the two codes, that describes the diagnosis in the best form. Regarding the allotments of the best form. Regarding the allotments of the catalogue of orthopedic and traumatic diagnosis including the possibility of double-coding we got a special better version of ID DIACOS than before.

Data Collection

Which measures of skewness and kurtosis are best?

Indices of distributional shape based on linear combinations of order statistics have recently been described by Hosking. Their usefulness as tools for practical data analysis is examined. They are found to have several advantages over the conventional indices of skewness and kurtosis (square root of b1 and b2) and no serious drawbacks. It is proposed, therefore, that they should replace square root of b1 and b2 in routine data analysis. To implement this suggestion, action by the developers of standard statistical software is needed.

Bias