PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Comparison of the capability of peak functions in describing real chromatographic peaks.

This paper describes the results of a comparison of four peak functions in describing real chromatographic peaks. They are the empirically transformed Gaussian, polynomial modified Gaussian, generalized exponentially modified Gaussian and hybrid function of Gaussian and truncated exponential functions. Real chromatographic peaks of different shapes (fronting. symmetric, and tailing) are obtained by various separation conditions of reversed-phase liquid chromatography. They are then fitted to the peak functions via the Marquardt-Levenberg algorithm, a nonlinear least-squares curve-fitting procedure, by Microsoft Solver. The qualities of the fits are evaluated by the sum of the squares of the residuals. It is concluded in the study that the empirically transformed Gaussian function offers the highest flexibility (best fits) to all shapes of chromatographic peaks, including extremely asymmetric tailing peaks with a peak asymmetry of up to 8. The flexibility of this function should improve our ability to process chromatographic peaks such as deconvolution of overlapped peaks and smoothing noisy peaks for the determination of statistical moments.

Chromatography, Liquid↗

Emerging tandem-mass-spectrometry techniques for the rapid identification of proteins.

State-of-the-art techniques such as liquid-chromatography-electrospray-ionisation tandem mass spectrometry have, in conjunction with database-searching computer algorithms, revolutionised the analysis of biochemical species from complex biological mixtures. With these techniques, it is now possible to perform high-throughput protein identification at picomolar to subpicomolar levels from protein mixtures. This article provides an overview of the techniques and methodologies available for the structural elucidation and identification of proteins and peptides from complex biological samples.

Algorithms↗

Detection of linear and nonlinear dependencies in time series using the method of surrogate data in S-PLUS.

A general implementation of the method of surrogate data in the S programming language for use with the S-PLUS statistical package is presented. We illustrate the application of the S functions to testing hypotheses about a human heart rate time series and demonstrate that there is evidence for both linear and nonlinear dependencies. We expect these S functions will be useful for the application of the method of surrogate data to the analysis of biomedical time series using the S-PLUS statistical software package.

Data Interpretation, Statistical↗

Statistical analysis of pair-wise compatibility of spatially nearest neighbor and adjacent residues in alpha-helix and beta-strands: application to a minimal model for secondary structure prediction.

Secondary structural elements like alpha-helix and beta-strands possess distinctly different structural features and thus the relative positioning of the nearest neighbor residues, and also the sequence-wise adjacent residues is important in determining the structural preference. In the present work we have statistically examined the pair-wise compatibility pattern of physically nearest neighbors and separately the adjacent residue pairs along the sequence in between the nearest neighbor partners in alpha-helices and beta-strands. It has been demonstrated that the patterns and hence, the physical basis of the compatibility of adjacent residue pairs and the spatially nearest neighbors are significantly different in most cases. The influence of tertiary contacts on the pair-wise compatibility is shown to be significant for beta-strands while it is small for alpha-helices. Based on the compatibility of physically nearest neighbors and the sequence-wise adjacent residue pairs, a minimal model has been constructed to predict the alpha-helices, beta-strands and coils of a protein from its sequence. Application of this method to 100 sequences shows that it has a predictive capability comparable to that of other more sophisticated statistical methods.

Base Pairing↗

The significance of statistical significance.

Currently, much nursing practice is based on limited evidence, for example, small-scale research, case studies and clinical experience. In a mature science this would be undesirable, but nursing is in the early stages of development as a science, and many of its practices depend on relatively informal knowledge. To encourage the spread of potentially valuable ideas, nurses must be willing to share their clinical experience and journal editors should consider publishing this information. High-quality research is essential to the long-term development of 'evidence-based practice', but it is crucial at the present stage of nursing science that we do not become too concerned with perfect research methodology at the expense of good ideas. This particularly applies to tests of statistical significance. If we accept only information that has demonstrated statistical significance, we risk the dismissal of qualitative research and other information which may be extremely valuable but which have not yet been fully investigated. The aim of this paper is to convince practitioners and journal editors that statistical significance is not the only way to judge clinical importance and to suggest that decisions on what should be submitted and accepted for publication should be based on potential clinical relevance as well as statistical analysis.

Bias↗

Nonexperimental data systems in surgery.

This article reviews nonexperimental data bases, emphasizing the present uses and future opportunities of routinely collected information. Data bases are discussed in terms of appropriate research designs. Possibilities for expanding available information through new data collection and through record linkage are stressed. The relationship of nonexperimental data systems to randomized trials and to clinical decision-making is examined.

Cohort Studies↗

Application of pattern recognition techniques to mass spectrometric data for sequencing C-terminal peptide residue series.

The application of pattern recognition to sequence elucidation from fast atom bombardment (FAB), collisionally activated dissociation (CAD), and tandem mass spectrometric data of peptides was investigated. Learning machine techniques for pattern recognition were applied to detect C-terminal series amino acid sequences up to the pentapeptide Try-Gly-Gly-Phe-Leu (YGGFL). The approach conditions the data by building upon known fragmentation pathways of peptides in FAB/CAD-related analysis. The intensities of critical sequence ion peaks are used to describe each pattern in the training set. A well-defined training set is then made use of to classify unknown species. The FORTRAN-77 program is adapted from one first developed by P.C. Jurs. It requires only the input of the critical peak intensities and the length of the unit to be tested. The method has potential in applications to larger peptides provided a database for the training set can be constructed.

Amino Acid Sequence↗

Statistical approaches to experimental design and data analysis of in vivo studies.

The objective of any experiment is to obtain an unbiased and precise estimate of a treatment effect in an efficient manner. Statistical aspects of the design, conduct, and analysis of the experiment play a major role in determining whether this goal is met. We highlight some of the more important statistical issues that pertain to in vivo studies. Particular emphasis is placed on the role of randomization, the number of animals, the utilization of repeated measures data, adjustments for missing data, and dealing with multiple causes of death or treatment failure. The discussion is not intended to be a comprehensive guide to all the statistical issues that can occur in animal experiments. Rather, the objective is to acquaint researchers with components of the experiment that will require careful statistical thought.

Animals↗

The anisotropic Hooke's law for cancellous bone and wood.

A method of data analysis for a set of elastic constant measurements is applied to data bases for wood and cancellous bone. For these materials the identification of the type of elastic symmetry is complicated by the variable composition of the material. The data analysis method permits the identification of the type of elastic symmetry to be accomplished independent of the examination of the variable composition. This method of analysis may be applied to any set of elastic constant measurements, but is illustrated here by application to hardwoods and softwoods, and to an extraordinary data base of cancellous bone elastic constants. The solid volume fraction or bulk density is the compositional variable for the elastic constants of these natural materials. The final results are the solid volume fraction dependent orthotropic Hooke's law for cancellous bone and a bulk density dependent one for hardwoods and softwoods.

Anisotropy↗

Statistical power and effect sizes of clinical neuropsychology research.

Cohen, in a now classic paper on statistical power, reviewed articles in the 1960 issue of one psychology journal and determined that the majority of studies had less than a 50-50 chance of detecting an effect that truly exists in the population, and thus of obtaining statistically significant results. Such low statistical power, Cohen concluded, was largely due to inadequate sample sizes. Subsequent reviews of research published in other experimental psychology journals found similar results. We provide a statistical power analysis of clinical neuropsychological research by reviewing a representative sample of 66 articles from the Journal of Clinical and Experimental Neuropsychology, the Journal of the International Neuropsychology Society, and Neuropsychology. The results show inadequate power, similar to that for experimental research, when Cohen's criterion for effect size is used. However, the results are encouraging in also showing that the field of clinical neuropsychology deals with larger effect sizes than are usually observed in experimental psychology and that the reviewed clinical neuropsychology research does have adequate power to detect these larger effect sizes. This review also reveals a prevailing failure to heed Cohen's recommendations that researchers should routinely report a priori power analyses, effect sizes and confidence intervals, and conduct fewer statistical tests.

Data Interpretation, Statistical↗

Making bootstrap statistical inferences: a tutorial.

Bootstrapping is a computer-intensive statistical technique in which extensive computational procedures are heavily dependent on modern high-speed digital computers. The payoff for such intensive computations is freedom from two major limiting factors that have dominated classical statistical theory since its beginning: the assumption that the data conform to a bell-shaped curve, and the need to focus on statistical measures whose theoretical properties can be analysed mathematically. The name "bootstrap" was derived from an old saying about pulling oneself up by one's bootstraps. In this case, bootstrapping means redrawing samples randomly from the original sample with replacement. The key idea, computations, advantages, limitations, and application potential of bootstrapping in the field of physical education and exercise science are introduced and illustrated using a set of national physical fitness testing data. Finally, an example of a bootstrapping application is provided. Through a step-by-step approach, the development and implementation of the bootstrap statistical inference are illustrated.

Data Interpretation, Statistical↗

Sequencing by hybridization with the generic 6-mer oligonucleotide microarray: an advanced scheme for data processing.

DNA sequencing by hybridization was carried out with a microarray of all 4(6) = 4,096 hexadeoxyribonucleotides (the generic microchip). The oligonucleotides immobilized in 100 x 100 x 20-microm polyacrylamide gel pads of the generic microchip were hybridized with fluorescently labeled ssDNA, providing perfect and mismatched duplexes. Melting curves were measured in parallel for all microchip duplexes with a fluorescence microscope equipped with CCD camera. This allowed us to discriminate the perfect duplexes formed by the oligonucleotides, which are complementary to the target DNA. The DNA sequence was reconstructed by overlapping the complementary oligonucleotide probes. We developed a data processing scheme to heighten the discrimination of perfect duplexes from mismatched ones. The procedure was united with a reconstruction of the DNA sequence. The scheme includes the proper definition of a discriminant signal, preprocessing, and the variational principle for the sequence indicator function. The effectiveness of the procedure was confirmed by sequencing, proofreading, and nucleotide polymorphism (mutation) analysis of 13 DNA fragments from 31 to 70 nucleotides long.

Algorithms↗

Ribosomal RNA as molecular barcodes: a simple correlation analysis without sequence alignment.

MOTIVATION: We explored the feasibility of using unaligned rRNA gene sequences as DNA barcodes, based on correlation analysis of composition vectors (CVs) derived from nucleotide strings. We tested this method with seven rRNA (including 12, 16, 18, 26 and 28S) datasets from a wide variety of organisms (from archaea to tetrapods) at taxonomic levels ranging from class to species. RESULT: Our results indicate that grouping of taxa based on CV analysis is always in good agreement with the phylogenetic trees generated by traditional approaches, although in some cases the relationships among the higher systemic groups may differ. The effectiveness of our analysis might be related to the length and divergence among sequences in a dataset. Nevertheless, the correct grouping of sequences and accurate assignment of unknown taxa make our analysis a reliable and convenient approach in analyzing unaligned sequence datasets of various rRNAs for barcoding purposes. AVAILABILITY: The newly designed software (CVTree 1.0) is publicly available at the Composition Vector Tree (CVTree) web server http://cvtree.cbi.pku.edu.cn.

Algorithms↗

The platypus is in its place: nuclear genes and indels confirm the sister group relation of monotremes and Therians.

Morphological data supports monotremes as the sister group of Theria (extant marsupials + eutherians), but phylogenetic analyses of 12 mitochondrial protein-coding genes have strongly supported the grouping of monotremes with marsupials: the Marsupionta hypothesis. Various nuclear genes tend to support Theria, but a comprehensive study of long concatenated sequences and broad taxon sampling is lacking. We therefore determined sequences from six nuclear genes and obtained additional sequences from the databases to create two large and independent nuclear data sets. One (data set I) emphasized taxon sampling and comprised five genes, with a concatenated length of 2,793 bp, from 21 species (two monotremes, six marsupials, nine placentals, and four outgroups). The other (data set II) emphasized gene sampling and comprised eight genes and three proteins, with a concatenated length of 10,773 bp or 3,669 amino acids, from five taxa (a monotreme, a marsupial, a rodent, human, and chicken). Both data sets were analyzed by parsimony, minimum evolution, maximum likelihood, and Bayesian methods using various models and data partitions. Data set I gave bootstrap support values for Theria between 55% and 100%, while support for Marsupionta was at most 12.3%. Taking base compositional bias into account generally increased the support for Theria. Data set II exclusively supported Theria, with the highest possible values and significantly rejected Marsupionta. Independent phylogenetic evidence in support of Theria was obtained from two single amino acid deletions and one insertion, while no supporting insertions and deletions were found for Marsupionta. On the basis of our data sets, the time of divergence between Monotremata and Theria was estimated at 231-217 MYA and between Marsupialia and Eutheria at 193-186 MYA. The morphological evidence for a basal position of Monotremata, well separated from Theria, is thus fully supported by the available molecular data from nuclear genes.

Amino Acid Sequence↗

DNA sequence elements located immediately upstream of the -10 hexamer in Escherichia coli promoters: a systematic study.

We have made a systematic study of how the activity of an Escherichia coli promoter is affected by the base sequence immediately upstream of the -10 hexamer. Starting with an activator-independent promoter, with a 17 bp spacing between the -10 and -35 hexamer elements, we constructed derivatives with all possible combinations of bases at positions -15 and -14. Promoter activity is greatest when the 'non-template' strand carries T and G at positions -15 and -14, respectively. Promoter activity can be further enhanced by a second T and G at positions -17 and -16, respectively, immediately upstream of the first 'TG motif'. Our results show that the base sequence of the DNA segment upstream of the -10 hexamer can make a significant contribution to promoter strength. Using published collections of characterised E.coli promoters, we have studied the frequency of occurrence of 'TG motifs' upstream of the promoters' -10 elements. We conclude that correctly placed 'TG motifs' are found at over 20% of E.coli promoters.

Base Sequence↗

Quantitative evaluation of multiplicity in epidemiology and public health research.

Epidemiologic and public health researchers frequently include several dependent variables, repeated assessments, or subgroup analyses in their investigations. These factors result in multiple tests of statistical significance and may produce type 1 experimental errors. This study examined the type 1 error rate in a sample of public health and epidemiologic research. A total of 173 articles chosen at random from 1996 issues of the American Journal of Public Health and the American Journal of Epidemiology were examined to determine the incidence of type 1 errors. Three different methods of computing type 1 error rates were used: experiment-wise error rate, error rate per experiment, and percent error rate. The results indicate a type 1 error rate substantially higher than the traditionally assumed level of 5% (p < 0.05). No practical or statistically significant difference was found between type 1 error rates across the two journals. Methods to determine and correct type 1 errors should be reported in epidemiologic and public health research investigations that include multiple statistical tests.

Bias↗

Reduction of protein sequence complexity by residue grouping.

It is well known that there are some similarities among various naturally occurring amino acids. Thus, the complexity in protein systems could be reduced by sorting these amino acids with similarities into groups and then protein sequences can be simplified by reduced alphabets. This paper discusses how to group similar amino acids and whether there is a minimal amino acid alphabet by which proteins can be folded. Various reduced alphabets are obtained by reserving the maximal information for the simplified protein sequence compared with the parent sequence using global sequence alignment. With these reduced alphabets and simplified similarity matrices, we achieve recognition of the protein fold based on the similarity score of the sequence alignment. The coverage in dataset SCOP40 for various levels of reduction on the amino acid types is obtained, which is the number of homologous pairs detected by program BLAST to the number marked by SCOP40. For the reduced alphabets containing 10 types of amino acids, the ability to detect distantly related folds remains almost at the same level as that by the alphabet of 20 types of amino acids, which implies that 10 types of amino acids may be the degree of freedom for characterizing the complexity in proteins.

Amino Acid Sequence↗

Structure refinement against synchrotron Laue data: strategies for data collection and reduction.

The synchrotron Laue technique has been applied to high-resolution structure refinement of the ribotoxin, restrictocin [Yang & Moffat (1996). Structure, 4, 837-852]. By employing carefully designed data-collection strategies and the data-reduction algorithms incorporated in the software system LaueView [Ren & Moffat (1995a). J. Appl. Cryst. 28, 461-481; Ren & Moffat (1995b). J. Appl. Cryst. 28, 482-493], a set of high-resolution Laue data with a completeness and accuracy comparable to excellent monochromatic data was obtained. Through detailed comparison with the monochromatic data and electron-density maps derived from the Laue data, optimum data-collection and reduction strategies were identified and the application of Laue diffraction techniques to conventional crystallographic refinement was demonstrated.

Algorithms↗