PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Predicting outcome in coronary disease. Statistical models versus expert clinicians.

To study the accuracy with which long-term prognosis can be predicted in patients with coronary artery disease, prognostic predictions from a data-based multivariable statistical model were compared with predictions from senior clinical cardiologists. Test samples of 100 patients each were selected from a large series of medically treated patients with significant coronary disease. Using detailed case summaries, five senior cardiologists each predicted one- and three-year survival and infarct-free survival probabilities for 100 patients. Fifty patients appeared in multiple samples for assessing interphysician variability. Cox regression models, developed using patients not in the test samples, predicted corresponding outcome probabilities for each test patient. Overall, model predictions correlated better with actual patient outcomes than did the doctors' predictions. For three-year survival, rank correlations were 0.61 (model) and 0.49 (doctors). For three-year infarct-free survival predictions, correlations with outcome were 0.48 (model) and 0.29 (doctors). Comparisons by individual doctor revealed Cox model three-year survival predictions were better than those of four of five doctors (model predictions added significant [p less than 0.05] prognostic information to the doctor's predictions, whereas the converse was not true). For infarct-free survival, the Cox model was superior to all five doctors. Where predictions were made by multiple doctors, the interphysician variability was substantial. In coronary artery disease, statistical models developed from carefully collected data can provide prognostic predictions that are more accurate than predictions of experienced clinicians made from detailed case summaries.

Coronary Disease↗

A statistical model to optimize indirect sandwich enzyme-linked immunosorbent assay parameters of antigen and antibody: a microcomputer program.

A new computer software program "AVCRV" was developed using a statistical model to analyze the data from the indirect sandwich enzyme-linked immunosorbent assay (ELISA). The software program calculates a sigmoid type of regression analysis and can be run on most microcomputers in the laboratory. The program permits those who are not familiar with computers to complete this type of analysis in a few seconds without a large mainframe computer or complicated software. This statistical model for a sigmoid type of regression analysis of ELISA data may improve the analysis of research data for various avian pathogens from several different experiments.

Animals↗

The frequency of ion-pair substructures in proteins is quantitatively related to electrostatic potential: a statistical model for nonbonded interactions.

A statistical analysis of ion pairs in protein crystal structures shows that their abundance with respect to uncharged controls is accurately predicted by a Boltzmann-like function of electrostatic potential. It appears that the mechanisms of protein folding and/or evolution combine to produce a "thermal" distribution of local nonbonded interactions, as has been suggested by statistical-mechanical theories. Using this relationship, we develop a maximum likelihood methodology for estimation of apparent energetic parameters from the data base of known structures, and we derive electrostatic potential functions that lead to optimal agreement of observed and predicted ion-pair frequencies. These are similar to potentials of mean force derived from electrostatic theory, but departure from Coulombic behavior is less than has been suggested.

Biological Evolution↗

Statistical models for discerning protein structures containing the DNA-binding helix-turn-helix motif.

A method for discerning protein structures containing the DNA-binding helix-turn-helix (HTH) motif has been developed. The method uses statistical models based on geometrical measurements of the motif. With a decision tree model, key structural features required for DNA binding were identified. These include a high average solvent-accessibility of residues within the recognition helix and a conserved hydrophobic interaction between the recognition helix and the second alpha helix preceding it. The Protein Data Bank was searched using a more accurate model of the motif created using the Adaboost algorithm to identify structures that have a high probability of containing the motif, including those that had not been reported previously.

Binding Sites↗

A statistical model for estimating donor postdonation platelet counts after plateletpheresis.

BACKGROUND: To avoid the need, in serial apheresis donors, either to delay plateletpheresis until a predonation platelet count is completed or to obtain a postdonation count after each procedure, a statistical model has been developed to predict the postdonation platelet count from the donor predonation platelet count, weight, and hematocrit. STUDY DESIGN AND METHODS: Predonation and postdonation platelet counts were measured in two groups of approximately 100 consecutive donors (Group A to test the model and Group B to validate it), and the postdonation counts were calculated with the model. Using stepwise multiple linear regression from donor data, estimated postdonation platelet counts were found to be comparable to the postdonation platelet counts actually measured. RESULTS: Estimated postdonation platelet counts x 10(9) per L (mean +/- SD) for each group, respectively, were Group A, 195 +/- 35, versus actual platelet counts of 195 +/- 39 (p = 0.43), and Group B, 183 +/- 36, versus actual platelet counts of 189 +/- 34 (p = 0.14). Sensitivity and specificity, respectively, were Group A, 57 and 99 percent and Group B, 62 and 99 percent. CONCLUSION: For most serial apheresis donors, application of this predictor model should preclude the need to obtain an extra postdonation platelet count.

Blood Donors↗

Gene-environment interaction and the mapping of complex traits: some statistical models and their implications.

The manifestation of many complex diseases or traits is very likely the result of an inextricable interplay of the biological and the environmental. Yet the role of environmental effect has traditionally been played down, for various reasons. In this paper, some simple statistical models that incorporate gene-environment interaction (GEI) have been proposed and their behavior and implications investigated. These implications concern the conditional independence assumption in likelihood calculation of pedigree data, the fine-tuning of the sib pair method for mapping quantitative traits, apportioning of disease or trait variation due to specific causes. In addition, they concern properties of gene mapping methods that do not take GEI into account, and they bring into question the utility of commonly used measures of genetic effects such as recurrence risk ratio for relative pairs, twin concordance rates, and heritability coefficients. In the presence of GEI, all these measures are functions not only of genetic effects and gene frequency, but also of environmental effects, the distribution of environmental factors in the population, and of GEI. Above all, these measures are all measures of familial aggregation, since they can be significant even in the absence of any genetic component of the disease. Thus their use as indicators of the genetic basis of complex diseases is cast into doubt.

Chromosome Mapping↗

Statistical model of amino acid code of protein secondary structure.

In the previous paper (Shestopalov, 2003) we presented the amino acid code of protein secondary structure as a partial solution of the fundamental problem of the protein three-dimensional structure calculation from the amino acid sequence. Here a statistical model of the code is described. The model is based on the structural data from 2258 protein chains (417,112 amino acid residues used). 60 and 61% of the secondary structure, calculated using the model, coincide, respectively, with the observed secondary structure in the training subset and test subset (104 protein chains and 21,166 residues used). This is equal to the threshold value for all the secondary structure calculations, based on the models, where, similarly as here, only the nearest and middle-range interactions are considered. Therefore the constructed model can be applied for the protein structure prediction from the amino acid sequence, especially when additional information is used along with expert analysis, as in the most successful prediction methods. The model can be used for analysis of the secondary structure changes during protein folding by comparison of the calculated and observed secondary structures. The information about the conformationally invariant segments can serve for the simulation of the supersecondary structure formation. One can try to obtain and examine the protein subset, in which the calculated and observed secondary structures are very similar.

Amino Acid Sequence↗

Angioarchitecture associated with haemorrhage in cerebral arteriovenous malformations: a prognostic statistical model.

The overall haemorrhagic risk of a cerebral arteriovenous malformation (cAVM) is 2-4% per year. However, the individual risk of haemorrhage has never been determined. This study was undertaken to assess the haemorrhage risk of an individual cAVM. Neuroangiographic findings of 160 cAVM were analysed retrospectively, looking at 30 angiographic features. A statistical model was established by logistic regression to evaluate the risk of an individual cAVM. We statistically correlated 15 parameters with the haemorrhage risk. The statistical model includes five independent parameters. Four are unfavourable: exclusively deep drainage, venous stenoses, venous reflux and the radio of afferent to efferent systems; one is favourable: venous recruitment. This model quantifies the individual risk of haemorrhage. When this model is applied to the population studied, the error rate is 5%. This model can contribute to therapeutic strategy, and to a better understanding of the natural history of cAVM.

Adolescent↗

Predictive accuracy study: comparing a statistical model to clinicians' estimates of outcomes after coronary bypass surgery.

BACKGROUND: The purpose of this study was to compare clinicians' prior probability estimates of operative mortality (OM) and prolonged intensive care unit stay (ICU) length of stay greater than 48 hours after coronary artery bypass graft surgery (CABG) with estimates derived from statistical models alone. METHODS: Nine clinicians estimated the predicted probability of OM and ICU stay greater than 48 hours from an abstract of information for each of 100 patients selected from the 1996 to 1997 database of 1,904 patients who underwent isolated CABG. Logistic regression models were used to calculate the predicted probability of OM and ICU stay greater than 48 hours for each patient. The study sample was split into two parts; clinicians were randomly given access to a predictive rule to guide their judgements for one part of the study. RESULTS: Clinicians' estimates were similar with or without access to the rule, and both parts of the study were therefore pooled. Clinicians significantly overestimated the probability of OM (model 6.3% +/- 1%, clinicians 7.6% +/- 3%, p = 0.0001) and ICU stay greater than 48 hours (model 25% +/- 2%, clinicians 28% +/- 1%, p = 0.0012). Clinicians' estimates of OM were not significantly higher than the model's for nonsurvivors (0.8% +/- 0.7%, p = 0.2), but were significantly higher for survivors (1.4% +/- 0.3%, p = 0.039). CONCLUSIONS: Clinicians trusted their own empiric estimates rather than a predictive rule and overestimated the probability of OM and ICU stay greater than 48 hours.

Aged↗

Statistical modelling for clinical mastitis in the dairy cow: problems and solutions.

Modelling case occurrence and risk factors for clinical mastitis, as a key multifactorial disease in the dairy cow, requires statistical models. The type of model used depends on the choice of perception or the study level: herd, lactation, animal, udder and quarter. The validity of the tests that are performed through these models is especially ensured when hypotheses of independence between statistical units are respected, and when the model adjustments do not involve overdispersion faced with the observed data. In the article, the main sources of overdispersion are identified according to the different levels of perception of mastitis risk. Then, the proposed solutions to control for overdispersion at each study level are discussed and the difficulty to compare the study results is highlighted through a variety of methodological choices of the authors. Two main categories of models are used for modelling clinical mastitis, i.e. generalist exploratory models and explanatory designed models. The contribution of the explanatory models to improve modelling accuracy and relevance is documented through the two main published methodological approaches, the first one being based on a states model, and the second on a survival model. The integration and optimisation of such explanatory modelling methods should be possible in the future in order to develop a more global explanatory model including herd risk factors, which could pertinently predict udder infections (both clinical and subclinical) at the cow, lactation, or even udder and quarter levels.

Animals↗

Validity of linear regression in method comparison studies: is it limited by the statistical model or the quality of the analytical input data?

We compared the application of ordinary linear regression, Deming regression, standardized principal component analysis, and Passing-Bablok regression to real-life method comparison studies to investigate whether the statistical model of regression or the analytical input data have more influence on the validity of the regression estimates. We took measurements of serum potassium as an example for comparisons that cover a narrow data range and measurements of serum estradiol-17beta as an example for comparisons that cover a wide data range. We demonstrate that, in practice, it is not the statistical model but the quality of the analytical input data that is crucial for interpretation of method comparison studies. We show the usefulness of ordinary linear regression, in particular, because it gives a better estimate of the standard deviation of the residuals than the other procedures. The latter is important for distinguishing whether the observed spread across the regression line is caused by the analytical imprecision alone or whether sample-related effects also contribute. We further demonstrate the usefulness of linear correlation analysis as a first screening test for the validity of linear regression data. When ordinary linear regression (in combination with correlation analysis) gives poor estimates, we recommend investigating the analytical reason for the poor performance instead of assuming that other linear regression procedures add substantial value to the interpretation of the study. This investigation should address whether (a) the x and y data are linearly related; (b) the total analytical imprecision (s(a,tot)) is responsible for the poor correlation; (c) sample-related effects are present (standard deviation of the residuals >> s(a,tot)); (d) the samples are adequately distributed over the investigated range; and (e) the number of samples used for the comparison is adequate.

Chromatography, Ion Exchange↗

PyEvolve: a toolkit for statistical modelling of molecular evolution.

BACKGROUND: Examining the distribution of variation has proven an extremely profitable technique in the effort to identify sequences of biological significance. Most approaches in the field, however, evaluate only the conserved portions of sequences - ignoring the biological significance of sequence differences. A suite of sophisticated likelihood based statistical models from the field of molecular evolution provides the basis for extracting the information from the full distribution of sequence variation. The number of different problems to which phylogeny-based maximum likelihood calculations can be applied is extensive. Available software packages that can perform likelihood calculations suffer from a lack of flexibility and scalability, or employ error-prone approaches to model parameterisation. RESULTS: Here we describe the implementation of PyEvolve, a toolkit for the application of existing, and development of new, statistical methods for molecular evolution. We present the object architecture and design schema of PyEvolve, which includes an adaptable multi-level parallelisation schema. The approach for defining new methods is illustrated by implementing a novel dinucleotide model of substitution that includes a parameter for mutation of methylated CpG's, which required 8 lines of standard Python code to define. Benchmarking was performed using either a dinucleotide or codon substitution model applied to an alignment of BRCA1 sequences from 20 mammals, or a 10 species subset. Up to five-fold parallel performance gains over serial were recorded. Compared to leading alternative software, PyEvolve exhibited significantly better real world performance for parameter rich models with a large data set, reducing the time required for optimisation from approximately 10 days to approximately 6 hours. CONCLUSION: PyEvolve provides flexible functionality that can be used either for statistical modelling of molecular evolution, or the development of new methods in the field. The toolkit can be used interactively or by writing and executing scripts. The toolkit uses efficient processes for specifying the parameterisation of statistical models, and implements numerous optimisations that make highly parameter rich likelihood functions solvable within hours on multi-cpu hardware. PyEvolve can be readily adapted in response to changing computational demands and hardware configurations to maximise performance. PyEvolve is released under the GPL and can be downloaded from http://cbis.anu.edu.au/software.

Animals↗

[Variability in scabies mites Sarcoptes scabiei De Geer (Acariformes, Sarcoptidae) in relation to scabies epidemiology. 1. A statistical model of female variability].

Individual variability of 235 grain mite females from Moscow and the Moscow Province was studied. The data were processed on a computer. Body proportions were used as a representative dimension index. The size of proterosomal scutellum was stable. A map of the chaetoid cover of notum is given, and its variability was studied. Statistical analysis of all signs was made. Frontal chaetoid and sejugal sections and the number of caudal chaetoids are stable. Naked, middle and abdominal chaetoid sections are the most variable. Linear sizes of naked section and the number of chaetoid sets on middle and abdominal sections correlate with the number of undeveloped chaetoids. The statistical model of variability may be useful for populational analysis of Sarcoptes forms in connection with scab epidemiology.

Animals↗

Statistical modelling of the determinants of historical exposure to bitumen and polycyclic aromatic hydrocarbons among paving workers.

INTRODUCTION: An industrial hygiene database has been constructed for the exposure assessment in a study of cancer risk among asphalt workers. AIM: To create models of bitumen and polycyclic aromatic hydrocarbons (PAH) exposure intensity among paving workers. METHODS: Individual exposure measurements from pavers (N = 1581) were collected from 8 countries. Correlation patterns between exposure measures were examined and factors affecting exposure were identified using statistical modelling. RESULTS: Inhalable dust appeared to be a good proxy of bitumen fume exposure. Bitumen fume and vapour levels were not correlated. Benzo(a)pyrene level appeared to be a good indicator of PAH exposure. All exposures steadily declined over the last 20 years. Mastic laying, re-paving, surface dressing, oil gravel paving and asphalt temperature were significant determinants of bitumen exposure. Coal tar use dictated PAH exposure levels. DISCUSSION: Bitumen fume, vapour and PAH have different determinants of exposure. For paving workers, exposure intensity can be assessed on the basis of time period and production characteristics.

Databases, Factual↗

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans↗

Statistical model for prediction of retrospective exposure to ethylene oxide in an occupational mortality study.

Since direct measures of individual exposure seldom exist for the entire period of an occupational mortality study, retrospective exposure estimates are necessary. This is often done in a subjective manner involving a consensus of opinion from a panel of epidemiologists and industrial hygienists. An alternative method utilizing a statistical model provides a more objective procedure for retrospective exposure assessment. The development of a weighted multiple regression model is presented for estimation of exposure levels to ethylene oxide (ETO) for inclusion in a cohort mortality study of workers in the sterilization industry. Three steps in development of the model are described: (1) data acquisition and assessment, (2) model building, and (3) evaluation of the model. The final model explained a remarkable 85% of the variability in 205 average measurements of ETO levels. Exposure factors included in the model were exposure category, product type, size of the sterilization unit, selected engineering controls, days after sterilization, and calendar year. The model was evaluated in two ways: against a set of measurement data not used to develop the model and a panel of 11 industrial hygienists representing the sterilization industry. The model predicted ETO exposures within 1.1 ppm of the validation data set with a standard deviation of 3.7 ppm. The arithmetic and geometric means of the 46 measurements in the validation data set were 4.6 and 2.2 ppm, respectively. The model also outperformed the panel of industrial hygienists relative to the validation data in terms of both bias and precision.

Ethylene Oxide↗

A hierarchical statistical model for estimating population properties of quantitative genes.

BACKGROUND: Earlier methods for detecting major genes responsible for a quantitative trait rely critically upon a well-structured pedigree in which the segregation pattern of genes exactly follow Mendelian inheritance laws. However, for many outcrossing species, such pedigrees are not available and genes also display population properties. RESULTS: In this paper, a hierarchical statistical model is proposed to monitor the existence of a major gene based on its segregation and transmission across two successive generations. The model is implemented with an EM algorithm to provide maximum likelihood estimates for genetic parameters of the major locus. This new method is successfully applied to identify an additive gene having a large effect on stem height growth of aspen trees. The estimates of population genetic parameters for this major gene can be generalized to the original breeding population from which the parents were sampled. A simulation study is presented to evaluate finite sample properties of the model. CONCLUSIONS: A hierarchical model was derived for detecting major genes affecting a quantitative trait based on progeny tests of outcrossing species. The new model takes into account the population genetic properties of genes and is expected to enhance the accuracy, precision and power of gene detection.

Algorithms↗

A statistical model that takes into account patient heterogeneity in decision making.

Statistical evaluation of clinical treatments or preventive medicine has profoundly contributed to decision making in medical fields such as with the acceptance of new treatment methods and health promotion policies. It is crucial in such decision making to find a correct statistical model to treat a surprisingly large variety of patients, or a heterogeneous group of patients, even with the same diagnosis. In diseases such as cancer, cardiovascular disease or diabetes, patients are often followed up to certain endpoints and these data are frequently analyzed by logrank tests or Cox-models to evaluate treatment effects. Although these methods have been widely accepted and extensively studied, we are sometimes faced with problems in applying these methods when the heterogeneity of patients is large and a lot of prognostic factors affecting the endpoints have to be considered. Based on the results of the analyses of survival data from more than 6,000 gastric cancer patients, it is revealed that the stratified logrank test may suffer serious power loss, even though primary prognostic factors are used as stratified factors. A so-called 'piecewise linear Cox regression method' for properly treating the heterogeneity of patients is introduced and extensively studied. This method is shown to be appropriate for patient groups with a high degree of heterogeneity such as the gastric cancer patients. The same method is, in principle, applicable to patients of other diseases, too, using statistical software such as SAS, BMDP and etc.

Algorithms↗