PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Statistical models for prevalent cohort data.

In prospective cohort studies individuals are sometimes recruited according to a certain cross-sectional sampling criterion. A prevalent cohort is defined as a group of individuals who have a certain disease at enrollment into the study. Statistical models for the analysis of prevalent cohort data are considered when the onset or diagnosis time of the disease is known. The incident proportional hazards model, where the time scale is duration with disease, is compared to the prevalent proportional hazards model, where the fundamental time scale is follow-up time. In certain cases the time of enrollment may coincide with another event (such as the initiation of treatment). This situation is also considered and its limitations highlighted. To illustrate the methodological ideas discussed in the paper, the analysis of data from an observational study of zidovudine (ZVD) in patients with the acquired immunodeficiency syndrome (AIDS) is presented.

Acquired Immunodeficiency Syndrome↗

Different statistical models used in the calculation of the prevalence of insulin-dependent diabetes mellitus according to the polymorphism of the HLA-DQ region.

The use of three statistical models yielded different estimates of the odds ratio relative to the association between the polymorphism in the HLA-DQ region and insulin-dependent diabetes mellitus (IDDM). The models used were: (1) the allele-dosage model which assumes that the number of susceptibility alleles has a linear effect on the logarithm of the odds; (2) the reference cell coding method used with alleles of susceptibility as a risk factor; or (3) a model that uses a classification of alpha/beta heterodimers as a susceptibility factor. We suggest that models which imply a log-linear relationship between a susceptibility marker and disease such as the first model are not appropriate in the assessment of the HLA-IDDM association. In contrast, although both latter models are valid, the third model is more compatible with current hypotheses of the pathological process of the disease. Once an estimation of the odds ratio is chosen, we use such an estimation to calculate an approximation of the prevalence of IDDM according to the polymorphism in HLA-DQ region using the iterative procedure of Newton-Raphson. These approaches are illustrated with data from a case-control study previously conducted in the city of Santiago, Chile.

Alleles↗

A statistical model examining repetitive criminal behavior in acts of violence.

A simple statistical model for examining repetitive criminal behavior in acts of violence is described. The units, called "parameters," are nonquestionable data concerning environment of the crime, personal properties, and postmortem findings of the victim, obtained by double-blind investigation performed by two forensic pathologists. Parameters shared by two or more criminal acts allegedly committed by the same assailant were compared with the same parameters recorded from 50 or 100 other mutually independent criminal acts committed by other known assailants. This allowed an evaluation of the probability (p) of a crime pattern expressed as a parameter score to recur in mutually independent cases. The distribution of the score, when plotted on a logarithmic scale in all examples, showed an approximately normal distribution. The relation between probability (p), the estimated mean (means), and standard deviation (SD) yielded a normal curve. Different patterns of action by different perpetrators and patterns indicating repetitive behavior could be obtained. The method is applicable during investigation of crimes in which the perpetrator acts in a repetitive manner, as in serial murders.

Crime↗

Multivariate statistical model for 3D image segmentation with application to medical images.

In this article we describe a statistical model that was developed to segment brain magnetic resonance images. The statistical segmentation algorithm was applied after a pre-processing stage involving the use of a 3D anisotropic filter along with histogram equalization techniques. The segmentation algorithm makes use of prior knowledge and a probability-based multivariate model designed to semi-automate the process of segmentation. The algorithm was applied to images obtained from the Center for Morphometric Analysis at Massachusetts General Hospital as part of the Internet Brain Segmentation Repository (IBSR). The developed algorithm showed improved accuracy over the k-means, adaptive Maximum Apriori Probability (MAP), biased MAP, and other algorithms. Experimental results showing the segmentation and the results of comparisons with other algorithms are provided. Results are based on an overlap criterion against expertly segmented images from the IBSR. The algorithm produced average results of approximately 80% overlap with the expertly segmented images (compared with 85% for manual segmentation and 55% for other algorithms).

Algorithms↗

A statistical model of infant mortality.

"We have developed here a statistical model for describing infant deaths. Even though the model is tested with Canadian data, it will be a good approximation of the relationship between infant deaths and age in any population. The model needs improvement when high risk populations are studied."

Age Factors↗

Predicting outcome in coronary disease. Statistical models versus expert clinicians.

To study the accuracy with which long-term prognosis can be predicted in patients with coronary artery disease, prognostic predictions from a data-based multivariable statistical model were compared with predictions from senior clinical cardiologists. Test samples of 100 patients each were selected from a large series of medically treated patients with significant coronary disease. Using detailed case summaries, five senior cardiologists each predicted one- and three-year survival and infarct-free survival probabilities for 100 patients. Fifty patients appeared in multiple samples for assessing interphysician variability. Cox regression models, developed using patients not in the test samples, predicted corresponding outcome probabilities for each test patient. Overall, model predictions correlated better with actual patient outcomes than did the doctors' predictions. For three-year survival, rank correlations were 0.61 (model) and 0.49 (doctors). For three-year infarct-free survival predictions, correlations with outcome were 0.48 (model) and 0.29 (doctors). Comparisons by individual doctor revealed Cox model three-year survival predictions were better than those of four of five doctors (model predictions added significant [p less than 0.05] prognostic information to the doctor's predictions, whereas the converse was not true). For infarct-free survival, the Cox model was superior to all five doctors. Where predictions were made by multiple doctors, the interphysician variability was substantial. In coronary artery disease, statistical models developed from carefully collected data can provide prognostic predictions that are more accurate than predictions of experienced clinicians made from detailed case summaries.

Coronary Disease↗

A statistical model to optimize indirect sandwich enzyme-linked immunosorbent assay parameters of antigen and antibody: a microcomputer program.

A new computer software program "AVCRV" was developed using a statistical model to analyze the data from the indirect sandwich enzyme-linked immunosorbent assay (ELISA). The software program calculates a sigmoid type of regression analysis and can be run on most microcomputers in the laboratory. The program permits those who are not familiar with computers to complete this type of analysis in a few seconds without a large mainframe computer or complicated software. This statistical model for a sigmoid type of regression analysis of ELISA data may improve the analysis of research data for various avian pathogens from several different experiments.

Animals↗

The frequency of ion-pair substructures in proteins is quantitatively related to electrostatic potential: a statistical model for nonbonded interactions.

A statistical analysis of ion pairs in protein crystal structures shows that their abundance with respect to uncharged controls is accurately predicted by a Boltzmann-like function of electrostatic potential. It appears that the mechanisms of protein folding and/or evolution combine to produce a "thermal" distribution of local nonbonded interactions, as has been suggested by statistical-mechanical theories. Using this relationship, we develop a maximum likelihood methodology for estimation of apparent energetic parameters from the data base of known structures, and we derive electrostatic potential functions that lead to optimal agreement of observed and predicted ion-pair frequencies. These are similar to potentials of mean force derived from electrostatic theory, but departure from Coulombic behavior is less than has been suggested.

Biological Evolution↗

Statistical models for discerning protein structures containing the DNA-binding helix-turn-helix motif.

A method for discerning protein structures containing the DNA-binding helix-turn-helix (HTH) motif has been developed. The method uses statistical models based on geometrical measurements of the motif. With a decision tree model, key structural features required for DNA binding were identified. These include a high average solvent-accessibility of residues within the recognition helix and a conserved hydrophobic interaction between the recognition helix and the second alpha helix preceding it. The Protein Data Bank was searched using a more accurate model of the motif created using the Adaboost algorithm to identify structures that have a high probability of containing the motif, including those that had not been reported previously.

Binding Sites↗

A statistical model for estimating donor postdonation platelet counts after plateletpheresis.

BACKGROUND: To avoid the need, in serial apheresis donors, either to delay plateletpheresis until a predonation platelet count is completed or to obtain a postdonation count after each procedure, a statistical model has been developed to predict the postdonation platelet count from the donor predonation platelet count, weight, and hematocrit. STUDY DESIGN AND METHODS: Predonation and postdonation platelet counts were measured in two groups of approximately 100 consecutive donors (Group A to test the model and Group B to validate it), and the postdonation counts were calculated with the model. Using stepwise multiple linear regression from donor data, estimated postdonation platelet counts were found to be comparable to the postdonation platelet counts actually measured. RESULTS: Estimated postdonation platelet counts x 10(9) per L (mean +/- SD) for each group, respectively, were Group A, 195 +/- 35, versus actual platelet counts of 195 +/- 39 (p = 0.43), and Group B, 183 +/- 36, versus actual platelet counts of 189 +/- 34 (p = 0.14). Sensitivity and specificity, respectively, were Group A, 57 and 99 percent and Group B, 62 and 99 percent. CONCLUSION: For most serial apheresis donors, application of this predictor model should preclude the need to obtain an extra postdonation platelet count.

Blood Donors↗

Statistical model for characterizing epistatic control of triploid endosperm triggered by maternal and offspring QTLs.

To study the effects of maternal and endosperm quantitative trait locus (QTL) interaction on endosperm development, we derive a two-stage hierarchical statistical model within the maximum-likelihood context, implemented with an expectation-maximization algorithm. A model incorporating both maternal and offspring marker information can improve the accuracy and precision of genetic mapping. Extensive simulations under different sampling strategies, heritability levels and gene action modes were performed to investigate the statistical properties of the model. The QTL location and parameters are better estimated when two QTLs are located at different intervals than when they are located at the same interval. Also, the additive effect of the offspring QTLs is better estimated than the additive effect of the maternal QTLs. The implications of our model for agricultural and evolutionary genetic research are discussed.

Algorithms↗

A statistical model for unwarping of 1-D electrophoresis gels.

A statistical model is proposed which relates density profiles in 1-D electrophoresis gels, such as those produced by pulsed-field gel electrophoresis (PFGE), to databases of profiles of known genotypes. The warp in each gel lane is described by a trend that is linear in its parameters plus a first-order autoregressive process, and density differences are modelled by a mixture of two normal distributions. Maximum likelihood estimates are computed efficiently by a recursive algorithm that alternates between dynamic time warping to align individual lanes and generalised-least-squares regression to ensure that the warp is smooth between lanes. The method, illustrated using PFGE of Escherichia coli O157 strains, automatically unwarps and classifies gel lanes, and facilitates manual identification of new genotypes.

Algorithms↗

Statistical model with a standard Gamma distribution.

We study a statistical model consisting of N basic units which interact with each other by exchanging a physical entity, according to a given microscopic random law, depending on a parameter lambda. We focus on the equilibrium or stationary distribution of the entity exchanged and verify through numerical fitting of the simulation data that the final form of the equilibrium distribution is that of a standard Gamma distribution. The model can be interpreted as a simple closed economy in which economic agents trade money and a saving criterion is fixed by the saving propensity lambda. Alternatively, from the nature of the equilibrium distribution, we show that the model can also be interpreted as a perfect gas at an effective temperature T(lambda), where particles exchange energy in a space with an effective dimension D(lambda).

Journal Article↗

Statistical modeling of biomedical corpora: mining the Caenorhabditis Genetic Center Bibliography for genes related to life span.

BACKGROUND: The statistical modeling of biomedical corpora could yield integrated, coarse-to-fine views of biological phenomena that complement discoveries made from analysis of molecular sequence and profiling data. Here, the potential of such modeling is demonstrated by examining the 5,225 free-text items in the Caenorhabditis Genetic Center (CGC) Bibliography using techniques from statistical information retrieval. Items in the CGC biomedical text corpus were modeled using the Latent Dirichlet Allocation (LDA) model. LDA is a hierarchical Bayesian model which represents a document as a random mixture over latent topics; each topic is characterized by a distribution over words. RESULTS: An LDA model estimated from CGC items had better predictive performance than two standard models (unigram and mixture of unigrams) trained using the same data. To illustrate the practical utility of LDA models of biomedical corpora, a trained CGC LDA model was used for a retrospective study of nematode genes known to be associated with life span modification. Corpus-, document-, and word-level LDA parameters were combined with terms from the Gene Ontology to enhance the explanatory value of the CGC LDA model, and to suggest additional candidates for age-related genes. A novel, pairwise document similarity measure based on the posterior distribution on the topic simplex was formulated and used to search the CGC database for "homologs" of a "query" document discussing the life span-modifying clk-2 gene. Inspection of these document homologs enabled and facilitated the production of hypotheses about the function and role of clk-2. CONCLUSION: Like other graphical models for genetic, genomic and other types of biological data, LDA provides a method for extracting unanticipated insights and generating predictions amenable to subsequent experimental validation.

Animals↗

Statistical model for predicting non-heme iron bioavailability from vegetarian meals.

Availability of non-heme iron has been extensively discussed when meals comprise heme as well as non-heme iron, but seldom so for exclusively vegetarian meals. The present study aimed to develop a statistical model for predicting non-heme iron availability from a composite vegetarian meal. Radioisotopic measurements of in vitro iron dialyzability of 208 out of 274 meals representing vegetarian diets from Asia, Africa, Europe and Latin America and the meal contents of iron, zinc, copper, ascorbic acid, beta-carotene, riboflavin, thiamin, folic acid, tannic acid, fiber and degraded phytate forms (IP6-IP1) were used for development of the model. A multiple regression model weighted for calorie contents was developed for the percentage iron dialyzability with the possible predictors as meal contents along with plausible interaction terms. The model was validated with in vitro iron dialyzability of 66 meals and in vivo iron absorption in five ileostomized adults. Application of the model was demonstrated using data on the daily dietary intake of 215 young adults whose hemoglobin levels were estimated twice in 3 weeks. Weighted multiple regression model was: ln(% Fe dialyzability)=1.340-0.259xln(IP2 [mg])+0.188xln(IP3 [mg])-0.278xln(IP5 [mg])+0.0912xln(ascorbic acid [mg])+0.06693xln(tannins [mg])+0.09552xln(beta-carotene [microg])+0.137xln(hemicellulose [g]) (P<0.01, R2=0.51). Good agreement was seen between observed and predicted dialyzability (r=0.90) and human absorption (r=0.89). The model would be useful to estimate bioavailable iron intakes of vegetarian populations and to identify at-risk individuals.

Adult↗

Gene-environment interaction and the mapping of complex traits: some statistical models and their implications.

The manifestation of many complex diseases or traits is very likely the result of an inextricable interplay of the biological and the environmental. Yet the role of environmental effect has traditionally been played down, for various reasons. In this paper, some simple statistical models that incorporate gene-environment interaction (GEI) have been proposed and their behavior and implications investigated. These implications concern the conditional independence assumption in likelihood calculation of pedigree data, the fine-tuning of the sib pair method for mapping quantitative traits, apportioning of disease or trait variation due to specific causes. In addition, they concern properties of gene mapping methods that do not take GEI into account, and they bring into question the utility of commonly used measures of genetic effects such as recurrence risk ratio for relative pairs, twin concordance rates, and heritability coefficients. In the presence of GEI, all these measures are functions not only of genetic effects and gene frequency, but also of environmental effects, the distribution of environmental factors in the population, and of GEI. Above all, these measures are all measures of familial aggregation, since they can be significant even in the absence of any genetic component of the disease. Thus their use as indicators of the genetic basis of complex diseases is cast into doubt.

Chromosome Mapping↗

A simple statistical model for prediction of acute coronary syndrome in chest pain patients in the emergency department.

BACKGROUND: Several models for prediction of acute coronary syndrome (ACS) among chest pain patients in the emergency department (ED) have been presented, but many models predict only the likelihood of acute myocardial infarction, or include a large number of variables, which make them less than optimal for implementation at a busy ED. We report here a simple statistical model for ACS prediction that could be used in routine care at a busy ED. METHODS: Multivariable analysis and logistic regression were used on data from 634 ED visits for chest pain. Only data immediately available at patient presentation were used. To make ACS prediction stable and the model useful for personnel inexperienced in electrocardiogram (ECG) reading, simple ECG data suitable for computerized reading were included. RESULTS: Besides ECG, eight variables were found to be important for ACS prediction, and included in the model: age, chest discomfort at presentation, symptom duration and previous hypertension, angina pectoris, AMI, congestive heart failure or PCI/CABG. At an ACS prevalence of 21% and a set sensitivity of 95%, the negative predictive value of the model was 96%. CONCLUSION: The present prediction model, combined with the clinical judgment of ED personnel, could be useful for the early discharge of chest pain patients in populations with a low prevalence of ACS.

Acute Disease↗

Statistical model of amino acid code of protein secondary structure.

In the previous paper (Shestopalov, 2003) we presented the amino acid code of protein secondary structure as a partial solution of the fundamental problem of the protein three-dimensional structure calculation from the amino acid sequence. Here a statistical model of the code is described. The model is based on the structural data from 2258 protein chains (417,112 amino acid residues used). 60 and 61% of the secondary structure, calculated using the model, coincide, respectively, with the observed secondary structure in the training subset and test subset (104 protein chains and 21,166 residues used). This is equal to the threshold value for all the secondary structure calculations, based on the models, where, similarly as here, only the nearest and middle-range interactions are considered. Therefore the constructed model can be applied for the protein structure prediction from the amino acid sequence, especially when additional information is used along with expert analysis, as in the most successful prediction methods. The model can be used for analysis of the secondary structure changes during protein folding by comparison of the calculated and observed secondary structures. The information about the conformationally invariant segments can serve for the simulation of the supersecondary structure formation. One can try to obtain and examine the protein subset, in which the calculated and observed secondary structures are very similar.

Amino Acid Sequence↗