PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Statistical modeling and visualization of molecular profiles in cancer.

Current cancer classifications using morphological criteria produce heterogeneous classes with variable prognosis and clinical course. By measuring gene expression for thousands of genes in a single hybridization experiment, microarrays have the potential to contribute to more effective classifications based on molecular information. This gives hope to improve both prognosis and treatment. Statistical methods for molecular classification have focused on using high dimensional representations of molecular profiles to identify subclasses. These can be noisy, unstable, and highly platform-specific. In this article, we emphasize the notion of molecular profiles based on latent categories signifying under-, over-, and baseline expression. Following this approach, we can generate results that are more easily interpretable, more easily translated into clinical tools, more robust to noise, and less platform-dependent. We illustrate both the methods and the associated software for molecular class discovery on a data set of 244 microarrays comprising six known leukemia classes.

Child↗

Statistical modeling of carcinogenic risks in dogs that inhaled 238PuO2.

Combined analyses of data on 260 life-span beagle dogs that inhaled 238PuO2 at the Inhalation Toxicology Research Institute (ITRI) and at Pacific Northwest National Laboratory (PNNL) were conducted. The hazard functions (age-specific risks) for incidence of lung, bone and liver tumors were modeled as a function of cumulative radiation dose, and estimates of lifetime risks based on the combined data were developed. For lung tumors, linear-quadratic functions provided an adequate fit to the data from both laboratories, and linear functions provided an adequate fit when analyses were restricted to doses less than 20 Gy. The estimated risk coefficients for these functions were significantly larger when based on ITRI data compared to PNNL data, and dosimetry biases are a possible explanation for this difference. There was also evidence that the bone tumor response functions differed for the two laboratories, although these differences occurred primarily at high doses. These functions were clearly nonlinear (even when restricted to average skeletal doses less than 1 Gy), and evidence of radiation-induced bone tumors was found for doses less than 0.5 Gy in both laboratories. Liver tumor risks were similar for the two laboratories, and linear functions provided an adequate fit to these data. Lifetime risk estimates for lung and bone tumors derived from these data had wide confidence intervals, but were consistent with estimates currently used in radiation protection. The dog-based lifetime liver tumor risk estimate was an order of magnitude larger than that used in radiation protection, but the latter also carries large uncertainties. The application of common statistical methodology to data from two studies has allowed the identification of differences in these studies and has provided a basis for common risk estimates based on both data sets.

Administration, Inhalation↗

Genetic Interaction Motif Finding by expectation maximization--a novel statistical model for inferring gene modules from synthetic lethality.

BACKGROUND: Synthetic lethality experiments identify pairs of genes with complementary function. More direct functional associations (for example greater probability of membership in a single protein complex) may be inferred between genes that share synthetic lethal interaction partners than genes that are directly synthetic lethal. Probabilistic algorithms that identify gene modules based on motif discovery are highly appropriate for the analysis of synthetic lethal genetic interaction data and have great potential in integrative analysis of heterogeneous datasets. RESULTS: We have developed Genetic Interaction Motif Finding (GIMF), an algorithm for unsupervised motif discovery from synthetic lethal interaction data. Interaction motifs are characterized by position weight matrices and optimized through expectation maximization. Given a seed gene, GIMF performs a nonlinear transform on the input genetic interaction data and automatically assigns genes to the motif or non-motif category. We demonstrate the capacity to extract known and novel pathways for Saccharomyces cerevisiae (budding yeast). Annotations suggested for several uncharacterized genes are supported by recent experimental evidence. GIMF is efficient in computation, requires no training and automatically down-weights promiscuous genes with high degrees. CONCLUSION: GIMF effectively identifies pathways from synthetic lethality data with several unique features. It is mostly suitable for building gene modules around seed genes. Optimal choice of one single model parameter allows construction of gene networks with different levels of confidence. The impact of hub genes the generic probabilistic framework of GIMF may be used to group other types of biological entities such as proteins based on stochastic motifs. Analysis of the strongest motifs discovered by the algorithm indicates that synthetic lethal interactions are depleted between genes within a motif, suggesting that synthetic lethality occurs between-pathway rather than within-pathway.

Algorithms↗

Monitoring and analysis of bovine spongiform encephalopathy (BSE) testing in Denmark using statistical models.

The evolution of monitoring and surveillance for bovine spongiform encephalopathy (BSE) from the phase of passive surveillance that began in the United Kingdom in 1988 until the present is described. Currently, surveillance for BSE in Europe consists of mass testing of cattle slaughtered for human consumption and cattle from certain groups considered to be at higher risk of having clinical or detectable BSE. The results of the ongoing BSE testing in Denmark have been analyzed using two statistical approaches: the "classical" frequentist and the Bayesian that is widely used in quantitative risk analysis. The analyses were intended to provide information for decision-makers, the media and the public as well as to provide inputs for future BSE surveillance models. The results to date suggest that the total number of BSE cases that will be found in Denmark in 2001 will not exceed 16.

Age Distribution↗

Statistical modeling of protein spray drying at the lab scale.

The objective of this study was to examine the effects of formulation and process variables on particle size and other characteristics of a spray-dried model protein, bovine serum albumin (BSA), using a partial factorial design for experiments. Formulation variables tested include concentration and zinc:protein complexation ratio. Process variables explored were inlet temperature, liquid feed rate, drying air flow rate, and atomizing nitrogen pressure on a lab-scale spray dryer. Statistical data analysis was used to determine F ratios for each of the inputs, which provided a means of ranking the importance of variables relative to one another for each powder characteristic of interest. It was found that protein concentration and atomizing nitrogen pressure had the greatest effects on the particle size of the protein powder. For determining product yield, results showed that protein concentration was the critical variable. Finally, the outlet temperature was mostly influenced by inlet temperature and liquid feed rate. Mathematical models based on these input-output relationships were constructed; these models provide insight into some of the controllable variables of the spray-drying process.

Chemistry, Pharmaceutical↗

A statistical model to evaluate analyte homogeneity for a material.

An underlying assumption for collaborative studies is that the analyte variation among test samples of the material (i.e., matrix and analyte concentration combination) under study has a negligible influence on the estimates of precision for the method. This assumption is expected to be fulfilled when the material under study is prepared (i.e., thoroughly mixed) such that the analyte is distributed uniformly throughout the matrix. Statistical design and intra-class correlation analysis procedures are proposed to assess the similarity or agreement among analytical results among- and within-containers for single and multiple occasions of use (e.g., collaborative and proficiency studies).

Chemistry Techniques, Analytical↗

Scaling theory and spreading dynamics in systems with one absorbing state derived from an equilibrium statistical model.

We show that for systems with one absorbing state, the widely assumed scaling properties of the survival probability and of the probability density of the size of activity avalanches cannot be true in the asymptotic limit. Trying to answer the question, what is the true asymptotic limit of these quantities, we study Domany-Kinzel probabilistic cellular automata using an equilibrium statistical mechanic model (ESM). We are able to express important quantities of the avalanche dynamics by correlation functions of the ESM. The application of scaling theory to the ESM allows for the derivation of the scaling properties of quantities of the avalanche dynamics in the form of infinite series. From these results we can obtain possible solutions for the apparent scaling problem, but cannot decide definitely which one is true. The most appealing solution, for which some evidence is given, states that there is a narrow range around the critical point in which, for example, the survival probability has the same power-law behavior as on the critical point. Outside this narrow range, the usually assumed scaling should be approximately valid.

Journal Article↗

Dose-response and threshold-mediated mechanisms in mutagenesis: statistical models and study design.

The objective of this paper is to review the use, in mutagenesis, of various mathematical models to describe the dose-response relationship and to try to identify thresholds. It is often taken as axiomatic that genotoxic carcinogens could damage DNA at any level of exposure, leading to a mutation, and that this could ultimately result in tumour development. This has led to the assumption that for genotoxic chemicals, there is no discernible threshold. This assumption is increasingly being challenged in the case of aneugens. The distinction between 'absolute' and 'pragmatic' thresholds is made and the difficulties in determining 'absolute' thresholds using hypothesis testing approaches are described. The potential of approaches, based upon estimation rather than statistical significance for the characterization of dose-response relationships, is stressed. The achievement of a good fit of a mathematical model to experimental data is not proof that the mechanism supposedly underlying this model is operating. It has been argued, in the case of genotoxic chemicals, that any effects produced by a genotoxic chemical which augments that producing a background incidence in unexposed individuals will lead to a dose-response relationship that is non-thresholded and is linear at low doses. The assumptions underlying this presumption are explored in the context of the increasing knowledge of the mechanistic basis of mutagenicity and carcinogenicity. The possibility that exposure to low levels of genotoxic chemicals may induce and enhance defence and repair mechanisms is not easily incorporated into many of the existing mathematical models and should be an objective in the development of the next generation of biologically based dose-response (BB-DR) models. Studies aimed at detecting or characterizing non-linearities in the dose-response relationship need appropriate experimental designs with careful attention to the choice of biomarker, number and selection of dose levels, optimum allocation of experimental units and appropriate levels of replication within and repetition of experiments. The characterization of dose-response relationships with appropriate measures of uncertainty can help to identify 'pragmatic' thresholds based upon biologically relevant criteria which can help in the regulatory process.

Aneuploidy↗

Statistical modeling for selecting housekeeper genes.

There is a need for statistical methods to identify genes that have minimal variation in expression across a variety of experimental conditions. These 'housekeeper' genes are widely employed as controls for quantification of test genes using gel analysis and real-time RT-PCR. Using real-time quantitative RT-PCR, we analyzed 80 primary breast tumors for variation in expression of six putative housekeeper genes (MRPL19 (mitochondrial ribosomal protein L19), PSMC4 (proteasome (prosome, macropain) 26S subunit, ATPase, 4), SF3A1 (splicing factor 3a, subunit 1, 120 kDa), PUM1 (pumilio homolog 1 (Drosophila)), ACTB (actin, beta) and GAPD (glyceraldehyde-3-phosphate dehydrogenase)). We present appropriate models for selecting the best housekeepers to normalize quantitative data within a given tissue type (for example, breast cancer) and across different types of tissue samples.

ATPases Associated with Diverse Cellular Activitie↗

Recognition of a category of responders to group II, slow-grower associated antigens amongst Kuwaiti senior school children, using a statistical model.

A mathematical model previously developed to test the validity of categorisation of skin test responders has been applied to data obtained from 3 age groups of Kuwaiti school children. Two specially designed sets of 4 new tuberculins were tested on senior school children to determine whether extra categories of responders might exist amongst them. Strong statistical evidence has been obtained that a proportion of the children respond to group ii, slow-grower associated antigen, creating a fourth responder category, but no evidence was found for responses to group iii, fast-grower associated antigen. The significance of group ii antigens in immune protection from tuberculosis has never been considered specifically. It is of special interest to note that responders to these antigens have been readily found in Kuwait, a country where BCG is thought to be effective, whereas no such category could be found in India or Sri Lanka, where the efficacy of the vaccine is less certain.

Adolescent↗

Approach to erythrocyte aggregation through erythrocyte sedimentation rate: application of a statistical model in pathology.

Erythrocyte sedimentation rate (ESR) is mainly used in clinical practice as a screening test for inflammatory diseases and sometimes in the follow-up of patients. However, ESR is highly dependent on erythrocyte aggregation. In this study, using a Sediscan (Becton) automatic device measuring the kinetics of ESR, these results are compared with the measurement of erythrocyte aggregation as determined by laser light backscattering (Erythroaggregometer Affibio). A series of 188 samples from in-patients were tested. Statistical analysis of 13 parameters indicates that 82% of ESR variance may be explained by fibrinogen level, haematocrit and a parameter characterizing erythrocyte aggregation: the aggregation index at 10 s. This correlation was then validated prospectively in 128 other patients and seems to be independent of the underlying disease. Thus ESR in combination with fibrinogen assay and haematocrit may be considered as a simple and economic method to assess erythrocyte aggregation.

Blood Sedimentation↗

Children with class III malocclusion: development of multivariate statistical models to predict future need for orthognathic surgery.

Until now, the literature does not provide an accurate model to predict the future need for orthognathic surgery in prepubertal patients with class III malocclusion. Because not all of these patients are candidates for later surgical correction, patient assessment and selection remain arbitrary with respect to diagnosis and treatment planning. The purpose of the present investigation was to analyze the value of classifying class III children before puberty into patients who can be effectively treated by orthopedic/orthodontic therapy alone and those who require orthognathic surgery. To obtain a robust model, the study design was multicentric (University Orthodontic Departments of Frankfurt, Heidelberg, and Würzburg). A total of 88 patients with class III malocclusion were grouped into orthopedic/orthodontic (n = 65) and surgery patients (n = 23), according to their records after puberty (mean age, 17 years three months). Discriminant analysis (DA) and logistic regression (LogR) were applied to 20 landmarks of the patients' cephalograms before puberty (mean age, nine years eight months) to identify the dentoskeletal variables that provide the best group separation and the best predictability of group membership, respectively. Both models were highly significant (P < .001), classifying 93.3% (DA) and 94.3% (LogR) of the patients correctly. The extracted variables were identical for both procedures: Wits appraisal, palatal plane angle, and individualized inclination of the lower incisors. The resulting equation of LogR was individual score = -7.968 - 1.323Wits - 0.363NL-NSL + 0.153[180 - (LI-ML) - (L1-ML(ind))]. We concluded that by means of multivariate statistics, prepubertal children with class III malocclusions may be classified into nonsurgery and surgery patients with high accuracy.

Adolescent↗

A statistical model for species extrapolation using categorical response data.

Predictions of human health risk for single chemicals are often based on animal studies and hence require some sort of adjustment for species differences in toxic susceptibility. In the past, either the animal dose has been divided by an uncertainty factor or the dose has been transformed by a mathematical model into a human equivalent dose. A generalization of the allometric model previously used for carcinogens, the so-called "surface area model," is investigated here for use with graded severity response data for noncarcinogenic systemic toxicity. Statistical methods for estimating one of the model's parameters, the power of body weight, are proposed and tested on simulated and actual toxicity data. Early results indicate reasonable accuracy if data are available for a large number of dose groups.

Animals↗

[Control of the efficacy of anti-arrhythmia drug therapy with the ambulatory electrocardiogram. Proposal for a new analytical statistical model].

Ambulatory electrocardiography is used for evaluating antiarrhythmic drug effectiveness. Statistical methods based on the analysis of the number of ventricular ectopic beats are currently employed. These techniques are not useful to compare groups of patients with different therapies, due to the wide spontaneous variability of the ectopic beats. We propose a new statistical method, based on the likelihood function. The new method has been tested both retrospectively on 102 patients treated with different antiarrhythmic drugs and prospectively on 12 patients subjected to three consecutive control ambulatory electrocardiograms and to a fourth one after treatment with propafenone. This new statistical method was found to be useful for comparing therapeutic effectiveness between groups of patients, whereas the traditional quantitative methods are to be preferred when drug effectiveness is evaluated in the single patient.

Adrenergic beta-Antagonists↗

Statistical modeling of complex backgrounds for foreground object detection.

This paper addresses the problem of background modeling for foreground object detection in complex environments. A Bayesian framework that incorporates spectral, spatial, and temporal features to characterize the background appearance is proposed. Under this framework, the background is represented by the most significant and frequent features, i.e., the principal features, at each pixel. A Bayes decision rule is derived for background and foreground classification based on the statistics of principal features. Principal feature representation for both the static and dynamic background pixels is investigated. A novel learning method is proposed to adapt to both gradual and sudden "once-off" background changes. The convergence of the learning process is analyzed and a formula to select a proper learning rate is derived. Under the proposed framework, a novel algorithm for detecting foreground objects from complex environments is then established. It consists of change detection, change classification, foreground segmentation, and background maintenance. Experiments were conducted on image sequences containing targets of interest in a variety of environments, e.g., offices, public buildings, subway stations, campuses, parking lots, airports, and sidewalks. Good results of foreground detection were obtained. Quantitative evaluation and comparison with the existing method show that the proposed method provides much improved results.

Algorithms↗

Structure and evolution of protein interaction networks: a statistical model for link dynamics and gene duplications.

BACKGROUND: The structure of molecular networks derives from dynamical processes on evolutionary time scales. For protein interaction networks, global statistical features of their structure can now be inferred consistently from several large-throughput datasets. Understanding the underlying evolutionary dynamics is crucial for discerning random parts of the network from biologically important properties shaped by natural selection. RESULTS: We present a detailed statistical analysis of the protein interactions in Saccharomyces cerevisiae based on several large-throughput datasets. Protein pairs resulting from gene duplications are used as tracers into the evolutionary past of the network. From this analysis, we infer rate estimates for two key evolutionary processes shaping the network: (i) gene duplications and (ii) gain and loss of interactions through mutations in existing proteins, which are referred to as link dynamics. Importantly, the link dynamics is asymmetric, i.e., the evolutionary steps are mutations in just one of the binding parters. The link turnover is shown to be much faster than gene duplications. Both processes are assembled into an empirically grounded, quantitative model for the evolution of protein interaction networks. CONCLUSIONS: According to this model, the link dynamics is the dominant evolutionary force shaping the statistical structure of the network, while the slower gene duplication dynamics mainly affects its size. Specifically, the model predicts (i) a broad distribution of the connectivities (i.e., the number of binding partners of a protein) and (ii) correlations between the connectivities of interacting proteins, a specific consequence of the asymmetry of the link dynamics. Both features have been observed in the protein interaction network of S. cerevisiae.

Biological Evolution↗

Statistical modelling of the differences between successive R-R intervals.

Understanding the behaviour of R-R interval data and its successive differences is critical to the dynamics of cardiac control. Several time domain measures that quantify R-R interval variability have important clinical significance in terms of risk stratification and evaluating the effectiveness of treatment procedures. The present approach at examining the distributions of successive beat-to-beat differences of R-R interval data from different populations and under different conditions (baseline and reaction times) provides a valuable insight into their previously unexplored distributional properties. In particular, our analysis reveals that the successive differences have non-normal statistical distributions (a contradiction to the commonly used assumption of normality), and the absolute successive R-R interval differences approximately follows a Weibull distribution. As an illustration of the utility of this approach, we explore the statistical properties of the time domain measure: root mean square successive difference, study the association between the Weibull scale parameter estimate and respiratory sinus arrhythmia, and propose improvements in artifact detection algorithms.

Adult↗