[Accuracy in predicting drug preparation demand using mathematical and statistical models].
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We applied factor analysis to indicators of health and living conditions from Brazilian states in order to define geographic areas at potential risk for neonatal tetanus. Two factors, namely 'health profile in the rural area' and 'proportional mortality on neonatal tetanus' were selected and plotted against each other. A cluster composed of States from the Northeast plus Para and Amapa was found to include most neonatal tetanus risk areas and a low case-reporting rate. Another cluster included States from the Southeast and South and displayed a neonatal tetanus reporting rate that was compatible with that for other indicators. Espirito Santo, however, was found to be a silent productive area. The Federal District appeared alone, showing the best health conditions. Finally, the States of the Middle West and Roraima constituted the last cluster, characterized by intermediate health status and high neonatal tetanus case-reporting rate. Our results were consistent with the overall Brazilian health profile, distinguishing the North and Northeast from the South and Southeast, with the Middle West in an intermediate position.
OBJECTIVE: To predict the risk of extracapsular extension and postoperative recurrence before radical prostatectomy (RP) for prostate cancer. METHODS: We performed multivariate Cox regression analysis on preoperative variables in 260 clinically localized prostate cancer patients who underwent RP. With these data, we constructed a relative risk of recurrence (Rr) equation and an equation to predict the probability of extracapsular extension (PECE) before RP. RESULTS: Rr is calculated as exp[(0.47 x race + 0.14 x PSAST) + (0.13 x worst biopsy Gleason sum) + (1.03 x stage T1c) + (1.55 x stage T2b,c)], where PSAST indicates a sigmoidal transformation of prostate-specific antigen. PECE is calculated as 1/[1 + exp(-Z)], where Z = -2.47 + 0.15 (PSAST) + 0.31 (worst biopsy Gleason sum) + 0.18 (race) + 0.16 (stage T1c) + 0.38 (stage T2b,c). CONCLUSION: These two equations can be used preoperatively to predict the probability of extracapsular disease and the risk of prostate-specific antigen recurrence in patients undergoing RP.
Statistical models were developed to help understand the relationship between the driver age and several important accident-related factors and circumstances such as injury severity, collision types, average daily traffic (ADT), roadway character, speed ratio, alcohol involvement, and accident location. By using techniques of categorical analysis on the 1994 and 1995 Florida accident database, four long-linear models with three variables in each model with all possible two-way interactions were developed. In order to compare the differences in response between the age groups and a particular accident-related variable, odds multipliers were computed. The effects of age and accident-related factors were examined, and interactions among them were considered. The results indicated significant relationships between the driver age and ADT, injury severity, manner of collision, speed, alcohol involvement, and roadway character. The findings' contribution to the understanding of the effect of age on accident involvement is addressed. A discussion of how log-linear and logit modeling with estimation of 'odds multipliers' may contribute to traffic safety studies is also provided.
Statistical modeling of links between genetic profiles with environmental and clinical data to aid in medical diagnosis is a challenge. Here, we present a computational approach for rapidly selecting important clinical data to assist in medical decisions based on personalized genetic profiles. What could take hours or days of computing is available on-the-fly, making this strategy feasible to implement as a routine without demanding great computing power. The key to rapidly obtaining an optimal/nearly optimal mathematical function that can evaluate the "disease stage" by combining information of genetic profiles with personal clinical data is done by querying a precomputed solution database. The database is previously generated by a new hybrid feature selection method that makes use of support vector machines, recursive feature elimination and random sub-space search. Here, to evaluate the method, data from polymorphisms in the renin-angiotensin-aldosterone system genes together with clinical data were obtained from patients with hypertension and control subjects. The disease "risk" was determined by classifying the patients' data with a support vector machine model based on the optimized feature; then measuring the Euclidean distance to the hyperplane decision function. Our results showed the association of renin-angiotensin-aldosterone system gene haplotypes with hypertension. The association of polymorphism patterns with different ethnic groups was also tracked by the feature selection process. A demonstration of this method is also available online on the project's web site.
Statistical models have been developed to delineate the major-gene and non-major-gene factors accounting for the familial aggregation of complex diseases. The mixed model assumes an underlying liability to the disease, to which a major gene, a multifactorial component, and random environment contribute independently. Affection is defined by a threshold on the liability scale. The regressive logistic models assume that the logarithm of the odds of being affected is a linear function of major genotype, phenotypes of antecedents and other covariates. An equivalence between these two approaches cannot be derived analytically. I propose a formulation of the regressive logistic models on the supposition of an underlying liability model of disease. Relatives are assumed to have correlated liabilities to the disease; affected persons have liabilities exceeding an estimable threshold. Under the assumption that the correlation structure of the relatives' liabilities follows a regressive model, the regression coefficients on antecedents are expressed in terms of the relevant familial correlations. A parsimonious parameterization is a consequence of the assumed liability model, and a one-to-one correspondence with the parameters of the mixed model can be established. The logits, derived under the class A regressive model and under the class D regressive model, can be extended to include a large variety of patterns of family dependence, as well as gene-environment interactions.
The data obtained from measurements of regional rCMRglu using [18F]fluorodeoxyglucose (FDG)/positron emission tomographic (PET) data contain more structure than can be identified with group mean rCMRglu profiles or regional correlation coefficients. This additional structure is revealed by a novel mathematical-statistical model of regional metabolic interactions that explicitly represents rCMRglu profiles as a combination of region-independent global effects, a group mean pattern and a mosaic of interacting networks. In its application to FDG/PET data, this model removes global subject effects [global scaling factors (GSFs)] and a group mean pattern (profile) so as to maximize statistical power for the detection and simultaneous discovery of all networks of two or more regions that form a significant and consistent linearly covarying pattern. The model approach presented here was applied to the combined rCMRglu data from 12 demented AIDS patients and 18 normal controls: Two significant metabolic covariance pattern descriptors that together accounted for 71 to 96% of the rCMRglu/GSF variation across subjects for 22/28 regions in the AIDS group were extracted. Each descriptor was found to be highly correlated with performance on several neuropsychological tests, providing independent validation of the analysis technique as a means of discovering and describing behaviorally related components of group rCMRglu profiles.
In the US, single-vehicle run-off-roadway accidents result in a million highway crashes with roadside features every year and account for approximately one third of all highway fatalities. Despite the number and severity of run-off-roadway accidents, quantification of the effect of possible countermeasures has been surprisingly limited due to the absence of data (particularly data on roadside features) needed to rigorously analyze factors affecting the frequency and severity of run-off-roadway accidents. This study provides some initial insight into this important problem by combining a number of databases, including a detailed database on roadside features, to analyze run-off-roadway accidents on a 96.6-km section of highway in Washington State. Using zero-inflated count models and nested logit models, statistical models of accident frequency and severity are estimated and the findings isolate a wide range of factors that significantly influence the frequency and severity of run-off-roadway accidents. The marginal effects of these factors are computed to provide an indication on the effectiveness of potential countermeasures. The findings show significant promise in applying new methodological approaches to run-off-roadway accident analysis.
BACKGROUND: Given the pressure on healthcare budgets, assessing the cost of managing a disease has become a major research focus; yet collection of these data are labor intensive and difficult. Understanding the predictors of cost provides an efficient means of incorporating such information in decision-making concerning new therapies. METHODS: Data from two 12-week multinational trials that collected information on a variety of neurological, functional, and cost parameters for 1341 ischemic stroke patients were examined by means of multiple linear regression. Because the intent is for the model to be predictive, only patient characteristics that can be known at the time of patient presentation or shortly thereafter were evaluated for inclusion in the model. RESULTS: The Barthel Index was the strongest predictor of cost in all models evaluated. Other major predictors, either directly or through their impact on survival, were stroke subtype, neurological impairment, congestive heart failure, and country. A good model fit was obtained, judging by the model statistics (model F:=84, 3 df, P:<0.0001) and the accuracy of the predictions (<3% difference between mean actual and predicted cost). CONCLUSIONS: Through the use of key patient characteristics, this regression model allows for prediction of the cost of stroke care, which may be helpful in the context of therapeutic decisions and budgetary planning purposes. It also provides insight into how specific treatments, through their impact on clinical characteristics, can modify the cost of stroke treatment.
MOTIVATION: Statistical models of protein families, such as position-specific scoring matrices, profiles and hidden Markov models, have been used effectively to find remote homologs when given a set of known protein family members. Unfortunately, training these models typically requires a relatively large set of training sequences. Recent work (Grundy, J. Comput. Biol., 5,<479-492, 1998) has shown that, when only a few family members are known, several theoretically justified statistical modeling techniques fail to provide homology detection performance on a par with Family Pairwise Search (FPS), an algorithm that combines scores from a pairwise sequence similarity algorithm such as BLAST. RESULTS: The present paper provides a model-based algorithm that improves FPS by incorporating hybrid motif-based models of the form generated by Cobbler (Henikoff and Henikoff, Protein Sci., 6, 698-705, 1997). For the 73 protein families investigated here, this cobbled FPS algorithm provides better homology detection performance than either Cobbler or FPS alone. This improvement is maintained when BLAST is replaced with the full Smith-Waterman algorithm. AVAILABILITY: http://fps.sdsc.edu
We studied the time of onset of chest pain in 1099 patients admitted to a coronary care unit with myocardial infarction using a statistical model. Statistical analysis demonstrated an excess of infarcts with time of onset of chest pain at 0700 hours (14%) and at midnight (11%), with the remaining infarct population (75%) forming a background distribution over the 24 hr.
If our goal in Artificial Intelligence in Medicine (AIM) is to engineer systems health-care providers will both use and, in the process, improve their performance, we must concentrate on the development of causal theories of knowledge and problem solving. One broad direction in pursuing this goal is understanding the relationships between existing models of rationality and bounded rationality for similar tasks. Models of rationality refer to those approaches in which the optimal properties of the models are deductively provable, i.e. in which the processing is rational. Representative models of rationality used in AIM are deductive logical models, statistical models such as Bayesian inference models, and decision-analytic models. Models of bounded rationality are those which do not guarantee such optimal properties nor yield to deductive correctness proofs. These models have their roots in cognitive psychology. In this article we show how explicating the relationship between models of rationality and bounded rationality might be done in the case of abductive tasks in medicine. This is done by positioning these modeling approaches within the same framework (an abstract computational model) and interpreting in this context both computational complexity results concerning the nature of the task and empirical results studies of human problem-solving behavior.
The statistics of base-pair choice in individual recognition sites on DNA is shown to be determined by the functional binding requirements for recognition and a selection parameter. This selection parameter can be identified as a generalized external force required to deform a random-choice base-pair distribution into the observed specific-choice distribution. This external force is balanced by the randomization pressure which--driven by mutations--always tends to increase randomness in the base-pair choices. The model makes it possible to predict relative binding constants of particular recognition sequences based primarily on the statistics of base-pair usage. A further consequence of this formulation is that the randomization pressure appears explicitly as an important force shaping the evolutionary selection not only of DNA sites, but also of other properties involving macromolecular design.
In this paper, a new statistical model for representing the amplitude statistics of ultrasonic images is presented. The model is called the Rician inverse Gaussian (RiIG) distribution, due to the fact that it is constructed as a mixture of the Rice distribution and the Inverse Gaussian distribution. The probability density function (pdf) of the RiIG model is given in closed form as a function of three parameters. Some theoretical background on this new model is discussed, and an iterative algorithm for estimating its parameters from data is given. Then, the appropriateness of the RiIG distribution as a model for the amplitude statistics of medical ultrasound images is experimentally studied. It is shown that the new distribution can fit to the various shapes of local histograms of linearly scaled ultrasound data better than existing models. A log-likelihood cross-validation comparison of the predictive performance of the RiIG, the K, and the generalized Nakagami models turns out in favor of the new model. Furthermore, a maximum a posteriori (MAP) filter is developed based on the RiIG distribution. Experimental studies show that the RiIG MAP filter has excellent filtering performance in the sense that it smooths homogeneous regions, and at the same time preserves details.
MOTIVATION: Li and Wong have described some useful statistical models for probe-level, oligonucleotide array data based on a multiplicative parametrization. In earlier work, we proposed similar analysis-of-variance-style mixed models fit on a log scale. With only subtle differences in the specification of their mean and stochastic error components, a question arises as to whether these models could lead to varying conclusions in practical application. RESULTS: In this paper, we provide an empirical comparison of the two models using a real data set, and find the models perform quite similarly across most genes, but with some interesting and important distinctions. We also present results from a simulation study designed to assess inferential properties of the models, and propose a modified test statistic for the Li-Wong model that provides an improvement in Type 1 error control. Advantages of both methods include the ability to directly assess and account for key sources of variability in the chip data and a means to automate statistical quality control.
There has been considerable research conducted over the last 20 years focused on predicting motor vehicle crashes on transportation facilities. The range of statistical models commonly applied includes binomial, Poisson, Poisson-gamma (or negative binomial), zero-inflated Poisson and negative binomial models (ZIP and ZINB), and multinomial probability models. Given the range of possible modeling approaches and the host of assumptions with each modeling approach, making an intelligent choice for modeling motor vehicle crash data is difficult. There is little discussion in the literature comparing different statistical modeling approaches, identifying which statistical models are most appropriate for modeling crash data, and providing a strong justification from basic crash principles. In the recent literature, it has been suggested that the motor vehicle crash process can successfully be modeled by assuming a dual-state data-generating process, which implies that entities (e.g., intersections, road segments, pedestrian crossings, etc.) exist in one of two states-perfectly safe and unsafe. As a result, the ZIP and ZINB are two models that have been applied to account for the preponderance of "excess" zeros frequently observed in crash count data. The objective of this study is to provide defensible guidance on how to appropriate model crash data. We first examine the motor vehicle crash process using theoretical principles and a basic understanding of the crash process. It is shown that the fundamental crash process follows a Bernoulli trial with unequal probability of independent events, also known as Poisson trials. We examine the evolution of statistical models as they apply to the motor vehicle crash process, and indicate how well they statistically approximate the crash process. We also present the theory behind dual-state process count models, and note why they have become popular for modeling crash data. A simulation experiment is then conducted to demonstrate how crash data give rise to "excess" zeros frequently observed in crash data. It is shown that the Poisson and other mixed probabilistic structures are approximations assumed for modeling the motor vehicle crash process. Furthermore, it is demonstrated that under certain (fairly common) circumstances excess zeros are observed-and that these circumstances arise from low exposure and/or inappropriate selection of time/space scales and not an underlying dual state process. In conclusion, carefully selecting the time/space scales for analysis, including an improved set of explanatory variables and/or unobserved heterogeneity effects in count regression models, or applying small-area statistical methods (observations with low exposure) represent the most defensible modeling approaches for datasets with a preponderance of zeros.
Statistical inference is considered for a two-state Markov model of a single ion channel, when time interval omission is incorporated. A simple method of obtaining confidence sets for the mean open and closed sojourn times for the underlying single channel, based on the method-of-moments estimators, is presented. Time interval omission induces non-identifiability, in that the method-of-moments usually leads to two distinct estimates of the mean open and closed sojourn times, one corresponding to the true values and the other being an artefact of time interval omission. A new method of overcoming such non-identifiability on the basis of one single channel record is described. The methodology is illustrated by a numerical example.
Some new nonlinear models for the relationship between the fraction of drug dose dissolved (absorbed) in vivo and that dissolved in vitro are described. The models are empirical in nature and are generalizations of the linear model that, at present, is the most commonly used model. The modeling approach is based on considering the time at which a drug molecule goes into solution (in vitro or in vivo) to be a random variable and relating the distribution functions using proportional odds, proportional hazards, and proportional reversed hazards models. The models are further extended by allowing the parameter that relates in vivo and in vitro to be a function of time. A statistical model for the data is developed and used as the basis for a statistical methodology for fitting these models. The methods are shown to be generalized linear mixed effects model (GLMM) methods. The models are fitted to some data sets, and the results demonstrate that these models have potential.