PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Phylogenetic position of Salinibacter ruber based on concatenated protein alignments.

A total of 22 genes from the genome of Salinibacter ruber strain M31 were selected in order to study the phylogenetic position of this species based on protein alignments. The selection of the genes was based on their essential function for the organism, dispersion within the genome, and sufficient informative length of the final alignment. For each gene, an individual phylogenetic analysis was performed and compared with the resulting tree based on the concatenation of the 22 genes, which rendered a single alignment of 10,757 homologous positions. In addition to the manually chosen genes, an automatically selected data set of 74 orthologous genes was used to reconstruct a tree based on 17,149 homologous positions. Although single genes supported different topologies, the tree topology of both concatenated data sets was shown to be identical to that previously observed based on small subunit (SSU) rRNA gene analysis, in which S. ruber was placed together with Bacteroidetes. In both concatenated data sets the bootstrap was very high, but an analysis with a gradually lower number of genes indicated that the bootstrap was greatly reduced with less than 12 genes. The results indicate that tree reconstructions based on concatenating large numbers of protein coding genes seem to produce tree topologies with similar resolution to that of the single 16S rRNA gene trees. For classification purposes, 16S rRNA gene analysis may remain as the most pragmatic approach to infer genealogic relationships.

Algorithms↗

Development and internal validation of preoperative transition zone prostate cancer nomogram.

OBJECTIVES: Up to 20% of men may harbor a transition zone (TZ) prostate cancer (PCa) at radical prostatectomy (RP). TZ PCa may be associated with more favorable RP pathologic findings than peripheral zone (PZ) PCa. To identify these men, we developed a model capable of predicting the probability of TZ PCa at RP. METHODS: The study cohort consisted of 945 consecutive men treated with RP, with clinical stage, prostate-specific antigen (PSA) level, and detailed biopsy and RP pathology data available. The preoperative variables were used as predictors in the multivariate logistic regression models to predict the rate of TZ PCa at RP. PCa was defined as a TZ tumor when more than 50% of the planimetrically measured tumor volume was situated within the TZ. Regression coefficients were used to develop nomograms, which were subjected to 200 bootstrap resamples to reduce overfit bias. RESULTS: TZ PCa at the final pathologic examination was recorded in 110 patients (11.6%). After 200 bootstraps, the most parsimonious and most accurate nomogram was 77.3% accurate in predicting the probability of TZ PCa. CONCLUSIONS: This nomogram is ideally suited to identify patients with markedly elevated, nearly metastatic serum PSA levels who harbor a TZ tumor that is highly treatable by RP.

Adult↗

Molecular systematics of Salmonidae: combined nuclear data yields a robust phylogeny.

The phylogeny of salmonid fishes has been the focus of intensive study for many years, but some of the most important relationships within this group remain unclear. We used 269 Genbank sequences of mitochondrial DNA (from 16 genes) and nuclear DNA (from nine genes) to infer phylogenies for 30 species of salmonids. We used maximum parsimony and maximum likelihood to analyze each gene separately, the mtDNA data combined, the nuclear data combined, and all of the data together. The phylogeny with the best overall resolution and support from bootstrapping and Bayesian analyses was inferred from the combined nuclear DNA data set, for which the different genes reinforced and complemented one another to a considerable degree. Addition of the mitochondrial DNA degraded the phylogenetic signal, apparently as a result of saturation, hybridization, selection, or some combination of these processes. By the nuclear-DNA phylogeny: (1) (Hucho hucho, Brachymystax lenok) form the sister group to (Salmo, Salvelinus, Oncorhynchus, H. perryi); (2) Salmo is the sister-group to (Oncorhynchus, Salvelinus); (3) Salvelinus is the sister-group to Oncorhynchus; and (4) Oncorhynchus masou forms a monophyletic group with O. mykiss and O. clarki, with these three taxa constituting the sister-group to the five other Oncorhynchus species. Species-level relationships within Oncorhynchus and Salvelinus were well supported by bootstrap levels and Bayesian analyses. These findings have important implications for understanding the evolution of behavior, ecology and life-history in Salmonidae.

Animals↗

Characterization of angiosperm nrDNA polymorphism, paralogy, and pseudogenes.

Many early reports of ITS region (ITS 1, 5.8S, and ITS 2) variation in flowering plants indicated that nrDNA arrays within individuals are homogeneous. However, both older and more recent studies have found intra-individual nrDNA polymorphism across a range of plant taxa including presumed non-hybrid diploids. In addition, polymorphic individuals often contain potentially non-functional nrDNA copies (pseudogenes). These findings suggest that complete concerted evolution should not be assumed when embarking on phylogenetic studies using nrDNA sequences. Here we (1). discuss paralogy in relation to species tree reconstruction and conclude that a priori determinations of orthology and paralogy of nrDNA sequences should not be made based on the functionality or lack of functionality of those sequences; (2). discuss why systematists might be particularly interested in identifying and including pseudogene sequences as a test of gene tree sampling; (3). examine the various definitions and characterizations of nrDNA pseudogenes as well as the relative merits and limitations of a subset of pseudogene detection methods and conclude that nucleotide substitution patterns are particularly appropriate for the identification of putative nrDNA pseudogenes; and (4). present and discuss the advantages of a tree-based approach to identifying pseudogenes based on comparisons of sequence substitution patterns from putatively conserved (e.g., 5.8S) and less constrained (e.g., ITS 1 and ITS 2) regions. Application of this approach, through a method employing bootstrap hypothesis testing, and the issues discussed in the paper are illustrated through reanalysis of two previously published matrices. Given the apparent robustness of the test developed and the ease of carrying out percentile bootstrap hypothesis tests, we urge researchers to employ this statistical tool. While our discussion and examples concern the literature on plant systematics, the issues addressed are relevant to studies of nrDNA and other multicopy genes in other taxa.

DNA, Ribosomal↗

Relationships of the temperate Australasian labrid fish tribe Odacini (Perciformes; Teleostei).

The labrid tribe Odacini comprises four genera and 12 species of fishes that inhabit shallow kelp forest and seagrass areas in temperate waters of Australia and New Zealand. Odacines are morphologically disparate, but share synapomorphies in fin structure and fusion of teeth into a beak-like oral jaw. A phylogenetic analysis of odacines was conducted to investigate their relationships to other labrid fishes, the relationships of species within the tribe, and the evolution of herbivory within the group. Fragments from two mitochondrial genes, 12S rDNA and 16S rDNA, and two nuclear genes, Tmo4C4 and RAG2, were sequenced for seven odacine species (representing all four genera), eight species representing the other major labrid lineages, and three outgroup species. Maximum likelihood and maximum parsimony analyses on the resulting 2338 bp of DNA sequence produced nearly identical topologies differing only in the placement of a clade containing the cheiline Cheilinus fasciatus and the scarine Cryptotomus roseus. The remaining clades received strong bootstrap support under maximum parsimony, and all clades in the maximum likelihood analysis received high bootstrap proportions and high posterior probabilities. The hypsigenyine labrid Choerodon anchorago formed the sister group to the odacines. Within the odacines, Odax cyanoallix+Odax pullus formed the sister to the remaining odacines, with Odax acroptilus, Odax cyanomelas, and Siphonognathus argyrophanes forming successively closer sister groups to the clade Haletta semifasciatus+Neoodax balteatus. Either herbivory evolved twice in the odacines, or herbivory evolved once with two reversions to carnivory. The latter hypothesis appears more likely in the light of odacine feeding biology.

Animals↗

Strange bayes indeed: uniform topological priors imply non-uniform clade priors.

While Bayesian analysis has become common in phylogenetics, the effects of topological prior probabilities on tree inference have not been investigated. In Bayesian analyses, the prior probability of topologies is almost always considered equal for all possible trees, and clade support is calculated from the majority rule consensus of the approximated posterior distribution of topologies. These uniform priors on tree topologies imply non-uniform prior probabilities of clades, which are dependent on the number of taxa in a clade as well as the number of taxa in the analysis. As such, uniform topological priors do not model ignorance with respect to clades. Here, we demonstrate that Bayesian clade support, bootstrap support, and jackknife support from 17 empirical studies are significantly and positively correlated with non-uniform clade priors resulting from uniform topological priors. Further, we demonstrate that this effect disappears for bootstrap and jackknife when data sets are free from character conflict, but remains pronounced for Bayesian clade supports, regardless of tree shape. Finally, we propose the use of a Bayes factor to account for the fact that uniform topological priors do not model ignorance with respect to clade probability.

Bayes Theorem↗

Biodiversity hotspots: evolutionary origins of biodiversity in wrasses (Halichoeres: Labridae) in the Indo-Pacific and new world tropics.

Halichoeres is a widely distributed coral reef fish genus with high levels of biodiversity in both the Indo-Pacific and New World tropics. This study employed molecular phylogenetic techniques and biogeographic analyses on 1700-1800 bp of mitochondrial CO1, 16s, and 12s to test competing hypotheses regarding the origins of biodiversity in this genus in these two biodiversity hotspots. Analyses indicate that Halichoeres is polyphyletic with distinct New World and Indo-Pacific Ocean components. The Halichoeres in the New World tropics formed a strongly supported clade (99% MP, 100% ML bootstrap values) that diverged 21.2-18.1 mya, suggesting that this lineage may represent a relictual fauna of the ancient Tethys Sea. The closure of the Isthmus of Panama contributed to the creation of Halichoeres biodiversity, but diversification across the Isthmus prior to its closure and within the W. Atlantic after the closure 3.1 mya were also important processes creating biodiversity in the New World tropics. Within the Indonesian Australian Archipelago (IAA) analysis of age vs. geographic distribution supported neither Center of Origin, Center of Accumulation or Center of Overlap hypotheses, and molecular clock estimates indicated that the role of Pleistocene sea level changes in the origins of IAA marine biodiversity may be less important than previously thought. Ancestral distribution reconstructions within the Indo-West Pacific (IWP) clade (99% ML bootstrap value) also failed to support these hypotheses as the reconstructions were highly sensitive to the inclusion of missing taxa. Results suggest plueralistic origins of biodiversity, but that vast amounts of habitat may favor the survival of biodiversity in the IAA biodiversity hotspot.

Animals↗

Phylogenetic relationships, host affinity, and geographic structure of boreal and arctic endophytes from three major plant lineages.

Although associated with all plants, fungal endophytes (microfungi that live within healthy plant tissues) represent an unknown proportion of fungal diversity. While there is a growing appreciation of their ecological importance and human uses, little is known about their host specificity, geographic structure, or phylogenetic relationships. We surveyed endophytic Ascomycota from healthy photosynthetic tissues of three plant species (Huperzia selago, Picea mariana, and Dryas integrifolia, representing lycophytes, conifers, and angiosperms, respectively) in northern and southern boreal forest (Québec, Canada) and arctic tundra (Nunavut, Canada). Endophytes were recovered from all plant species surveyed, and were present in <1-41% of 2 mm2 tissue segments examined per host species. Sequence data from the nuclear ribosomal internal transcribed spacer region (ITS) were obtained for 280 of 558 isolates. Species-accumulation curves based on ITS genotypes remained non-asymptotic, and bootstrap analyses indicated that a large number of genotypes remain to be found. The majority of genotypes were recovered from only a single host species, and only 6% of genotypes were shared between boreal and arctic communities. Two independent Bayesian analyses and a neighbor-joining bootstrapping analysis of combined data from the nuclear large and small ribosomal subunits (LSUrDNA, SSUrDNA; 2.4 kb) showed that boreal and arctic endophytes represent Dothideomycetes, Sordariomycetes, Chaetothyriomycetidae, Leotiomycetes, and Pezizomycetes. Many well-supported phylotypes contained only endophytes despite exhaustive sampling of available sequences of Ascomycota. Together, these data demonstrate greater than expected diversity of endophytes at high-latitude sites and provide a framework for assessing the evolution of these poorly known but ubiquitous symbionts of living plants.

Ascomycota↗

Combined mitochondrial and nuclear sequences support the monophyly of forcipulatacean sea stars.

Previous molecular phylogenetic analyses of forcipulatacean sea stars (Echinodermata: Asteroidea) have reconstructed a non-monophyletic order Forcipulatida, provided that two or more forcipulate families are included. This result could mean that one or more assumptions of the reconstruction method was violated, or else the traditional classification could be erroneous. The present molecular phylogenetic analysis included 12 non-forcipulatacean and 39 forcipulatacean sea stars, with multiple representatives of all but one of the forcipulate families and/or subfamilies. Bayesian analysis of approximately 4.2kb of sequence data representing seven partitions (nuclear 18S rRNA and 28S rRNA, mitochondrial 12S rRNA, 16S rRNA, 5 tRNAs and cytochrome oxidase I with first and second codon positions analyzed separately from third codon positions) recovered a consensus tree with three well-supported clades (78%-100% bootstrap support) that corresponded at least approximately to traditional taxonomic ranks: the superorder Forcipulatacea (Forcipulatida + Brisingida) + Pteraster, the Brisingida/Brisingidae and Asteriidae + Rathbunaster + Pycnopodia. When a molecular clock was enforced, the partitioned Bayesian analysis recovered the traditional Forcipulatacea. Five of six genera represented by two or more species were monophyletic with 100% bootstrap support. Most of the traditional subfamilial and familial groupings within the Forcipulatida were either unresolved or non-monophyletic. The separate partitions differed considerably in estimates of model parameters, mainly between nuclear sequences (with high GC content, low rates of sequence substitution and high transition/transversion rate ratios) and mitochondrial sequences.

Animals↗

On the calculation of intestinal schistosome infection intensity.

The effects of using different methods to calculate individual infection intensities on the age-infection distribution of Schistosoma mansoni field data are demonstrated. Methods are tested on a maximum of three stool samples per person collected on three consecutive days; the methods considered for the calculation of individual infection intensities are the geometric mean (GM), arithmetic mean (AM) and pseudo geometric mean (GM of stool samples instead of replicates). In addition, the effects of calculating the infection intensity for each age group using either AMs or GMs are compared. Differences occur in the shape of the age-infection profiles obtained by using either the arithmetic or geometric group mean. When using the AM, peak infection intensity occurs in a younger age group compared to using the GM, and all three methods of calculating individual infection intensity give the first peak of infection in the same age group. However, differences occur in the position of the second peak which occurs earlier with the two GMs than with the AM. Bootstrapping procedures show that the individual AM, gives a different age group for the first peak of infection at least 25% of the time when compared to either of the GMs, and 31% of the time for the second peak, while the two GMs give the same peak age groups around 90-92% of the time for both peaks. When using the GM, to calculate infection intensity for each age group, there are no differences between the three methods used to calculate individual infection intensity. This is confirmed by bootstrapping procedures. The results are discussed in relation to the distribution of parasites and levels of parasite aggregation. The implications of the results for field studies are also discussed.

Adolescent↗

Development and validation of a prediction model for strokes after coronary artery bypass grafting.

BACKGROUND: A prospective study of patients undergoing coronary artery bypass graft surgery (CABG) was conducted to identify patient and disease factors related to the development of a perioperative stroke. A preoperative risk prediction model was developed and validated based on regionally collected data. METHODS: We performed a regional observational study of 33,062 consecutive patients undergoing isolated CABG surgery in northern New England between 1992 and 2001. The regional stroke rate was 1.61% (532 strokes). We developed a preoperative stroke risk prediction model using logistic regression analysis, and validated the model using bootstrap resampling techniques. We assessed the model's fit, discrimination, and stability. RESULTS: The final regression model included the following variables: age, gender, presence of diabetes, presence of vascular disease, renal failure or creatinine greater than or equal to 2 mg/dL, ejection fraction less than 40%, and urgent or emergency. The model significantly predicted (chi(2) [14 d.f.] = 258.72, p < 0.0001) the occurrence of stroke. The correlation between the observed and expected strokes was 0.99. The risk prediction model discriminated well, with an area under the relative operating characteristic curve of 0.70 (95% CI, 0.67 to 0.72). In addition, the model had acceptable internal validity and stability as seen by bootstrap techniques. CONCLUSIONS: We developed a robust risk prediction model for stroke using seven readily obtainable preoperative variables. The risk prediction model performs well, and enables a clinician to estimate rapidly and accurately a CABG patient's preoperative risk of stroke.

Adult↗

Comparing non-hierarchical models: application to non-linear mixed effects modeling.

There is no method available to compare the fit of two non-hierarchical non-linear mixed effects models, although the common practice is to select the model with the lower objective function. Bootstrapping the log-likelihood differences (LLDs) of non-hierarchical models and constructing a bootstrap confidence interval on the LLDs is proposed for comparing the goodness-of-fit of such models. This is illustrated with different parameterizations of clearance models for an anti-infective agent in a longitudinal pharmacokinetic study which are compared. Additive and exponential models of creatinine clearance as a predictor of clearance are used as examples.

Adult↗

Aortic valve replacement: is valve size important?

OBJECTIVE: We sought to determine whether aortic prosthesis size adversely influences survival after aortic valve replacement. METHODS: A total of 892 adults receiving a mechanical (n = 346), pericardial (n = 463), or allograft (n = 83) valve for aortic stenosis were observed for up to 20 years (mean, 5.0 +/- 3.9 years) after primary isolated aortic valve replacement. We used multivariable propensity scores to adjust for valve selection factors, multivariable hazard function analyses to identify risk factors for all-cause mortality, and bootstrap resampling to quantify the reliability of the results. RESULTS: Twenty-five percent of patients had indexed internal orifice areas of less than 1.5 cm(2)/m(2) and more than 2 SDs (Z-value) below predicted normal aortic valve size. Mechanical valve orifices were smaller (1.3 +/- 0. 29 cm(2)/m(2), Z = -2.2 +/- 1.16) than pericardial (1.9 +/- 0.36 cm(2)/m(2), Z = -0.40 +/- 1.01) or allograft valves (2.1 +/- 0.50, Z = 0.24 +/- 1.17). The overall survival was 98%, 96%, 86%, 69%, and 49% at 30 days and 1, 5, 10, and 15 years postoperatively. Univariably, survival was weakly and inversely related to manufacturer valve size (P =.16) and internal orifice diameter (P =. 2) but completely unrelated to indexed valve area (P =.6) or Z-value (P =.8). These, and univariable differences among valve types (P =. 004), were accounted for by different prevalences in patient risk factors and not by valve size or type per se. Bootstrap resampling indicated that these findings had a less than 15% chance of being incorrect. CONCLUSIONS: Survival after aortic valve replacement is strongly related to patient risk factors but appears not to be adversely affected by moderate patient-prosthesis mismatch (down to about 4 SDs below normal). Aortic root enlargement to accommodate a large prosthesis may be required in few situations.

Adolescent↗

Validation of a nomogram for prediction of side specific extracapsular extension at radical prostatectomy.

PURPOSE: We have previously have reported a tree structured regression model for predicting SS-ECE. Others recently reported a logistic regression based SS-ECE nomogram. We developed a nomogram and compared the performance and discriminant properties of the tree regression and the nomogram in a contemporary cohort of European patients treated with radical retropubic prostatectomy. MATERIALS AND METHODS: The cohort consisted of 1,118 patients with pretreatment prostate specific antigen 0.1 to 73.2 ng/ml (median 6.6). Each of the 2,236 prostate lobes was considered separately. Clinical stage, pretreatment PSA, biopsy Gleason sum, percent positive cores and percent cancer in the biopsy specimen were used as predictors in a logistic regression model predicting SS-ECE. Regression coefficients were then used to generate an SS-ECE nomogram. Performance characteristics and discriminant properties of the previously published tree regression were also tested in the same cohort. For internal validation and to decrease overfit bias 200 bootstrap re-samples were applied to accuracy estimates for each method. RESULTS: ECE was present in 303 of 1,118 radical retropubic prostatectomy specimens (27%) and in 385 lobes (17%). In logistic regression models all variables were statistically significant multivariate predictors of SS-ECE except the percent of positive biopsy cores (p = 0.7). Bootstrap corrected predictive accuracy of the SS-ECE nomogram was 0.840 vs 0.700 for the tree regression model. CONCLUSIONS: Logistic regression based nomogram predictions of SS-ECE are highly accurate and represent a valuable aid for assessing the risk of ECE prior to surgery.

Calibration↗

A statistical procedure for the analysis of microbial communities based on phenotypic properties of isolates.

A novel statistical procedure for the analysis of microbial communities based on phenotypic properties of randomly collected isolates is presented and discussed. The procedure allows the representation of the microbial communities as a set of ellipses in a bidimensional graph. This representation is obtained by the following steps: (a) measurement of a set of binary phenotypic properties for n isolates belonging to k samples, each representing a different community; (b) repeated sampling by bootstrapping of the m samples, thus obtaining, for each community, i subsamples of j isolates; (c) calculation of the frequency of positive results for each test for each subsample; (d) calculation of the matrix of Euclidean distances between the k x i frequency vectors; (e) use of multidimensional scaling (MDS) to obtain a representation in two dimensions of the distance relationships between the frequency vectors; (f) plotting of the 95% confidence ellipses for the i frequency vectors of each of the k communities. By using both simple, synthetic microbial communities, and samples of lactic acid bacteria isolated from natural microbial communities (sourdoughs, compressed yeast, fermented sausages), it was demonstrated that the position and shape of the ellipses are clearly related to the composition of the community, while the relationship between the size of the ellipses and the phenotypical diversity of the community is less straightforward: while communities with very different diversity (measured with the Functional Evenness index and the mean taxonomic distance) had ellipses that were very different in size, there was no strict proportionality between the size of the ellipse and the diversity of the community. Nevertheless, the representation of microbial communities obtained by bootstrapping and multidimensional scaling appears to be superior to the more usual representation based on tabulation of the frequencies of isolates belonging to different clusters.

Animals↗

Sensitivity analysis for high quantiles of ochratoxin A exposure distribution.

Using available data from a consumption survey and contamination data on ochratoxin A (OA) in food, a sensitivity analysis (SA) for high quantiles (95th and 99th quantiles) of OA exposure distribution was carried out, obtained by a Monte Carlo simulation in French children. Exposure assessment for food contaminants is important to control the risk of foodborne diseases. Risk assessors are interested in high quantiles of contaminant exposure distributions. As these exposure distributions are generally very asymmetrical, it is difficult to obtain relevant and stable high quantiles in such a context. Determining OA exposure distribution is complex because it is based on the sum of elementary exposure distributions (eight foodstuffs are analysed here), and each one of these is the product of a consumption distribution and a contamination distribution. The SA enables us to quantify the influences of the parameter variability of the consumption and contamination probability density functions (pdf) which have been fitted to the data, our simulation model inputs, on the 95th and 99th quantiles of the output exposure distribution. After some preliminary trials, we have postulated a quadratic polynomial regression model for the quantiles of OA exposure distribution in view of undertaking this SA. This regression model comprises 32 main factors, their 496 two-factor interactions and their 32 quadratic terms. The 32 factors are the parameters of the fitted pdf: 16 parameters of Gamma distributions relative to the eight consumed foods and 16 parameters of Gamma distributions relative to the eight food OA contaminations. For an optimal parameter estimation of such a large model, we used an experimental design approach depending on a resolution-V fractional factorial design of 6561 experiments. The factor ranges are established by a preliminary study of bootstrap sampling. From the bootstrap samples, the factor ranges are obtained taking into account the correlation between the two parameters of the fitted Gamma pdf. A full exposure distribution is simulated for each of the 6561 experiments. The consumption dependencies are taken into account by the Iman and Conover method. On the basis of this analysis, validated and useful models for each desired quantile are obtained showing a major influence of the parameters of "Cereals" (consumption and contamination) and slightly less so for parameter of "Pork" consumption in the sensitivity of the quantiles.

Adolescent↗

Artificial neural networks for infant mortality modelling.

This work aims to investigate a simple to use and easy to interpret methodology for assessing the relative importance of input variables in artificial neural networks (ANNs) applied to epidemiological modelling. The independent variables were 43 variables of the social, economic, environmental and health sector of 59 Brazilian municipalities, and the outcomes were infant mortality rates from these municipalities. Two assays were developed for the ANN modelling. On the first, all 43 variables were taken as input; and on the second, input variables were chosen with the help of factor analysis (FA). The relative importance of the input variables was investigated by means of bootstrap replications of the ANN model on the second assay. Further, multiple linear regression models (LRMs) were developed with the same data set and compared to the ANN models. The FA analysis allowed the selection of eight variables for the second assay. The percent of explained variance R(2) on the ANNs was in the range 0.74-0.80, while linear models had R(2)=0.4-0.5. These findings were validated by the bootstrap replications, in which the ANN models remained with higher R(2) and lower mean square error than the LRMs. The analysis of the best (second) ANN model indicated the highest ranking of importance for the variables literacy, agricultural and livestock sector jobs, number of commercial establishments and telephones. The approach presented here successfully integrated a data-oriented model with expert knowledge, indicating the potentiality of ANN modelling in the prediction, planning and assessment of public health actions.

Brazil↗

COMPROC and CHECKNORM: computer programs for comparing accuracies of diagnostic tests using ROC curves in the presence of verification bias.

To assess relative accuracies of two diagnostic tests, we often compare the areas under the receiver operating characteristic (ROC) curves of these two tests in a paired design. Standard methods for analyzing data from a paired design require that every patient tested has the known disease status. In practice, however, some of the patients with test results may not have verified disease status. Any analysis using only verified cases may result in verification bias. COMPROC is an easy to use program for comparing the effectiveness of two diagnostic tests based on the area under the ROC curve in the presence of verification bias. COMPROC compensates for verification bias by implementing the maximum likelihood (ML) estimation of the areas and covariance matrix of two ROC curves under the missing at random (MAR) assumption as described by Zhou (Biometrics 54 (1998) 349-366). This method assumes normality of the difference of the two ROC curve area estimators. We also describe a program CHECKNORM that does a bootstrap analysis to test this normality assumption (B. Efron, R.J. Tibshirani, An Introduction to the Bootstrap, Chapman and Hall, London, 1993). COMPROC allows for the inclusion of observed covariates that may influence the decision to verify the disease status of a patient. The program computes the estimates of the area under the ROC curve for the two diagnostic tests along with the variance of each area, the covariance between the two areas, a two-sided p-value, and a confidence interval for the difference of the areas. The programs COMPROC and CHECKNORM require the scripting language Perl and the statistical software SAS and can be run on both UNIX machines as well as PCs. The use of COMPROC and CHECKNORM is illustrated in a clinical study designed to compare relative accuracies of MRI and CT in detecting pancreatic cancer.

Area Under Curve↗