PubMed Health⌕ Search

Biomedical subjects

Paul Joyce

Publications and source records attributed to Paul Joyce.

16 recordsLinked to original sources

The population biology of bacterial plasmids: a hidden Markov model approach.

Horizontal plasmid transfer plays a key role in bacterial adaptation. In harsh environments, bacterial populations adapt by sampling genetic material from a horizontal gene pool through self-transmissible plasmids, and that allows persistence of these mobile genetic elements. In the absence of selection for plasmid-encoded traits it is not well understood if and how plasmids persist in bacterial communities. Here we present three models of the dynamics of plasmid persistence in the absence of selection. The models consider plasmid loss (segregation), plasmid cost, conjugative plasmid transfer, and observation error. Also, we present a stochastic model in which the relative fitness of the plasmid-free cells was modeled as a random variable affected by an environmental process using a hidden Markov model (HMM). Extensive simulations showed that the estimates from the proposed model are nearly unbiased. Likelihood-ratio tests showed that the dynamics of plasmid persistence are strongly dependent on the host type. Accounting for stochasticity was necessary to explain four of seven time-series data sets, thus confirming that plasmid persistence needs to be understood as a stochastic process. This work can be viewed as a conceptual starting point under which new plasmid persistence hypotheses can be tested.

Bacteria↗

Properties of adaptive walks on uncorrelated landscapes under strong selection and weak mutation.

We examine properties of adaptive walks on uncorrelated (i.e. random) fitness landscapes starting from moderately fit genotypes under strong selection weak mutation. As an extension of Orr's model for a single step in an adaptive walk under these conditions, we show that the fitness rank of the dominant genotype in a population after the fixation of a beneficial mutation is, on average, (i+6)/4, where i is the fitness rank of the starting genotype. This accounts for the change in rank due to acquiring a new set of single-mutation neighbors after fixing a new allele through natural selection. Under this scenario, adaptive walks can be modeled as a simple Markov chain on the space of possible fitness ranks with an absorbing state at i = 1, from which no beneficial mutations are accessible. We find that these walks are typically short and are often completed in a single step when starting from a moderately fit genotype. As in Orr's original model, these results are insensitive to both the distribution of fitness effects and most biological details of the system under consideration.

Adaptation, Biological↗

Statistical methods for characterizing diversity of microbial communities by analysis of terminal restriction fragment length polymorphisms of 16S rRNA genes.

The analysis of terminal restriction fragment length polymorphisms (T-RFLP) of 16S rRNA genes has proven to be a facile means to compare microbial communities and presumptively identify abundant members. The method provides data that can be used to compare different communities based on similarity or distance measures. Once communities have been clustered into groups, clone libraries can be prepared from sample(s) that are representative of each group in order to determine the phylogeny of the numerically abundant populations in a community. In this paper methods are introduced for the statistical analysis of T-RFLP data that include objective methods for (i) determining a baseline so that 'true' peaks in electropherograms can be identified; (ii) a means to compare electropherograms and bin fragments of similar size; (iii) clustering algorithms that can be used to identify communities that are similar to one another; and (iv) a means to select samples that are representative of a cluster that can be used to construct 16S rRNA gene clone libraries. The methods for data analysis were tested using simulated data with assumptions and parameters that corresponded to actual data. The simulation results demonstrated the usefulness of these methods in their ability to recover the true microbial community structure generated under the assumptions made. Software for implementing these methods is available at http://www.ibest.uidaho.edu/tools/trflp_stats/index.php.

Bacteria↗

Developing the electronic health record: what about patient safety?

This paper examines the development of electronic health records within the National Health Service (NHS) by an analysis of a series of pilot projects funded by the Electronic Record Development and Implementation Project (ERDIP), one aspect of the work of the NHS Information Authority (NHSIA) (As of 1 April 2005, the NHSIA ceased to operate. Much of its work is continued by Connecting for Health and the Health and Social Care Information Centre.) The focus of the analysis is on the extent to which identifying and correcting error within health records was explored through these projects. The inherent potential for error and resultant impact on patient safety is highlighted, by considering the context of the record, the content of the record and the process of change from paper-based or piecemeal electronic health records to integrated electronic health records. While the process of change highlights issues of data security and access, it is the variability in starting points for different organizations that possibly poses most risk to patient safety. Issues relating to the content of the record can to some extent be minimized by the effective use of technology, but the tension between coding and qualitative data requires further consideration in terms of its impact on patient safety. This paper concludes that the development of electronic health records has to be viewed within the context of governance and patient safety, and the implications articulated.

Diffusion of Innovation↗

An empirical test of the mutational landscape model of adaptation using a single-stranded DNA virus.

The primary impediment to formulating a general theory for adaptive evolution has been the unknown distribution of fitness effects for new beneficial mutations. By applying extreme value theory, Gillespie circumvented this issue in his mutational landscape model for the adaptation of DNA sequences, and Orr recently extended Gillespie's model, generating testable predictions regarding the course of adaptive evolution. Here we provide the first empirical examination of this model, using a single-stranded DNA bacteriophage related to phiX174, and find that our data are consistent with Orr's predictions, provided that the model is adjusted to incorporate mutation bias. Orr's work suggests that there may be generalities in adaptive molecular evolution that transcend the biological details of a system, but we show that for the model to be useful as a predictive or inferential tool, some adjustments for the biology of the system will be necessary.

Adaptation, Biological↗

Evaluating the performance of a successive-approximations approach to parameter optimization in maximum-likelihood phylogeny estimation.

Almost all studies that estimate phylogenies from DNA sequence data under the maximum-likelihood (ML) criterion employ an approximate approach. Most commonly, model parameters are estimated on some initial phylogenetic estimate derived using a rapid method (neighbor-joining or parsimony). Parameters are then held constant during a tree search, and ideally, the procedure is repeated until convergence is achieved. However, the effectiveness of this approximation has not been formally assessed, in part because doing so requires computationally intensive, full-optimization analyses. Here, we report both indirect and direct evaluations of the effectiveness of successive approximations. We obtained an indirect evaluation by comparing the results of replicate runs on real data that use random trees to provide initial parameter estimates. For six real data sets taken from the literature, all replicate iterative searches converged to the same joint estimates of topology and model parameters, suggesting that the approximation is not starting-point dependent, as long as the heuristic searches of tree space are rigorous. We conducted a more direct assessment using simulations in which we compared the accuracy of phylogenies estimated using full optimization of all model parameters on each tree evaluated to the accuracy of trees estimated via successive approximations. There is no significant difference between the accuracy of the approximation searches relative to full-optimization searches. Our results demonstrate that successive approximation is reliable and provide reassurance that this much faster approach is safe to use for ML estimation of topology.

Algorithms↗

Managing risk: a taxonomy of error in health policy.

This paper discusses the current initiatives on error and adverse events within healthcare, with a particular focus on the NHS, within the context of health policy. One of the key features of the paper is the proposal for an emergent taxonomy of the medical error literature, developed from the ideologies and rationales that underpin their approaches. This taxonomy provides details of three categories--empiricists, organisational rationalists and reformers of professional culture--and these act as an organising framework for the exploration of the potential consequences of current policy on errors and adverse events. This discussion highlights the tension between optimising health outcomes for patients and managing the health system as effectively as possible. In particular, the inherent tension between explicit managerial formulations of risk and implicit risk management strategies associated with medical professionalism are considered.

Classification↗

A new method for estimating the size of small populations from genetic mark-recapture data.

The use of non-invasive genetic sampling to estimate population size in elusive or rare species is increasing. The data generated from this sampling differ from traditional mark-recapture data in that individuals may be captured multiple times within a session or there may only be a single sampling event. To accommodate this type of data, we develop a method, named capwire, based on a simple urn model containing individuals of two capture probabilities. The method is evaluated using simulations of an urn and of a more biologically realistic system where individuals occupy space, and display heterogeneous movement and DNA deposition patterns. We also analyse a small number of real data sets. The results indicate that when the data contain capture heterogeneity the method provides estimates with small bias and good coverage, along with high accuracy and precision. Performance is not as consistent when capture rates are homogeneous and when dealing with populations substantially larger than 100. For the few real data sets where N is approximately known, capwire's estimates are very good. We compare capwire's performance to commonly used rarefaction methods and to two heterogeneity estimators in program capture: Mh-Chao and Mh-jackknife. No method works best in all situations. While less precise, the Chao estimator is very robust. We also examine how large samples should be to achieve a given level of accuracy using capwire. We conclude that capwire provides an improved way to estimate N for some DNA-based data sets.

Computer Simulation↗

Use of stochastic models to assess the effect of environmental factors on microbial growth.

We present a novel application of a stochastic ecological model to the study and analysis of microbial growth dynamics as influenced by environmental conditions in an extensive experimental data set. The model proved to be useful in bridging the gap between theoretical ideas in ecology and an applied problem in microbiology. The data consisted of recorded growth curves of Escherichia coli grown in triplicate in a base medium with all 32 possible combinations of five supplements: glucose, NH(4)Cl, HCl, EDTA, and NaCl. The potential complexity of 2(5) experimental treatments and their effects was reduced to 2(2) as just the metal chelator EDTA, the presumed osmotic pressure imposed by NaCl, and the interaction between these two factors were enough to explain the variability seen in the data. The statistical analysis showed that the positive and negative effects of the five chemical supplements and their combinations were directly translated into an increase or decrease in time required to attain stationary phase and the population size at which the stationary phase started. The stochastic ecological model proved to be useful, as it effectively explained and summarized the uncertainty seen in the recorded growth curves. Our findings have broad implications for both basic and applied research and illustrate how stochastic mathematical modeling coupled with rigorous statistical methods can be of great assistance in understanding basic processes in microbial ecology.

Edetic Acid↗

Modeling the impact of periodic bottlenecks, unidirectional mutation, and observational error in experimental evolution.

Antibiotic resistant bacteria are a constant threat in the battle against infectious diseases. One strategy for reducing their effect is to temporarily discontinue the use of certain antibiotics in the hope that in the absence of the antibiotic the resistant strains will be replaced by the sensitive strains. An experiment where this strategy is employed in vitro produces data which showed a slow accumulation of sensitive mutants. Here we propose a mathematical model and statistical analysis to explain this data. The stochastic model elucidates the trend and error structure of the data. It provides a guide for developing future sampling strategies, and provides a framework for long term predictions of the effects of discontinuing specific antibiotics on the dynamics of resistant bacterial populations.

Bacteria↗

Accounting for uncertainty in the tree topology has little effect on the decision-theoretic approach to model selection in phylogeny estimation.

Currently available methods for model selection used in phylogenetic analysis are based on an initial fixed-tree topology. Once a model is picked based on this topology, a rigorous search of the tree space is run under that model to find the maximum-likelihood estimate of the tree (topology and branch lengths) and the maximum-likelihood estimates of the model parameters. In this paper, we propose two extensions to the decision-theoretic (DT) approach that relax the fixed-topology restriction. We also relax the fixed-topology restriction for the Bayesian information criterion (BIC) and the Akaike information criterion (AIC) methods. We compare the performance of the different methods (the relaxed, restricted, and the likelihood-ratio test [LRT]) using simulated data. This comparison is done by evaluating the relative complexity of the models resulting from each method and by comparing the performance of the chosen models in estimating the true tree. We also compare the methods relative to one another by measuring the closeness of the estimated trees corresponding to the different chosen models under these methods. We show that varying the topology does not have a major impact on model choice. We also show that the outcome of the two proposed extensions is identical and is comparable to that of the BIC, Extended-BIC, and DT. Hence, using the simpler methods in choosing a model for analyzing the data is more computationally feasible, with results comparable to the more computationally intensive methods. Another outcome of this study is that earlier conclusions about the DT approach are reinforced. That is, LRT, Extended-AIC, and AIC result in more complicated models that do not contribute to the performance of the phylogenetic inference, yet cause a significant increase in the time required for data analysis.

Computational Biology↗

Evaluating the performance of likelihood methods for detecting population structure and migration.

A plethora of statistical models have recently been developed to estimate components of population genetic history. Very few of these methods, however, have been adequately evaluated for their performance in accurately estimating population genetic parameters of interest. In this paper, we continue a research program of evaluation of population genetic methods through computer simulation. Specifically, we examine the software MIGRATEE-N 1.6.8 and test the accuracy of this software to estimate genetic diversity (Theta), migration rates, and confidence intervals. We simulated nucleotide sequence data under a neutral coalescent model with lengths of 500 bp and 1000 bp, and with three different per site Theta values of (0.00025, 0.0025, 0.025) crossed with four different migration rates (0.0000025, 0.025, 0.25, 2.5) to construct 1000 evolutionary trees per-combination per-sequence-length. We found that while MIGRATEE-N 1.6.8 performs reasonably well in estimating genetic diversity (Theta), it does poorly at estimating migration rates and the confidence intervals associated with them. We recommend researchers use this software with caution under conditions similar to those used in this evaluation.

Computer Simulation↗

Combining mathematical models and statistical methods to understand and predict the dynamics of antibiotic-sensitive mutants in a population of resistant bacteria during experimental evolution.

Temporarily discontinuing the use of antibiotics has been proposed as a means to eliminate resistant bacteria by allowing sensitive clones to sweep through the population. In this study, we monitored a tetracycline-sensitive subpopulation that emerged during experimental evolution of E. coli K12 MG1655 carrying the multiresistance plasmid pB10 in the absence of antibiotics. The fraction of tetracycline-sensitive mutants increased slowly over 500 generations from 0.1 to 7%, and loss of resistance could be attributed to a recombination event that caused deletion of the tet operon. To help understand the population dynamics of these mutants, three mathematical models were developed that took into consideration recurrent mutations, increased host fitness (selection), or a combination of both mechanisms (full model). The data were best explained by the full model, which estimated a high mutation frequency (lambda = 3.11 x 10(-5)) and a significant but small selection coefficient (sigma = 0.007). This study emphasized the combined use of experimental data, mathematical models, and statistical methods to better understand and predict the dynamics of evolving bacterial populations, more specifically the possible consequences of discontinuing the use of antibiotics.

Anti-Bacterial Agents↗

Performance-based selection of likelihood models for phylogeny estimation.

Phylogenetic estimation has largely come to rely on explicitly model-based methods. This approach requires that a model be chosen and that that choice be justified. To date, justification has largely been accomplished through use of likelihood-ratio tests (LRTs) to assess the relative fit of a nested series of reversible models. While this approach certainly represents an important advance over arbitrary model selection, the best fit of a series of models may not always provide the most reliable phylogenetic estimates for finite real data sets, where all available models are surely incorrect. Here, we develop a novel approach to model selection, which is based on the Bayesian information criterion, but incorporates relative branch-length error as a performance measure in a decision theory (DT) framework. This DT method includes a penalty for overfitting, is applicable prior to running extensive analyses, and simultaneously compares all models being considered and thus does not rely on a series of pairwise comparisons of models to traverse model space. We evaluate this method by examining four real data sets and by using those data sets to define simulation conditions. In the real data sets, the DT method selects the same or simpler models than conventional LRTs. In order to lend generality to the simulations, codon-based models (with parameters estimated from the real data sets) were used to generate simulated data sets, which are therefore more complex than any of the models we evaluate. On average, the DT method selects models that are simpler than those chosen by conventional LRTs. Nevertheless, these simpler models provide estimates of branch lengths that are more accurate both in terms of relative error and absolute error than those derived using the more complex (yet still wrong) models chosen by conventional LRTs. This method is available in a program called DT-ModSel.

Bayes Theorem↗

Assessing allelic dropout and genotype reliability using maximum likelihood.

A growing number of population genetic studies utilize nuclear DNA microsatellite data from museum specimens and noninvasive sources. Genotyping errors are elevated in these low quantity DNA sources, potentially compromising the power and accuracy of the data. The most conservative method for addressing this problem is effective, but requires extensive replication of individual genotypes. In search of a more efficient method, we developed a maximum-likelihood approach that minimizes errors by estimating genotype reliability and strategically directing replication at loci most likely to harbor errors. The model assumes that false and contaminant alleles can be removed from the dataset and that the allelic dropout rate is even across loci. Simulations demonstrate that the proposed method marks a vast improvement in efficiency while maintaining accuracy. When allelic dropout rates are low (0-30%), the reduction in the number of PCR replicates is typically 40-50%. The model is robust to moderate violations of the even dropout rate assumption. For datasets that contain false and contaminant alleles, a replication strategy is proposed. Our current model addresses only allelic dropout, the most prevalent source of genotyping error. However, the developed likelihood framework can incorporate additional error-generating processes as they become more clearly understood.

Alleles↗