PubMed Health⌕ Search

Biomedical subjects

G A Churchill

Publications and source records attributed to G A Churchill.

At least 19 recordsLinked to original sources

Statistical aspects of genetic mapping in autopolyploids.

Many plant species of agriculture importance are polyploid, having more than two copies of each chromosome per cell. In this paper, we describe statistical methods for genetic map construction in autopolyploid species with particular reference to the use of molecular markers. The first step is to determine the dosage of each DNA fragment (electrophoretic band) from its segregation ratio. Fragments present in a single dose can be used to construct framework maps for individual chromosomes. Fragments present in multiple doses can often be used to link the single chromosome maps into homologous groups and provide additional ordering information. Marker phenotype probabilities were calculated for pairs of markers arranged in different configurations among the homologous chromosomes. These probabilities were used to compute a maximum likelihood estimator of the recombination fraction between pairs of markers. A likelihood ratio test for linkage of multidose markers was derived. The information provided by each configuration and power and sample size considerations are also discussed. A set of 294 RFLP markers scored on 90 plants of the species Saccharum spontaneum L. was used to illustrate the construction of an autopolyploid map. Previous studies conducted on the same data revealed that this species of sugar cane is an autooctaploid with 64 chromosomes arranged into eight homologous groups. The methodology described permitted consolidation of 54 linkage groups into ten homologous groups.

Chromosomes↗

Genetic analysis of susceptibility to dextran sulfate sodium-induced colitis in mice.

The genetic basis for differential sensitivity of inbred mice to inflammatory bowel disease induced by dextran sulfate sodium (DSS) is unknown. Susceptible C3H/HeJ were outcrossed to partially resistant C57BL/6J mice. F2 and N2 progeny were phenotyped by evaluating histopathologic lesions in large intestine detected 16 days after a 5-day period of feeding 3.5% DSS. Screening for DSS colitis (Dssc) loci revealed quantitative trait loci (QTL) on Chr 5 (Dssc1) and Chr 2 (Dssc2). These traits contributed additively, explaining 17.5% of the variation in total colonic lesions. Additional QTL on Chr 18 and 1 that collectively explained 11% of the variation in total colon lesions were indicated. In the cecum, only a putative QTL on Chr 11 was associated with pathology (lesion severity) in the cecum. Reduced DSS susceptibility was observed in congenic stocks in which the highly susceptible NOD/Lt strain carried putative resistance alleles from either B6 on Chr 2 or from the highly resistant NON/Lt strain on Chr 9. We conclude that multiple genes control susceptibility to DSS colitis in mice. Possible Dssc candidate genes are discussed in terms of current knowledge of inflammatory bowel disease susceptibility loci in humans.

Animals↗

Quantitative trait loci for bone density in C57BL/6J and CAST/EiJ inbred mice.

Genetic analyses for loci regulating bone mineral density have been conducted in a cohort of F(2) mice derived from intercross matings of (C57BL/6J x CAST/EiJ)F(1) parents. Femurs were isolated from 714 4-month-old females when peak adult bone density had been achieved. Bone mineral density (BMD) data were obtained by peripheral quantitative computed tomography (pQCT), and genotype data were obtained by Polymerase Chain Reaction (PCR) assays for polymorphic markers carried in genomic DNA of each mouse. Genome-wide scans for co-segregation of genetic marker data with high or low BMD revealed loci on eight different chromosomes, four of which (Chrs 1, 5, 13, and 15) achieved conservative statistical criteria for suggestive, significant, or highly significant linkage with BMD. These four quantitative trait loci (QTLs) were confirmed by a linear regression model developed to describe the main effects; none of the loci exhibited significant interaction effects by ANOVA. The four QTLs have been named Bmd1 (Chr 1), Bmd2 (Chr 5), Bmd3 (Chr 13), and Bmd4 (Chr 15). Additive effects were observed for Bmd1, recessive for Bmd3, and dominant effects for Bmd2 and Bmd4. The current large size of the QTL regions (6-->31 cM) renders premature any discussion of candidate genes at this time. Fine mapping of these QTLs is in progress to refine their genetic positions and to evaluate human homologies.

Age Factors↗

Bayesian restoration of a hidden Markov chain with applications to DNA sequencing.

Hidden Markov models (HMMs) are a class of stochastic models that have proven to be powerful tools for the analysis of molecular sequence data. A hidden Markov model can be viewed as a black box that generates sequences of observations. The unobservable internal state of the box is stochastic and is determined by a finite state Markov chain. The observable output is stochastic with distribution determined by the state of the hidden Markov chain. We present a Bayesian solution to the problem of restoring the sequence of states visited by the hidden Markov chain from a given sequence of observed outputs. Our approach is based on a Monte Carlo Markov chain algorithm that allows us to draw samples from the full posterior distribution of the hidden Markov chain paths. The problem of estimating the probability of individual paths and the associated Monte Carlo error of these estimates is addressed. The method is illustrated by considering a problem of DNA sequence multiple alignment. The special structure for the hidden Markov model used in the sequence alignment problem is considered in detail. In conclusion, we discuss certain interesting aspects of biological sequence alignments that become accessible through the Bayesian approach to HMM restoration.

Algorithms↗

Multigenic and imprinting control of ovarian granulosa cell tumorigenesis in mice.

Spontaneous juvenile ovarian granulosa cell (GC) tumors that occur in young girls are similar to GC carcinomas that develop in SWR-derived inbred mice. We analyzed female offspring from a series of matings among SWR and SJL inbred mice for chromosomal loci underlying tumor susceptibility. Intercross F2 female mice were produced by reciprocal matings of (SWR x SJL)F1 and (SJL x SWR)F1 parents. Tumorigenesis in these F2 mice as well as in SWXJ recombinant inbred and congenic strains of mice derived from SWR and SJL showed significant (P < 0.001) association with Gct1, a dominant susceptibility locus on chromosome (CHR) 4 and with Gct2 on CHR 12. Suggestive (P < 0.01) association was found with Gct3 on CHR 15. A fourth susceptibility locus, Gct4 on CHR X, was demonstrated with a strong parent-of-origin effect associated with the paternal genotype. Imprinting and complex interactions among these four loci combine to establish the probability for GC tumorigenesis in this mouse model.

Alleles↗

Biases in amino acid replacement matrices and alignment scores due to rate heterogeneity.

Empirically derived amino acid replacement matrices are widely used in sequence comparison and database searches. We consider an extension of the usual Markov process model of protein evolution that admits site to site rate heterogeneity and demonstrates that rate heterogeneity can introduce a bias in estimated replacement probabilities and the corresponding alignment scores derived from these matrices. We suggest an approach to obtain unbiased estimates of replacement probabilities and alignment scores and derive the details for the case where rates are assumed to vary according to a gamma distribution.

Amino Acid Sequence↗

Permutation tests for multiple loci affecting a quantitative character.

The problem of detecting minor quantitative trait loci (QTL) responsible for genetic variation not explained by major QTL is of importance in the complete dissection of quantitative characters. Two extensions of the permutation-based method for estimating empirical threshold values are presented. These methods, the conditional empirical threshold (CET) and the residual empirical threshold (RET), yield critical values that can be used to construct tests for the presence of minor QTL effects while accounting for effects of known major QTL. The CET provides a completely nonparametric test through conditioning on markers linked to major QTL. It allows for general nonadditive interactions among QTL, but its practical application is restricted to regions of the genome that are unlinked to the major QTL. The RET assumes a structural model for the effect of major QTL, and a threshold is constructed using residuals from this structural model. The search space for minor QTL is unrestricted, and RET-based tests may be more powerful than the CET-based test when the structural model is approximately true.

Alleles↗

Heterogeneity in rates of recombination across the mouse genome.

If loci are randomly distributed on a physical map, the density of markers on a genetic map will be inversely proportional to recombination rate. First, proposed by Mary Lyon, we have used this idea to estimate recombination rates from the Drosophila melanogaster linkage map. These results were compared with results of two other studies that estimated regional recombination rates in D. melanogaster using both physical and genetic maps. The three methods were largely concordant in identifying large-scale genomic patterns of recombination. The marker density method was then applied to the Mus musculus microsatellite linkage map. The distribution of microsatellites provided evidence for heterogeneity in recombination rates. Centromeric regions for several mouse chromosomes had significantly greater numbers of markers than expected, suggesting that recombination rates were lower in these regions. In contrast, most telomeric regions contained significantly fewer markers than expected. This indicates that recombination rates are elevated at the telomeres of many mouse chromosomes and is consistent with a comparison of the genetic and cytogenetic maps in these regions. The density of markers on a genetic map may provide a generally useful way to estimate regional recombination rates in species for which genetic, but not physical, maps are available.

Animals↗

A Hidden Markov Model approach to variation among sites in rate of evolution.

The method of Hidden Markov Models is used to allow for unequal and unknown evolutionary rates at different sites in molecular sequences. Rates of evolution at different sites are assumed to be drawn from a set of possible rates, with a finite number of possibilities. The overall likelihood of phylogeny is calculated as a sum of terms, each term being the probability of the data given a particular assignment of rates to sites, times the prior probability of that particular combination of rates. The probabilities of different rate combinations are specified by a stationary Markov chain that assigns rate categories to sites. While there will be a very large number of possible ways of assigning rates to sites, a simple recursive algorithm allows the contributions to the likelihood from all possible combinations of rates to be summed, in a time proportional to the number of different rates at a single site. Thus with three rates, the effort involved is no greater than three times that for a single rate. This "Hidden Markov Model" method allows for rates to differ between sites and for correlations between the rates of neighboring sites. By summing over all possibilities it does not require us to know the rates at individual sites. However, it does not allow for correlation of rates at nonadjacent sites, nor does it allow for a continuous distribution of rates over sites. It is shown how to use the Newton-Raphson method to estimate branch lengths of a phylogeny and to infer from a phylogeny what assignment of rates to sites has the largest posterior probability. An example is given using beta-hemoglobin DNA sequences in eight mammal species; the regions of high and low evolutionary rates are inferred and also the average length of patches of similar rates.

Animals↗

Properties of statistical tests of neutrality for DNA polymorphism data.

A class of statistical tests based on molecular polymorphism data is studied to determine size and power properties. The class includes Tajima's D statistic as well as the D* and F* tests proposed by Fu and Li. A new method of constructing critical values for these tests is described. Simulations indicate that Tajima's test is generally most powerful against the alternative hypotheses of selective sweep, population bottleneck, and population subdivision, among tests within this class. However, even Tajima's test can detect a selective sweep or bottleneck only if it has occurred within a specific interval of time in the recent past or population subdivision only when it has persisted for a very long time. For greatest power against the particular alternatives studied here, it is better to sequence more alleles than more sites.

Computer Simulation↗

Estimation and reliability of molecular sequence alignments.

The problem of estimating the relatedness of a pair of biological sequences is addressed. A stochastic model of sequence evolution is described that allows insertion and deletion as well as replacement of amino acid residues (or substitution of nucleotides) over time. An expectation-maximization (EM) algorithm that obtains maximum likelihood estimates of the model parameters is introduced. The method assumes that the sequences are related by descent from a common ancestor but the alignment (i.e., the precise evolutionary correspondence between residues in each sequence) is unknown. Results from the E-step of the EM algorithm are used to assess the likelihood that any two residues are related by direct descent from a common ancestor.

Algorithms↗

Empirical threshold values for quantitative trait mapping.

The detection of genes that control quantitative characters is a problem of great interest to the genetic mapping community. Methods for locating these quantitative trait loci (QTL) relative to maps of genetic markers are now widely used. This paper addresses an issue common to all QTL mapping methods, that of determining an appropriate threshold value for declaring significant QTL effects. An empirical method is described, based on the concept of a permutation test, for estimating threshold values that are tailored to the experimental data at hand. The method is demonstrated using two real data sets derived from F(2) and recombinant inbred plant populations. An example using simulated data from a backcross design illustrates the effect of marker density on threshold values.

Chromosome Mapping↗

Pooled-sampling makes high-resolution mapping practical with DNA markers.

A pooled-sample approach to the construction of high-resolution genetic maps is described. The strategy depends on the existence of an easily selectable target locus and the ability to produce large segregating populations. If these requirements are met, the pooled-sample mapping approach allows tightly linked markers (e.g., restriction fragment length polymorphisms) to be mapped relative to the target with a great economy of effort. The recombination fractions among loci can be estimated by the maximum likelihood method and a simple approximate estimator is derived. The order of loci is deduced using a Bayesian statistical framework to yield posterior probabilities for all possible orderings of a marker set. Optimal pooling strategies and the effects of misclassification of selected individuals are discussed and studied by computer simulation. The feasibility of this method is demonstrated by the high-resolution mapping of a region on chromosome 5 of tomato that contains a gene regulating fruit ripening.

Animals↗

Network models for sequence evolution.

We introduce a general class of models for sequence evolution that includes network phylogenies. Networks, a generalization of strictly tree-like phylogenies, are proposed to model situations where multiple lineages contribute to the observed sequences. An algorithm to compute the probability distribution of binary character-state configurations is presented and statistical inference for this model is developed in a likelihood framework. A stepwise procedure based on likelihood ratios is used to explore the space of models. Starting with a star phylogeny, new splits (nontrivial bipartitions of the sequence set) are successively added to the model until no significant change in the likelihood is observed. A novel feature of our approach is that the new splits are not necessarily constrained to be consistent with a treelike mode of evolution. The fraction of invariable sites is estimated by maximum likelihood simultaneously with other model parameters and is essential to obtain a good fit to the data. The effect of finite sequence length on the inference methods is discussed. Finally, we provide an illustrative example using aligned VP1 genes from the foot and mouth disease viruses (FMDV). The different serotypes of the FMDV exhibit a range of treelike and network evolutionary relationships.

Aphthovirus↗

Phylogenetic inference: linear invariants and maximum likelihood.

We develop a new statistical method for inferring phylogenies, based on a likelihood ratio test. This method does not require parameter constraints but does require identical evolutionary processes in the sites considered. Another method of phylogenetic inference is the method of linear invariants, described by Cavender (1989, Molecular Biology and Evolution 6, 301-316), based on a notion of Lake (1987, Molecular Biology and Evolution 4, 167-191). We describe a sound mathematical basis for the use of linear invariants. We show that the validity of the method requires parameter constraints, but does not require that the evolutionary processes in differing sites be identical. We show that the method of linear invariants is asymptotically equivalent to a less powerful version of our likelihood ratio test, and is thus essentially a maximum likelihood technique.

Animals↗

The accuracy of DNA sequences: estimating sequence quality.

In this paper we describe a method for the statistical reconstruction of a large DNA sequence from a set of sequenced fragments. We assume that the fragments have been assembled and address the problem of determining the degree to which the reconstructed sequence is free from errors, i.e., its accuracy. A consensus distribution is derived from the assembled fragment configuration based upon the rates of sequencing errors in the individual fragments. The consensus distribution can be used to find a minimally redundant consensus sequence that meets a prespecified confidence level, either base by base or across any region of the sequence. A likelihood-based procedure for the estimation of the sequencing error rates, which utilizes an iterative EM algorithm, is described. Prior knowledge of the error rates is easily incorporated into the estimation procedure. The methods are applied to a set of assembled sequence fragments from the human G6PD locus. We close the paper with a brief discussion of the relevance and practical implications of this work.

Algorithms↗

Sample size for a phylogenetic inference.

The objective of this work is to describe sample-size calculations for the inference of a nonzero central branch length in an unrooted four-species phylogeny. Attention is restricted to independent binary characters, such as might be obtained from an alignment of the purine-pyrimidine sequences of a nucleic acid molecule. A statistical test based on a multinomial model for character-state configurations is described. The importance of including invariable sites in models for sequence change is demonstrated, and their effect on sample size is quantified. The methods are applied to a four-species alignment of small-subunit rRNA sequences derived from two archaebacteria, a eubacteria and a eukaryote. We conclude that the information in these sequences is not sufficient to resolve the branching order of this tree. Estimates of the number of aligned nucleotide positions required to provide a reasonably powerful test are given.

Archaea↗

Methods for inferring phylogenies from nucleic acid sequence data by using maximum likelihood and linear invariants.

Likelihood methods and methods using invariants are procedures for inferring the evolutionary relationships among species through statistical analysis of nucleic acid sequences. A likelihood-ratio test may be used to determine the feasibility of any tree for which the maximum likelihood can be computed. The method of linear invariants described by Cavender, which includes Lake's method of evolutionary parsimony as a special case, is essentially a form of the likelihood-ratio method. In the case of a small number of species (four or five), these methods may be used to find a confidence set for the correct tree. An exact version of Lake's asymptotic chi 2 test has been mentioned by Holmquist et al. Under very general assumptions, a one-sided exact test is appropriate, which greatly increases power.

Animals↗