PubMed HealthSearch

SEARCH · PubMed Health

Results for “Binomial Distribution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho = 0), the beta-binomial distribution (Rho = 0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5 K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus

Segregation ratios within Segregation Distorter lines of Drosophila melanogaster conform to a beta-binomial distribution.

Segregation Distorter (SD) chromosomes are preferentially recovered from SD/SD+ males due to the dysfunction of sperm bearing the SD+ chromosome. The proportion of offspring bearing the SD chromosome is given the symbol k. The nature of the frequency distribution of k was examined by comparing observed k distributions produced by six different SD chromosomes, each with a different mean, with k distributions predicted by two different statistical models. The first model was one where the k of all males with a given SD chromosome were considered to be equal prior to the determination of those gametes which produce viable zygotes. In this model the only source of variation of k would be binomial sampling. The results rigorously demonstrated for the first time that the observed k distributions did not fit the prediction that the only source of variation was binomial sampling. The next model tested was that the prior distribution of segregation ratios conformed to a beta distribution, such that the distribution of k would be a beta-binomial distribution. The predicted distributions of this model did not differ significantly from the observed distributions of k in five of the six cases examined. The sixth case probably failed to fit a beta-binomial distribution due to a major segregating modifier. The demonstration that the prior distribution of segregation ratios of SD lines can generally be approximated with a beta distribution is crucial for the biometrical analysis of segregation distortion.

Animals

Use of the beta-binomial distribution in dominant-lethal testing for "weak mutagenic activity: part 2.

Experiments in Dominant-Lethal Testing have been simulated on the computer to estimate the type I error rates and the power of the Beta-Binomial test under various models. (1) The mating ratio is one; and p, the probability that an implant will die, is distributed over the couples. (2) The mating ratio is larger than one; and p is distributed over the males, the females mated to the same male being binomial observations of the value p supplied by the male. (3) The mating ratio is larger than one; and p is distributed over the females. The average rates of dead implants have been set at 0.08 and 0.10 for the control and treatment groups, respectively, and a nominal level of significance equal to 0.05 has been chosen. The type I error rate of the traditional chi-square test has also been estimated. A by-product of these simulations is the behaviour of the estimates alpha and beta of the beta-distribution parameters, which discloses that, in the actual experiments with mice, p is distributed over the females. Our results lead to the recommendations that, for a given number of animals per group, a mating ratio larger than one should be adopted and that the males should be considered as the experimental units for the calculations. With 300 and 450 animals per group, average powers of 0.72 and 0.85 are reached, respectively, for the chosen increment of 2% in the rate of dead implants. Under these models, the type I error rate of the traditional chi-square test may grow to 0.30 for the nominal level of 0.05.

Genes, Dominant

[The result-sequence method (sequence analysis) in leptospira research. 3. Communication: The bilateral sequential test (de Boer, Armitage) for testing the difference of mean values of two binomial distributions (author's transl)].

The influence of the liver-activated and non-activated cytostatic drug Cyclophosphamid on the respiration of Leptospira biflexa semaranga Veldrat S 173 was tested in three different concentrations by group-sequential testing of two relative frequencies in a two-sided test-reading. The conditioned probability theta = 0.82 results from thetan = 0.40 and thetaa = 0.75. On account of a practicable sample size, we choose theta' = 0.80 (activated form superior) and theta'' = 0.20 (non-activated form superior). The test-adjusted level of significance is alpha = 0.05 less than beta = 0.10. The expected values for the number of discordant pairs are Etheta' = Etheta'' = 13, Etheta = 1/1 = 14 and Ethetamax = 16. The acceptance inspection performed by control chart and tabulated reference figures results - in comparison with conventional procedures - in savings of discordant pairs of 65 per cent (concentration 10(-4) g/ml) in a decision in favour of activated form, of 65 per cent (concentration 10(-8) g/ml) in an equal efficiency of both states of Cyclophosphamid and of 10 per cent (concentration 10(-11) g/ml) in a non-significant difference. Taking into consideration all unrestrictedly selected pairs, the duration of the experiment takes only 4.25 months instead of 6.75 if the statistical data analysis is not performed conventionally, but group-sequentially instead.

Cyclophosphamide

An approach to numerical identification of bacterial species.

The distribution of matching coefficients (M values) for strains of a species to their own 'hypothetical mean organism' (HMO) or to HMO patterns of other organisms was studied in 754 strains of 19 mycobacterial species, testing for 91 discriminating characters. The M values of strains of a species to the HMO of other species usually showed a normal distribution, and M values to their own HMO showed either a normal distribution of a binomial distribution, depending on the mean of M values. If the number of test characters was large, the binomial distribution usually resembled the normal distribution. After preparation of the HMO for every species and estimation of the mean of the M values (M) and the standard deviation (s), numerical identification could be carried out: if a test strain had an M value to the HMO of species chi that only fell within the range (M +/2s) for species chi, the strain would be identified as a member of that species.

Bacteria

Multichannel 18-test panels: are 60% of panels abnormal by chance?

Current teaching concerning the frequency of abnormal results secondary to chance alone in a multichannel panel is theoretically based on the binomial distribution. However, this distribution can be used only when the probability of an abnormal result (pi) is the same for each test in the panel. In modern-day multichannel testing, pi varies from test to test and most often is less than the usually reported 0.05. On the other hand, a test such as cholesterol may have a pi level as high as 0.55. Theoretically the only distribution that can take this variability into consideration is the Lexis distribution, a form of the binomial distribution that allows for varying pi s. Since no formula is available to calculate this distribution, we wrote a computer program to generate it. We arranged 18-test panels from 203 normal patients in a frequency distribution. This was then compared with the theoretical Lexis and binomial distributions. This analysis showed that although there was a 50% chance of having one abnormality per panel and a 16% chance of having two abnormalities per panel, there was less than 4% chance of having three or more abnormalities per 18-test panel. In addition, most of the abnormalities noted were minor and were thought to be clinically unimportant.

Adult

The micronucleus test: statistical design and analysis.

Alternative statistical procedures are discussed which may be employed to compare the incidences among treatment groups of micronucleated polychromatic and normochromatic erythrocytes and their ratios. Comparison of incidences of micronucleated polychromatic erythrocytes using a sequential sampling strategy based on the negative binomial distribution is shown to require fewer animals for the same sensitivity of test than a similar procedure based on the binomial distribution. The sequential test is superior, both in power and number of animals required, to an alternative 1-stage test based on the same distribution. The procedure described permits the investigator to optimize the number of animals in each test group and the number of cells counted per animal to detect a predetermined increase in the incidence of micronucleated cells over that observed in the control population within chosen limits of type I and type II error. An alternative sequential approach based on the binomial distribution is presented, which is applicable when the number of cells analyzed per animal is variable.

Animals

Frequency of elevated urinary beta 2-microglobulin levels in relatives of patients with asymptomatic low-molecular-weight proteinuria.

We studied urinary beta 2-microglobulin levels in a total of 29 apparently healthy relatives (aged 0.8-70 years) of 8 male patients with asymptomatic low-molecular-weight proteinuria in six families. The frequency of levels above the age- and sex-associated 95% confidence limit was 7 of 29 (24%), 4 of 12 (33%) in first-degree relatives, 2 of 6 (33%) fathers, and 2 of 6 (33%) mothers. These frequencies were significantly above those in the general population (P less than 0.01, by a normal distribution test, a binomial distribution, and Poisson distribution test for the sample proportion). The increased frequency in fathers argues against an X-linked pattern of inheritance for this entity, suggesting that there is heterogeneity in the inheritance.

Adolescent

Predictive probability early termination plans for phase II clinical trials.

A phase II clinical trial is designed to gather data to help decide whether an experimental treatment has sufficient effectiveness to justify further study. In a one-arm trial with dichotomous outcome, we wish to test a simple null hypothesis on the Bernoulli parameter against a one-sided alternative in a sample of N patients. It is advisable to have a rule to terminate the trial early when evidence accumulates that the treatment is ineffective. Predictive probabilities based on the binomial distribution and beta and uniform prior distributions for the binomial parameter are found to be useful as the basis of group sequential designs. Size, power and average sample size for these designs are discussed. A process for the specification of an early termination plan, advice on the quantification of prior beliefs, and illustrative examples are included.

Antineoplastic Agents

On the regularities of distribution of Hypoderma bovis De Geer larvae parasitizing cattle herds in different parts of the range of this warble fly.

The type and parameters of the distribution of the second and third instar larvae of Hypoderma bovis in cattle herds in Czechoslovakia (54 herds, 7233 head), Mongolia (20 herds, 1809 head) and the USSR (48 herds, 4978 head) were studied. A statistical analysis showed a) that in the majority of cases negative binomial distribution serves as a model of the distribution of larvae with sufficient reliability; b) that the regularity of dependence of the distribution exponent k of the negative binomial distribution on the incidence of infestation remains constant in different parts of the range of this warble fly. The latter fact indicates that the regulatory systems limiting the population numbers of this warble fly are associated with host-parasite relationships and do not depend on a complex of conditions specific for various natural zones. In order to understand the regulatory processes in parasite populations it is necessary to study equally the regulatory mechanisms operating primarily on the level of specimens as well as the regulatory systems operating on the level of populations.

Animals

On the regularities of distribution of Hypoderma bovis De Geer larvae parasitizine cattle herds in different parts of the range of this warble fly.

The type and parameters of the distribution of the second and third instar larvae of Hypoderma bovis in cattle herds in Czechoslovakia (54 herds, 7233 head), Mongolia (20 herds, 1809 head) and the USSR (48 herds, 4978 head) were studied. A statistical analysis showed a) that in the majority of cases negative binomial distribution serves as a model of the distribution of larvae with sufficient reliability; b) that the regularity of dependence of the distribution exponent k of the negative binomial distribution on the incidence of infestation remains constant in different parts of the range of thes warble fly. The latter fact indicates that the regulatory systems limiting the population numbers of this warble fly are associated with host-parasite relationships and do not depend on a complex of conditions specific for various natural zones. In order to understand the regulatory processes in parasite populations it is necessary to study equally the regulatory mechanisms operating primarily on the level of specimens as well as the regulatory systems operating on the level of populations.

Animals

The distribution of fetal death in control mice and its implications on statistical tests for dominant lethal effects.

In dominant lethal testing fetal death is generally assumed to follow either a Poisson or binomial distribution. However, both of these models were found to be inappropriate when three large sets of mouse control data and other data sets from the literature were examined. The validity of statistical test procedures based on these inappropriate models was then studied in detail. It was found that chi-square tests (which assume an underlying binomial distribution) may seriously exaggerate the level of significance and hence should not be used. In contrast, the inappropriateness of the underlying Poisson or binomial model appeared to have little effect on the validity of pairwise comparisons by analysis of variance procedures. Unlike chi-square, these procedures regard the pregnant female rather than the individual implant as the experimental unit. However, a statistical analysis of dominant lethal data generally involves more than a series of pairwise comparisons, and it is unclear how an invalid underlying model may affect statistical test procedures in this more complex situation. Moreover, it is difficult to justify the use of statistical models that are demonstrably invalid when a reasonable alternative exists. Thus, until a satisfactory parametric model can be found and appropriate test procedures derived, we prefer to analyze dominant lethal data by non-parametric (distribution-free) methods. Proportion of dead implants per female appears to be a more meaningful measure of fetal death than number of dead implants per female for several seasons which include (1) analyses based on proportions take the total number of implants per female into account and (2) analyses based on proportions make more reasonable assumptions concerning pre-implantation losses and are more powerful when such losses occur. Despite our concern with the appropriateness of the underlying model, in practice we have found few instances in which non-parametric and analysis of variance procedures have led to markedly different conclusions.

Animals

[Falsely positive values in multi-channel analysis: An iniquiry into reference and patient groups (author's transl)].

The number of falsely positive values occurring in 12-channel analysis was determined in two groups of patients and reference individuals. It revealed that the portion of falsely positive values actually found was statistically significant beyond that calculated on the assumption of a binomial distribution. Partly distinct correlations of the parameters combined to a profile as well as clear deviations from the normal distribution have to be taken into consideration as reasons for this discrepancy between theory and reality. The results show that the application of the binomial distribution leads to statements which significantly differ from the conditions actually present.

Autoanalysis

[Use of various accordance measures in the study of sick leave frequency distribution per man-year].

In the studies of a wide number of phenomena very often an attempt to describe their distributions with the help of certain theoretical distributions is undertaken. Pearson's chi2 and Kolmogorovs lambda traditional tests, when applied to the evaluation of accordance between the empirical distribution and the distribution calculated on the basis of theoretical probability function, depend directly on the dimensions of population studied, and they are not useful in the case of big quantities. In the paper the author presents a great number of measures for two (empirical and theoretical) distributions accordance and their application to the study of sick-leave frequency distribution per man per year. On the basis of received results a high level of distributions accordance between the sick leave cases distribution and the negative binomial distribution has been observed. The distribution of sick leave shows a high accordance with the logarithmically-normal distribution. In the light of the above mentioned distribution study, the imperfection of Pearson's chi2 and Kolmogorov's lambda tests can be ascertained. In the case of big quantities the evaluation of distributions accordance seems to require the application of the accordance measures based on relative quantities.

Absenteeism

Frequency distribution of Wuchereria bancrofti microfilariae in human populations and its relationships with age and sex.

This paper examines the effects of host age and sex on the frequency distribution of Wuchereria bancrofti infections in the human host. Microfilarial counts from a large data base on the epidemiology of bancroftian filariasis in Pondicherry, South India are analysed. Frequency distributions of microfilarial counts divided by age are successfully described by zero-truncated negative binomial distributions, fitted by maximum likelihood. Parameter estimates from the fits indicate a significant trend of decreasing overdispersion with age in the distributions above age 10; this pattern provides indirect evidence for the operation of density-dependent constraints on microfilarial intensity. The analysis also provides estimates of the proportion of mf-positive individuals who are identified as negative due to sampling errors (around 5% of the total negatives). This allows the construction of corrected mf age-prevalence curves, which indicate that the observed prevalence may underestimate the true figures by between 25% and 100%. The age distribution of mf-negative individuals in the population is discussed in terms of current hypotheses about the interaction between disease and infection.

Adolescent