PubMed Health⌕ Search

PubMed · 2242405

Sampling strategies for distances between DNA sequences.

Abstract

An international effort is now underway to obtain the DNA sequence for the entire human genome (Watson and Jordan, 1989, Genomics 5, 654-656; Barnhart, 1989, Genomics 5, 657-660). This Human Genome Initiative will generate sequence data from several species other than humans, and will result in several copies per species of at least some regions of the genome. Although the project has generated much interest, it is but one aspect of the widespread effort to generate DNA sequence data. Published sequences are collected in common databases, and release 63 of GenBank in March 1990 contained 40,127,752 bases from 33,337 reported sequences (News from GenBank 3; Mountain View, California: Intelligenetics, Inc., 1990). Large though this database is, it is only about 1% of the number of bases in the human genome. Interpretations of data of such magnitude are going to require the collaborative efforts of biometricians and molecular biologists, and an aim of this paper is to show that there is also a role for readers of this journal in the design of surveys of DNA sequences. Discussion here will center on the use of sequence data in evolutionary studies, where some region of DNA is sequenced in several different species. The object is to infer the evolutionary history of that particular region, or of the species themselves. Statistical issues in the very important studies on sequences to locate and characterize regions responsible for human diseases will not be addressed here. We will discuss appropriate ways of measuring distances between DNA sequences and of predicting the sampling properties of the distances. There are procedures for inferring evolutionary histories for a set of elements that depend on a matrix of distances between each pair of elements, and the precision of resulting trees must be influenced by the precision of the distances. We will show that account needs to be taken of two sampling processes--the sampling of sequences by the investigator ("statistical sampling"), and the sampling of genetic material involved in the formation of offspring from a parental population ("genetic sampling").

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B S Weir, C J Basten. 1990. Sampling strategies for distances between DNA sequences.. https://pubmed.ncbi.nlm.nih.gov/2242405/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Molecular heterochrony and the evolution of sociality in bumblebees (Bombus terrestris).

Sibling care is a hallmark of social insects, but its evolution remains challenging to explain at the molecular level. The hypothesis that sibling care evolved from ancestral maternal care in primitively eusocial insects has been elaborated to involve heterochronic changes in gene expression. This elaboration leads to the prediction that workers in these species will show patterns of gene expression more similar to foundress queens, who express maternal care behaviour, than to established queens engaged solely in reproductive behaviour. We tested this idea in bumblebees (Bombus terrestris) using a microarray platform with approximately 4500 genes. Unlike the wasp Polistes metricus, in which support for the above prediction has been obtained, we found that patterns of brain gene expression in foundress and queen bumblebees were more similar to each other than to workers. Comparisons of differentially expressed genes derived from this study and gene lists from microarray studies in Polistes and the honeybee Apis mellifera yielded a shared set of genes involved in the regulation of related social behaviours across independent eusocial lineages. Together, these results suggest that multiple independent evolutions of eusociality in the insects might have involved different evolutionary routes, but nevertheless involved some similarities at the molecular level.

Analysis of Variance↗

Confidence intervals for the standardized effect arising in the comparison of two normal populations.

Confidence intervals for a standardized effect are derived after stabilizing the variance of the Welch t-statistic. Simulation studies demonstrate the viability of the resulting intervals for a wide range of parameter values and sample sizes as small as five. The methodology is extended to the combination of results from several studies, so as to obtain a confidence interval for a representative standardized effect for all the studies. The methods are illustrated on a recent meta-analytic study of systolic blood pressure reduction during a weight reducing regime, as well as the classical Mumford data on psychological intervention and hospital length of stay.

Analysis of Variance↗