PubMed Health⌕ Search

Biomedical subjects

D N Stivers

Publications and source records attributed to D N Stivers.

16 recordsLinked to original sources

Microarrays: handling the deluge of data and extracting reliable information.

Application of powerful, high-throughput genomics technologies is becoming more common and these technologies are evolving at a rapid pace. Genomics facilities are being established in major research institutions to produce inexpensive, customized cDNA microarrays that are accessible to researchers in a broad range of fields. These high-throughput platforms have generated a massive onslaught of data, which threatens to overwhelm researchers. Although microarrays show great promise, the technology has not matured to the point of consistently generating robust and reliable data when used in the average laboratory. This article addresses several aspects related to the handling of the deluge of microarray data and extracting reliable information from these data. We review the essential elements of data acquisition, data processing and data analysis, and briefly discuss issues related to the quality, validation and storage of data. Our goal is to point out some of the problems that must be overcome before this promising technology can achieve its full potential.

DNA, Complementary↗

Identifying differentially expressed genes in cDNA microarray experiments.

A major goal of microarray experiments is to determine which genes are differentially expressed between samples. Differential expression has been assessed by taking ratios of expression levels of different samples at a spot on the array and flagging spots (genes) where the magnitude of the fold difference exceeds some threshold. More recent work has attempted to incorporate the fact that the variability of these ratios is not constant. Most methods are variants of Student's t-test. These variants standardize the ratios by dividing by an estimate of the standard deviation of that ratio; spots with large standardized values are flagged. Estimating these standard deviations requires replication of the measurements, either within a slide or between slides, or the use of a model describing what the standard deviation should be. Starting from considerations of the kinetics driving microarray hybridization, we derive models for the intensity of a replicated spot, when replication is performed within and between arrays. Replication within slides leads to a beta-binomial model, and replication between slides leads to a gamma-Poisson model. These models predict how the variance of a log ratio changes with the total intensity of the signal at the spot, independent of the identity of the gene. Ratios for genes with a small amount of total signal are highly variable, whereas ratios for genes with a large amount of total signal are fairly stable. Log ratios are scaled by the standard deviations given by these functions, giving model-based versions of Studentization. An example is given.

Analysis of Variance↗

Microsatellites and intragenic polymorphisms of transforming growth factor beta and platelet-derived growth factor and their receptor genes in Native Americans with systemic sclerosis (scleroderma): a preliminary analysis showing no genetic association.

OBJECTIVE: Abnormalities of transforming growth factor beta (TGFbeta) and platelet-derived growth factor (PDGF) alpha and beta and/or their receptors have been demonstrated in systemic sclerosis (SSc). This study aimed to determine whether genetic polymorphisms in or near the TGFbeta and PDGF gene families were associated with susceptibility to SSc in a Native American population with a high disease prevalence. METHODS: Genotyping of 5 intragenic polymorphisms within the TGFbeta1 gene and mapping of 35 microsatellites near the genes for TGFbeta1, latent TGFbeta1 binding protein (LTBP1), TGFbeta receptors I and II, PDGFalpha, PDGFbeta, PDGF receptor alpha, and PDGF receptor beta was performed in 19 SSc patients, 76 controls, and 42 family members. Allele distributions and frequencies were examined between SSc patients and controls, and marker haplotypes were examined in families when allele frequencies appeared to be different between patients and controls. RESULTS: Although 1 polymorphism within the TGFbeta1 gene (TGFbeta1) was modestly increased in the SSc patients, this did not maintain statistical significance after correction. Similarly, 1 microsatellite (D9S120) near the TGFbeta receptor I gene (TGFBR1) showed a significant disturbance of allele frequencies between patients and controls; however, it did not form a disease-associated haplotype with other nearby markers. Weak disturbances of markers near PDGFalpha (PDGFA) and PDGFbeta, (PDGFB) also failed to maintain significance after correction. Both PDGF receptor genes (PDGFRA and PDGFRB) also showed no disease associations. CONCLUSION: The results of these preliminary analyses suggest that genetic anomalies of the TGFbeta1 and PDGF gene families are not likely to explain the dysregulation seen in SSc or to account for the susceptibility to SSc in this population.

Alleles↗

The utility of short tandem repeat loci beyond human identification: implications for development of new DNA typing systems.

Since the first characterization of the population genetic properties of repeat polymorphisms, the number of short tandem repeat (STR) loci validated for forensic use has now grown to at least 13. Worldwide variations of allele frequencies at these loci have been studied, showing that variations of interpopulation diversity at these loci do not compromise the power of identification of individuals. However, data collected for validation of these loci for forensic use has utility beyond human identification; the origin and past migration history of modern humans can be reconstructed from worldwide variations at these loci. Furthermore, complex forensic cases previously unresolvable can now be investigated with the help of the validated STR loci. Here, we provide the absolute power of the validated set of 13 STR loci for addressing these issues using multilocus genotype data on 1,401 individuals belonging to seven populations (US European-American, US African-American, Jamaican, Italian, Swiss, Chinese and Apache Native-American). Genomic research is discovering new classes of polymorphic loci (such as the single nucleotide polymorphisms, SNPs) and lineage markers (such as the mitochondrial DNA and Y-chromosome markers); our aim, therefore, was to determine how many SNP loci are needed to match the power of this set of 13 STR loci. We conclude that the current set of STR loci is adequate for addressing most problems of human identification (including interpretations of DNA mixtures). However, if suitable number of SNPs are used that would match the power of the STR loci, they alone cannot resolve more complex cases unless they are supplemented by the validated STR loci.

Chromosome Mapping↗

HLA haplotypes and microsatellite polymorphisms in and around the major histocompatibility complex region in a Native American population with a high prevalence of scleroderma (systemic sclerosis).

Choctaw Native Americans in southeastern Oklahoma have the highest prevalence of scleroderma or systemic sclerosis yet found (469/100,000). An Amerindian HLA DR2 haplotype (DRB1*1602) was significantly associated with scleroderma in this population in a previous study. It is not known, however, if other disease genes are linked to this HLA haplotype. The regions flanking the HLA loci were studied with polymorphic microsatellite markers. An extended HLA DR2 (DRB1*1602, DQA1*0501, DQB1*0301, DPB1*1301) haplotype that includes the class I and III regions was identified which was significantly associated with scleroderma in the Oklahoma Choctaw. No other significant associations with microsatellite marker alleles immediately flanking the HLA region were found.

Autoimmune Diseases↗

Association of microsatellite markers near the fibrillin 1 gene on human chromosome 15q with scleroderma in a Native American population.

OBJECTIVE: To localize disease genes for scleroderma, or systemic sclerosis (SSc), in a population of Choctaw Native Americans with a high prevalence of SSc, in which there is evidence of a possible founder effect. METHODS: A candidate gene approach was used in which microsatellite alleles on human chromosomes 15q and 2q, homologous to the murine tight skin 1 (tsk1) and tsk2 loci, respectively, were analyzed in Choctaw SSc cases and race-matched normal controls for possible disease association. Genotyping first-degree relatives of the cases identified potential disease haplotypes, and haplotype frequencies were obtained by expectation-maximization and maximum-likelihood estimation methods. Simultaneously, the ancestral origins of contemporary Choctaw SSc cases were ascertained using census and historical records. RESULTS: A multilocus 2-cM haplotype was identified on human chromosome 15q homologous to the murine tsk1 region, which showed a significantly increased frequency in SSc cases compared with controls. This haplotype contains 2 intragenic markers for the fibrillin 1 (FBN1) gene. Genealogical studies demonstrated that the SSc cases were distantly related, and their ancestry could be traced back to 5 founding families in the mid-eighteenth century. The probability that the SSc cases share this haplotype due to familial aggregation effects alone was calculated and found to be very low. There was no evidence of any microsatellite allele disturbances on chromosome 2q in the region homologous to the tsk2 locus or the region containing the interleukin-1 family. CONCLUSION: A 2-cM haplotype on chromosome 15q that contains FBN1 is associated with scleroderma in Choctaw Native Americans from Oklahoma. This haplotype may have been inherited from common founders about 10 generations ago and may contribute to the high prevalence of SSc that is now seen.

Alleles↗

Relative mutation rates at di-, tri-, and tetranucleotide microsatellite loci.

Using the generalized stepwise mutation model, we propose a method of estimating the relative mutation rates of microsatellite loci, grouped by the repeat motif. Applying ANOVA to the distributions of the allele sizes at microsatellite loci from a set of populations, grouped by repeat motif types, we estimated the effect of population size differences and mutation rate differences among loci. This provides an estimate of motif-type-specific mutation rates up to a multiplicative constant. Applications to four different sets of di-, tri-, and tetranucleotide loci from a number of human populations reveal that, on average, the non-disease-causing microsatellite loci have mutation rates inversely related to their motif sizes. The dinucleotides appear to have mutation rates 1.5-2 times higher than the tetranucleotides, and the non-disease-causing trinucleotides have mutation rates intermediate between the di- and tetranucleotides. In contrast, the disease-causing trinucleotides have mutation rates 3.9-6.9 times larger than the tetranucleotides. Comparison of these estimates with the direct observations of mutation rates at microsatellites indicates that the earlier suggestion of higher mutation rates of tetranucleotides in comparison with the dinucleotides may stem from a nonrandom sampling of tetranucleotide loci in direct mutation assays.

Analysis of Variance↗

A discrete-time, multi-type generational inheritance branching process model of cell proliferation.

Mammalian cell populations, such as tumors, may contain subpopulations differing in parameters such as cell lifetimes, even if the populations are derived from single cells. The mode of inheritance of cell lifetimes has previously been the subject of experimental and mathematical investigation. To obtain data on cell lifetimes over more cell generations then previously available, Axelrod et al. [Cell Prolif. 26:235-249(1988)] measured the number of cells in primary colonies and secondary colonies derived form the primary colonies. The experimental results indicated large variance of cells per colony and highly significant correlations between the numbers of cells in primary and secondary colonies. To mathematically model these results we derive, for previously uninvestigated multi-type Galton-Watson branching process models, the covariance of the cell counts in the primary and secondary colonies. As a result, we are able to successfully model the data with two subpopulations having differing proliferation rates, in which the proliferation rate of a daughter cell is primarily determined by the proliferation rate of its mother. Interestingly, simulations display a trade-off between high values of variances and correlation coefficients. The values obtained from experiment are located on the boundary of the region attainable by simulation.

Animals↗

Estimation of mutation rates from parentage exclusion data: applications to STR and VNTR loci.

Nonpaternity is a common source of bias in estimating mutation rates when they are obtained from family data showing discordance of parental and children's genotypes. With the availability of hypervariable DNA markers, this source of bias can be largely eliminated. However, the proportion of cases where parentage exclusion is caused by presumed mutation(s) of parental alleles must be adjusted to obtain a valid mutation rate estimate. The present work derives the basis of this adjustment factor, called the proportional bias. This proportional bias depends upon the allele frequency distribution at the locus. The maximum and minimum bounds of the proportional bias depend on the number of alleles at the locus. Using data from Caucasian populations at tandem repeat loci commonly used for parentage testing and forensic identification purposes, we show that when mutation rates are estimated at these loci, the proportional bias is generally very close to the maximum possible value for the observed number of alleles (or binned fragment sizes) at each locus. The expected proportional bias decreases with increasing mutation rate at a locus. For the short tandem repeat loci, without bias correction, the direct count method can result in an underestimation of up to 60% of their true value. In contrast, for the minisatellite VNTR loci, even with crude measurements on allele sizes, we show that the absolute proportional bias is generally below the coefficient of variation of the direct estimates.

Chromosome Mapping↗

Dynamics of repeat polymorphisms under a forward-backward mutation model: within- and between-population variability at microsatellite loci.

Suggested molecular mechanisms for the generation of new tandem repeats of simple sequences indicate that the microsatellite loci evolve via some of forward-backward mutation. We provide a mathematical basis for suggesting a measure of genetic distance between populations based on microsatellite variation. Our results indicate that such a genetic distance measure can remain proportional to the divergence time of populations even when the forward-backward mutations produce variable and/or directionally biased alleles size changes. If the population size and the rate of mutation remain constant, then the measure will be proportional to the time of divergence of populations. This genetic distance is expressed in terms of a ratio of components of variance of allele sizes, based on expressions developed for studying population dynamics of quantitative traits. Application of this measure to data on 18 microsatellite loci in the nine human populations leads to evolutionary trees consistent with the known ethnohistory of the populations.

Alleles↗

Distribution and evolution of CTG repeats at the myotonin protein kinase gene in human populations.

We have analyzed the CTG repeat length and the neighboring Alu insertion/deletion (+/-) polymorphism in DNA samples from 16 ethnically and geographically diverse human populations to understand the evolutionary dynamics of the myotonic dystrophy-associated CTG repeat. Our results show that the CTG repeat length is variable in human populations. Although the (CTG)5 repeat is the most common allele in the majority of populations, this allele is absent among Costa Ricans and New Guinea highlanders. We have detected a (CTG)4 repeat allele, the smallest CTG known allele, in an American Samoan individual. (CTG) > or = 19 alleles are the most frequent in Europeans followed by the populations of Asian origin and are absent or rare in Africans. To understand the evolution of CTG repeats, we have used haplotype data from the CTG repeat and Alu(+/-) locus. Our results are consistent with previous studies, which show that among individuals of Caucasian and Japanese origin, the association of the Alu(+) allele with CTG repeats of 5 and > or = 19 is complete, whereas the Alu(-) allele is associated with (CTG)11-16 repeats. However, these associations are not exclusive in non-Caucasian populations. Most significantly, we have detected the (CTG)5 repeat allele on an Alu(-) background in several populations including Native Africans. As no (CTG)5 repeat allele on an Alu(-) background was observed thus far, it was proposed that the Alu(-) allele arose on a (CTG)11-13 background. Our data now suggest that the most parsimonious evolutionary model is (1) (CTG)5-Alu(+) is the ancestral haplotype; (2) (CTG)5-Alu(-) arose from a (CTG)5-Alu(+) chromosome later in evolution; and (3) expansion of CTG alleles occurred from (CTG)5 alleles on both Alu(+) and Alu(-) backgrounds.

Biological Evolution↗

Segregation distortion of the CTG repeats at the myotonic dystrophy locus.

Myotonic dystrophy (DM), an autosomal dominant neuromuscular disease, is caused by a CTG-repeat expansion, with affected individuals having > or = 50 repeats of this trinucleotide, at the DMPK locus of human chromosome 19q13.3. Severely affected individuals die early in life; the milder form of this disease reduces reproductive ability. Alleles in the normal range of CTG repeats are not as unstable as the (CTG)(> or = 50) alleles. In the DM families, anticipation and parental bias of allelic expansions have been noted. However, data on mechanism of maintenance of DM in populations are conflicting. We present a maximum-likelihood model for examining segregation distortion of CTG-repeat alleles in normal families. Analyzing 726 meiotic events in 95 nuclear families from the CEPH panel pedigrees, we find evidence of preferential transmission of larger alleles (of size < or = 29 repeats) from females (the probability of transmission of larger alleles is .565 +/- 0.03, different from .5 at P approximately equal .028). There is no evidence of segregation distortion during male meiosis. We propose a hypothesis that preferential transmission of larger CTG-repeat alleles during female meiosis can compensate for mutational contraction of repeats within the normal allelic size range, and reduced viability and fertility of affected individuals. Thus, the pool of premutant alleles at the DM locus can be maintained in populations, which can subsequently mutate to the full mutation status to give rise to DM.

Alleles↗

Paternity exclusion by DNA markers: effects of paternal mutations.

In parentage testing when one parent is excluded, the distribution of the number of loci showing exclusion due to mutations of the transmitting alleles is derived, and it is contrasted with the expected distribution when the exclusion is caused by nonpaternity. This theory is applied to allele frequency data on short tandem repeat loci scored by PCR analysis, and VNTR data scored by Southern blot RFLP analysis that are commonly used in paternity analysis. For such hypervariable loci, wrongly accused males should generally be excluded based two or more loci, while a true father is unlikely to be excluded based on multiple loci due to mutations of paternal alleles. Thus, when these DNA markers are used for parentage analysis, the decision to infer non-paternity based on exclusions at two or more loci has a statistical support. Our approach places a reduced weight on the combined exclusion probability. Even with this reduced power of exclusion, the probability of exclusion based on combined tests on STR and VNTR loci is sufficiently large to resolve most paternity dispute cases in general populations.

Adult↗

Time-continuous branching walk models of unstable gene amplification.

We consider a stochastic mechanism of the loss of resistance of cancer cells to cytotoxic agents, in terms of unstable gene amplification. Two models being different versions of a time-continuous branching random walk are presented. Both models assume strong dependence in replication and segregation of the extrachromosomal elements. The mathematical part of the paper includes the expression for the expected number of cells with a given number of gene copies in terms of modified Bessel functions. This adds to the collection of rare explicit solutions to branching process models. Original asymptotic expansions are also demonstrated. Fitting the model to experimental data yields estimates of the probabilities of gene amplification and deamplification. The thesis of the paper is that purely stochastic mechanisms may explain the dynamics of reversible drug resistance of cancer cells. Various stochastic approaches and their limitations are discussed.

Animals↗

A test of allelic independence based on distributions of allele size differences at microsatellite loci.

Thousands of polymorphic microsatellite loci are now mapped in the human genome, most of which exhibit a large number of segregating alleles. For forensic, evolutionary and gene mapping applications, it is important to establish intralocus allelic independence at these loci. We develop a test for intralocus allelic independence based on repeat size differences between alleles within genotypes of individuals, and show its relationship with the intraclass correlation of allele sizes within loci. Applications of this test to 13 short tandem repeat loci in 4 subpopulations indicate that the distribution of size differences of alleles has adequate power to detect intralocus dependence of alleles caused by population substructure. In addition, size differences between randomly chosen alleles provide information regarding heterozygosity and contiguity of allele sizes at the locus, which by themselves do not necessarily indicate the presence of population substructure.

Alleles↗