Experimental DNA melting behavior of the three major Schistosoma species.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to R D Blake.
Explore the source record for details and available documents.
We demonstrate that differential scanning calorimetry (DSC) can be used to yield high-resolution melting profiles for DNA plasmids that agree in all major features with the corresponding plasmid melting profiles derived using more traditional optical techniques. We further demonstrate that by combining information derived from both calorimetric and optical melting profiles one can glean insights that are unavailable from either melting curve alone. By using both optical and calorimetric observables, we show how one can resolve, identify, and measure the thermodynamic properties of particular sequences/domains of interest within a plasmid. We also show that complementary DSC and optical melting studies on plasmids with and without specifically designed inserts can provide fundamental advantages over the corresponding melting studies on other model system constructs for thermodynamically characterizing nucleic acid sequences/structures.
MOTIVATION: MELTSIM is a windows-based statistical mechanical program for simulating melting curves of DNAs of known sequence and genomic dimensions under different conditions of ionic strength with great accuracy. The program is useful for mapping variations of base compositions of sequences, conducting studies of denaturation, establishing appropriate conditions for hybridization and renaturation, determinations of sequence complexity, and sequence divergence. RESULTS: Good agreement is achieved between experimental and calculated melting curves of plasmid, bacterial, yeast and human DNAs. Denaturation maps that accompany the calculated curves indicate non-coding regions have a significantly lower (G+C) composition than coding regions in all species examined. Curves of partially sequenced human DNA suggest the current database may be heavily biased with coding regions, and excluding large (A+T)-rich elements. AVAILABILITY: MELTSIM 1.0 is available at: //www.uml.edu/Dept/Chem/UMLBIC/Apps/MEL TSIM/MELTSIM-1.0-Win/meltsim. zip. Melting curve plots in this paper were made with GNUPLOT 3.5, available at: http://www.cs.dartmouth.edu/gnuplot_inf o.html Contact : blake@maine.maine.edu;
Tij and Delta Hij for stacking of pair i upon j in DNA have been obtained over the range 0.034-0.114 M Na+from high-resolution melting curves of well-behaved synthetic tandemly repeating inserts in recombinant pN/MCS plasmids. Results are consistent with neighbor-pair thermodynamic additivity, where the stability constant, sij , for different domains of length N depend quantitatively on the product of stability constants for each individual pair in domains, sijN . Unit transition enthalpies with average errors less than +/-5%, were determined by analysis of two-state equilibria associated with the melting of internal domains and verified from variations of Tij with [Na+]. Enthalpies increase with Tij , in close agreement with the empirical function: Delta Hij = 52.78@ Tij - 9489, and in parallel with a smaller increase in Delta Sij . Delta Hij and Delta Sij are in good agreement with the results of an extensive compilation of published Delta Hcal and Delta Scal for synthetic and natural DNAs. Neighbor-pair additivity was also observed for (dA@dT)-tracts at melting temperatures; no evidence could be detected of the familiar and unusual structural features that characterize tracts at lower temperatures. The energetic effects of loops were determined from the melting behavior of repeating inserts installed between (G+C)-rich barrier domains in the pN/MCS plasmids. A unique set of values for the cooperativity, loop exponent and stiffness parameters were found applicable to internal domains of all sizes and sequences. Statistical mechanical curves calculated with values of Tij([Na+]) , Delta Hij and these loop parameters are in good agreement with observation.
The slime mold, Dictyostelium discoideum, possesses an (A+T) rich eukaryotic genome that is being sequenced in the Human Genome Project. High resolution melting curves of isolated total and fractionated nuclear D. discoideum DNA(AX3 strain) were determined experimentally and are compared to melting curves calculated from GENBANK sequences (1.59% of genome) by the statistical thermodynamics program MELTSIM (1), parameterized for long DNA sequences (2,3). The lower and upper temperature limits of calculated melting agree well with the observed melting of total DNA. The experimental curve is unusual in that it contains a number of sharp peaks. MELTSIM allowed us to calculate positional denaturation maps of D. discoideum GENBANK sequence documents containing the 26S, 5.8S and 17S rDNA gene sequences, a major satellite DNA and repetitive sequence family present in 100-200 copies/nucleus. These denaturation maps contain subtransitions that correspond with a number of the experimentally observed peaks, some of which we show to correspond with rDNA gene enriched CsCl gradient fractions of D. discoideum DNA. MELTSIM calculated curves of coding, intron and flanking sequences indicate that both intron and flanking sequences are extremely (A+T) rich and account for most of the low temperature melting. There is no temperature overlap between thermal stabilities of these sequence domains and those of coding DNA. The latter must satisfy triplet codon constraints of higher (G+C) content. These large stability property differences enable a denaturation mapping feature of MELTSIM to clearly distinguish exon positions from those of introns and flanking DNA in long D. discoideum gene containing sequences.
High-resolution derivative melting was used to obtain detailed distributions of local (G + C) contents in a number of ruminant DNAs. Profiles over low (G + C) regions [20-36% (G + C)] are congruent for all ruminants. This region represents 45-50% of the nuclear DNA content and primarily contains intergenic and intron sequences. The high (G + C) region, where most coding sequences are found [38-68% (G + C)], is marked by satellite bands denoting the presence of transcriptionally inert, tandemly repetitive sequence families. These bands can be analyzed for the abundance, base composition, and sequence divergence of satellite families with relatively high precision. Band patterns are unique to each species; even closely related species can be readily distinguished by their base distribution profiles. Variations in nuclear DNA contents in ruminants, determined by flow cytometry, are primarily due to variations in abundances of these repetitive sequence families. Thus, A. alces (moose) is found to have 8.85 +/- 0.2 pg DNA/cell, 25% more than the average in ruminants, while the base distribution curve indicates the presence of an unusually abundant satellite of 52.6% (G + C). The size (1 kb) and sequence of this satellite corresponds to satellite-I of other cervids, and in consequence it is designated Alces-I. The sequence of a cloned repeat of Alces-I has a length of 968 bp, a (G + C) content of 52.6%, and contributes 35%, or almost 3 million copies to the nuclear DNA, exceeding by approximately 300% the average array size of this repeat family in related cervids. In situ hybridization indicates the repeat is distributed throughout centromeric regions of all 62 acrocentric autosomes. Alces-I has much greater-than-expected numbers of GG, GA, and AG and far fewer numbers of TA and CG duplets, characteristics of all tandem repeats. The sequence is judged to be orthologous with satellite-I sequences from Rangifer tarandus (caribou), Capreolus capreolus (roe deer), Muntiacus muntjac (Chinese muntjac) and Muntiacus reevesi (Indian muntjac), as well as Antilocapra americana (pronghorn), and the bovids Bos taurus and Ovis aries. A tentative tree for the five cervids is in excellent agreement with one proposed on the basis of morphological characteristics. Differences from a consensus sequence indicate transversions exceed transitions by almost twofold, suggesting that substitutions occur randomly, or nearly so.
Formamide lowers melting temperatures (Tm) of DNAs linearly by 2.4-2.9 degrees C/mole of formamide (C(F)) depending on the (G+C) composition, helix conformation and state of hydration. The inherent cooperativity of melting is unaffected by the denaturant. dTm/dC(F)for 11 plasmid domains of 0.23 < (G+C)<0.71 generally fit to a linear dependence on (G+C)-content, which, however, is consistent with a (G+C)-independent alteration in the apparent equilibrium constant for thermally induced helix <--> coil transitions. Results indicate that formamide has a destabilizing effect on the helical state, and that sequence-dependent variations in hydration patterns are primarily responsible for small variations in sensitivity to the denaturant. The average unit transition enthalpy delta H(m)[see text for complete expression], exhibits a biphasic dependence on formamide concentration. The initial drop of -0.8 kcal/mol bp at low formamide concentrations is attributable to a delta delta H(m)[see text for complete expression], for exchange of solvent in the vicinity of the helix: displacement by formamide of weakly bound hydrate or counterion. The phenomenological effects are equivalent to lowering the bulk counterion concentration. Poly(dA.dT) exhibits a much lower sensitivity to formamide, due to the specific pattern of tightly bound, immobilized water bridges that buttress the helix from within the narrow minor groove. Tracts of three (A.T)-pairs behave normally, but tracts of six exhibit the same level of reduced sensitivity as the polymer, suggesting a conformational shift as tracts are elongated beyond some critical length [McCarthy J.G. and Rich,A. (1991) Nucleic Acids Res. 19, 3421-3429].
We report the results of a theoretical study, combining the results of sequence analysis and integral equation structural methods for nucleic acids in aqueous solutions, on the effects of nearest neighbors on the (T.G) mispair in solution, for 12 nearest neighbor contexts. Attempts have been made to classify the structural and energetic effects of the 5' and 3' neighbors with respect to the observed spontaneous mutation rates in vertebrates. It is found that 5' nearest neighbor is probably the most critical structural factor in facilitating or discouraging mutations. Local conformational states correlate with discrimination of bases to be excised in mispairs. Our study confirms the role of the flexibility of the DNA molecule in governing the rates of spontaneous mutations.
The (G + C) distribution and the presence and amounts of repetitive sequence families in the white-tailed deer (Odocoileus virginianus) have been examined. The distribution ranges from 20 to 70% (G+C) and shows four distinct repeat families. A 0.7-kb family, DII, corresponds to satellite II in domestic bovids--ox, sheep, and goat--and was singled out for detailed characterization. DII has a prototypic repeat of 67% (G + C), consists of 25,000 tandem copies, and contributes 1.7% to the genomic DNA. Sequencing and electrophoretic analysis indicate a repeat length of 691 bp. These characteristics are similar to those of the bovid satellite II families as well as to those of other cervids that we have examined. The intraspecific sequence divergence within this family has a variance of only 2.5 +/- 0.3%.
The pattern of 20,200 point substitutions in the 16 unique neighbor-pair environments has been determined from aligned gene/pseudogene sequences in the current database of human DNA sequences. Substitution rates, representing averages over those for different regions of the genome, are distributed over a 60-fold range with strong biases in particular neighbor-pair environments. The rates for substitutions involving the CG doublet are the most rapid overall, where changes of the C.G pair vary over a tenfold range depending on the type of substitution and the 5' neighbor-pair. In general, the rates are fastest in alternating purine-pyrimidine sequences and slowest in purine.pyrimidine tracts, suggesting that the frequencies of one or both key molecular misadventures that can occur during replication, dNTP misinsertion and transient misalignment, may be associated with structural alternations and flexibility of the backbone. By contrast, purine.pyrimidine tracts are less flexible, less prone to substitution, and therefore their proportions accumulate in sequences over time. Characteristic biases of the content and arrangement of oligonucleotide strings or tuples in all sequence elements, but particularly in non-coding regions, appear to be due to the pattern of different neighbor-dependent substitution rates. Computer simulations of numerous replicative cycles have been carried out with substitutions occurring on the same schedule found in this study for pseudogenes. Statistical analyses of tuple frequencies at periodic intervals during the simulation experiment indicate that sequences slowly change in lexical complexity toward a quasi-equilibrium state that corresponds to that for introns.
It has been shown that the frequency versus size distribution of A and T overlapping and non-overlapping homopolymer tracts of N > 5 in D. discoideum gene flanking and intron regions are significantly greater than in coding regions(1). In the present report, we demonstrate, that a spatial periodicity exists in long A and T tracts (N > 10) in long flanking sequences by scored alignments of those tracts (N > 10) with the nucleosomal repeat. A tract spacing was found at 185-190 bp that corresponds to a maximum alignment score. This is exactly the average spacing of D. discoideum nucleosomes determined experimentally. A majority of A and T tracts in flanking sequences are often spaced by short DNA stretches and the total length of adjacent A and T tracts plus the interrupting short DNA stretch corresponds closely to the average experimentally measured nucleosomal linker DNA size in D. discoideum-42 bp. These data suggest a model which has A and T runs of N > 10 bp in flanking DNA of D. discoideum organized in a regular phase with nonhomopolymer sequences along the DNA. This model has functional implications for A and T tracts, suggesting that they are found in nucleosomal linker DNA regions of chromatin during some necessary portion(s) of the life of the cell.
The results of this theoretical study combining sequence analysis and minimization with integral equation liquid structural methods indicate that the local sequence context of a T-G wobble mismatch influences the local conformation of the helix, and that conformational alterations are correlated with mutational activity. Studies on the mismatch in four different 5' and 3' neighbor contexts indicate that the nature of the 5' base to the thymine of the mispair is probably the single most critical factor in determining the structural features that facilitate or discourage mutations. When cytosine is the 5' neighbor, the helix adopts a mostly BII conformation, whereas a 5' guanine preserves the canonical BI. Structures that vary little from the BI structure on the incorporation of the mismatch have sequences that correspond to lower rates of transition, whereas those with mostly BII conformations, have sequences with high mutation rates. Subtle variations in stacking patterns around the mismatch precipitate a structural Domino-effect, with a variety of changes in conformation. The helix opens at the mismatch with increased roll angle and propeller twist, causing the thymine to migrate into the major groove and the guanine into the minor groove, exposing the heteroatomic groups to the solvent in the major and minor grooves, respectively, and allowing for some unusual hydrogen bonds. These alterations show a tentative correlation with mutation rates, implying that stacking and structure around the mismatch are important features in the discrimination by proofreading activities of canonical W-C and wobble mismatch base pairs during replication-repair. Variations in the C1'-C1' distances, high propeller twists, changes in the electrostatic complementarity leading to unusual hydrogen bonding patterns probably all correlate with detectability.
D. discoideum, the slime mold, is one of the most AT rich eukaryotic genomes known. In this paper we examine this organism's database for overlapping N-tuples of high frequency and find A and T tracts possess among the highest frequencies in flanking sequences but not in coding sequences. We examined both overlapping and non-overlapping frequencies of the A, T, G and C homopolymer tracts of 2 < N < 6. Overlapping (dG).(dC) and (dA).(dT) tracts occur at greater frequencies than expected, based on random occurrence. Long (dA).(dT) tracts of N > 10 occur at well above expected frequencies in flanking and intron regions, while (dG).(dC) tracts above N = 5 are rarely found. Some of the implications of these findings for tract origins in slip-strand replication and for chromatin structure are discussed.
The numbers and local sequence environments of the two types of substitution mutation plus additions and deletions have been obtained directly in this study from differences between a large number of extant primate gene and pseudogene sequences. A total of 3786 mutations were scored in regions where similarities between pseudogene and corresponding gene sequences is greater than or equal to 85%, comprising approximately 30% of the pseudogene database of 80,584 bp. The pattern of mutations obtained in this fashion is almost identical to that obtained by Li et al. (1984) using a slightly different, more direct approach and with a smaller database. When mutations were scored, the neighbor pairs on the 5' and 3' sides were also noted, leading to a large 16 x 12 matrix of transitions and transversions. Biases of varying magnitude are found in the rates of substitution of the same base pair in different local sequence environments. The overall order for the effect of the 5' neighbor on the rates of substitution mutation of a pyrimidine is A greater than C much greater than T greater than G, and G greater than A greater than T greater than C for the 3' neighbor; where these results represent the average of substitution rates for the complement purine with complement neighbors of bases ordered above. The order for the 3' neighbor is essentially the same for the two transitions and most of the four transversions as well; however, the order for the 5' neighbor is more variable. The overall rate for the C.G----T.A transition is not unusual, however the presence of a 3' neighboring G.C pair boosts the rate substantially, presumably due to specific cytosine methylation of the CG doublet in primate DNAs. The rate of the T.A----C.G transition is also well above average when the 3' neighbor is an A.T, and to a lesser extent a G.C, pair. The latter bias is typical in that it reflects the association of alternating pyrimidine-purine sequences with increasing mutation rates. The substitution of the pyrimidine in a 5'purine-pyrimidine-purine3' sequence generally occurs much faster than in a pyrimidine tract and points to the local conformation as a major determining factor of the substitution rate. An apparent inverse relationship is found between starting and product doublet frequencies of base pairs undergoing mutations with specific 3' neighbors, indicating that differences in intrinsic substitution rates of base pairs with specific neighbors are a key factor in producing the familiar biases of nearest-neighbor frequencies.
Variations in base mono- and dipoles result in variations in stacking energies for the 10 unique neighbor pairs in DNA. Stacking energies for pair M on N, expressed as TMN, were derived by matrix decomposition of a large set of linear algebraic expressions relating the measured Tm for subtransitions emanating from large polymeric DNAs, and the fractional neighbor frequencies, fMN, for the domains responsible for the transitions, Tm = sigma fMNTMN. Tm were determined for subtransitions that dissociate in approximately all-or-none fashion in high resolution melting profiles of partially deleted and recombinant forms of pBR322 DNA. Three different analytical maneuvers were undertaken to resolve subtransitions: site-specific cleavage of domains; deletion of domains; and addition of domains. Three dozen domains of widely divergent, quasi-random neighbor frequencies were identified and assigned, resulting in a unique set of values for TMN with standard deviation, sigma = +/- 0.23 degree C. The average difference between calculated and experimental Tm for domains is only +/- 0.17 degree C, indicating that the thermodynamic properties of these domains are not in any way unusual. Assuming delta S to be constant for all pairs, the corresponding delta HMN are found to have a precision of +/- 10 calories.mol-1 and an accuracy of +/- 606 calories.mol-1. TMN used to calculate melting curves by statistical mechanical analysis of sequences of the different plasmid specimens in this study were in quantitative agreement with observed curves for most sequences. These TMN differ significantly from those determined previously and also correlate poorly with values determined by quantum chemical analysis. Stabilities of neighbor pairs, expressed as the difference in free energy between that for a given pair (MN) and that for the average of like pairs (M, N), depend on the relationship of stacked purines and pyrimidines as follows. delta delta Gpu-py(-466 cal) greater than delta delta Gpu-pu(+52 cal) greater than delta delta Gpy-pu(+335 cal) Differences between experimental Tm and Tm calculated with TMN for the isolated neighbor pairs in the B-conformation are useful in the identification of altered structures and unusual modes of dissociation of helixes. A significantly higher Tm is observed for the highly biased repeated sequence synthetic helixes dA.dT, d(AGC).d(GCT), and d(GAT).d(ATC), reflecting auxiliary sources of stability such as bifurcated hydrogen bonds and/or altered structures for these helixes.
The Tm of internal loop-forming (dA.dT)N domains in pBR322 DNA has been measured over a tenfold range of [Na+]. The slopes SN = dTm/d log [Na+] are linear and decrease in magnitude with decreasing loop size N, signaling a reduction in Na+ released during the transition of these domains to the coil state. Values of SN decrease linearly with increasing N-1 in accordance with the expectation of a simple model for the occurrence of a gradient of long-range electrostatic forces at helix-coil boundaries, and extrapolate almost precisely to the value of S infinity observed for (dA.dT) infinity. These results indicate (1) less counterion is released per phosphate residue from the finite loop than from the infinite-sized loop, and (2) the difference in binding is constant for each boundary formed and independent of the size of the loop within the range examined: approximately 350 base pair (bp) greater than N greater than 71 bp. The slope of the dependence of SN on N-1 indicates the region of higher charge density at the boundary extends at least 18 A into the coil and probably 40-50 A before dropping to a value characteristic of the unperturbed coil. The free energy for excess counterion binding at boundaries can be expressed by -delta G/RT = 10.47 log[Na+] + 5.234 When the loop entropy function in a statistical mechanical algorithm for the dissociation of DNA is weighted by this quantity, calculated Tm are seen to vary by only +/- 0.09 degrees C from observed.
Explore the source record for details and available documents.
Explore the source record for details and available documents.