PubMed Health⌕ Search

Biomedical subjects

Douglas Poland

Publications and source records attributed to Douglas Poland.

15 recordsLinked to original sources

Enthalpy distribution functions for protein-DNA complexes: example of the binding of AT-hooks to target DNA.

In this article we use the published heat capacity data of Dragan et al. [A.I. Dragan, et al., The energetics of specific binding of AT-hooks from HMGA1 to target DNA, J. Mol. Biol. 327 (2003) 393-411] on the association of proteins with DNA duplexes to construct enthalpy probability distributions for the protein/DNA complexes formed in these systems. We first analyze the multistep equilibrium that determines the species concentrations in this system to determine whether or not the DNA-peptide complex goes cleanly to DNA single-strands and peptide. Using the heat capacity data for this case we employ the maximum-entropy method to construct enthalpy probability distribution functions for the species involved in this equilibrium. We find that the distribution functions for this system clearly show bimodal behavior indicating a two-state transition from complex to non-complex form.

AT-Hook Motifs↗

Enthalpy distribution functions for the unwinding of a short DNA duplex.

In this article we use the published heat capacity data of Dragan et al. (J Mol Biol 2003, 327, 293-411) for a short DNA duplex to calculate the enthalpy probability distribution for this species as a function of temperature. Our approach is based on a procedure that we developed (Poland, D. J Chem Phys 2000, 112, 6554) whereby one obtains moments of the enthalpy distribution from the temperature dependence of the heat capacity. One then uses the maximum-entropy method to construct the enthalpy probability distribution from the set of enthalpy moments. For the DNA duplex treated here the heat capacity goes through a maximum as a function of temperature reflecting the unwinding of the duplex structure. In the neighborhood of the heat capacity maximum, the enthalpy distribution functions show a clear bimodal structure, indicating the coexistence of two distinct states, the duplex and the single-strand state. The probabilities of theses two states can be estimated from the enthalpy distribution functions and can be used to calculate the temperature dependence of the equilibrium constant for the unwinding of the DNA duplex. This example illustrates that the temperature dependence of the heat capacity can be used to give a detailed picture of conformational transitions in biological macromolecules. In particular, the structure of the enthalpy distribution in this case allows one to see the temperature evolution of the two-state distribution in detail.

DNA↗

Universal scaling of the C--G distribution of genes.

Using our previous result that the C--G distribution in genomes is very broad, varying as a power law of the size of the block of genome considered, we examine the C--G distribution in genes themselves. We show that the widths of the C--G distributions for the genes of several simple organisms also vary as power laws. This suggests that the power law behavior gives a universal scaling whereby the distributions for the C--G content of the genes from all species are mapped onto a single function.

Algorithms↗

Energy distributions of gallium nanoclusters.

Starting with the heat-capacity data of Breaux et al., [J. Am. Chem. Soc. 126, 8629 (2004)] we use the maximum-entropy method to calculate energy distribution functions for gallium-ion nanoclusters over a wide temperature range (100-1050 K). Specifically, we calculate energy distributions for clusters containing n = 39 and n = 45 gallium atoms. For the case of n = 39 clusters the energy distribution gets systematically broader as a function of temperature with no indication of any marked structural change in the cluster. On the other hand, the energy distribution for the n = 45 cluster first gets broader as a function of temperature but then gets narrower again as the temperature is further increased, indicating that there is some kind of structural transition taking place in this cluster species.

Journal Article↗

The phylogeny of persistence in DNA.

We continue our study, Poland [Biophysical Chemistry 110 (2004) 59-2], of the distribution of C or G (C-G for short) in the DNA of select organisms, in particular, the tendency for C-G to cluster on all scales with respect to the number of bases considered. We previously found that if we counted the number of C-G bases in consecutive, nonoverlapping boxes containing a total of m bases, then the width of the distribution function describing how many C-G bases are in a box increases with respect to m dramatically relative to the width expected for a random distribution. The relative width of the C-G composition distribution function was found to vary accurately as a power law with respect to m, the size of the box, over a very wide range of m values. We express the power law in terms of a characteristic exponent gamma, that is, the relative widths of the distributions vary as m(gamma). The enhanced relative width of the distribution functions is a direct consequence of the tendency for boxes of similar composition to follow one another. This tendency represents persistence in composition from box to box and hence we refer to gamma as the persistence exponent. The occurrence of a power law means that the tendency for C-G to cluster is present on all scales of sequence length (box size) up to the total length of the chromosome which for bacteria is the entire genome. The persistence exponent gamma that characterizes the power law is thus an important parameter describing the distribution of C-G on all scales from individual base pairs up to the total length of the DNA sample considered. In the present paper, we determine the characteristic exponent gamma and the associated fractal dimension of DNA samples for a selection of species representing all of the major types of organism, that is, we explore the phylogeny of the exponent gamma. Here we treat six prokaryotes and six eukaryotes which, together with the species we have previously treated, brings the total number of species we have examined to 15. We find the power law form for the C-G distribution for all of the species treated and hence this behavior seems to be ubiquitous. The values of the characteristic exponent gamma that we find tend to cluster around the value gamma=0.20 with no obvious pattern with respect to phylogeny. The extreme values that we obtain are gamma=0.057 (yeast) and gamma=0.386 (human). We conclude by showing that the persistence of C-G clustering on the scale of the length of a chromosome is dramatically illustrated by interpreting the C-G distribution as a random walk.

Animals↗

The persistence exponent of DNA.

Using the complete genome of Thermoplasma volcanium, as an example, we have examined the distribution functions for the amount of C or G in consecutive, non-overlapping blocks of m bases in this system. We find that these distributions are very much broader (by many factors) than those expected for a random distribution of bases. If we plot the widths of the C-G distributions relative to the widths expected for random distributions, as a function of the block size used, we obtain a power law with a characteristic exponent. The broadening of the C-G distributions follows from the empirical finding that blocks containing a given C-G content tend to be followed by blocks of similar C-G content thus indicating a statistical persistence of composition. The exponent associated with the power law thus measures the strength of persistence in a given DNA. This behavior can be understood using Mandelbrot's model of a fractional Brownian walk. In this model there is a hierarchy of persistence (correlation between blocks) between all parts of the system. The model gives us a way to scale the C-G distributions such that all these functions are collapsed onto a master curve. For a fractional Brownian walk, the fractal dimension of the C-G distribution is simply related to the persistence exponent for the power law. The persistence exponent for T. volcanium is found to be gamma = 0.29 while for a 10 million base segment of the human genome we obtain gamma = 0.39, similar to but not identical with the value found for the microbe.

Base Composition↗

DNA melting profiles from a matrix method.

In this article we give a new method for the calculation of DNA melting profiles. Based on the matrix formulation of the DNA partition function, the method relies for its efficiency on the fact that the required matrices are very sparse, essentially reducing matrix multiplication to vector multiplication and thus making the computer time required to treat a DNA molecule containing N base pairs proportional to N(2). A key ingredient in the method is the result that multiplication by the inverse matrix can also be reduced to vector multiplication. The task of calculating the melting profile for the entire genome is further reduced by treating regions of the molecule between helix-plateaus, thus breaking the molecule up into independent parts that can each be treated individually. The method is easily modified to incorporate changes in the assignment of statistical weights to the different structural features of DNA. We illustrate the method using the genome of Haemophilus influenzae.

Algorithms↗

Long-range correlations in the helix free energy distribution in DNA.

In this paper we explore the free energy distribution in the helical form of DNA using the genome of the virus Rickettsia prowazekii Madrid E as an example. The genome of this organism has been determined by Andersson et al. (Nature 396 (1998) 133) and is available on the World Wide Web (www.tigr.org). Using the helix statistical weights based on nearest-neighbor base pairs of SantaLucia (Proc. Natl. Acad. Sci. USA 95 (1998) 1460), we calculate the free energy in consecutive blocks of m base pairs in the DNA sequence and then construct the free energy distribution for these values. Using the maximum-entropy method we can fit the distribution curves with a function based on the moments of the distribution. For blocks containing 10-20 base pairs the distribution is slightly skewed and we require four moments to accurately fit the function. For blocks containing 100 base pairs or more, the distribution is well approximated by a Gaussian function based on the first two moments of the distribution. We find that the free energy distribution for m=20 can be reproduced using random sequences that have the local (singlet, doublet or triplet) statistics of Rickettsia. However, for much larger blocks, for example m=500, the width of the free energy distribution based on the actual Rickettsia genome is broader by almost a factor of 3 than the distributions based on random local statistics. We find that the distribution functions for the C or G content in blocks of m base pairs have almost the same behavior as a function of block size as do the free energy distributions. In order to duplicate the width of the distribution functions based on the actual Rickettsia sequence, we need to introduce tables (matrices) that correlate the states of consecutive blocks hundreds of base pairs long. This indicates that correlations on the order of the number of base pairs contained in the average gene are required to give the actual widths for either the C or G content or the helix free energy distributions. Above a certain m value, the distributions for larger m can be accurately expressed in terms of the distribution functions for smaller m. Thus, for example, the distribution for m=5000 can be expressed in terms of the generating function for m=1000.

Base Composition↗

DNA probability profiles: examples from the Treponema pallidum genome.

In this paper we apply an algorithm developed by Poland (Biopolymers 13 (1974) 1859) to treat the statistical mechanics of the thermal unwinding of DNA to the genome of Treponema pallidum, the syphilis spirochete. We calculate probability profiles (giving the probability that each unit in the molecule is in the helix-state) and other statistical distributions for genes and sequences of genes, the longest containing 100 genes and 107,139 base pairs (approximately 10% of the genome).

Algorithms↗

Free energy of proton binding in proteins.

In this article we use literature data on the titration of denatured ribonuclease to test the accuracy of proton-binding distributions obtained using our recent approach employing moments. We find that using only the local slope of the titration curve at a small number of points (five, for example) we can reproduce the detailed proton-binding distribution at all pH values. Our method gives the complete proton-binding polynomial for a given protein and each coefficient in this polynomial in turn yields the free energy for binding a given number of protons in all ways to the protein. Using these net free energies, we can then compute the average proton-binding free energy per proton as a function of the fraction of protons bound. We find that this function is remarkably similar for different proteins, even for proteins that exhibit quite different titration behavior. For the special case of binding to independent sites, we obtain simple relations for the first and last terms in the free energy per-proton function. For this special case we also can calculate the distribution functions giving the probability that a molecule has a given number of positive or negative charges and the joint distribution that a molecule simultaneously has a given number of positive and negative charge.

Animals↗

Maximum-entropy calculation of free energy distributions in tRNAs.

We have previously shown that the distribution function describing the probability that a biological macromolecule picked at random has a particular enthalpy value can be calculated from the temperature dependence of the heat capacity of the macromolecule. The free energy as a function of enthalpy (free energy distribution) can then be determined from the enthalpy probability distribution. In addition, the free energy distribution at an arbitrary temperature can be calculated from the free energy distribution at a reference temperature. Here we apply this approach to a family of similar macromolecules, namely a set of transfer RNAs, specifically tRNA(Phe), tRNA(Val), tRNA(Met), tRNA(Ser), tRNA(Asp) and tRNA(Ile). Using published heat-capacity data, we calculate the enthalpy probability distribution functions for all of these molecules at five different temperatures covering the range from 30 to 80 degrees C. We then use these distributions to give a reference free-energy distribution, from which the thermodynamics at any temperature can be calculated for each species. We compare the reference free-energy distribution for the five tRNAs and find that, while the overall form of the distributions is similar, the local behavior of the functions varies considerably between the species.

Entropy↗

Contribution of secondary structure to the heat capacity and enthalpy distribution of the unfolded state in proteins.

We have recently shown that one can construct the enthalpy distribution for protein molecules from experimental knowledge of the temperature dependence of the heat capacity. For many proteins the enthalpy distribution evaluated at the midpoint of the denaturation transition (corresponding to the maximum in the heat capacity vs temperature curve) is broad and biphasic, indicating two different populations of molecules (native and unfolded) with distinctly different enthalpies. At temperatures above the denaturation point, the heat capacity for the unfolded state in many proteins is quite large and using the analysis just mentioned, we obtain a gaussian-like enthalpy distribution that is very broad. A large value of the heat capacity indicates that there are structural changes going on in the unfolded state above the transition temperature. In the present paper we investigate the origin of this large heat capacity by considering the presence of changing amounts of secondary structure (specifically, alpha-helix) in the unfolded state. For this purpose we use the empirical estimates of the Zimm-Bragg sigma and s factors for all of the native amino acids in water as determined by Scheraga and co-workers. Using myoglobin as an example, we calculate probability profiles and distribution functions for the total number of helix states in the specific-sequence molecule. Given the partition function for the specific-sequence molecule, we can then calculate a set of enthalpy moments for the molecule from which we obtain a good estimate of the enthalpy distribution in the unfolded state. This distribution turns out to be quite narrow when compared with the distribution obtained from the raw heat capacity data. We conclude that there must be other major structural changes (backbone and solvent) that are not accounted for by the inclusion of alpha-helix in the unfolded state.

Animals↗

Maximum-entropy calculation of free energy distributions for two forms of myoglobin.

The temperature dependence of the heat capacity of myoglobin depends dramatically on pH. At low pH (near 4.5), there are two weak maxima in the heat capacity at low and intermediate temperatures, respectively, whereas at high pH (near 10.7), there is one strong maximum at high temperature. Using literature data for the low-pH form (Hallerbach and Hinz, 1999) and for the high-pH form (Makhatadze and Privalov, 1995), we applied a recently developed technique (Poland, 2001d) to calculate the free energy distributions for the two forms of the protein. In this method, the temperature dependence of the heat capacity is used to calculate moments of the protein enthalpy distribution function, which in turn, using the maximum-entropy method, are used to construct the actual distribution function. The enthalpy distribution function for a protein gives the fraction of protein molecules in solution having a given value of the enthalpy, which can be interpreted as the probability that a molecule picked at random has a given enthalpy value. Given the enthalpy distribution functions at several temperatures, one can then construct a master free energy function from which the probability distributions at all temperatures can be calculated. For the high-pH form of myoglobin, the enthalpy distribution function that is obtained exhibits bimodal behavior at the temperature corresponding to the maximum in the heat capacity (Poland, 2001a), reflecting the presence of two populations of molecules (native and unfolded). For this form of myoglobin, the temperature evolution of the relative probabilities of the two populations can be obtained in detail from the master free energy function. In contrast, the enthalpy distribution function for the low-pH form of myoglobin does not show any special structure at any temperature. In this form of myoglobin the enthalpy distribution function simply exhibits a single maximum at all temperatures, with the position of the maximum increasing to higher enthalpy values as the temperature is increased, indicating that in this case there is a continuous evolution of species rather than a shift between two distinct population of molecules.

Calorimetry↗

Protein denaturant binding polynomials.

We show how moments of the denaturant binding distribution function can be extracted from experimental data on the denaturation of a protein as a function of the concentration of denaturant and how in turn these moments can be used to construct the denaturant binding distribution function. This approach is similar to our recent work on using the maximum-entropy method to construct ligand-binding distributions from moments obtained from titration curves for nucleic acids and proteins. As an example we take literature data on the denaturation of ferro- and ferricytochrome c by guanidine hydrochloride and from it construct the denaturant binding polynomial and binding distribution function for the unfolded protein.

Cytochrome c Group↗