PubMed HealthSearch

Biomedical subjects

D L Brutlag

Publications and source records attributed to D L Brutlag.

At least 19 recordsLinked to original sources

Identification of protein motifs using conserved amino acid properties and partitioning techniques.

Analyzing a set of protein sequences involves a fundamental relationship between the coherency of the set and the specificity of the motif that describes it. Motifs may be obscured by training sets that contain incoherent sequences, in part due to protein subclasses, contamination, or errors. We develop an algorithm for motif identification that systematically explores possible patterns of coherency within a set of protein sequences. Our algorithm constructs alternative partitions of the training set data, where one subset of each partition is presumed to contain coherent data and is used for forming a motif. The motif is represented by multiple overlapping amino acid groups based on evolutionary, biochemical, or physical properties. We demonstrate our method on a training set of reverse transcriptases that contains subclasses, sequence errors, misalignments, and contaminating sequences. Despite these complications, our program identifies a novel motif for the subclass of retroviral and retrovirus-related reverse transcriptases. This motif has a much higher specificity than previously reported motifs and suggests the importance of conserved hydrophilic and hydrophobic residues in the structure of reverse transcriptases.

Algorithms

Discovering structural correlations in alpha-helices.

We have developed a new representation for structural and functional motifs in protein sequences based on correlations between pairs of amino acids and applied it to alpha-helical and beta-sheet sequences. Existing probabilistic methods for representing and analyzing protein sequences have traditionally assumed conditional independence of evidence. In other words, amino acids are assumed to have no effect on each other. However, analyses of protein structures have repeatedly demonstrated the importance of interactions between amino acids in conferring both structure and function. Using Bayesian networks, we are able to model the relationships between amino acids at distinct positions in a protein sequence in addition to the amino acid distributions at each position. We have also developed an automated program for discovering sequence correlations using standard statistical tests and validation techniques. In this paper, we test this program on sequences from secondary structure motifs, namely alpha-helices and beta-sheets. In each case, the correlations our program discovers correspond well with known physical and chemical interactions between amino acids in structures. Furthermore, we show that, using different chemical alphabets for the amino acids, we discover structural relationships based on the same chemical principle used in constructing the alphabet. This new representation of 3-dimensional features in protein motifs, such as those arising from structural or functional constraints on the sequence, can be used to improve sequence analysis tools including pattern analysis and database search.

Amino Acid Sequence

Discovering side-chain correlation in alpha-helices.

Using a new representation for interactions in protein sequences based on correlations between pairs of amino acids, we have examined alpha-helical segments from known protein structures for important interactions. Traditional techniques for representing protein sequences usually make an explicit assumption of conditional independence of residues in the sequences. Protein structure analyses, however, have repeatedly demonstrated the importance of amino acid interactions for structural stability. We have developed an automated program for discovering sequence correlations in sets of aligned protein sequences using standard statistical tests and for representing them with Bayesian networks. In this paper, we demonstrate the power of our discovery program and representation by analyzing pairs of residues from alpha-helices. The sequence correlations we find represent physical and chemical interactions among amino-acid side chains in helical structures. Furthermore, these local interactions are likely to be important for stabilizing and packing alpha-helices. Lastly, we have also detect correlations in side-chain comformations that indicate important structural interactions but which don't appear as sequence correlations.

Computer Simulation

Detection of correlations in tRNA sequences with structural implications.

Using an flexible representation of biological sequences, we have performed a comparative analysis of 1208 known tRNA sequences. We believe we our technique is a more sensitive method for detecting structural and functional relationships in sets of aligned sequences because we use a flexible representation (for sequences), as well as a general statistical method that can detect a wide range of relationships between positions in a sequence. Our method utilizes functional classifications of the sequence building-blocks (nucleotide bases and amino acids) based on physical or chemical properties. This flexibility in sequence representation improves the significance of finding sequence relationships mediated by the defining property. For example, using a purine/pyrimidine classification, we can detect base-stacking interactions in sets of nucleotide sequences that form base-paired helices. We use several statistical measures, including chi 2-tests, Monte Carlo simulations and an information measure to detect significant correlations in sequences. In this paper we illustrate our method by analyzing a set of tRNA sequences and showing that the correlations our program discovers, in each case, correspond to the known base-pairing and higher order interactions observed in tRNA crystal structures. Furthermore, we show that novel and interesting features of tRNAs are detected when sequence correlations with the charged amino acid (and anticodon) are evaluated. This technique is a powerful method for predicting the structure of RNAs and for analyzing specific functional characteristics.

Base Composition

Knowledge-based simulation of DNA metabolism: prediction of enzyme action.

We have developed a knowledge-based simulation of DNA metabolism that accurately predicts the actions of enzymes on DNA under a large number of environmental conditions. Previous simulations of enzyme systems rely predominantly on mathematical models. We use a frame-based representation to model enzymes, substrates and conditions. Interactions between these objects are expressed using production rules and an underlying truth maintenance system. The system performs rapid inference and can explain its reasoning. A graphical interface provides access to all elements of the simulation, including object representations and explanation graphs. Predicting enzyme action is the first step in the development of a large knowledge base to envision the metabolic pathways of DNA replication and repair.

Adenosine Monophosphate

Improved sensitivity of biological sequence database searches.

We have increased the sensitivity of DNA and protein sequence database searches by allowing similar but non-identical amino acids or nucleotides to match. In addition, one can match k-tuples or words instead of matching individual residues in order to speed the search. A matching matrix species which k-tuples match each other. The matching matrix can be calculated from a similarity matrix of amino acids and a threshold of similarity required for matching. This permits amino acid similarity matrices or replacement matrices (PAM matrices) to be used in the first step of a sequence comparison rather than in a secondary scoring phase. The concept of matching non-identical k-tuples also increases the power of DNA database searches. For example, a matrix that specifies that any 3-tuple in a DNA sequence can match any other 3-tuple encoding the same amino acid permits a DNA database search using a DNA query sequence for regions that would encode a similar amino acid sequence.

Amino Acid Sequence

Conversion and reciprocal exchange between tandem repeats in Drosophila melanogaster.

We have developed an experimental system to assay conversion and reciprocal exchange between tandem repeats in Drosophila melanogaster. In this system, the recombining markers map 0.76 kb apart within the Adh gene, and the length of the repeated unit is 4.75 kb. Our results provide a preliminary record of germline frequencies of gene conversion and unequal exchange between these markers. Conversions involving dispersed repeats were not observed, and may be less frequent. This work demonstrates that conversion takes place at an appreciable frequency between tandem repeats in metazoan germline. It confirms that gene conversion can mediate homogenization of reiterated sequences in higher eukaryotes.

Alcohol Dehydrogenase

Is there a relationship between DNA sequences encoding peptide ligands and their receptors?

It has been suggested that the coding for a ligand and its receptor may have originated in inverse complementary strands of the same DNA. This would imply a deficiency of stop codons in the complementary strand of the ligand message sequence. We have sought evidence of such deficiencies by an analysis of the usage of selected codons in 23 human neuropeptide and hormone mRNA sequences. We have also searched directly for similarities between substance K or substance P and the substance K receptor. Although bovine proopiomelanocortin has an open reading frame for the full extent of the inverse complement of the coding region, this seems to be a unique case. The data as a whole do not support the hypothesis.

Amino Acid Sequence

Expression of the Drosophila type II topoisomerase is developmentally regulated.

The expression of the type II topoisomerase from Drosophila melanogaster was studied during development and in tissue culture cells. RNA blot and protein blot analyses using probes specific for Drosophila topoisomerase II show that the enzyme is developmentally regulated. Levels of both RNA transcript and protein appear highest during early embryogenesis and pupation, periods which are known to have the highest mitotic activity. Tissue culture analysis using Drosophila Kc cells supports these results as levels of topoisomerase II message are higher in rapidly dividing cells than in quiescent cells. Analysis of topoisomerase II levels in early embryos suggests that levels are adequately high for the enzyme to act in DNA replication or segregation at termination of replication. Apparent in vivo proteolysis of topoisomerase II is seen throughout the life cycle, in spite of careful precautions. Whether these proteolytic fragments are important in vivo is still uncertain.

Animals

Identical satellite DNA sequences in sibling species of Drosophila.

The evolution of simple satellite DNAs was examined by DNA-DNA hybridization of ten Drosophila melanogaster satellite sequences to DNAs of the sibling species, Drosophila simulans and Drosophila erecta. Seven of these repeat types are present in tandem arrays in D. simulans and each of the ten sequences is repeated in D. erecta. In thermal melts, six of the seven satellite sequences in D. simulans and seven of the ten sequences in D. erecta melted within 1 deg.C of the corresponding values in D. melanogaster. The remaining sequences melted within 3 deg.C of the homologous hybrids. Therefore, there is little or no alteration in those satellite sequences held in common, despite a period of about ten million years since the divergence of D. melanogaster and D. simulans from a common ancestor. Simple satellite sequences appear to be more highly conserved than coding regions of the genome, on a per nucleotide basis. Since multiple copies of three satellite sequences could not be detected in D. simulans yet are present in D. erecta, a species more distantly related to D. melanogaster than is D. simulans, these sequences show discontinuities in evolution. There were major quantitative variations between species, showing that satellite DNAs are prone to massive amplification or diminution events over timespans as short as those separating sibling species. In D. melanogaster, these sequences amount to 21% of the genome but only 5% in D. simulans and 0.4% in D. erecta. There was a general trend of lower abundance with evolutionary distance for most satellites, suggesting that the amounts of different satellite sequences do not vary independently during evolution.

Animals

Adjacent satellite DNA segments in Drosophila structure of junctions.

The structure of eight satellite DNA molecules containing a junction between tandem arrays of different repeated sequences is described. In one class of junctions there was an abrupt switch with the juxtaposition of two satellite arrays. These arrays were closely related and the periodicity of repeats was maintained in phase across the junction. These arrays usually showed extreme homogeneity in their repeating sequences. A second class of junctions was more complex, and in two cases may have arisen by the insertion of a mobile element into a satellite array. A novel mechanism of satellite formation is proposed to explain the precision of junctions and sequence similarities of neighboring satellite arrays. Homogeneous satellite arrays would be generated enzymatically by synthesis of a repeat using the preceding repeat as template. Occasional errors in copying of the template, either single base changes or misreading the length of the repeat unit, would lead to abrupt switches in the repeating sequence.

Animals

Multiple forms and cellular localization of Drosophila DNA topoisomerase II.

Purified type II topoisomerase from Drosophila melanogaster embryos was reported earlier to contain a major polypeptide of 166,000 daltons and several smaller peptides between 132,000 and 145,000 daltons (Shelton, E. R., Osheroff, N. and Brutlag, D. L. (1983) J. Biol. Chem. 258, 9530-9535). Using purified topoisomerase II we have raised antibodies against the 132,000-166,000-dalton cluster of polypeptides. In this paper we demonstrate that at least three of these polypeptides are also present in embryos immediately upon lysis. Using antigen-affinity purified antibody from the cluster of purified topoisomerase II antigens, we have also discovered several smaller polypeptides in the molecular size range of 30,000-40,000 daltons in embryo extracts. These observations suggest the presence of multiple forms of DNA topoisomerases in the cell. In addition, we demonstrate that purified Drosophila topoisomerase II antibody recognizes yeast topoisomerase II antigens expressed by lambda gt 11-yeast topoisomerase II recombinants (Goto, T. and Wang, J. C. (1984) Cell 36, 1073-1080) establishing a structural homology between yeast and Drosophila enzymes. Antibody preparations were also used to localize the distribution of topoisomerase II in polytene nuclei. In contrast with the distribution of topoisomerase I which is located primarily at puffs, the Drosophila topoisomerase II is distributed generally along the chromosomes paralleling the distribution of DNA itself.

Animals

Multiplicity of satellite DNA sequences in Drosophila melanogaster.

Three Drosophila melanogaster satellite DNAs (1.672, 1.686, and 1.705 g/ml in CsCl), each containing a simple sequence repeated in tandem, were cloned in pBR322 as small fragments about 500 base pairs long. This precaution minimized deletions, since inserts of the same size as the fragments used for cloning were recovered in a stable form. A homogeneous tandem array of one sequence type usually extended the length of the insert. Eleven distinct repeat sequences were discovered, but only one sequence was predominant in each satellite preparation. The remaining classes were minor in amount. The repeat unit lengths were restricted to 5, 7, or 10 base pairs, with sequences closely related. Each sequence conforms to the expression (RRN)m(RN)n, where R is A or G. The multiplicity of simple repeated sequences revealed despite the small sample size suggests that numerous repeat sequences reside in heterochromatin and that particular rules apply to the structure of the repeating sequence.

Animals

Similarities in structure and function of calf thymus and Drosophila casein kinase II.

Both calf and Drosophila contain a type II casein kinase with similar molecular structure and catalytic activity. Purified calf thymus casein kinase II is composed of three subunits of Mr = 44,000 (alpha), 40,000 (alpha'), and 26,000 (beta) (Dahmus, M.E. (1981) J. Biol. Chem. 256, 3319-3325), whereas the Drosophila enzyme is composed of two subunits of Mr = 36,700 (alpha) and 28,200 (beta) (Glover, C. V. C., Shelton, E. R., and Brutlag, D. L. (1983) J. Biol. Chem. 258, 3258-3265). The native form of the enzyme is an alpha 2 beta 2 tetramer. Polyclonal antibodies prepared against each enzyme react with both the alpha and beta subunits of the homologous enzyme and cross-react with both subunits of the heterologous enzyme. Reaction of polyclonal antibodies with proteins resolved by sodium dodecyl sulfate-polyacrylamide gel electrophoresis establishes that no significant difference in subunit molecular weight exists between the purified enzymes and the enzyme present in initial cell extracts. Each antibody effectively inhibits the in vitro activity of the homologous enzyme and causes a slight inhibition in the activity of the heterologous enzyme. Peptide maps derived from purified subunits indicate that the alpha and beta subunits are unique and that there is extensive primary sequence homology between the corresponding subunits of the calf and Drosophila enzyme. Casein kinase II from both sources phosphorylates the same subunits of calf thymus RNA polymerase II and an identical set of proteins in a complex mixture of acid-soluble proteins from Drosophila tissue culture cells. The striking similarity in molecular structure and catalytic activity between the calf and Drosophila enzyme suggests that casein kinase II has been highly conserved in evolution.

Animals

Rapid searches for complex patterns in biological molecules.

The intrinsic redundancy of genetic information makes searching for patterns in biological sequences a difficult task. We have designed an interactive self-documenting computer program called QUEST that allows rapid searching of large DNA and protein data banks for highly redundant consensus sequences or character patterns. QUEST uses a concise language for specifying character patterns containing several levels of ambiguity and pattern arrangement. Examples of the use of this program for sequence data are given. Details of the algorithm and pattern optimization are explained.

Amino Acid Sequence

DNA topoisomerase II from Drosophila melanogaster. Purification and physical characterization.

A type II DNA topoisomerase has been purified from the nuclei of Drosophila melanogaster 6- to 18-h-old embryos. The enzyme, as assayed by its ability to catenate supercoiled DNA, behaved as a single homogeneous species throughout the procedure and the yield was approximately 0.5 mg of protein/100 g of dechorionated embryos. The final product was entirely ATP-dependent and free of topoisomerase I, endonuclease and protease activities. The purified topoisomerase II had a Stokes radius of 69 A and a sedimentation coefficient (S20,w) of 9.2 S, leading to a calculated native molecular weight of approximately 261,000. The protein consists of a single polypeptide of molecular weight 166,000, as determined by electrophoresis on sodium dodecyl sulfate-polyacrylamide gels. Taken together with the above hydrodynamic studies, the Drosophila enzyme is probably a homodimer, as has been observed for other eukaryotic type II enzymes. Thus, it appears that during the course of evolution the heterologous subunits which comprise bacterial type II topoisomerases have been combined into a single polypeptide chain in eukaryotes.

Animals

DNA topoisomerase II from Drosophila melanogaster. Relaxation of supercoiled DNA.

In order to study the double-strand DNA passage reaction of eukaryotic type II topoisomerases, a quantitative assay to monitor the enzymic conversion of supercoiled circular DNA to relaxed circular DNA was developed. Under conditions of maximal activity, relaxation catalyzed by the Drosophila melanogaster topoisomerase II was processive and the energy of activation was 14.3 kcal . mol-1. Removal of supercoils was accompanied by the hydrolysis of either ATP or dATP to inorganic phosphate and the corresponding nucleoside diphosphate. Apparent Km values were 200 microM for pBR322 plasmid DNA, 140 microM for SV40 viral DNA, 280 microM for ATP, and 630 microM for dATP. The turnover number for the Drosophila enzyme was at least 200 supercoils of DNA relaxed/min/molecule of topoisomerase II. The enzyme interacts preferentially with negatively supercoiled DNA over relaxed molecules, is capable of removing positive superhelical twists, and was found to be strongly inhibited by single-stranded DNA. Kinetic and inhibition studies indicated that the beta and gamma phosphate groups, the 2'-OH of the ribose sugar, and the C6-NH2 of the adenine ring are important for the interaction of ATP with the enzyme. While the binding of ATP to Drosophila topoisomerase II was sufficient to induce a DNA strand passage event, hydrolysis was required for enzyme turnover. The ATPase activity of the topoisomerase was stimulated 17-fold by the presence of negatively supercoiled DNA and approximately 4 molecules of ATP were hydrolyzed/supercoil removed. Finally, a kinetic model describing the switch from a processive to a distributive relaxation reaction is presented.

Animals

Purification and characterization of a type II casein kinase from Drosophila melanogaster.

A cyclic nucleotide-independent protein kinase has been isolated from Drosophila melanogaster by chromatography on phosphocellulose and hydroxylapatite followed by gel filtration and glycerol gradient sedimentation. As determined by sodium dodecyl sulfate gel electrophoresis, the purified enzyme is greater than 95% homogeneous and is composed of two distinct subunits, alpha and beta, having Mr = 36,700 and 28,200, respectively. The native form of the enzyme is an alpha 2 beta 2 tetramer having a Stokes radius of 48 A, a sedimentation coefficient of 6.4 S, and Mr approximately 130,000. The purified kinase undergoes an autocatalytic reaction resulting in the specific phosphorylation of the beta subunit, exhibits a low apparent Km for both ATP and GTP as nucleoside triphosphate donor (17 and 66 microM, respectively), phosphorylates both casein and phosvitin but neither histones nor protamine, modifies both serine and threonine residues in casein, and is strongly inhibited by heparin (I50 = 21 ng/ml). These properties are remarkably similar to those of casein kinase II, an enzyme previously described in several mammalian and avian species. The strong similarities among the insect, avian, and mammalian enzymes suggest that casein kinase II has been highly conserved during evolution.

Animals