PubMed HealthSearch

Biomedical subjects

A F Neuwald

Publications and source records attributed to A F Neuwald.

14 recordsLinked to original sources

Extracting protein alignment models from the sequence database.

Biologists often gain structural and functional insights into a protein sequence by constructing a multiple alignment model of the family. Here a program called Probe fully automates this process of model construction starting from a single sequence. Central to this program is a powerful new method to locate and align only those, often subtly, conserved patterns essential to the family as a whole. When applied to randomly chosen proteins, Probe found on average about four times as many relationships as a pairwise search and yielded many new discoveries. These include: an obscure subfamily of globins in the roundworm Caenorhabditis elegans ; two new superfamilies of metallohydrolases; a lipoyl/biotin swinging arm domain in bacterial membrane fusion proteins; and a DH domain in the yeast Bud3 and Fus2 proteins. By identifying distant relationships and merging families into superfamilies in this way, this analysis further confirms the notion that proteins evolved from relatively few ancient sequences. Moreover, this method automatically generates models of these ancient conserved regions for rapid and sensitive screening of sequences.

Algorithms

A superfamily of conserved domains in DNA damage-responsive cell cycle checkpoint proteins.

Computer analysis of a conserved domain, BRCT, first described at the carboxyl terminus of the breast cancer protein BRCA1, a p53 binding protein (53BP1), and the yeast cell cycle checkpoint protein RAD9 revealed a large superfamily of domains that occur predominantly in proteins involved in cell cycle checkpoint functions responsive to DNA damage. The BRCT domain consists of approximately 95 amino acid residues and occurs as a tandem repeat at the carboxyl terminus of numerous proteins, but has been observed also as a tandem repeat at the amino terminus or as a single copy. The BRCT superfamily presently includes approximately 40 nonorthologous proteins, namely, BRCA1, 53BP1, and RAD9; a protein family that consists of the fission yeast replication checkpoint protein Rad4, the oncoprotein ECT2, the DNA repair protein XRCC1, and yeast DNA polymerase subunit DPB11; DNA binding enzymes such as terminal deoxynucleotidyltransferases, deoxycytidyl transferase involved in DNA repair, and DNA-ligases III and IV; yeast multifunctional transcription factor RAP1; and several uncharacterized gene products. Another previously described domain that is shared by bacterial NAD-dependent DNA-ligases, the large subunits of eukaryotic replication factor C, and poly(ADP-ribose) polymerases appears to be a distinct version of the BRCT domain. The retinoblastoma protein (a universal tumor suppressor) and related proteins may contain a distant relative of the BRCT domain. Despite the functional diversity of all these proteins, participation in DNA damage-responsive checkpoints appears to be a unifying theme. Thus, the BRCT domain is likely to perform critical, yet uncharacterized, functions in the cell cycle control of organisms from bacteria to humans. The carboxyterminal BRCT domain of BRCA1 corresponds precisely to the recently identified minimal transcription activation domain of this protein, indicating one such function.

Amino Acid Sequence

Detection of likely transmembrane beta strand regions in sequences of mitochondrial pore proteins using the Gibbs sampler.

The mitochondrial channel VDAC is presumed to fold as a beta-barrel although the number and identity of transmembrane beta-strands in the protein are controversial. Previously, a novel multiple alignment algorithm called the Gibbs sampler was used to detect a residue-frequency motif in sequences of bacterial outer-membrane proteins that corresponds to transmembrane beta-strands in bacterial porins of known structure (Neuwald et al., 1995, Protein Science, 4, 1618. In the present study, this bacterial motif has been used to screen sets of mitochondrial membrane protein sequences, with matches occurring in only two classes of proteins: VDACs and the outer-membrane protein import pore (1SP42, M0M38). These results suggest a structural (and perhaps evolutionary) relatedness between the bacterial and mitochondrial pore proteins, with the mitochondrial subsequences that match the bacterial motif corresponding to transmembrane beta-strands as in the porins.

Amino Acid Sequence

Gibbs motif sampling: detection of bacterial outer membrane protein repeats.

The detection and alignment of locally conserved regions (motifs) in multiple sequences can provide insight into protein structure, function, and evolution. A new Gibbs sampling algorithm is described that detects motif-encoding regions in sequences and optimally partitions them into distinct motif models; this is illustrated using a set of immunoglobulin fold proteins. When applied to sequences sharing a single motif, the sampler can be used to classify motif regions into related submodels, as is illustrated using helix-turn-helix DNA-binding proteins. Other statistically based procedures are described for searching a database for sequences matching motifs found by the sampler. When applied to a set of 32 very distantly related bacterial integral outer membrane proteins, the sampler revealed that they share a subtle, repetitive motif. Although BLAST (Altschul SF et al., 1990, J Mol Biol 215:403-410) fails to detect significant pairwise similarity between any of the sequences, the repeats present in these outer membrane proteins, taken as a whole, are highly significant (based on a generally applicable statistical test for motifs described here). Analysis of bacterial porins with known trimeric beta-barrel structure and related proteins reveals a similar repetitive motif corresponding to alternating membrane-spanning beta-strands. These beta-strands occur on the membrane interface (as opposed to the trimeric interface) of the beta-barrel. The broad conservation and structural location of these repeats suggests that they play important functional roles.

Algorithms

Detecting patterns in protein sequences.

The detection of conserved sequence patterns (motifs) in related proteins often yields valuable structural and functional insights. We describe a method that utilizes rigorous statistics and a depth-first search procedure to efficiently and exhaustively search a set of proteins for significant patterns up to a specified length. Additional procedures classify related patterns into groups and identify protein segments most likely to share a common motif. The utility of the method was demonstrated on several difficult test problems; detection of motifs among 56 proteins in the acyltransferase family, detection of a dinucleotide-binding fold present within a small subset of a set of 91 distantly related and unrelated proteins, detection of the helix-turn-helix motif in 15 distantly related proteins and detection of subtle internal repeats in a prenyltransferase. In a search of a large set of sequences for internal repeats, the method detected novel ankyrin-like repeats in an Escherichia coli protein.

Acetyltransferases

Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.

A wealth of protein and DNA sequence data is being generated by genome projects and other sequencing efforts. A crucial barrier to deciphering these sequences and understanding the relations among them is the difficulty of detecting subtle local residue patterns common to multiple sequences. Such patterns frequently reflect similar molecular structures and biological properties. A mathematical definition of this "local multiple alignment" problem suitable for full computer automation has been used to develop a new and sensitive algorithm, based on the statistical method of iterative sampling. This algorithm finds an optimized local alignment model for N sequences in N-linear time, requiring only seconds on current workstations, and allows the simultaneous detection and optimization of multiple patterns and pattern repeats. The method is illustrated as applied to helix-turn-helix proteins, lipocalins, and prenyltransferases.

Algorithms

Conditional dihydrofolate reductase deficiency due to transposon Tn5tac1 insertion downstream from the folA gene in Escherichia coli.

Transposon Tn5tac1 can generate conditional mutations by virtue of an outward-facing tac promoter, which is regulated by the lac repressor and isopropyl-beta-D-thiogalactopyranoside (IPTG). We report here on a Tn5tac1 insertion in Escherichia coli that results in a conditional (IPTG-elicited) folA mutant phenotype: During aerobic growth, IPTG caused decreased synthesis of dihydrofolate reductase (DHFR; encoded by the folA gene) and hypersensitivity to trimethoprim (a DHFR inhibitor); during anaerobic growth, IPTG elicited auxotrophy that was satisfied by thymine or glycine or threonine. The Tn5tac1 insertion was downstream from folA, with the tac promoter pointing into the gene (antisense direction). Complementation tests indicated that the conditional folA deficiency was a cis effect of transcription from the tac promoter, perhaps due to head-to-head collision between converging RNA polymerases.

Alleles

Mutational analysis of the Escherichia coli serB promoter region reveals transcriptional linkage to a downstream gene.

Genes encoding proteins with unrelated functions can be cotranscribed, and this may be used by cells to coordinate different metabolic pathways during growth. We describe a gene, designated sms, which is downstream from the serine biosynthetic gene serB in Escherichia coli but does not appear to be involved in amino acid (aa) biosynthesis. The sms gene is 1380 bp long. The Sms product migrates at 55 kDa on sodium dodecyl sulfate(SDS)-polyacrylamide gels and has a M(r) of 49472 (460 aa residues) calculated from the nucleotide sequence. The deduced Sms aa sequence shares regions of similarity with two ATP-dependent proteases, Lon and RecA, and contains two motifs: a C-x(2)-C-x(n)-C-x(2)-C motif, which is found in some nucleic acid binding proteins, and an ATP/GTP binding site motif. Insertional inactivation of sms led to increased sensitivity to the alkylating agent methylmethane sulfonate, but not to a requirement for serine or other metabolites. Several promoter mutations were isolated and characterized, which suggest that serB has a typical promoter recognized by sigma 70. After the serB coding sequence there is a 48-bp region with no obvious promoter sequence preceding the sms translation start codon. Analyses using sms'-lacZ fusions cloned downstream from wild-type and mutant serB promoters showed that sms is cotranscribed with serB.

ATP-Dependent Proteases

cysQ, a gene needed for cysteine synthesis in Escherichia coli K-12 only during aerobic growth.

The initial steps in assimilation of sulfate during cysteine biosynthesis entail sulfate uptake and sulfate activation by formation of adenosine 5'-phosphosulfate, conversion to 3'-phosphoadenosine 5'-phosphosulfate, and reduction to sulfite. Mutations in a previously uncharacterized Escherichia coli gene, cysQ, which resulted in a requirement for sulfite or cysteine, were obtained by in vivo insertion of transposons Tn5tac1 and Tn5supF and by in vitro insertion of resistance gene cassettes. cysQ is at chromosomal position 95.7 min (kb 4517 to 4518) and is transcribed divergently from the adjacent cpdB gene. A Tn5tac1 insertion just inside the 3' end of cysQ, with its isopropyl-beta-D-thiogalactopyranoside-inducible tac promoter pointed toward the cysQ promoter, resulted in auxotrophy only when isopropyl-beta-D-thiogalactopyranoside was present; this conditional phenotype was ascribed to collision between converging RNA polymerases or interaction between complementary antisense and cysQ mRNAs. The auxotrophy caused by cysQ null mutations was leaky in some but not all E. coli strains and could be compensated by mutations in unlinked genes. cysQ mutants were prototrophic during anaerobic growth. Mutations in cysQ did not affect the rate of sulfate uptake or the activities of ATP sulfurylase and its protein activator, which together catalyze adenosine 5'-phosphosulfate synthesis. Some mutations that compensated for cysQ null alleles resulted in sulfate transport defects. cysQ is identical to a gene called amtA, which had been thought to be needed for ammonium transport. Computer analyses, detailed elsewhere, revealed significant amino acid sequence homology between cysQ and suhB of E. coli and the gene for mammalian inositol monophosphatase. Previous work had suggested that 3'-phosphoadenoside 5'-phosphosulfate is toxic if allowed to accumulate, and we propose that CysQ helps control the pool of 3'-phosphoadenoside 5'-phosphosulfate, or its use in sulfite synthesis.

Aerobiosis

Diverse proteins homologous to inositol monophosphatase.

Bovine inositol monophosphatase (IMP) and several homologous proteins were found to share two sequence motifs with bovine inositol polyphosphate 1-phosphatase (IPP). These motifs may correspond to binding sites within IMP and IPP for inositol phosphates or for lithium, since both substances are bound by these proteins. This suggests that the proteins homologous to IMP, which have diverse biological roles but whose function is not clear, may act by enhancing the synthesis or degradation of phosphorylated compounds.

Amino Acid Sequence

IS30 activation of an smp'-lacZ gene fusion in Escherichia coli.

The Escherichia coli serB gene is divergently transcribed from the gene. Six independently isolated IS30 insertions within the 5' end of the serB structural gene resulted in increased expression of an smp'-lacZ gene fusion. DNA sequence analysis of two of the insertions suggests that a promoter was created upon IS30 insertion.

Base Sequence

An Escherichia coli membrane protein with a unique signal sequence.

The smp gene, whose promoter divergently overlaps with the Escherichia coli serB promoter, encodes a 24-kDa membrane protein of unknown function. An smp fusion to the lacZ gene was constructed to select for smp promoter up-mutations. The mutations isolated, however, were located within the smp structural gene rather than within the promoter region. A 15-bp deletion and two different point mutations, corresponding to the hydrophobic region of the Smp signal sequence, resulted in a Lac+ phenotype. The point mutations introduced a positively charged amino acid into the hydrophobic region. The Smp signal sequence is unique in that it is encoded by an mRNA which has the potential to form a 9-bp hairpin structure with a 39-nt loop, that conforms to the idealized mirror symmetric sequence GC(RGUGR)(YUGUY)5 (RGUGR)CG. This highly symmetrical region may encode additional intragenic information important for the expression of smp. A computer search revealed that the smp gene product shares homology with elongation factor Ts and ribosomal subunit L4, two components of the E. coli translational apparatus.

Amino Acid Sequence

DNA sequence and characterization of the Escherichia coli serB gene.

We have determined the sequence of a DNA fragment containing the Escherichia coli serB gene. An open reading frame of 966 nucleotides was identified that encodes a polypeptide of 322 amino acids with a molecular weight of 35,002 daltons. The transcription start site was determined by Mung Bean nuclease mapping. The -10 and -35 regions of the serB promoter lack homology to the consensus sequences. In addition, the -35 region of the serB promoter overlaps the -35 region of a second divergent promoter. Frameshift mutations were constructed at three different sites within the serB gene. When plasmids carrying these mutations were used as templates in a minicell system, mutations closer to the proposed transcription and translation start sites resulted in smaller polypeptides than those further away, confirming the proposed direction of transcription and translation. The observed sizes of the truncated and native polypeptides were in agreement with those predicted from the DNA sequence. A very stable stem and loop structure (delta G= -32 kcal/mole) that does not fit the criteria of known transcription terminators was found one nucleotide downstream from the putative UAA translation stop codon.

Amino Acid Sequence