PubMed Health⌕ Search

Biomedical subjects

A Marchler-Bauer

Publications and source records attributed to A Marchler-Bauer.

13 recordsLinked to original sources

Characterization of p40/GPR69A as a peripheral membrane protein related to the lantibiotic synthetase component C.

The 40 kDa erythrocyte membrane protein p40/GPR69A, previously assigned to the G-protein-coupled receptor superfamily, was now identified by peptide-antibodies and characterized as a loosely associated peripheral membrane protein. This result is in striking contrast to the proposed seven-transmembrane protein structure and function and therefore we wish to correct our previous proposal. p40 is located at the cytoplasmic side of the membrane and is neither associated with the cytoskeleton nor lipid rafts. Refined sequence analysis revealed that p40 is related to the LanC family of bacterial membrane-associated proteins which are involved in the biosynthesis of antimicrobial peptides. Therefore, we rename p40 to LanC-like protein 1 (LANCL1) and suggest that it may play a similar role as a peptide-modifying enzyme component in eukaryotic cells.

Amino Acid Sequence↗

Combination of threading potentials and sequence profiles improves fold recognition.

Using a benchmark set of structurally similar proteins, we conduct a series of threading experiments intended to identify a scoring function with an optimal combination of contact-potential and sequence-profile terms. The benchmark set is selected to include many medium-difficulty fold recognition targets, where sequence similarity is undetectable by BLAST but structural similarity is extensive. The contact potential is based on the log-odds of non-local contacts involving different amino acid pairs, in native as opposed to randomly compacted structures. The sequence profile term is that used in PSI-BLAST. We find that combination of these terms significantly improves the success rate of fold recognition over use of either term alone, with respect to both recognition sensitivity and the accuracy of threading models. Improvement is greatest for targets between 10 % and 20 % sequence identity and 60 % to 80 % superimposable residues, where the number of models crossing critical accuracy and significance thresholds more than doubles. We suggest that these improvements account for the successful performance of the combined scoring function at CASP3. We discuss possible explanations as to why sequence-profile and contact-potential terms appear complementary.

Algorithms↗

MMDB: 3D structure data in Entrez.

Three-dimensional structures are now known for roughly half of all protein families. It is thus quite likely, in searching sequence databases, that one will encounter a homolog with known structure and be able to use this information to infer structure-function properties. The goal of Entrez's 3D structure database is to make this information accessible and useful to molecular biologists. To this end, Entrez's search engine provides three powerful features: (i) Links between databases; one may search by term matching in Medline((R)), for example, and link to 3D structures reported in these articles. (ii) Sequence and structure neighbors; one may select all sequences similar to one of interest, for example, and link to any known 3D structures. (iii) Sequence and structure visualization; identifying a homolog with known structure, one may view a combined molecular-graphic and alignment display, to infer approximate 3D structure. Entrez's MMDB (Molecular Modeling DataBase) may be accessed at: http://www.ncbi.nlm.nih.gov/Entrez/structure.html

Amino Acid Sequence↗

Domain size distributions can predict domain boundaries.

MOTIVATION: The sizes of protein domains observed in the 3D-structure database follow a surprisingly narrow distribution. Structural domains are furthermore formed from a single-chain continuous segment in over 80% of instances. These observations imply that some choices of domain boundaries on an otherwise uncharacterized sequence are more likely than others, based solely on the size and segment number of predicted domains. This property might be used to guess the locations of protein domain boundaries. RESULTS: To test this possibility we enumerate putative domain boundaries and calculate their relative likelihood under a probability model that considers only the size and segment number of predicted domains. We ask, in a cross-validated test using sequences with known 3D structure, whether the most likely guesses agree with the observed domain structure. We find that domain boundary predictions are surprisingly successful for sequences up to 400 residues long and that guessing domain boundaries in this way can improve the sensitivity of threading analysis.

Algorithms↗

MMDB: Entrez's 3D structure database.

The three dimensional structures for representatives of nearly half of all protein families are now available in public databases. Thus, no matter which protein one investigates, it is increasingly likely that the 3D structure of a homolog will be known and may reveal unsuspected structure-function relationships. The goal of Entrez's 3D-structure database is to make this information accessible and usable by molecular biologists (http://www.ncbi.nlm.nih.gov/Entrez). To this end Entrez provides two major analysis tools, a search engine based on sequence and structure 'neighboring' and an integrated visualization system for sequence and structure alignments. From a protein's sequence 'neighbors' one may rapidly identify other members of a protein family, including those where 3D structure is known. By comparing aligned sequences and/or structures in detail, using the visualization system, one may identify conserved features and perhaps infer functional properties. Here we describe how these analysis tools may be used to investigate the structure and function of newly discovered proteins, using the PTEN gene product as an example.

Amino Acid Sequence↗

Threading with explicit models for evolutionary conservation of structure and sequence.

We have attempted to predict the three-dimensional structures of 19 proteins for the CASP3 experiment, each showing less than 25% sequence identity with known structures. Predictions were based on a threading method that aligns the target sequence with the conserved cores of structural templates, as identified from structure-structure alignments of the template with homologous neighbors. Alternative alignments were scored using contact potentials and a position-specific score matrix derived from sequence neighbors of the template. We find that this method identified the correct structural family for 11 of the 19 targets and predicted the remaining 8 targets to be similar to "none" of the templates, avoiding false positives. Threading alignments are relatively accurate for 10 of the 11 targets, including alignments for 6 of 7 identified at CASP3 as fold-recognition targets. These predictions were ranked "first place" by the CASP3 assessor when compared to fold-recognition predictions made by other methods. It appears that threading with family-specific models for structure and sequence conservation has improved threading prediction accuracy.

Algorithms↗

A measure of progress in fold recognition?

We present a retrospective analysis of CASP3 threading predictions, applying evaluation and assessment criteria used at CASP2. Our purpose is twofold. First, we wish to ask whether measures of model accuracy are comparable between CASP3 and CASP2, even though they have been calculated differently. We find that these quantities are effectively the same, and that either may be used to compare model accuracy. Secondly, we wish to assess progress in fold recognition by comparing the numbers of CASP2 and CASP3 models that cross specific accuracy thresholds. We find that the number of accurate models at CASP3 drops sharply as the targets become more difficult, with less extensive similarity to known structures, exactly the pattern seen at CASP2. CASP3 teams do not seem to have predicted accurate models for targets of greater difficulty, and for a given difficulty range the best CASP3 models seem no more accurate than the best models at CASP2. At CASP3, however, we find greater numbers of accurate models for medium-difficulty targets, with extensive similarity to a known structure but no shared sequence motifs. Threading methods would appear to have become more reliable for modeling based on remote evolutionary relationships.

Algorithms↗

Isolation, molecular characterization, and tissue-specific expression of a novel putative G protein-coupled receptor.

We isolated a 40 kDa integral membrane protein (p40) from human erythrocyte ghosts by affinity chromatography, using a C-terminal peptide of stomatin, and obtained partial sequences which enabled us to isolate two full-length cDNAs from human bone marrow and fetal brain cDNA libraries. The cDNA sequences were identical and encoded a novel putative G protein-coupled receptor (399 amino acids). Northern and RNA dot blot analyses demonstrated that the major 4.8 kb-transcript is predominantly expressed in brain. In situ hybridization studies of tissue sections revealed high expression in neurons of the brain and spinal cord, in thymocytes, megakaryocytes, and macrophages.

Amino Acid Sequence↗

Measures of threading specificity and accuracy.

Threading predictions for CASP2 target proteins were compared to their true structures using a series of precisely defined measures of agreement, calculated in a fully automatic way. Fold recognition specificity was calculated as the proportion of a predictor's "bet" that was placed on previously-known structures similar to the prediction target, as identified by a "jury" of well-tested structure-structure comparison methods. Values approaching 100% indicate that a prediction correctly identified the structural and/or evolutionary family to which a target belongs. Alignment specificity was calculated as the proportion of aligned residue paris in the predicted target-to-known-structure alignment that also occur in the structure-structure alignments produced by the "jury" methods. Contact specificity was calculated as the proportion of nonlocal residue contacts in the molecular model implied by threading alignment, that also occur in the experimental structure of the target. Alignment specificity and contact specificity measure the accuracy of a predicted 3-dimensional model. Values approaching 100% indicate that target residues have been assigned to the correct spatial locations and that the model is as accurate as possible for a threading prediction.

Amino Acid Sequence↗

A retrospective analysis of CASP2 threading predictions.

Analysis of CASP2 protein threading results shows that the success rate of structure predictions varies widely among prediction targets. We set "critical" thresholds in fold recognition specificity and threading model accuracy at the points where "incorrect" CASP2 predictions just outnumber "correct" predictions. Using these thresholds we find that correct predictions were made for all of those targets and for only those targets where more than 50% of target residues may be superimposed on previously known structures. Three-fourths of these correct predictions were furthermore made for targets with greater than 12% residue identity in structural alignment, where characteristic sequence motifs are also present. Based on these observations we suggest that the sustained performance of threading methods is best characterized by counting the numbers of correct predictions for targets of increasing "difficulty." We suggest that target difficulty may be assigned, once the true structure of the target is known, according to the fraction of residues superimposable onto previously known structures and the fraction of identical residues in those structural alignments.

Models, Molecular↗

A measure of success in fold recognition.

Prediction of protein structure by fold recognition, or threading, was recently put to the test in a 'blind' structure prediction experiment, CASP2. Thirty-two teams from around the world participated, preparing predictions for 22 different 'target' proteins whose structures were soon to be determined. As experimental structures became available, we, as organizers of the threading competition, computed objective measures of fold-recognition specificity and model accuracy, to identify and characterize successful predictions. Here, we present a brief summary of these prediction evaluations, a tally of 'correct' predictions and a discussion of factors associated with correct predictions. We find that threading produced specific recognition and accurate models whenever the structural database contained a template spanning a large fraction of target sequence. Presence of conserved sequence motifs was helpful, but not required, and it would appear that threading can succeed whenever similarity to a known structure is sufficiently extensive.

Computer Simulation↗

The Saccharomyces cerevisiae zinc finger proteins Msn2p and Msn4p are required for transcriptional induction through the stress response element (STRE).

The MSN2 and MSN4 genes encode homologous and functionally redundant Cys2His2 zinc finger proteins. A disruption of both MSN2 and MSN4 genes results in a higher sensitivity to different stresses, including carbon source starvation, heat shock and severe osmotic and oxidative stresses. We show that MSN2 and MSN4 are required for activation of several yeast genes such as CTT1, DDR2 and HSP12, whose induction is mediated through stress-response elements (STREs). Msn2p and Msn4p are important factors for the stress-induced activation of STRE dependent promoters and bind specifically to STRE-containing oligonucleotides. Our results suggest that MSN2 and MSN4 encode a DNA-binding component of the stress responsive system and it is likely that they act as positive transcription factors.

Base Sequence↗

Structural requirements for low-pH-induced rearrangements in the envelope glycoprotein of tick-borne encephalitis virus.

The exposure of the flavivirus tick-borne encephalitis (TBE) virus to an acidic pH is necessary for virus-induced membrane fusion and leads to a quantitative and irreversible conversion of the envelope protein E dimers to trimers. To study the structural requirements for this oligomeric rearrangement, the effect of low-pH treatment on the oligomeric state of different isolated forms of protein E was investigated. Full-length E dimers obtained by solubilization of virus with the detergent Triton X-100 formed trimers at low pH, whereas truncated E dimers lacking the stem-anchor region underwent a reversible dissociation into monomers without forming trimers. These data suggest that the low-pH-induced rearrangement in virions is a two-step process involving a reversible dissociation of the E dimers followed by an irreversible formation of trimers, a process which requires the stem-anchor portion of the protein. This region contains potential amphipathic alpha-helical and conserved structural elements whose interactions may contribute to the rearrangements which initiate the fusion process.

Amino Acid Sequence↗