Protein sequence database.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The side-chain-side-chain interaction between Phe residues and sulfur-containing residues (Cis and Met) in the two possible orientations at positions i, i + 4 of alpha-helices is described. We have analyzed the contribution to helical stability of the above interactions by studying eight polyalanine-based peptides differing at the residues at positions 9 and 13. These two positions were independently mutated from Ala (AA), to Cys (AC and CA), Met (AM and MA), and Phe (AF and FA) and to the pairs Phe-Met (FM), Met-Phe (MF), Phe-Cys (FC), and Cys-Phe (CF). The intrinsic helical propensities of Cys, Met, and Phe were found to be those previously described in the algorithm AGADIR. NMR analysis of the FM, MF, FC, and CF peptides showed the formation in aqueous solution of contacts between the aromatic ring and the side chains of Cys or Met, at the two i, i + 4 orientations. CD studies demonstrated the important contribution of two of these interactions (FM and FC) to alpha-helix stability (up to 2 kcal mol-1 in the Phe-Cys pair). Statistical analysis of the protein database provides a rationale for the stereospecificity and free energies of the interactions. The very favorable interaction between an aromatic ring and a sulfur-containing amino acid explains why in the protein database around 50% of the sulfur atoms are contacting aromatic rings (Reid et al., 1985).
A tool is presented that helps to find biological functions for new protein sequences. Running on any Macintosh computer system, MacPattern provides a unique combination of different algorithms, speed and user-friendliness. It supports searches for protein patterns using the PROSITE database, protein block searches with the BLOCKS database, and the identification of statistically significant protein segments. MacPattern allows batch processing of sequences and automatic translations of nucleotide sequence data. It is particularly suited for genome analysis or cDNA sequencing projects.
The advent of storage phosphor technology has been of considerable benefit to the imaging of gel-separated radiolabeled proteins due to the rapid and quantitative nature of the data acquisition process. Previously, times over one month were required to obtain fluorographs of the same gel to yield data of sufficient dynamic range for quantitative analysis of high-resolution two-dimensional (2-D) gels. As we are in the process of building a human 2-D gel protein database, and therefore have a high throughput of 2-D gels both to image and quantitate using the Quest II software, we undertook an evaluation of a storage phosphor imager, including an evaluation of signal fade. The results of this evaluation demonstrate the feasibility of using such a system, and we describe the procedures that allow us to use this technique for quantitative analysis of many complex 2-D gel patterns. These procedures include a useful batch printing program that allows printing of many images in a non-interactive mode. Examples will be presented of how autoradiography, using storage phosphor plates and the Quest II system, have enabled us to begin building a human 2-D gel protein database including posttranslational modification information, without the previous time constraints associated with such a project.
Explore the source record for details and available documents.
The sequences of two previously known tail genes, R and S, of the temperate bacteriophage P2 and the sequence of an additional open reading frame (orf-30) located between S and V, were determined. Amber mutations mapping within R and S, Ram3, Ram42, Ram23, Sam75, and Sam89 were sequenced and found to be within their corresponding open reading frames. We constructed overproducing plasmids for R and S and identified these proteins by SDS-PAGE of whole-cell lysates and Coomassie blue staining. The predicted molecular masses of proteins R and S were M(r) 17,400 and 17,300, respectively, although both polypeptides migrated more slowly during gel electrophoresis than would be expected from the sequence data. orf-30 occupies the strand opposite from RS and V and is preceded by several weak potential sigma 70-RNA polymerase promoters, some of which overlap with the V promoter. A construct that had the putative orf-30 promoter region upstream of the lacZ gene produced low levels of beta-galactosidase activity in vivo. Expression from the orf-30 promoter was not stimulated by the phage P4 transcriptional activator protein, delta, which acts at all the known P2 and P4 late promoters. Insertion mutagenesis showed that orf-30 was not an essential gene for P2 growth in Escherichia coli. None of the gene or protein sequences exhibited extensive homology to sequences in the nucleic acid and protein databases. However, the R protein contains a small region homologous to one in the phage T4 tail protein gp15, which is required for T4 tails to bind heads. We propose that R and S are tail completion proteins that are essential for stable head joining.
We have developed and tested a systematic method for the location and statistical evaluation of potential DNA-binding regions of the lambda Cro type in protein sequences. Using this approach to examine proteins expected to contain such regions, we have been able to compile a statistically homogeneous master set of 37 lambda Cro-like DNA-binding domains. Examination of a protein database revealed other prokaryotic proteins that are similar to this lambda Cro-like group. There are also many DNA-binding proteins that are not found to be significantly similar to the lambda Cro group, consistent with previous suggestions that different types of protein sequence may be able to achieve a similar mode of binding and that there exist other modes of sequence-specific DNA-binding. A useful feature of the method is that it can be applied without a computer.
Explore the source record for details and available documents.
A systematic comparison of the protein synthesis patterns of cultured normal and transformed human fibroblasts and epithelial cells, using two-dimensional gel protein analysis combined with computerized imaging and data acquisition, identified a 90-kD protein (SSP 5714) as one of the most striking downregulated markers typical of the transformed state. Using the information stored in the comprehensive human cellular protein database, we found this protein strongly expressed in several fetal tissues and one of them, epidermis, served as a source for preparative two-dimensional gel electrophoresis. Partial amino acid sequences were generated from peptides obtained by in situ digestion of the electroblotted protein. These sequences identified the marker protein as gelsolin, a finding that was confirmed by two-dimensional immunoblotting of human MRC-5 fibroblast proteins using specific antibodies and by coelectrophoresis with purified human gelsolin. These results suggest that an important regulatory protein of the microfilament system may play a role in defining the phenotype of transformed human fibroblast and epithelial cells in culture.
We have cloned a full-length cDNA for the B cell membrane protein CD22, which is referred to as B lymphocyte cell adhesion molecule (BL-CAM). Using subtractive hybridization techniques, several B lymphocyte-specific cDNAs were isolated. Northern blot analysis with one of the clones, clone 66, revealed expression in normal activated B cells and a variety of B cell lines, but not in normal activated T cells, T cell lines, Hela cells, or several tissues, including brain and placenta. One major transcript of approximately 3.3 kb was found in B cells although several smaller transcripts were also present in low amounts (approximately 2.6, 2.3, and 1.6 kb). Sequence analysis of a full-length cDNA clone revealed an open reading frame of 2,541 bases coding for a predicted protein of 847 amino acids with a molecular mass of 95 kD. The BL-CAM cDNA is nearly identical to a recently isolated cDNA clone for CD22, with the exception of an additional 531 bases in the coding region of BL-CAM. BL-CAM has a predicted transmembrane spanning region and a 140-amino acid intracytoplasmic domain. Search of the National Biological Research Foundation protein database revealed that this protein is a member of the immunoglobulin super family and that it had significant homology with three homotypic cell adhesion proteins: carcinoembryonic antigen (29% identity over 460 amino acids), myelin-associated glycoprotein (27% identity over 425 amino acids), and neural cell adhesion molecule (21.5% over 274 amino acids). Northern blot analysis revealed low-level BL-CAM mRNA expression in unactivated tonsillar B cells, which was rapidly increased after B cell activation with Staphylococcus aureus Cowan strain 1 and phorbol myristate acetate, but not by various cytokines, including interleukin 4 (IL-4), IL-6, and gamma interferon. In situ hybridization with an antisense BL-CAM RNA probe revealed expression in B cell-rich areas in tonsil and lymph node, although the most striking hybridization was in the germinal centers. COS cells transfected with a BL-CAM expression vector were immunofluorescently stained positively with two different CD22 antibodies, each of which recognizes a different epitope. Additionally, both normal tonsil B cells and a B cell line were found to adhere to COS transfected with BL-CAM in the sense but not the antisense direction.(ABSTRACT TRUNCATED AT 400 WORDS)
A comprehensive DNA analysis computer program was described in the second special issue of Nucleic Acids Research on the applications of computers to research on nucleic acids by Stone and Potter (1). Criteria used in designing the program were user friendliness, ability to handle large DNA sequences, low storage requirement, migratability to other computers and comprehensive analysis capability. The program has been used extensively in an industrial-research environment. This paper talks about improvements to that program. These improvements include testing for methylation blockage of restriction enzyme recognition sites, homology analysis, RNA folding analysis, integration of a large DNA database (GenBank), a site specific mutagenesis analysis, a protein database and protein searching programs. The original design of the DNA analysis program using a command executive from which any analytical programs can be called, has proven to be extremely versatile in integrating both developed and outside programs to the file management system employed.
A large-scale effort to map and sequence the human genome is now under way. Crucial to the success of this research is a group of computer programs that analyze and compare data on molecular sequences. This article describes the classic algorithms for similarity searching and sequence alignment. Because good performance of these algorithms is critical to searching very large and growing databases, we analyze the running times of the algorithms and discuss recent improvements in this area.
The BLAST sequence comparison programs have been ported to a variety of parallel computers-the shared memory machine Cray Y-MP 8/864 and the distributed memory architectures Intel iPSC/860 and nCUBE. Additionally, the programs were ported to run on workstation clusters. We explain the parallelization techniques and consider the pros and cons of these methods. The BLAST programs are very well suited for parallelization for a moderate number of processors. We illustrate our results using the program blastp as an example. As input data for blastp, a 799 residue protein query sequence and the protein database PIR were used.
A rapid method for similarity searches (FASTP program) was used to identify similarities between a protein database and the human basic proteins from myelin [P2 protein and 17.2K, 18.5K, and 21.5K variants of myelin basic protein (MBP)]. From similarity scores, we concluded that none of the presently known proteins are in a family containing the MBPs. No new members were found for the lipid-binding family of which P2 is a member. Sequence similarities deemed relevant to the molecular mimicry hypothesis for virus-induced autoimmunity were identified in FASTP data with the aid of microcomputer programs. Several MBP/viral protein similarities were found that have not been reported previously. Of note because of their association with demyelinating conditions were proteins from visna and vaccinia. Similarity with visna was specific to the 21.5K and 20.2K MBPs. The most interesting new possibility for mimicry involving the P2 protein was between the influenza A NS2 protein and a sequence region of P2 thought to be neuritogenic in animals and mitogenic for lymphocytes from some patients with Guillain-Barré syndrome (GBS). This may have relevance for some cases of GBS associated with the 1976 U.S.A. swine flu vaccination program. Because FASTP reports only the best similarities between proteins, searches with FASTP may not have detected all the examples of mimicry present in the database. Searches might also be more effective if similarities could be scored on immunological rather than structural relatedness.
A plasmid containing the bacterial chloramphenicol acetyltransferase (CAT) gene under the control of an Autographa california nuclear polyhedrosis virus (AcNPV) late gene promoter was constructed. This plasmid (pL2cat) also contained the AcNPV hr5 enhancer element. Transient-expression assay experiments indicated that the late promoter was active in Spodoptera frugiperda cells cotransfected with pL2cat and AcNPV DNA but not when pL2cat was transfected alone. Low levels of CAT activity were observed in cells cotransfected with pL2cat and pIE-1 DNAs. However, CAT activity was not induced in a similar plasmid which lacked the cis-linked enhancer element, indicating that the enhancer was required for expression of the late gene. Cotransfection mapping of pPstI clones of AcNPV DNA indicated that the pPstI-G clone of viral DNA contained a factor which further stimulated late gene expression 3- to 10-fold. Transient-expression assay analysis of subclones of pPstI-G localized the trans-active factor to a 3.0-kilobase XbaI fragment. The nucleotide sequence of this fragment was determined and found to contain three potential open reading frames. A computer-assisted search of a protein database revealed no closely related proteins. One of the predicted amino acid sequences contained potential metal-binding domains similar to those found in nucleic acid-binding proteins. Subcloning and subsequent CAT assay indicated that two of the open reading frames were required for the activation of pL2cat. Nuclease S1 mapping of infected and transfected RNAs indicated that the two open reading frames were transcribed as delayed-early genes. Quantitative nuclease S1 analysis and differential DNA digestion of recovered plasmids indicated that the activation of pL2cat was not due to an increase in steady-state levels of mRNA replication of the viral DNA.
Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.