PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Systematic method for the detection of potential lambda Cro-like DNA-binding regions in proteins.

We have developed and tested a systematic method for the location and statistical evaluation of potential DNA-binding regions of the lambda Cro type in protein sequences. Using this approach to examine proteins expected to contain such regions, we have been able to compile a statistically homogeneous master set of 37 lambda Cro-like DNA-binding domains. Examination of a protein database revealed other prokaryotic proteins that are similar to this lambda Cro-like group. There are also many DNA-binding proteins that are not found to be significantly similar to the lambda Cro group, consistent with previous suggestions that different types of protein sequence may be able to achieve a similar mode of binding and that there exist other modes of sequence-specific DNA-binding. A useful feature of the method is that it can be applied without a computer.

Amino Acid Sequence

Comparative two-dimensional gel analysis and microsequencing identifies gelsolin as one of the most prominent downregulated markers of transformed human fibroblast and epithelial cells.

A systematic comparison of the protein synthesis patterns of cultured normal and transformed human fibroblasts and epithelial cells, using two-dimensional gel protein analysis combined with computerized imaging and data acquisition, identified a 90-kD protein (SSP 5714) as one of the most striking downregulated markers typical of the transformed state. Using the information stored in the comprehensive human cellular protein database, we found this protein strongly expressed in several fetal tissues and one of them, epidermis, served as a source for preparative two-dimensional gel electrophoresis. Partial amino acid sequences were generated from peptides obtained by in situ digestion of the electroblotted protein. These sequences identified the marker protein as gelsolin, a finding that was confirmed by two-dimensional immunoblotting of human MRC-5 fibroblast proteins using specific antibodies and by coelectrophoresis with purified human gelsolin. These results suggest that an important regulatory protein of the microfilament system may play a role in defining the phenotype of transformed human fibroblast and epithelial cells in culture.

Amino Acid Sequence

cDNA cloning of the B cell membrane protein CD22: a mediator of B-B cell interactions.

We have cloned a full-length cDNA for the B cell membrane protein CD22, which is referred to as B lymphocyte cell adhesion molecule (BL-CAM). Using subtractive hybridization techniques, several B lymphocyte-specific cDNAs were isolated. Northern blot analysis with one of the clones, clone 66, revealed expression in normal activated B cells and a variety of B cell lines, but not in normal activated T cells, T cell lines, Hela cells, or several tissues, including brain and placenta. One major transcript of approximately 3.3 kb was found in B cells although several smaller transcripts were also present in low amounts (approximately 2.6, 2.3, and 1.6 kb). Sequence analysis of a full-length cDNA clone revealed an open reading frame of 2,541 bases coding for a predicted protein of 847 amino acids with a molecular mass of 95 kD. The BL-CAM cDNA is nearly identical to a recently isolated cDNA clone for CD22, with the exception of an additional 531 bases in the coding region of BL-CAM. BL-CAM has a predicted transmembrane spanning region and a 140-amino acid intracytoplasmic domain. Search of the National Biological Research Foundation protein database revealed that this protein is a member of the immunoglobulin super family and that it had significant homology with three homotypic cell adhesion proteins: carcinoembryonic antigen (29% identity over 460 amino acids), myelin-associated glycoprotein (27% identity over 425 amino acids), and neural cell adhesion molecule (21.5% over 274 amino acids). Northern blot analysis revealed low-level BL-CAM mRNA expression in unactivated tonsillar B cells, which was rapidly increased after B cell activation with Staphylococcus aureus Cowan strain 1 and phorbol myristate acetate, but not by various cytokines, including interleukin 4 (IL-4), IL-6, and gamma interferon. In situ hybridization with an antisense BL-CAM RNA probe revealed expression in B cell-rich areas in tonsil and lymph node, although the most striking hybridization was in the germinal centers. COS cells transfected with a BL-CAM expression vector were immunofluorescently stained positively with two different CD22 antibodies, each of which recognizes a different epitope. Additionally, both normal tonsil B cells and a B cell line were found to adhere to COS transfected with BL-CAM in the sense but not the antisense direction.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

Methylation blockage and other improvements to a comprehensive DNA analysis program.

A comprehensive DNA analysis computer program was described in the second special issue of Nucleic Acids Research on the applications of computers to research on nucleic acids by Stone and Potter (1). Criteria used in designing the program were user friendliness, ability to handle large DNA sequences, low storage requirement, migratability to other computers and comprehensive analysis capability. The program has been used extensively in an industrial-research environment. This paper talks about improvements to that program. These improvements include testing for methylation blockage of restriction enzyme recognition sites, homology analysis, RNA folding analysis, integration of a large DNA database (GenBank), a site specific mutagenesis analysis, a protein database and protein searching programs. The original design of the DNA analysis program using a command executive from which any analytical programs can be called, has proven to be extremely versatile in integrating both developed and outside programs to the file management system employed.

Base Sequence

Searching gene and protein sequence databases.

A large-scale effort to map and sequence the human genome is now under way. Crucial to the success of this research is a group of computer programs that analyze and compare data on molecular sequences. This article describes the classic algorithms for similarity searching and sequence alignment. Because good performance of these algorithms is critical to searching very large and growing databases, we analyze the running times of the algorithms and discuss recent improvements in this area.

Algorithms

Implementations of BLAST for parallel computers.

The BLAST sequence comparison programs have been ported to a variety of parallel computers-the shared memory machine Cray Y-MP 8/864 and the distributed memory architectures Intel iPSC/860 and nCUBE. Additionally, the programs were ported to run on workstation clusters. We explain the parallelization techniques and consider the pros and cons of these methods. The BLAST programs are very well suited for parallelization for a moderate number of processors. We illustrate our results using the program blastp as an example. As input data for blastp, a 799 residue protein query sequence and the protein database PIR were used.

Computers

An approach to searching protein sequences for superfamily relationships or chance similarities relevant to the molecular mimicry hypothesis: application to the basic proteins of myelin.

A rapid method for similarity searches (FASTP program) was used to identify similarities between a protein database and the human basic proteins from myelin [P2 protein and 17.2K, 18.5K, and 21.5K variants of myelin basic protein (MBP)]. From similarity scores, we concluded that none of the presently known proteins are in a family containing the MBPs. No new members were found for the lipid-binding family of which P2 is a member. Sequence similarities deemed relevant to the molecular mimicry hypothesis for virus-induced autoimmunity were identified in FASTP data with the aid of microcomputer programs. Several MBP/viral protein similarities were found that have not been reported previously. Of note because of their association with demyelinating conditions were proteins from visna and vaccinia. Similarity with visna was specific to the 21.5K and 20.2K MBPs. The most interesting new possibility for mimicry involving the P2 protein was between the influenza A NS2 protein and a sequence region of P2 thought to be neuritogenic in animals and mitogenic for lymphocytes from some patients with Guillain-Barré syndrome (GBS). This may have relevance for some cases of GBS associated with the 1976 U.S.A. swine flu vaccination program. Because FASTP reports only the best similarities between proteins, searches with FASTP may not have detected all the examples of mimicry present in the database. Searches might also be more effective if similarities could be scored on immunological rather than structural relatedness.

Animals

Functional mapping of Autographa california nuclear polyhedrosis virus genes required for late gene expression.

A plasmid containing the bacterial chloramphenicol acetyltransferase (CAT) gene under the control of an Autographa california nuclear polyhedrosis virus (AcNPV) late gene promoter was constructed. This plasmid (pL2cat) also contained the AcNPV hr5 enhancer element. Transient-expression assay experiments indicated that the late promoter was active in Spodoptera frugiperda cells cotransfected with pL2cat and AcNPV DNA but not when pL2cat was transfected alone. Low levels of CAT activity were observed in cells cotransfected with pL2cat and pIE-1 DNAs. However, CAT activity was not induced in a similar plasmid which lacked the cis-linked enhancer element, indicating that the enhancer was required for expression of the late gene. Cotransfection mapping of pPstI clones of AcNPV DNA indicated that the pPstI-G clone of viral DNA contained a factor which further stimulated late gene expression 3- to 10-fold. Transient-expression assay analysis of subclones of pPstI-G localized the trans-active factor to a 3.0-kilobase XbaI fragment. The nucleotide sequence of this fragment was determined and found to contain three potential open reading frames. A computer-assisted search of a protein database revealed no closely related proteins. One of the predicted amino acid sequences contained potential metal-binding domains similar to those found in nucleic acid-binding proteins. Subcloning and subsequent CAT assay indicated that two of the open reading frames were required for the activation of pL2cat. Nuclease S1 mapping of infected and transfected RNAs indicated that the two open reading frames were transcribed as delayed-early genes. Quantitative nuclease S1 analysis and differential DNA digestion of recovered plasmids indicated that the activation of pL2cat was not due to an increase in steady-state levels of mRNA replication of the viral DNA.

Amino Acid Sequence

Spermatocytes and round spermatids of rat testis: protein patterns.

Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.

Adolescent

Harnessing deep learning for proteome-scale detection of amyloid signaling motifs.

MOTIVATION: Amyloid signaling sequences adopt the cross-β fold that is capable of self-replication in the templating process. Propagation of the amyloid fold from the receptor to the effector protein is used for signal transduction in the immune response pathways in animals, fungi, and bacteria. So far, a dozen of families of amyloid signaling motifs (ASMs) have been classified. Unfortunately, due to the wide variety of ASMs it is difficult to identify them in large protein databases available, which limits the possibility of conducting experimental studies. To date, various deep learning (DL) models have been applied across a range of protein-related tasks, including domain family classification and the prediction of protein structure and protein-protein interactions. RESULTS: In this study, we develop tailor-made bidirectional LSTM and BERT-based architectures to model ASM, and compare their performance against a state-of-the-art machine learning grammatical model. Our research is focused on developing a discriminative model of generalized ASMs, capable of detecting ASMs in large datasets. The DL-based models are trained on a diverse set of motif families and a global negative set, and used to identify ASMs from remotely related families. We analyze how both models represent the data and demonstrate that the DL-based approaches effectively detect ASMs, including novel motifs, even at the genome scale. AVAILABILITY AND IMPLEMENTATION: The models are provided as a Python package, asmscan-bilstm, and a Docker image at https://github.com/chrispysz/asmscan-proteinbert-run. The source code can be accessed at https://github.com/jakub-galazka/asmscan-bilstm and https://github.com/chrispysz/asmscan-proteinbert. Data and results are at https://github.com/wdyrka-pwr/ASMscan.

Deep Learning

Stochastic motif extraction using hidden Markov model.

In this paper, we study the application of an HMM (hidden Markov model) to the problem of representing protein sequences by a stochastic motif. A stochastic protein motif represents the small segments of protein sequences that have a certain function or structure. The stochastic motif, represented by an HMM, has conditional probabilities to deal with the stochastic nature of the motif. This HMM directly reflects the characteristics of the motif, such as a protein periodical structure or grouping. In order to obtain the optimal HMM, we developed the "ilerative duplication method" for HMM topology learning. It starts from a small fully-connected network and iterates the network generation and parameter optimization until it achieves sufficient discrimination accuracy. Using this method, we obtained an HMM for a leucine zipper motif. Compared to the accuracy of a symbolic pattern representation with accuracy of 14.8 percent, an HMM achieved 79.3 percent in prediction. Additionally, the method can obtain an HMM for various types of zinc finger motifs, and it might separate the mixed data. We demonstrated that this approach is applicable to the validation of the protein database; a constructed HMM has indicated that one protein sequence annotated as "leucine-zipper like sequence" in the database is quite different from other leucine-zipper sequences in terms of likelihood, and we found this discrimination is plausible.

Algorithms

Mapping cells and sub-cellular organelles on 2-D gels: 'new tricks for an old horse'.

Nowadays, investigators in all fields are faced with the identification of unknown, up- or down-regulated, modified proteins that they are trying to identify. Two-dimensional (2-D) gel electrophoresis, with its ability to resolve several thousand proteins, is an extremely powerful technique. The current resolution and reproducibility of 2-D gel technology and the establishment of computer assisted 2-D gel protein databases have paved new ways for the identification of proteins.

Cell Fractionation

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

The mouse male germ cell-specific gene Tpx-1: molecular structure, mode of expression in spermatogenesis, and sequence similarity to two non-mammalian genes.

Tpx-1 is a testis-specific gene that maps on mouse Chromosome (Chr) 17. The deduced TPX-1 protein shows 55% amino acid sequence similarity to acidic epididymal glycoprotein (AEG), assumed to be involved in sperm maturation. In the present study, we determined the genomic structure of the mouse Tpx-1 gene and the cellular localization of its transcripts. The gene was found to contain ten exons, with an unusually large intron (approximately 17.0 kilobase pairs) between exons 8 and 9. In situ hybridization of testicular sections showed that Tpx-1 is transcribed abundantly by haploid male germ cells. A computer search of protein databases revealed that deduced TPX-1/AEG proteins have significant sequence similarity (approximately 30%) to two non-mammalian proteins: "pathogenesis-related" proteins 1 of tobaccos, and venom sac proteins of white-face hornets, known as Dol m V. Amino acid residues encoded by exon 10 of the Tpx-1 gene and most of those encoded by exon 9 were absent in the non-mammalian proteins. This result suggests that the ancestor of Tpx-1 acquired exons 9 and 10 after its divergence from the ancestors of the plant and insect proteins.

Amino Acid Sequence

A homologue of the Drosophila female sterile homeotic (fsh) gene in the class II region of the human MHC.

The RING3 gene maps in the class II region of the human major histocompatibility complex, at a CpG island distal of the HLA-DNA gene. RING3 cDNAs were obtained from a T cell cDNA library and the longest (4 kb) was sequenced. The sequence contained an open reading frame encoding a protein of 754 amino acids. A screen of protein databases revealed striking homology between the RING3 protein and the Drosophila female sterile homeotic gene (fsh) which is implicated in the establishment of segments in the early embryo. Partial sequence homology was also observed with some other proteins involved in cell cycle control (CCG1), cell division (ftsA) and regulation of cell growth (gamma interferons). This highly conserved gene may play an important role in human development. In addition, its location in the MHC class II region may be related to some HLA-associated diseases.

Amino Acid Sequence

Secondary structure analysis identifies a putative mouse protein demonstrating similarity to the repeat units found in CDC4, the G protein beta subunits and related proteins.

The predicted protein product of an anonymous clone isolated from a cDNA library prepared from 12 day post coitum (p.c) embryonic mouse heart tissue demonstrated the same segmental repeats previously identified in the cell division control protein, CDC4 and the G protein beta 1 subunit. A search of the protein database subsequently identified three other classes of protein containing the repeat. Secondary structure analyses performed on the repeat sequences revealed a high degree of conservation suggesting that the repeat motif performs a specific function in a diverse range of proteins.

Amino Acid Sequence

Database algorithm for generating protein backbone and side-chain co-ordinates from a C alpha trace application to model building and detection of co-ordinate errors.

The problem of constructing all-atom model co-ordinates of a protein from an outline of the polypeptide chain is encountered in protein structure determination by crystallography or nuclear magnetic resonance spectroscopy, in model building by homology and in protein design. Here, we present an automatic procedure for generating full protein co-ordinates (backbone and, optionally, side-chains) given the C alpha trace and amino acid sequence. To construct backbones, a protein structure database is first scanned for fragments that locally fit the chain trace according to distance criteria. A best path algorithm then sifts through these segments and selects an optimal path with minimal mismatch at fragment joints. In blind tests, using fully known protein structures, backbones (C alpha, C, N, O) can be reconstructed with a reliability of 0.4 to 0.6 A root-mean-square position deviation and not more than 0 to 5% peptide flips. This accuracy is sufficient to identify possible errors in protein co-ordinate sets. To construct full co-ordinates, side-chains are added from a library of frequently occurring rotamers using a simple and fast Monte Carlo procedure with simulated annealing. In tests on X-ray structures determined at better than 2.5 A resolution, the positions of side-chain atoms in the protein core (less than 20% relative accessibility) have an accuracy of 1.6 A (r.m.s. deviation) and 70% of chi 1 angles are within 30 degrees of the X-ray structure. The computer program MaxSprout is available on request.

Algorithms