PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Tumorigenic poxviruses: genomic organization and DNA sequence of the telomeric region of the Shope fibroma virus genome.

Shope fibroma virus (SFV), a tumorigenic poxvirus, has a 160-kb linear double-stranded DNA genome and possesses terminal inverted repeats (TIRs) of 12.4 kb. The DNA sequence of the terminal 5.5 kb of the viral genome is presented and together with previously published sequences completes the entire sequence of the SFV TIR. The terminal 400-bp region contains no major open reading frames (ORFs) but does possess five related imperfect palindromes. The remaining 5.1 kb of the sequence contains seven tightly clustered and tandemly oriented ORFs, four larger than 100 amino acids in length (T1, T2, T4, and T5) and three smaller ORFs (T3A, T3B, and T3C). All are transcribed toward the viral hairpin and almost all possess the consensus sequence TTTTTNT near their 3' ends which has been implicated for the transcription termination of vaccinia virus early genes. Searches of the published DNA database revealed no sequences with significant homology with this region of the SFV genome but when the protein database was searched with the translation products of ORFs T1-T5 it was found that the N-terminus of the putative T4 polypeptide is closely related to the signal sequence of the hemagglutinin precursor from influenza A virus, suggesting that the T4 polypeptide may be secreted from SFV-infected cells. Examination of other SFV ORFs shows that T1 and T2 also possess signal-like hydrophobic amino acid stretches close to their N-termini. The protein database search also revealed that the putative T2 protein has significant homology to the insulin family of polypeptides. In terms of sequence repetitions, seven tandemly repeated copies of the hexanucleotide ATTGTT and three flanking regions of dyad symmetry were detected, all in ORF T3C. A search for palindromic sequences also revealed two clusters, one in ORF T3A/B and a second in ORF T2. ORF T2 harbors five short sequence domains, each of which consists of a 6-bp short palindrome and a 10- to 18-bp larger palindrome. The significance of these palindromic domains in this ORF is unclear but the coincidence of the end of one larger palindrome with the end of the translated protein sequence that has homology with the B chain of insulin suggests that the palindromes may divide the T2 protein into several functional units. The salient organizational features of the complete SFV TIR are also discussed in light of what is known about other poxviral TIRs.

Amino Acid Sequence

Comparison of methods for searching protein sequence databases.

We have compared commonly used sequence comparison algorithms, scoring matrices, and gap penalties using a method that identifies statistically significant differences in performance. Search sensitivity with either the Smith-Waterman algorithm or FASTA is significantly improved by using modern scoring matrices, such as BLOSUM45-55, and optimized gap penalties instead of the conventional PAM250 matrix. More dramatic improvement can be obtained by scaling similarity scores by the logarithm of the length of the library sequence (In()-scaling). With the best modern scoring matrix (BLOSUM55 or JO93) and optimal gap penalties (-12 for the first residue in the gap and -2 for additional residues), Smith-Waterman and FASTA performed significantly better than BLASTP. With In()-scaling and optimal scoring matrices (BLOSUM45 or Gonnet92) and gap penalties (-12, -1), the rigorous Smith-Waterman algorithm performs better than either BLASTP and FASTA, although with the Gonnet92 matrix the difference with FASTA was not significant. Ln()-scaling performed better than normalization based on other simple functions of library sequence length. Ln()-scaling also performed better than scores based on normalized variance, but the differences were not statistically significant for the BLOSUM50 and Gonnet92 matrices. Optimal scoring matrices and gap penalties are reported for Smith-Waterman and FASTA, using conventional or In()-scaled similarity scores. Searches with no penalty for gap extension, or no penalty for gap opening, or an infinite penalty for gaps performed significantly worse than the best methods. Differences in performance between FASTA and Smith-Waterman were not significant when partial query sequences were used. However, the best performance with complete query sequences was obtained with the Smith-Waterman algorithm and In()-scaling.

Algorithms

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing

Systematic method for the detection of potential lambda Cro-like DNA-binding regions in proteins.

We have developed and tested a systematic method for the location and statistical evaluation of potential DNA-binding regions of the lambda Cro type in protein sequences. Using this approach to examine proteins expected to contain such regions, we have been able to compile a statistically homogeneous master set of 37 lambda Cro-like DNA-binding domains. Examination of a protein database revealed other prokaryotic proteins that are similar to this lambda Cro-like group. There are also many DNA-binding proteins that are not found to be significantly similar to the lambda Cro group, consistent with previous suggestions that different types of protein sequence may be able to achieve a similar mode of binding and that there exist other modes of sequence-specific DNA-binding. A useful feature of the method is that it can be applied without a computer.

Amino Acid Sequence

Comparative two-dimensional gel analysis and microsequencing identifies gelsolin as one of the most prominent downregulated markers of transformed human fibroblast and epithelial cells.

A systematic comparison of the protein synthesis patterns of cultured normal and transformed human fibroblasts and epithelial cells, using two-dimensional gel protein analysis combined with computerized imaging and data acquisition, identified a 90-kD protein (SSP 5714) as one of the most striking downregulated markers typical of the transformed state. Using the information stored in the comprehensive human cellular protein database, we found this protein strongly expressed in several fetal tissues and one of them, epidermis, served as a source for preparative two-dimensional gel electrophoresis. Partial amino acid sequences were generated from peptides obtained by in situ digestion of the electroblotted protein. These sequences identified the marker protein as gelsolin, a finding that was confirmed by two-dimensional immunoblotting of human MRC-5 fibroblast proteins using specific antibodies and by coelectrophoresis with purified human gelsolin. These results suggest that an important regulatory protein of the microfilament system may play a role in defining the phenotype of transformed human fibroblast and epithelial cells in culture.

Amino Acid Sequence

cDNA cloning of the B cell membrane protein CD22: a mediator of B-B cell interactions.

We have cloned a full-length cDNA for the B cell membrane protein CD22, which is referred to as B lymphocyte cell adhesion molecule (BL-CAM). Using subtractive hybridization techniques, several B lymphocyte-specific cDNAs were isolated. Northern blot analysis with one of the clones, clone 66, revealed expression in normal activated B cells and a variety of B cell lines, but not in normal activated T cells, T cell lines, Hela cells, or several tissues, including brain and placenta. One major transcript of approximately 3.3 kb was found in B cells although several smaller transcripts were also present in low amounts (approximately 2.6, 2.3, and 1.6 kb). Sequence analysis of a full-length cDNA clone revealed an open reading frame of 2,541 bases coding for a predicted protein of 847 amino acids with a molecular mass of 95 kD. The BL-CAM cDNA is nearly identical to a recently isolated cDNA clone for CD22, with the exception of an additional 531 bases in the coding region of BL-CAM. BL-CAM has a predicted transmembrane spanning region and a 140-amino acid intracytoplasmic domain. Search of the National Biological Research Foundation protein database revealed that this protein is a member of the immunoglobulin super family and that it had significant homology with three homotypic cell adhesion proteins: carcinoembryonic antigen (29% identity over 460 amino acids), myelin-associated glycoprotein (27% identity over 425 amino acids), and neural cell adhesion molecule (21.5% over 274 amino acids). Northern blot analysis revealed low-level BL-CAM mRNA expression in unactivated tonsillar B cells, which was rapidly increased after B cell activation with Staphylococcus aureus Cowan strain 1 and phorbol myristate acetate, but not by various cytokines, including interleukin 4 (IL-4), IL-6, and gamma interferon. In situ hybridization with an antisense BL-CAM RNA probe revealed expression in B cell-rich areas in tonsil and lymph node, although the most striking hybridization was in the germinal centers. COS cells transfected with a BL-CAM expression vector were immunofluorescently stained positively with two different CD22 antibodies, each of which recognizes a different epitope. Additionally, both normal tonsil B cells and a B cell line were found to adhere to COS transfected with BL-CAM in the sense but not the antisense direction.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

Methylation blockage and other improvements to a comprehensive DNA analysis program.

A comprehensive DNA analysis computer program was described in the second special issue of Nucleic Acids Research on the applications of computers to research on nucleic acids by Stone and Potter (1). Criteria used in designing the program were user friendliness, ability to handle large DNA sequences, low storage requirement, migratability to other computers and comprehensive analysis capability. The program has been used extensively in an industrial-research environment. This paper talks about improvements to that program. These improvements include testing for methylation blockage of restriction enzyme recognition sites, homology analysis, RNA folding analysis, integration of a large DNA database (GenBank), a site specific mutagenesis analysis, a protein database and protein searching programs. The original design of the DNA analysis program using a command executive from which any analytical programs can be called, has proven to be extremely versatile in integrating both developed and outside programs to the file management system employed.

Base Sequence

Searching gene and protein sequence databases.

A large-scale effort to map and sequence the human genome is now under way. Crucial to the success of this research is a group of computer programs that analyze and compare data on molecular sequences. This article describes the classic algorithms for similarity searching and sequence alignment. Because good performance of these algorithms is critical to searching very large and growing databases, we analyze the running times of the algorithms and discuss recent improvements in this area.

Algorithms

An approach to searching protein sequences for superfamily relationships or chance similarities relevant to the molecular mimicry hypothesis: application to the basic proteins of myelin.

A rapid method for similarity searches (FASTP program) was used to identify similarities between a protein database and the human basic proteins from myelin [P2 protein and 17.2K, 18.5K, and 21.5K variants of myelin basic protein (MBP)]. From similarity scores, we concluded that none of the presently known proteins are in a family containing the MBPs. No new members were found for the lipid-binding family of which P2 is a member. Sequence similarities deemed relevant to the molecular mimicry hypothesis for virus-induced autoimmunity were identified in FASTP data with the aid of microcomputer programs. Several MBP/viral protein similarities were found that have not been reported previously. Of note because of their association with demyelinating conditions were proteins from visna and vaccinia. Similarity with visna was specific to the 21.5K and 20.2K MBPs. The most interesting new possibility for mimicry involving the P2 protein was between the influenza A NS2 protein and a sequence region of P2 thought to be neuritogenic in animals and mitogenic for lymphocytes from some patients with Guillain-Barré syndrome (GBS). This may have relevance for some cases of GBS associated with the 1976 U.S.A. swine flu vaccination program. Because FASTP reports only the best similarities between proteins, searches with FASTP may not have detected all the examples of mimicry present in the database. Searches might also be more effective if similarities could be scored on immunological rather than structural relatedness.

Animals

Functional mapping of Autographa california nuclear polyhedrosis virus genes required for late gene expression.

A plasmid containing the bacterial chloramphenicol acetyltransferase (CAT) gene under the control of an Autographa california nuclear polyhedrosis virus (AcNPV) late gene promoter was constructed. This plasmid (pL2cat) also contained the AcNPV hr5 enhancer element. Transient-expression assay experiments indicated that the late promoter was active in Spodoptera frugiperda cells cotransfected with pL2cat and AcNPV DNA but not when pL2cat was transfected alone. Low levels of CAT activity were observed in cells cotransfected with pL2cat and pIE-1 DNAs. However, CAT activity was not induced in a similar plasmid which lacked the cis-linked enhancer element, indicating that the enhancer was required for expression of the late gene. Cotransfection mapping of pPstI clones of AcNPV DNA indicated that the pPstI-G clone of viral DNA contained a factor which further stimulated late gene expression 3- to 10-fold. Transient-expression assay analysis of subclones of pPstI-G localized the trans-active factor to a 3.0-kilobase XbaI fragment. The nucleotide sequence of this fragment was determined and found to contain three potential open reading frames. A computer-assisted search of a protein database revealed no closely related proteins. One of the predicted amino acid sequences contained potential metal-binding domains similar to those found in nucleic acid-binding proteins. Subcloning and subsequent CAT assay indicated that two of the open reading frames were required for the activation of pL2cat. Nuclease S1 mapping of infected and transfected RNAs indicated that the two open reading frames were transcribed as delayed-early genes. Quantitative nuclease S1 analysis and differential DNA digestion of recovered plasmids indicated that the activation of pL2cat was not due to an increase in steady-state levels of mRNA replication of the viral DNA.

Amino Acid Sequence

Spermatocytes and round spermatids of rat testis: protein patterns.

Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.

Adolescent