PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Nucleic acid and protein sequence databases.

Nucleic acid and protein sequences contain a wealth of information of interest to molecular biologists. The advent of molecular sequence databases provides a unique opportunity for the computer analysis of all available sequences. Sequence databases serve two main functions: (i) to facilitate comparisons with newly determined sequences, and (ii) to act as a source of data for the generation and testing of hypotheses concerning molecular sequence organisation and evolution. The large amounts of sequence data now becoming available require that algorithms for database searching be fast and efficient and considerable progress is being made in this area.

Algorithms

EMBOPRO--an automatically generated protein sequence database.

For the identification of newly sequenced proteins it is necessary to have a large stock of known proteins for comparison. In this paper we present an automatically generated protein sequence database. The translation program introduced allows a periodical translation of every new release of the EMBL database. Possible errors of the translation are discussed as well as the reliability of the nucleotide sequence data, which turns out to be quite good. A comparison of our translated database with some established ones is given.

Amino Acid Sequence

A protein secondary structure database (PSS).

A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.

Amino Acid Sequence

A relational database of protein structures designed for flexible enquiries about conformation.

A relational database of protein structure has been developed to enable rapid and flexible enquiries about the occurrence of many aspects of protein architecture. The coordinates of 294 proteins from the Brookhaven Data Bank have been processed by standard computer programs to generate many additional terms that quantify aspects of protein structure. These terms include solvent accessibility, main-chain and side-chain dihedral angles, and secondary structure. In a relational database, the information is stored in tables with columns holding the different terms and rows holding the different entries for the terms. The different relational base tables store the information about the protein coordinate set, the different chains in the protein, the amino acid residues and ligands, the atomic coordinates, the salt bridges, the hydrogen bonds, the disulphide bridges and the close tertiary contacts. The database was established under ORACLE management system. Enquiries are constructed in ORACLE using SQL (structured query language) which is simple to use and alleviates the need for extensive computer programs. A single table can be searched for entries that meet various criteria, e.g. all protein solved to better than a given resolution. The power of the database occurs when several tables, or the entries in a single table, are cross-correlated. For example the dihedral angles of proline in the fourth position in an alpha-helix in high resolution structures can be rapidly obtained. The structural database provides a powerful tool to obtain empirical rules about protein conformation. This database of protein structures is part of a joint project between Birkbeck College and Leeds University to establish an integrated data resource of protein sequences and structures (ISIS) that encodes the complex patterns of residues and coordinates that define protein conformation. The entire data resource (ISIS) will provide a system to guide all areas of protein modelling including structure prediction, site-directed mutagenesis and de novo protein design. The availability of ISIS is described in the paper.

Computer Simulation

An object-oriented database for protein structure analysis.

An object-oriented database system has been developed which is being used to store protein structure data. The database can be queried using the logic programming language Prolog or the query language Daplex. Queries retrieve information by navigating through a network of objects which represent the primary, secondary and tertiary structures of proteins. Routines written in both Prolog and Daplex can integrate complex calculations with the retrieval of data from the database, and can also be stored in the database for sharing among users. Thus object-oriented databases are better suited to prototyping applications and answering complex queries about protein structure than relational databases. This system has been used to find loops of varying length and anchor positions when modelling homologous protein structures.

Amino Acid Sequence

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

Scrutineer: a computer program that flexibly seeks and describes motifs and profiles in protein sequence databases.

Scrutineer is an interactive, user-friendly program designed to search for motifs, patterns and profiles in the Swissprot, Protein Identification Resource (PIR) or SeqDb protein sequence databases. Basic capabilities include (i) searches for strings of amino acids with multiple choices at a given position; (ii) searches for strings including variable-length segments and delocalized constraints; (iii) searches over subsets of a database or particular regions within each sequence (e.g. N-terminal one-third); (iv) searches involving secondary structure predictions, physicochemical characteristics, and the like; and (v) searches using aligned sequences as targets with various optional weighting schemes. The various search criteria and hits can be combined and complex targets located. Once the data are loaded into virtual memory, all occurrences in PIR release 22.0 (3.7 x 10(6) amino acids) of a given short string of amino acids (e.g. a hexamer) are found in approximately 36 s. Scrutineer can also describe the entire database, user-specified hits, user-defined regions of sequence and all hits. The source code and accompanying manual are being freely distributed.

Algorithms

Molecular cloning and expression of a novel keratinocyte protein (psoriasis-associated fatty acid-binding protein [PA-FABP]) that is highly up-regulated in psoriatic skin and that shares similarity to fatty acid-binding proteins.

Analysis by means of two-dimensional (2D) gel electrophoresis of the protein patterns of normal and psoriatic unfractionated non-cultured keratinocytes has revealed a few low-molecular-weight proteins that are highly up-regulated in psoriatic skin. These include psoriasin; calgranulin B, also known as MRP 14, L1, or calprotectin; calgranulin A or MRP 8; and cystatin A or stefin A. Here, we have cloned and sequenced the cDNA (clone 1592) encoding a new member of this group of low-molecular-weight proteins [isoelectric focusing (IEF) SSP 3007 in the keratinocyte 2D gel protein database] that we have termed PA-FABP (psoriasis-associated fatty acid-binding protein). The deduced sequence predicted a protein with molecular weight of 15,164 daltons and a calculated pI of 6.96, values that are close to those recorded in the keratinocyte 2D gel protein database. The protein comigrated with PA-FABP as determined by 2D gel analysis of [35S]-methionine-labeled proteins expressed by transformed human amnion (AMA) cells transfected with clone 1592 using the vaccinia virus expression system and reacted with a rabbit polyclonal antibody raised against 2D gel purified PA-FABP. Structural analysis of the amino acid sequence revealed 48%, 52%, and 56% identity to known low-molecular-weight fatty acid-binding proteins belonging to the FABP family. Northern blot analysis showed that PA-FABP mRNA is indeed highly up-regulated in psoriatic keratinocytes. The transcript is present in human cell lines of epithelial and lymphoid (Molt 4) origin but cannot be detected in normal or SV40 transformed MRC-5 fibroblasts. 2D gel protein analysis of normal primary keratinocytes cultured for at least 8 d under conditions that promoted incomplete terminal differentiation [serum-free keratinocyte (SFK) medium supplemented with epidermal growth factor (EGF), pituitary extract, and 10% fetal calf serum] revealed a strong up-regulation of PA-FABP, psoriasin, calgranulins A and B, and a few other proteins that are highly expressed in psoriatic skin. The levels of these proteins exceeded by far those observed in non-cultured normal keratinocytes implying that the cultured cells have followed an altered pattern of differentiation that resembles--at least in part--that of non-cultured psoriatic keratinocytes. The implications of these results for the study of psoriasis are discussed.

Amino Acid Sequence

Secondary structure-based profiles: use of structure-conserving scoring tables in searching protein sequence databases for structural similarities.

The profile method, for detecting distantly related proteins by sequence comparison, has been extended to incorporate secondary structure information from known X-ray structures. The sequence of a known structure is aligned to sequences of other members of a given folding class. From the known structure, the secondary structure (alpha-helix, beta-strand or "other") is assigned to each position of the aligned sequences. As in the standard profile method, a position-dependent scoring table, termed a profile, is calculated from the aligned sequences. However, rather than using the standard Dayhoff mutation table in calculating the profile, we use distinct amino acid mutation tables for residues in alpha-helices, beta-strands or other secondary structures to calculate the profile. In addition, we also distinguish between internal and external residues. With this new secondary structure-based profile method, we created a profile for eight-stranded, antiparallel beta barrels of the insecticyanin folding class. It is based on the sequences of retinol-binding protein, insecticyanin and beta-lactoglobulin. Scanning the sequence database with this profile, it was possible to detect the sequence of avidin. The structure of streptavidin is known, and it appears to be distantly related to the antiparallel beta barrels. Also detected is the sequence of complement component C8, which we therefore predict to be a member of this folding class.

Amino Acid Sequence

Human genome protein function database.

A database which focuses on the normal functions of the currently-known protein products of the Human Genome was constructed. Information is stored as text, figures, tables, and diagrams. The program contains built-in functions to modify, update, categorize, hypertext, search, create reports, and establish links to other databases. The semi-automated categorization feature of the database program was used to classify these proteins in terms of biomedical functions.

Databases, Factual

Generalized protein tertiary structure recognition using associative memory Hamiltonians.

In previous papers, a method of protein tertiary structure recognition was described based on the construction of an associative memory Hamiltonian, which encoded the amino acid sequence and the C alpha co-ordinates of a set of database proteins. Using molecular dynamics with simulated annealing, the ability of the Hamiltonian to successfully recall the structure of a protein in the memory database was successfully demonstrated, as long as the total number of database proteins did not exceed a characteristic value, called the capacity of the Hamiltonian, equal to 0.5N to 0.7N, where N is the number of amino acid residues in the protein to be recalled. In this paper, we describe the development of additional methods to increase the capacity of the Hamiltonian, including use of a more complete representation of the protein backbone and the incorporation of contextual information into the Hamiltonian through the use of secondary structure prediction. In addition, we further extend the ability of associative memory models to predict the tertiary structures of proteins not present in the protein data set, by making the Hamiltonian invariant with respect to biological symmetries that represent site mutations and insertions and deletions. The ability of the Hamiltonian to generalize from homologous proteins to an unknown protein in the presence of other unrelated proteins in the data set is demonstrated.

Cytochromes

Tumorigenic poxviruses: genomic organization and DNA sequence of the telomeric region of the Shope fibroma virus genome.

Shope fibroma virus (SFV), a tumorigenic poxvirus, has a 160-kb linear double-stranded DNA genome and possesses terminal inverted repeats (TIRs) of 12.4 kb. The DNA sequence of the terminal 5.5 kb of the viral genome is presented and together with previously published sequences completes the entire sequence of the SFV TIR. The terminal 400-bp region contains no major open reading frames (ORFs) but does possess five related imperfect palindromes. The remaining 5.1 kb of the sequence contains seven tightly clustered and tandemly oriented ORFs, four larger than 100 amino acids in length (T1, T2, T4, and T5) and three smaller ORFs (T3A, T3B, and T3C). All are transcribed toward the viral hairpin and almost all possess the consensus sequence TTTTTNT near their 3' ends which has been implicated for the transcription termination of vaccinia virus early genes. Searches of the published DNA database revealed no sequences with significant homology with this region of the SFV genome but when the protein database was searched with the translation products of ORFs T1-T5 it was found that the N-terminus of the putative T4 polypeptide is closely related to the signal sequence of the hemagglutinin precursor from influenza A virus, suggesting that the T4 polypeptide may be secreted from SFV-infected cells. Examination of other SFV ORFs shows that T1 and T2 also possess signal-like hydrophobic amino acid stretches close to their N-termini. The protein database search also revealed that the putative T2 protein has significant homology to the insulin family of polypeptides. In terms of sequence repetitions, seven tandemly repeated copies of the hexanucleotide ATTGTT and three flanking regions of dyad symmetry were detected, all in ORF T3C. A search for palindromic sequences also revealed two clusters, one in ORF T3A/B and a second in ORF T2. ORF T2 harbors five short sequence domains, each of which consists of a 6-bp short palindrome and a 10- to 18-bp larger palindrome. The significance of these palindromic domains in this ORF is unclear but the coincidence of the end of one larger palindrome with the end of the translated protein sequence that has homology with the B chain of insulin suggests that the palindromes may divide the T2 protein into several functional units. The salient organizational features of the complete SFV TIR are also discussed in light of what is known about other poxviral TIRs.

Amino Acid Sequence

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing